Pith. sign in

Paper Citation Record · LEDGER

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding

As of 12 August 2026, this Paper Citation Record lists 100 of 136 outbound references and 0 inbound Pith citation observations for arXiv:2412.16158.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.16158 v2

Coverage vector

measured 100 of 136 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T10:49:08.294410Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 136 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved90
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cd6c5f04-9712-41e3-be0f-31ceb933e1b2 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:07.905038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:07.905038Z digest=sha256:e8d7eb5a77c3e279bf523bf8f8f1957d51ad58de9955a4fe869253db84d845a9

Observation 5ba4b669-ff7f-467a-a2a1-414704687d86 · outbound

This paper cites GPT-4 Technical Report.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:07.909674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:07.909674Z digest=sha256:b8e1e8e86bc26142c56e28ea6310d1e2a55a05d5d192679cb6b4d0a60e0049e8

Observation 0a5bb838-0f25-4d62-b1c5-f868ddb95507 · outbound

This paper cites MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:07.913530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:07.913530Z digest=sha256:629b3c8961ee65c0e6fa463c044408c1b0b9aa95e1d676df3fee1e54d08426f3

Observation d71830f1-b684-4fb7-ac77-5feae7eb21f5 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding The claude 3 model family: Opus, sonnet, haiku

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:07.917951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:07.917951Z digest=sha256:63bee8c2d780a934e182be184af2a2f76af3df6b430d8c3793380b241bf1e7e2

Observation 219df7c9-9f6a-4c83-aee1-bc117e7de5eb · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:07.922761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:07.922761Z digest=sha256:6acf2847c4d229dc3b5ea68515825343d20decad82241f166c795f9989d9fbee

Observation c79ec08e-c8a9-4a92-aa7a-5a39306077c5 · outbound

This paper cites Introducing our multimodal models.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Introducing our multimodal models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:07.927736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:07.927736Z digest=sha256:859201079db923d897e187db9a8b7c434de4acabe7cff009080579f476a88e3d

Observation 3084111e-c871-446c-8d82-fbb750e5f9cb · outbound

This paper cites Vqa-med: Overview of the medical visual question answering task at imageclef 2019.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Vqa-med: Overview of the medical visual question answering task at imageclef 2019

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:07.932167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:07.932167Z digest=sha256:4e21c8d01ce092b4b6ad4a4b58ee0fee6fca433b5b5a91362984530d922807af

Observation 4a365c0e-44d3-4631-8276-f6568993093e · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding PaliGemma: A versatile 3B VLM for transfer

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:07.936839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:07.936839Z digest=sha256:3f7c40082d245098376ed4c86f8bd3bbac44cdd943f1f33e39c863e1e07f1823

Observation f20bedfa-9a3a-4066-a5d7-cfb43aac0a9d · outbound

This paper cites Scene text visual question answering.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Scene text visual question answering

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:07.941064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:07.941064Z digest=sha256:9eaf14af643d53e401c7735779f59605bc6269e1e131285e14154cccf61a72df

Observation 3764b123-8dc2-4337-b946-525af5df5b4b · outbound

This paper cites Coyo-700m: Image-text pair dataset.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Coyo-700m: Image-text pair dataset

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:07.944784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:07.944784Z digest=sha256:8bdf1195a0d7d4a8566ea6ca1f8cc649570f6239c0b853848f1d02d9d0e3e7eb

Observation 7496fd43-beee-455c-b42b-bcf3dacea6c2 · outbound

This paper cites InternLM2 Technical Report.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding InternLM2 Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:07.948584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:07.948584Z digest=sha256:80f01269dbf81aeb33b43e0f89014597e4c766119bd53ae297a25ab08b3e91e7

Observation 6ab6b2d0-258b-4220-9d72-5ac19bca1a47 · outbound

This paper cites An augmented benchmark dataset for geometric question answering through dual parallel text encoding.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding An augmented benchmark dataset for geometric question answering through dual parallel text encoding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:07.953251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:07.953251Z digest=sha256:533432fc81a8e7763797a07ca83e1443562178c305624d7ddf4da56bcbbd4aa2

Observation 4ebc1937-cb48-4f23-8e8a-2ba54d57d6e3 · outbound

This paper cites Textocr-gpt4v.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Textocr-gpt4v

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:07.956769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:07.956769Z digest=sha256:74349a07df5dca3246f50ec78524a30401d9e06321a117852bbd26d9ed72c3ef

Observation 553352aa-705f-41c3-a5af-9da1daecb769 · outbound

This paper cites MapQA: A Dataset for Question Answering on Choropleth Maps.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding MapQA: A Dataset for Question Answering on Choropleth Maps

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:07.960195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:07.960195Z digest=sha256:3baffade29dfa60233c778770f78d413b55c512a74807d19c2eb9b1216ba608e

Observation 2903dba2-41f0-4bf7-b35d-9574f3f3fa5d · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:07.964031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:07.964031Z digest=sha256:c784653d8bf3f72300f579217ad244a3b824516c08ce263d05c3a9dea9d7687c

Observation a08d26df-b966-4971-bc91-a8c4295d12c0 · outbound

This paper cites UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:07.967954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:07.967954Z digest=sha256:044efdf6b8ef402e2b84f2a184bf4f6bc7031eebf2bb3ce007b07380545f2e6c

Observation 4d90f477-d841-46ea-a979-e7e2bc1a52c2 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:07.971952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:07.971952Z digest=sha256:e227618b172c8904302c3519deb4367a8cbebb434d3c4492d235129fc40a87cf

Observation 3b1b00db-8e89-435d-aa31-4194798f1758 · outbound

This paper cites SOLO: A Single Transformer for Scalable Vision-Language Modeling.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding SOLO: A Single Transformer for Scalable Vision-Language Modeling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:07.975874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:07.975874Z digest=sha256:1c39d7ba0f57bd05e5a253d5c50fb95ea7801a42891b79276e47487c3fe17759

Observation bb26d3d4-69e5-474e-9828-6aec5db85a51 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:07.980271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:07.980271Z digest=sha256:686ffa8715e4a2080a27bdea6fcffa5ca9ab8cbbdc3a8333125968a5ecabd1ca

Observation c9b53d88-4422-42c2-a76f-5ee0b18f2c88 · outbound

This paper cites Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:07.984241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:07.984241Z digest=sha256:cc16b15723f48003489a4d2bf596abbe83f0d7249723c12ea3f81675a89e9565

Observation 16c46b68-bfd5-4e09-91ae-f253d8fe9984 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:07.988105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:07.988105Z digest=sha256:75b97fb1b16d442a1b747d875c63673df08bf0b26a389f830be6d39754386aab

Observation 2be1e6d9-7317-492a-96cb-b4d5d3814ce7 · outbound

This paper cites Complicated Table Structure Recognition.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Complicated Table Structure Recognition

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:07.992386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:07.992386Z digest=sha256:f5d0cbbb9c78cf73f8b0f27d83d6d2d5de1bc85eef6236c92dd5987e4cd8ef28

Observation 950614fd-409f-4756-900a-3ae088d9d0ce · outbound

This paper cites Icdar2019 robust read- ing challenge on arbitrary-shaped text-rrc-art.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Icdar2019 robust read- ing challenge on arbitrary-shaped text-rrc-art

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:07.996407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:07.996407Z digest=sha256:c213f7c2e9b4eb030397b42c0d16216a4cc4f81ed3f1d98c7aa24dbdcc81fa76

Observation 6e3b402d-5f3c-4763-b0f0-70d90ec52b27 · outbound

This paper cites Simple and Effective Multi-Paragraph Reading Comprehension.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Simple and Effective Multi-Paragraph Reading Comprehension

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.000495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.000495Z digest=sha256:5d2084cd75279481a06d3732468287c932d1a8cb6ebe1c87566b6fab0abc3997

Observation 70e0a205-5cea-4259-94c7-09066d0dad3e · outbound

This paper cites Deep visual template-free form parsing.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Deep visual template-free form parsing

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.004724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.004724Z digest=sha256:b5c64db89eab4968215f9461285aad8b46ab54d75766b17e777f7250aee702d6

Observation 09609b9d-e3e4-4981-bc3b-3ab8bd6810fa · outbound

This paper cites Unveiling Encoder-Free Vision-Language Models.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Unveiling Encoder-Free Vision-Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.009009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.009009Z digest=sha256:92d3370c0687fc505364dbc8293880f2d71aa22b3bbc3b6f7654b3c1b3020525

Observation ec215a64-3e95-45c2-9d6d-ca8bca9819be · outbound

This paper cites Compressing visual- linguistic model via knowledge distillation.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Compressing visual- linguistic model via knowledge distillation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.013547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.013547Z digest=sha256:6b5531f161b8ffccf7254d99e284cf3fe8671344c8db568389a25feda81c5b1d

Observation 27dc2e1d-83fc-48c1-8f66-d62f52bf04a6 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.017524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.017524Z digest=sha256:ce001176443bef75f475ceb532262299e73f6c6f22f5a8a98d3d7b549f1cff43

Observation efe9e84f-84b3-4cce-9b1d-4a249d151906 · outbound

This paper cites Making the v in vqa matter: El- evating the role of image understanding in visual ques- tion answering.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Making the v in vqa matter: El- evating the role of image understanding in visual ques- tion answering

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.021980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.021980Z digest=sha256:076d08b22744e77f10dfcf0d60191135f503465376e6dc30ffc59ca20f945dbe

Observation 08586c5c-0b64-44b9-8ad0-57e846085f2d · outbound

This paper cites Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.025990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.025990Z digest=sha256:5aa60456c057db77d85f813fa43a3a3e865abdb239f91721b20666d10342476c

Observation bd9afb9a-37d4-4826-b452-1dd5b6af07a7 · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.030191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.030191Z digest=sha256:853618eaa57b8a81596f46c4eb2a0fa7a569f0ae63efacc700cf7b79e6bedfa0

Observation c35905a8-7175-4d19-b240-07edb1c118f7 · outbound

This paper cites Eaten: Entity-aware attention for sin- gle shot visual text extraction.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Eaten: Entity-aware attention for sin- gle shot visual text extraction

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.035321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.035321Z digest=sha256:59cd87dd853122d9ce7e4b2f84750d163979ccad9b0e30c7f30c932c14fe9e52

Observation 91afe247-306f-45d8-aeb2-ef2c6d29aee4 · outbound

This paper cites Don’t stop pretraining: Adapt language models to domains and tasks.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Don’t stop pretraining: Adapt language models to domains and tasks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.039338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.039338Z digest=sha256:7353019012013353ea93e90ed9f1562b56df6fd13b3a491701b2398322208d09

Observation 7b16ef0f-983d-4d5c-bf77-819bfe5c5843 · outbound

This paper cites Icpr2018 contest on robust read- ing for multi-type web images.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Icpr2018 contest on robust read- ing for multi-type web images

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.043425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.043425Z digest=sha256:2e11e410e1e1f108f7e46a1aa907d4ba24f11a48951a7215ac4c8329b1581e2b

Observation 318b72b3-a1c7-4c86-a704-5608702eef7b · outbound

This paper cites PathVQA: 30000+ Questions for Medical Visual Question Answering.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding PathVQA: 30000+ Questions for Medical Visual Question Answering

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.047377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.047377Z digest=sha256:9d429be340e9aef839b88af09608b35cc931e9602c2b5a8bd48366d3d52a7c42

Observation db154994-c89b-4c6e-9c64-cd612c122987 · outbound

This paper cites Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.051867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.051867Z digest=sha256:6496fe26077ebf54da6a521fb28595bba0edef70be63b82888e0e5b655e84e45

Observation 151511c6-132e-47b9-a5c9-77aed2643d2e · outbound

This paper cites mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.055871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.055871Z digest=sha256:42d2e164ac92a1f7752fb65dde86028527c7d765baf70be13a05f774649adf75

Observation ebf20f45-12ad-4ab3-87a8-d9b266e788b7 · outbound

This paper cites Medical-diff-vqa: a large-scale medical dataset for difference visual question answering on chest x-ray images, 2023.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Medical-diff-vqa: a large-scale medical dataset for difference visual question answering on chest x-ray images, 2023

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.059932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.059932Z digest=sha256:2c258e4268fbcde616bc6221cf0f98c2757a87e3e02e53330600a3dd2be7827e

Observation 68dde9d3-8582-44e4-99fe-dfb9e84cf9a7 · outbound

This paper cites Visual program distillation: Distill- ing tools and programmatic reasoning into vision-language models.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Visual program distillation: Distill- ing tools and programmatic reasoning into vision-language models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.063333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.063333Z digest=sha256:47d2da395ca1af1f02666da91c570b68d8d72d901b569f430b7ccd840a4e7d8d

Observation 7796b487-7650-45e4-96c9-996c021f5c68 · outbound

This paper cites Movienet: A holistic dataset for movie under- standing.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Movienet: A holistic dataset for movie under- standing

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.066809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.066809Z digest=sha256:08b725fd44e11341984c8a032388024a9357ae890674590dab1480bfff2298da

Observation 3b13cf06-d12d-429b-8c5a-50b5583407e0 · outbound

This paper cites Hires-llava: Restor- ing fragmentation input in high-resolution large vision- language models.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Hires-llava: Restor- ing fragmentation input in high-resolution large vision- language models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.070041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.070041Z digest=sha256:c97ac2df13e99869a38dabe92b588f132616d05fb90537699729c038b81fbcb8

Observation 30cc1899-fe7f-4b2e-a7a2-03cd107798c6 · outbound

This paper cites Icdar2019 com- petition on scanned receipt ocr and information extraction.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Icdar2019 com- petition on scanned receipt ocr and information extraction

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.073191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.073191Z digest=sha256:bec8dabd84f62399163abab79c6747a8d1894fb09268b4149c4494822dca6e31

Observation cbe1ef73-4e59-4946-9d6a-b0194e53a2bc · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.076253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.076253Z digest=sha256:7c1e70ff5611a22338c2175120eb8da5afaa9155e83befb535cfb063e6798144

Observation e6984120-16ce-4e69-a01b-f1b273bf083a · outbound

This paper cites Egotaskqa: Understanding human tasks in ego- centric videos.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Egotaskqa: Understanding human tasks in ego- centric videos

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.079486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.079486Z digest=sha256:c91caf63b88398d678c7804d93386b081ca10e2420e48495f302c57f7cce61a4

Observation cfa1c8fa-d7c2-4353-a39a-2bf1d5c572d1 · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Clevr: A diagnostic dataset for compositional language and elementary visual reasoning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.082744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.082744Z digest=sha256:68391e4afc318c217303be4c9304528e378e7296e8372348cdb35807f2418fed

Observation 28620bdf-2340-4720-9ac1-79731b88a07e · outbound

This paper cites Dvqa: Understanding data visualizations via ques- tion answering.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Dvqa: Understanding data visualizations via ques- tion answering

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.085944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.085944Z digest=sha256:b1d92bb0f3f5af67ea06d617d84b6bec0617c19daf1fa54d43c9795ccd31c11c

Observation 03c1d237-a189-48b7-8da5-92294ca0421a · outbound

This paper cites FigureQA: An Annotated Figure Dataset for Visual Reasoning.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding FigureQA: An Annotated Figure Dataset for Visual Reasoning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.089410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.089410Z digest=sha256:8bcf8909a0d87e8492b3e213416d0e41fb835e66c12ddf8f7bcd24cf0edfe775

Observation c0d0f8f9-24b0-4182-a9ae-f39dc7c599cf · outbound

This paper cites Chart-to-Text: A Large-Scale Benchmark for Chart Summarization.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Chart-to-Text: A Large-Scale Benchmark for Chart Summarization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.093786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.093786Z digest=sha256:27dac8dcaf8801dcaba2755a119c3b561a6d68236a65183fe56f2e3351df7315

Observation 80283ccd-36bc-4554-87e9-a114ef6a3331 · outbound

This paper cites Referitgame: Referring to objects in pho- tographs of natural scenes.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Referitgame: Referring to objects in pho- tographs of natural scenes

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.097976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.097976Z digest=sha256:d67a1c917657d3fc82db4ee77a46ccc2d98d451c651d520964e996ae39ae6830

Observation 0a0f73ef-57e7-4010-8415-a8e8d1f296c7 · outbound

This paper cites A diagram is worth a dozen images.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding A diagram is worth a dozen images

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.102029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.102029Z digest=sha256:5db8789af48ed261e783e5642237d8d9a94635dde3e15f35762e270fd5f710fd

Observation c279f269-b582-41e2-b927-1c78e92357e1 · outbound

This paper cites Are you smarter than a sixth grader? textbook question answer- ing for multimodal machine comprehension.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Are you smarter than a sixth grader? textbook question answer- ing for multimodal machine comprehension

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.105666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.105666Z digest=sha256:f85b5120272259342eb3993419c35b58dc7fa388c91a907647faf2767f7b9fa4

Observation 7134f05b-beec-48a2-a185-691bae37b15e · outbound

This paper cites Visual in- formation extraction in the wild: practical dataset and end- to-end solution.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Visual in- formation extraction in the wild: practical dataset and end- to-end solution

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.109365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.109365Z digest=sha256:194f7c94839f654fdba9a48158a3b174870b34c217e97c0c0f863a0e66a809db

Observation eaf396a2-b074-45ee-84c9-f8a36123baee · outbound

This paper cites Laion-gpt4v dataset.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Laion-gpt4v dataset

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.113186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.113186Z digest=sha256:de2aa4a94aeb84a6bfa16a98058fa1952fc7ca9f64f7de3bc324320b3522adb2

Observation 18371d88-e852-4c69-aaf9-b2848e659812 · outbound

This paper cites A dataset of clinically generated visual questions and answers about radiology images.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding A dataset of clinically generated visual questions and answers about radiology images

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.116964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.116964Z digest=sha256:8cab109d7fcb63caf445ba157bd33b9d31f999f43128038dba19bb28a9feb2c5

Observation 5346708f-c4c5-40d2-a45c-d12753262f82 · outbound

This paper cites Viquae, a dataset for knowledge-based visual question answering about named entities.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Viquae, a dataset for knowledge-based visual question answering about named entities

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.120859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.120859Z digest=sha256:98b427ef76f25ea6dd313eb02a8ee87fa2a14188d5f295ca841db5ed5f83e761

Observation 7ed100a3-649d-4370-a71d-7b87ffceba10 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.125301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.125301Z digest=sha256:35cd1251741a92abeed906bb566e070b3e5c5e5a3b14c8568ce47a9149f25288

Observation 4151d93a-f332-46eb-9cb1-a6363bf78675 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.129667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.129667Z digest=sha256:5d3078c44f8ffaad0d4b6f8e8df33d0a085540441df20a4893e68f7c4c933c03

Observation cbe43d26-fab1-471d-9190-0aa7d4622402 · outbound

This paper cites Chemvlm: Exploring the power of multimodal large language models in chemistry area.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Chemvlm: Exploring the power of multimodal large language models in chemistry area

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.133784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.133784Z digest=sha256:6ec65dbfcd7b2a8f2f6f170b4f1005714c2eb9e65dd5d323bc1a4628fa52d3cd

Observation 2a2d965d-ccce-46bb-a5ad-b58730986555 · outbound

This paper cites Mvbench: A comprehensive multi-modal video under- standing benchmark.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Mvbench: A comprehensive multi-modal video under- standing benchmark

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.137546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.137546Z digest=sha256:a03234fb18956c712c84facb54fdc8507bc68892e006fab4201221c4cce72ad8

Observation e671cbb1-0f09-43cf-a571-468d9ffa7b2a · outbound

This paper cites Evaluating object hallucination in large vision-language models.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Evaluating object hallucination in large vision-language models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.141261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.141261Z digest=sha256:5398639de1bd4e150a1951451e1dc88d89c8819687217d4276d9c216113ae6f4

Observation c4ac0e4e-d25d-4e78-946b-ab047510633d · outbound

This paper cites Super-clevr: A virtual benchmark to diagnose do- main robustness in visual reasoning.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Super-clevr: A virtual benchmark to diagnose do- main robustness in visual reasoning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.145017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.145017Z digest=sha256:762af05955c7493708477b420aa79b86199997f8cfb62ee840ec849f006e0511

Observation 010c99f5-cc3b-4171-80de-94cc76d4e775 · outbound

This paper cites Mon- key: Image resolution and text label are important things for large multi-modal models.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Mon- key: Image resolution and text label are important things for large multi-modal models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.148587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.148587Z digest=sha256:d6f049ce8d18b06df3721d0680b2dc9de0646410ba28d2ee9777d60edfea5b9c

Observation 9cf66104-80da-49c7-b15d-24469907c359 · outbound

This paper cites Vila: On pre-training for vi- sual language models.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Vila: On pre-training for vi- sual language models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.152181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.152181Z digest=sha256:fe8983952369d25fafc994e1a9cab3ab37190bd4a827dd56193f732b722994bc

Observation bda4c47d-4915-4028-8d63-e6bd1a850253 · outbound

This paper cites Microsoft coco: Common objects in context.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Microsoft coco: Common objects in context

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.155842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.155842Z digest=sha256:117e62aff72dfa96c005e38183fd716189ee762f89fe85972b5d3d47a79e3324

Observation 5fc4dfd6-c238-4b8c-86b9-44ddf1e760a5 · outbound

This paper cites Slake: A semantically-labeled knowledge- enhanced dataset for medical visual question answering.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Slake: A semantically-labeled knowledge- enhanced dataset for medical visual question answering

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.159415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.159415Z digest=sha256:8e129cb6c87ec7b063cbb5c22a82a5d6e0bd891f6ecefd38507309b30e0059d6

Observation 1a2ccfc3-e863-4ee6-99ce-b02541c4840e · outbound

This paper cites Casia online and offline chinese handwriting databases.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Casia online and offline chinese handwriting databases

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.162971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.162971Z digest=sha256:27af0e717651ef9714c37c20640c02413a0b00758ef063b8669bd06b6fd9ae97

Observation 69b4cebf-8ad5-4755-94e7-8ea27cb25f13 · outbound

This paper cites Visual spa- tial reasoning.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Visual spa- tial reasoning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.166664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.166664Z digest=sha256:cb8aa1bb193b05a94ec82dd743f8022b62f57d579786d17bef2ce0b051577249

Observation 9987486b-b6cd-4533-b172-f59f88e1e569 · outbound

This paper cites Mitigating hallucination in large multi-modal models via robust instruction tuning.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Mitigating hallucination in large multi-modal models via robust instruction tuning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.170195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.170195Z digest=sha256:fd34d7840743a9c643c0287e26e76fdce6eba63baabbfdbe579ded6f0b163c2a

Observation 8e398f10-bccb-4d5f-82ac-6275bcb36b51 · outbound

This paper cites MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.173803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.173803Z digest=sha256:3634b593928217b52853bfbdfc5246f2fba5aedccd599002c83682b0a46354f0

Observation d04af017-bdcf-41d4-b039-4eceb8be0b54 · outbound

This paper cites Improved baselines with visual instruction tuning.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Improved baselines with visual instruction tuning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.177610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.177610Z digest=sha256:3f5a8a19ab62e64651d0a21d6f5e1e908153a1d6f0617653c6aa5ea7207e1a4c

Observation 0cbafedf-256f-41ba-8c7b-458a19ad09aa · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Llava-next: Im- proved reasoning, ocr, and world knowledge

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.181331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.181331Z digest=sha256:e24887dfd200306fa72f0f3d4391b4d6c30c0dfac0ee7efdae37bfb326182c9b

Observation f119e9d5-9757-463c-b564-0e4f31c35e43 · outbound

This paper cites Visual instruction tuning.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Visual instruction tuning

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.185186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.185186Z digest=sha256:f7a311f1ed8addb1cefda98ce75b2b6bc43616e96c55ccd5227662bf2141e86b

Observation 4da4d5f6-e510-4afe-8194-14131a1f0c54 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.188971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.188971Z digest=sha256:5252759f8d0d5e34c527c54d3e34313fa2c6ce2dd92d0c93352e37ce70988b8a

Observation db7d8205-78c7-41da-91a1-0ced09a51ec2 · outbound

This paper cites Mmbench: Is your multi- modal model an all-around player? In ECCV, pages 216–.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Mmbench: Is your multi- modal model an all-around player? In ECCV, pages 216–

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.192881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.192881Z digest=sha256:7be04a1e3a87093aebf273848c8c8829cc74ad500eeb91970b2da22ebab854f4

Observation 9c9b2023-4620-4bd9-958a-e7b2eaacdf17 · outbound

This paper cites MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.199448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.199448Z digest=sha256:31d7099083fc986b80febe95b120b85beb06c98f1e61dc7bbe4836cf5bc0d7e2

Observation 224d0ffb-4d93-4b43-bedf-6d21db74f05f · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.202823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.202823Z digest=sha256:a33f8a3b898d234954be8ab1c9022998e4b7e908c2cccb43a20c027e1698c2ab

Observation 95b58660-03a8-4f0c-b686-16b5209b814b · outbound

This paper cites Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.206505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.206505Z digest=sha256:4d6f45cacaaea3f9273c5cc6fc44fd27cab02572b3af4755d563e422a706291d

Observation 40d830c6-af17-45df-b644-0eff71ae5228 · outbound

This paper cites Learn to explain: Multimodal reason- ing via thought chains for science question answering.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Learn to explain: Multimodal reason- ing via thought chains for science question answering

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.210595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.210595Z digest=sha256:8472c8c9f1b25a8da3bb770e790f49b21f8a680c211df492bf6f8ba7f50602d7

Observation f716cc29-6023-4004-ac70-5c53ffd6d5bd · outbound

This paper cites Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.214747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.214747Z digest=sha256:31b4670b404574bf00d9be812ef3aa6415452a3da6a6045fa723d6be4222bc36

Observation b48b5cc7-0962-44ff-ab5f-e4054d78f056 · outbound

This paper cites VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.218792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.218792Z digest=sha256:310d0b2234fda9e9c3fe323ef9565d849ac9a7b9bfbd471602542d5320a662f7

Observation 2a39fe23-80c4-4757-9ee0-5017b9b623e0 · outbound

This paper cites Deepart: Learning joint representations of visual arts.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Deepart: Learning joint representations of visual arts

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.223013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.223013Z digest=sha256:abd49b7fcd2a5159ecfa2ce82cb1ef29b2143b59efc1755318d7cd840cd197c0

Observation 2982c736-06af-4659-b210-ceb0321ed8fc · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.226714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.226714Z digest=sha256:43bb5b2d89dc2c1348b198fa1b4ef84c346f771edb68e19f8c4b3fca7a4b17cf

Observation e44e2131-99da-4a84-aa6a-7e5a6b5fbb1d · outbound

This paper cites The iam-database: an english sentence database for offline handwriting recognition.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding The iam-database: an english sentence database for offline handwriting recognition

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.230195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.230195Z digest=sha256:2811f1f8fc167cc9c93603bf22d7939fc0124ce8dcf10292a1f6dcdf3a9b85c9

Observation 23bbdee0-b510-4ad9-8d1d-5ce6f722210a · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.233888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.233888Z digest=sha256:fa59d1371c5652938ac6e8fafc66911510f4101c32ca4aded8b9550922913ce9

Observation 65e13ded-6c39-4d97-9faa-2e78de22ff4f · outbound

This paper cites UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.238683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.238683Z digest=sha256:104dfe769249c07cb01ad11c150074831b76de414a267a3207fce509429e902f

Observation af3059cb-c264-425a-b98c-2fef2c7f4691 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Docvqa: A dataset for vqa on document images

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.243283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.243283Z digest=sha256:665a21c4c24eb5856f4501817744e8a01b216aaa808d81e6434a9c1b72c62080

Observation b5dd9c2f-9d21-4bac-a888-6e7203069433 · outbound

This paper cites Infographicvqa.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Infographicvqa

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.247064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.247064Z digest=sha256:77cf674a67aa8fee66e7122a8d197911f2ce56ebcbf1dffb80787bd4af5555c9

Observation fcda2877-eaa5-4991-bc3a-709b18ed9086 · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.250605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.250605Z digest=sha256:45a5682c40a9af9c279ae067651eb82a17f59ff08e311b7d969c0c5ffa2f2afd

Observation 3f5dba71-6c98-455c-9d5f-7e7985cf1ea8 · outbound

This paper cites Plotqa: Reasoning over scientific plots.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Plotqa: Reasoning over scientific plots

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.254401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.254401Z digest=sha256:7dd8a5baca5c6a6a95df68497b77d956ca44280e6e258476787c550d20709024

Observation 90f291d5-e8bf-4899-9895-3c61536d3e8d · outbound

This paper cites Localized sym- bolic knowledge distillation for visual commonsense mod- els.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Localized sym- bolic knowledge distillation for visual commonsense mod- els

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:09.618290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:49:08.257990Z digest=sha256:aee631115e7ff5022616dd6033189f6366ca684e9dc3c3bee0f0576ff3e607e2

Observation 6df6201a-e410-43c9-b86e-7f74c9d7bb18 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.261497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.261497Z digest=sha256:162989fead46da8752ed068883f3cf957ad25b00b5f6d4cd5a4bcf4209972cd1

Observation 7fa6a765-a7d3-421e-8900-87353103a565 · outbound

This paper cites Efficiently scaling trans- former inference.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Efficiently scaling trans- former inference

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:09.607182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:49:08.265314Z digest=sha256:b5afeb19dd636e5a2788337efd3c84cb5c30fd842f0c05f9f23a64fe8725a114

Observation 81d33351-6e77-4c06-afa7-296895af5ae3 · outbound

This paper cites A dataset for movie description.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding A dataset for movie description

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:09.595623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:49:08.268871Z digest=sha256:b52e200a1329189bac93d7ae5862930f83ba4c20f75b71b9e1a5bb66dd6e2549

Observation 317103cd-3ba8-4db8-817e-448820126bac · outbound

This paper cites Laion-5b: An open large-scale dataset for train- ing next generation image-text models.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Laion-5b: An open large-scale dataset for train- ing next generation image-text models

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:09.584080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:49:08.272406Z digest=sha256:82f0ce5ab28453fdb723e6beee460429ac24598468f209930b32573eb8ea7a43

Observation 304cc45e-0549-4418-ae66-886a56839fbd · outbound

This paper cites Laion coco: 600m synthetic captions from laion2b-en.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Laion coco: 600m synthetic captions from laion2b-en

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:09.571488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:49:08.276344Z digest=sha256:5b712c0133a08bd985bc563911cf26a65cf7d2001ee6cf971b4f5e3281bfbe02

Observation ee3c7c9c-d14c-4a16-a9a2-810b9ada4996 · outbound

This paper cites Solving geometry problems: Combining text and diagram interpretation.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Solving geometry problems: Combining text and diagram interpretation

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:09.560158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:49:08.279893Z digest=sha256:cb577380ba0026869aaae69c1fc653a6dcb9cdb907230af871d2302a5e354776

Observation 58617fa4-785f-4792-b157-0b563d1883a8 · outbound

This paper cites Kvqa: Knowledge-aware visual question answering.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Kvqa: Knowledge-aware visual question answering

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:09.549517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:49:08.283532Z digest=sha256:471a57d95141c52b67b0f94e6b7da35dc30816eddc0fe973c032473a78eeaf4d

Observation a3859706-8d68-47a4-83d7-113ab49bce50 · outbound

This paper cites Ntu rgb+d: A large scale dataset for 3d human activity anal- ysis.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Ntu rgb+d: A large scale dataset for 3d human activity anal- ysis

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:09.538809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:49:08.287136Z digest=sha256:fe7f46bb181ae0018c3d4de56cc725e15776f377fd1e1db6941ec53fb0c08692

Observation 72537121-778e-4cd3-bf3c-1b233fa6f39a · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Objects365: A large-scale, high-quality dataset for object detection

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:09.527193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:49:08.290745Z digest=sha256:afc31a36f3fe3b5a6ea50c9cd5b5f9e43cfbfaa66389e15ce732bde3b79ffefe

Observation cd7ca002-83ff-4c81-9a90-310725577f72 · outbound

This paper cites Towards vqa models that can read.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Towards vqa models that can read

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:09.515161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T10:49:08.294410Z digest=sha256:3b25d956c6fe2ccd278b5df16944117b748f4d5eaef9dbc21f708227a56ed950

Pith citing papers

No inbound Pith citation observations are available.