Pith. sign in

Paper Citation Record · LEDGER

LLaVA-OneVision: Easy Visual Task Transfer

As of 11 August 2026, this Paper Citation Record lists 100 of 179 outbound references and 100 inbound Pith citation observations for arXiv:2408.03326.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.03326 v3

Coverage vector

measured 100 of 179 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T14:23:49.412830Z

measured 200 of 200 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 100 of 917 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:55:10.312257Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 179 outbound references displayed

  • verified exact16
  • verified fuzzy79
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

29
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 80aaa1d5-a0d1-4057-b0fb-c768eb867b2d · outbound

This paper cites Tallyqa: Answering complex counting questions.

LLaVA-OneVision: Easy Visual Task Transfer Tallyqa: Answering complex counting questions

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.792224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:2a15a052c854bad901b8cf1a28266493f8462e74aa52debcca4a8955fc4db883

Observation 8299c215-5244-4383-ad68-54ec3ce42dac · outbound

This paper cites Mathqa: Towards interpretable math word problem solving with operation-based formalisms.

LLaVA-OneVision: Easy Visual Task Transfer Mathqa: Towards interpretable math word problem solving with operation-based formalisms

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.796285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:596d1ba5ff8ae2ec239d609e47759c1500fb9f996c2b183291527a49426b3807

Observation 3980e647-3419-4456-88e6-1200d3f2b3f2 · outbound

This paper cites Claude-3.5.

LLaVA-OneVision: Easy Visual Task Transfer Claude-3.5

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.797959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:3f66580a42813de685f2a8b27904baf9795413e36a074604ab2b900e087a136f

Observation 49780143-6ee2-445d-9afd-48c018046913 · outbound

This paper cites Vqa: Visual question answering.

LLaVA-OneVision: Easy Visual Task Transfer Vqa: Visual question answering

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.799980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:b4f9b96c3e8ac148aa5dcdd42733262ea8676c7aabdba8105c05d1738271110c

Observation 939a577e-2304-43bd-a4e1-eb3fe5aba008 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

LLaVA-OneVision: Easy Visual Task Transfer Scanqa: 3d question answering for spatial scene understanding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.801802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:2699abbf612b08822f9574dea66dabeb6032e15ea8b50c940318fcb43fa02ee9

Observation fff6ee15-b520-4502-bea8-745273ed97a9 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

LLaVA-OneVision: Easy Visual Task Transfer Scanqa: 3d question answering for spatial scene understanding

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.803856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:012fe0154b959c3d41343f3fbebeaa4e313bde972a4a26e782d0dc8b5fe3b595

Observation 09b9f03d-3afb-4f80-b33b-2b88bcce453f · outbound

This paper cites Vision datasets: A benchmark for vision-based industrial inspection.

LLaVA-OneVision: Easy Visual Task Transfer Vision datasets: A benchmark for vision-based industrial inspection

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.805810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:ce50303dfab264ffe6d7c62a4331effc9416be717328cf8146b25fa38e01bb22

Observation 36401b18-0905-4bb6-a6b7-7c5f545b0bb7 · outbound

This paper cites Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond.

LLaVA-OneVision: Easy Visual Task Transfer Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.807667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:12eae689ff5c8425bdbf7ffb5331382cdcb814a3d07eb605c0714664920634fb

Observation 4cc965fd-d81a-4718-b465-ca12e151f77f · outbound

This paper cites Visual question answering on image sets.

LLaVA-OneVision: Easy Visual Task Transfer Visual question answering on image sets

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.809561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:9b9eb3ca29396600cc97fe0b20e0a098794d564be662772aaee264b94ab0bf3b

Observation 9977947e-a56d-4700-b6e6-ed7d5c078497 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

LLaVA-OneVision: Easy Visual Task Transfer PaliGemma: A versatile 3B VLM for transfer

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:21.987351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:5382afd10728583877b264666a4ab796635f0dc87fde37d1d8db65cd9267efd8

Observation 57fae0da-a23c-4282-8bc6-11e1d162d150 · outbound

This paper cites Scene text visual question answering.

LLaVA-OneVision: Easy Visual Task Transfer Scene text visual question answering

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.811541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:d1769feb524989154ae5269b7ad9d6061ad5bdec33d65af9f7bfa3a2ca663952

Observation b11632f6-694c-4c6d-84ad-1067a68c4a81 · outbound

This paper cites Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom.

LLaVA-OneVision: Easy Visual Task Transfer Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.813209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:f3b8cc22cd2f2284d48b4852fd7f29cfb1025591c9a2bd276f0615818103882e

Observation ba37d1e6-51c1-4e9d-a3ce-3614b774efa3 · outbound

This paper cites Textocr-gpt4v.

LLaVA-OneVision: Easy Visual Task Transfer Textocr-gpt4v

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.814830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:2d108106a4bc6d921257b2038a1c3b2de08b12b986bf13f5dd760aae9fa90993

Observation a19211ea-8f19-46ac-b957-f00b74c18c02 · outbound

This paper cites Mapqa: A dataset for question answering on choropleth maps.

LLaVA-OneVision: Easy Visual Task Transfer Mapqa: A dataset for question answering on choropleth maps

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.816854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:8d1f66431f5f6d02693f611871ba935ad46e9e8417023f9b2d92ba7a53eb2c28

Observation e467d038-9853-47e7-b9ba-17e74180245d · outbound

This paper cites WebQA: Multihop and Multimodal QA.

LLaVA-OneVision: Easy Visual Task Transfer WebQA: Multihop and Multimodal QA

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:23:49.561566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:06a63ae4c425f6c9151895815424e11441a04ef3ab41d6af0b7627fdfc12cb37

Observation 19822b22-57b7-44b1-9d38-ca6665b1f39d · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

LLaVA-OneVision: Easy Visual Task Transfer ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:66d0c0a304476ad50d86ba96390bfca6d54874d169950726acd2a0f3905c0fd8

Observation 62947ea0-eb12-401e-b3c2-43811352c8d0 · outbound

This paper cites Unigeo: Unifying geometry logical reasoning via reformulating mathematical expression.

LLaVA-OneVision: Easy Visual Task Transfer Unigeo: Unifying geometry logical reasoning via reformulating mathematical expression

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.818695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:4427c60cfa7f807f52fdce4018332d4c995ca416d3403b39de422693232dfe28

Observation 244068af-5d7e-4930-9116-5ed59e936f3d · outbound

This paper cites Xing, and Liang Lin.

LLaVA-OneVision: Easy Visual Task Transfer Xing, and Liang Lin

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.820565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:1278b09324bfe1138cba7a934db7b5263b36bbc0db3b9a4519c77f0067ff4b6c

Observation 8babf22d-2ce6-4233-9f01-2eb45017b6cb · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

LLaVA-OneVision: Easy Visual Task Transfer Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:41:44.612219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:ac44354c04107751588e2a1553c59f777c23cab201c9387d1dc653fdb50262a4

Observation 94c68049-8bc0-4ff0-94e3-41252d02cebd · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

LLaVA-OneVision: Easy Visual Task Transfer ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:08:13.182567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:b347573952d497775dde96e13a3f60eeab0ff5580cccac9928b0e3b2df258529

Observation 13c3cee9-ad9c-4bf5-9fc0-a5837391ad38 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

LLaVA-OneVision: Easy Visual Task Transfer ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:23:49.613523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:2886db334dc4871998ce8a4c23df42c52e2967c7db74afc08d51e48ea53f5ccd

Observation 447de910-863d-4632-9536-226cceb1fdb0 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

LLaVA-OneVision: Easy Visual Task Transfer InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:46:10.650334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:ccccdbc00bb0d96d7b59ca6c316b510161b6aff34da31c57274422b090ed3bb8

Observation 77253167-594a-4623-b5e2-128dcad032f3 · outbound

This paper cites Hitab: A hierarchical table dataset for question answering and natural language generation.

LLaVA-OneVision: Easy Visual Task Transfer Hitab: A hierarchical table dataset for question answering and natural language generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.822561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:2f2102a19ab545cf6c5fc3c23db0b725094a4afee4e75bc3d51c537f8e1455c7

Observation 04b06fc2-0556-4778-b1ca-40832d375154 · outbound

This paper cites PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns.

LLaVA-OneVision: Easy Visual Task Transfer PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:23:49.519210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:641579b82826864ff23cff46eaf472d45cfb22c7e8015f12690d265ad6ce2de6

Observation f22179d7-39a8-4176-a85d-f58bcb331c17 · outbound

This paper cites Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner.

LLaVA-OneVision: Easy Visual Task Transfer Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.824389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:5635d2d82dafd9ef765172f71fa83fa5f27de0a56a46aabb4d798062e8e85dd2

Observation 3e0c6430-5282-44e5-95a6-f0016fe55e8a · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

LLaVA-OneVision: Easy Visual Task Transfer Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.826614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:2838810ddf9b102e707be46a3e847741f2235a9ba01bcbcbd0457538911019cd

Observation c22a8d7d-9323-48d2-8ec6-a634e3a95fd8 · outbound

This paper cites Neural naturalist: Generating fine-grained image comparisons.

LLaVA-OneVision: Easy Visual Task Transfer Neural naturalist: Generating fine-grained image comparisons

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.828690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:6c5c443f356160d79c70ded0c026ee55fdb60eae8658718fda79d5d232bfe383

Observation db47c7bf-afaf-4fae-a598-b945cc5cf105 · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models.

LLaVA-OneVision: Easy Visual Task Transfer Mme: A comprehensive evaluation benchmark for multimodal large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.830608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:23c16660e66e014560b865aa7364528be124fafbb81110a343f71e0ff3ed8bd5

Observation 3011ee19-4410-42db-93f3-0292ec3202c3 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

LLaVA-OneVision: Easy Visual Task Transfer Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:58:42.298802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:dffe02459ad945ff9a8ede1b066adfba43290fca44909c0d45e5e73754cbabc1

Observation 0e6c63e6-ab24-49f0-85c4-c660d45bb2fe · outbound

This paper cites Dreamsim: Learning new dimensions of human visual similarity using synthetic data.

LLaVA-OneVision: Easy Visual Task Transfer Dreamsim: Learning new dimensions of human visual similarity using synthetic data

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.832521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:77e980667c44f6af2c3b3ab68fc0295da74566980ec9a4072a3b4f4984e85cf3

Observation d256730f-a2e3-4983-90d8-0043a8ba6f24 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

LLaVA-OneVision: Easy Visual Task Transfer BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:881a300d1ded01729f48ce40564b8a044622d3f92c6aa64e6a1a74e09460c01b

Observation c8550678-7331-49be-a443-8f261fe114cf · outbound

This paper cites G-llava: Solving geometric problem with multi-modal large language model.

LLaVA-OneVision: Easy Visual Task Transfer G-llava: Solving geometric problem with multi-modal large language model

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.834308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:02f764415134ff8558835070386af70207fa14c2928a8a519f1058ad8d26d78a

Observation 38e0cf1a-67e1-4c36-947d-c0b3eb208f91 · outbound

This paper cites an unresolved cited work.

LLaVA-OneVision: Easy Visual Task Transfer Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-05-10T14:23:49.835940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:c43b3dc72f14069f75faaa2eb14cc15001a35e01279905cc22972be9b0fc855e

Observation ef769a95-7c90-4816-89cb-1101658638a1 · outbound

This paper cites Sciverse.

LLaVA-OneVision: Easy Visual Task Transfer Sciverse

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.837972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:6e0533c7fc7e4052e071a1ce3e1d6313de497619091a96e37b7985114f751c45

Observation ee4d18be-d8a8-4d3a-8433-57f5ae56d34d · outbound

This paper cites Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following.

LLaVA-OneVision: Easy Visual Task Transfer Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:23:49.552473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:184b73628a345ff09febd8577db105f4a0fa5cd27556aa4868d2b589a5f23641

Observation 3c03def1-69a2-4cb7-82f1-7faa39ef1274 · outbound

This paper cites Imagine this! scripts to compositions to videos.

LLaVA-OneVision: Easy Visual Task Transfer Imagine this! scripts to compositions to videos

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.839916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:069aada995c68024d88a54b9f37f8ea82958cb7ae072aeb97b51bc06b7ff374a

Observation 2d47fe6a-6734-42ee-a1ba-b5c0909a299c · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

LLaVA-OneVision: Easy Visual Task Transfer Vizwiz grand challenge: Answering visual questions from blind people

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.842039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:d9be812d4d2e88aaa479579409aaa8d201227034f244cb15ec8a50fd33bcd821

Observation 877bde4c-0fbc-4d83-b9a9-3416de912c39 · outbound

This paper cites 3d-llm: Injecting the 3d world into large language models.

LLaVA-OneVision: Easy Visual Task Transfer 3d-llm: Injecting the 3d world into large language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.843897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:6b963b8dd5155f33787c1c4f01292a260b50a1e88420fd030ef744d75e982577

Observation 9f16cc92-efff-4b33-8e86-5e857d6aca25 · outbound

This paper cites Image change captioning by learning from an auxiliary task.

LLaVA-OneVision: Easy Visual Task Transfer Image change captioning by learning from an auxiliary task

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.845663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:590082be2321b320acc0b2744b41f3ef64e081afe994fd68b406ee610f81a5f6

Observation aecba392-37a3-4508-94c3-ce749f3a07ed · outbound

This paper cites Huang, Francis Ferraro, Nasrin Mostafazadeh, Ishan Misra, Jacob Devlin, Aish- warya Agrawal, Ross Girshick, Xiaodong He, Pushmeet Kohli, Dhruv Batra, et al.

LLaVA-OneVision: Easy Visual Task Transfer Huang, Francis Ferraro, Nasrin Mostafazadeh, Ishan Misra, Jacob Devlin, Aish- warya Agrawal, Ross Girshick, Xiaodong He, Pushmeet Kohli, Dhruv Batra, et al

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.847800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:710d7919ab9dc34e9f2c053599373a6be43c85a6e26887f00f1b3bf1fc4a1572

Observation c8e012eb-f212-4ee1-b62d-4391ebfea9ec · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

LLaVA-OneVision: Easy Visual Task Transfer Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.849676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:09a260a00be2683778dd3cb1756697923ef80d75dfab36579912297be86dc952

Observation a8891752-b795-4f50-8f6d-f7041698b6de · outbound

This paper cites Hq-edit: A high-quality dataset for instruction-based image editing.

LLaVA-OneVision: Easy Visual Task Transfer Hq-edit: A high-quality dataset for instruction-based image editing

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.851798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:111472eb356b1e6e411bbf72d78e56d89b09af07bc2428c1d17dc69a2f0af1d8

Observation 009f1ef6-5956-4b13-bc90-5d9d2262356f · outbound

This paper cites Lim, and Edward H.

LLaVA-OneVision: Easy Visual Task Transfer Lim, and Edward H

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.853720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:850f0b0c21a048f1f537d3557eddad5b74b61c15ac8de2f67877321e8ea5c72a

Observation 8a890935-c89e-4a05-a816-4319b646dab0 · outbound

This paper cites The amazing mysteries of the gutter: Drawing inferences between panels in comic book narratives.

LLaVA-OneVision: Easy Visual Task Transfer The amazing mysteries of the gutter: Drawing inferences between panels in comic book narratives

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.855712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:1e5c6d50dd380a58039fd84d6d3754a77257fe53f6e2038856972e2c4470360c

Observation edf95c37-ef1b-485f-934b-4277e539d0f2 · outbound

This paper cites Learning to Describe Differences Between Pairs of Similar Images.

LLaVA-OneVision: Easy Visual Task Transfer Learning to Describe Differences Between Pairs of Similar Images

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:23:49.516058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:663b2c18a59b9d2509fbde543378bed1e52c149515f261afae1a743d8c6dffba

Observation b187c262-76b1-44b1-9e40-22657ef7bd7b · outbound

This paper cites Learning to describe differences between pairs of similar images.

LLaVA-OneVision: Easy Visual Task Transfer Learning to describe differences between pairs of similar images

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.857718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:0bb0504897166fefd88195730fb578c133b2f64cab36bff0569d26ef061de62a

Observation 7c9a2926-243f-46e2-999e-7107c13e8884 · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

LLaVA-OneVision: Easy Visual Task Transfer MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:23:49.546122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:eb4286cdc861f2a9bf732d416e76b076190293b4f739b1d8e437c7eecb2e44fe

Observation 1a8ea4f3-11bf-4324-ad5a-3532696c0aef · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning.

LLaVA-OneVision: Easy Visual Task Transfer Clevr: A diagnostic dataset for compositional language and elementary visual reasoning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.859640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:8070e2172403936f936f25841c30e2e87fd5d27ff748be0c2e414e7cbf3ff102

Observation f24d1fd9-eea2-4afa-871c-acd1ac9ede1e · outbound

This paper cites Dvqa: Understanding data visualizations via question answering.

LLaVA-OneVision: Easy Visual Task Transfer Dvqa: Understanding data visualizations via question answering

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.861410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:bda2422ddb366fcbf226d61b6330598c933c400ad96f883f76c2af76cbbb36d5

Observation f52a5e4b-2ffe-4e2b-83fb-722adfdf2b32 · outbound

This paper cites Figureqa: An annotated figure dataset for visual reasoning.

LLaVA-OneVision: Easy Visual Task Transfer Figureqa: An annotated figure dataset for visual reasoning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.863246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:06046b8f6642787d14ca9fac75b1ffef736e975e382c19bc9db353cc9533147b

Observation a7731beb-ab5f-42e2-b610-b74dc1f08f26 · outbound

This paper cites Prismatic vlms: Investigating the design space of visually-conditioned language models.

LLaVA-OneVision: Easy Visual Task Transfer Prismatic vlms: Investigating the design space of visually-conditioned language models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.865179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:b1d7246ad4de392b3c66e096659688052403d1e1d24981e1172beaa7635b6306

Observation b56f4798-8e8e-43ea-8c81-c08d90c27881 · outbound

This paper cites GeomVerse: A Systematic Evaluation of Large Models for Geometric Reasoning.

LLaVA-OneVision: Easy Visual Task Transfer GeomVerse: A Systematic Evaluation of Large Models for Geometric Reasoning

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:23:49.589748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:ed2bb5862a0e021b9a49287ec70f7763bb2d0161ded90220596e66f415ea9357

Observation 4e69fa9d-8692-4b84-9521-c652b6d6df8a · outbound

This paper cites A diagram is worth a dozen images.

LLaVA-OneVision: Easy Visual Task Transfer A diagram is worth a dozen images

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.866875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:c688d92f087ef85976bcd16403eb5265c4f2cf689923eb780c750c014ece091b

Observation 910b65a7-0897-43f4-8873-c8c33340a03e · outbound

This paper cites A diagram is worth a dozen images.

LLaVA-OneVision: Easy Visual Task Transfer A diagram is worth a dozen images

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.868762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:2311fbb7961c284fb61c07a8d4a5f37bad350ec526ad062bdee887e40b7bf9dc

Observation 729ca3a5-b14d-40fa-aac9-6f9a80b095d4 · outbound

This paper cites Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension.

LLaVA-OneVision: Easy Visual Task Transfer Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.871064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:51f1a52681444907d5a3bb24e9da060995e1d9df88da430c616f1d2f315faafa

Observation 227e1777-e9f2-4c22-a978-4b7a29c8e5b3 · outbound

This paper cites Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension.

LLaVA-OneVision: Easy Visual Task Transfer Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.873260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:b48661813e1291484aa0a95f5590b050dff2cbe048b7affc4ad2ad5d267e2b32

Observation e21be52d-e899-4b21-a1fd-be0231c46d5b · outbound

This paper cites The hateful memes challenge: Detecting hate speech in multimodal memes.

LLaVA-OneVision: Easy Visual Task Transfer The hateful memes challenge: Detecting hate speech in multimodal memes

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.875174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:a7abf83a507c4ee1a2fc14d121454334ba8e178b219e63adea7b0ed9be1320ec

Observation 54881f83-74b0-44a8-8905-1a6b3d11b93c · outbound

This paper cites Ocr-free document understanding transformer.

LLaVA-OneVision: Easy Visual Task Transfer Ocr-free document understanding transformer

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.877024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:8342efb5a2f898e24a49e696638601bd23049a35971d03f2176108b1f4cef97f

Observation 78f3114f-fe43-484b-8dc0-97e79c7e5058 · outbound

This paper cites Shamma, Michael S.

LLaVA-OneVision: Easy Visual Task Transfer Shamma, Michael S

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.879063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:9772d2785bf56d68bef700401a29892c4a46296551666609fea49dfaa56e1342

Observation 8bb47279-fe45-4ad2-bbcb-9caa9c469f91 · outbound

This paper cites Image retrieval from contextual descriptions.

LLaVA-OneVision: Easy Visual Task Transfer Image retrieval from contextual descriptions

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.620840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:3429fa3d06df6e1d8648cfc5e429a061fdf0b979400f8c8eb87932d9a2582e34

Observation 9a059159-be65-4fe1-b559-1a0c37a94e5b · outbound

This paper cites Sharegpt-4o: Comprehensive multimodal annotations with gpt-4o.

LLaVA-OneVision: Easy Visual Task Transfer Sharegpt-4o: Comprehensive multimodal annotations with gpt-4o

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.623045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:071d99d6d081c5f111f5d18308b5e258d4d02ed6c31c9fd406976967eadb0821

Observation 00792893-240a-4162-a2bf-46c9319827f9 · outbound

This paper cites A dataset of clinically generated visual questions and answers about radiology images.

LLaVA-OneVision: Easy Visual Task Transfer A dataset of clinically generated visual questions and answers about radiology images

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.625193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:33a8e01038df96a6f72d3bcc17b2d8c82ef57bf06b4ab7bedf66bfb526c5fff3

Observation dfac752f-7085-48e8-9d86-61e7ad93252d · outbound

This paper cites What matters when building vision-language models? Technical Report.

LLaVA-OneVision: Easy Visual Task Transfer What matters when building vision-language models? Technical Report

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.627614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:8fc224dca9bfde0baff5832418f426ff9bc32315b38d5ca345717641ac7f05e7

Observation d02dd82a-a468-4742-8b8d-ce036e2b3e9a · outbound

This paper cites Llava-next: What else influences visual instruction tuning beyond data?, May 2024.

LLaVA-OneVision: Easy Visual Task Transfer Llava-next: What else influences visual instruction tuning beyond data?, May 2024

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.629811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:3627095e02c784ff878e024ba189149a679dd29d18b7e6805d3c8e69dd64b1a8

Observation 386fdd86-e05c-4908-8f51-5ee8d0da447e · outbound

This paper cites Llava-next: Stronger llms supercharge multimodal capabilities in the wild, May 2024.

LLaVA-OneVision: Easy Visual Task Transfer Llava-next: Stronger llms supercharge multimodal capabilities in the wild, May 2024

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.631785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:301babceb32e6b5be8d779a26ebe57fac4d21bb608a20d95f763db4e45f8d649

Observation b6d57fb6-77cc-422a-add0-a0574805be89 · outbound

This paper cites Seed-bench: Benchmarking multimodal llms with generative comprehension.

LLaVA-OneVision: Easy Visual Task Transfer Seed-bench: Benchmarking multimodal llms with generative comprehension

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.633881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:e6d88f4fe01d782d3de15c24b5d0847df07d089844b4b4b9edb58dbfa6fd27ac

Observation 50d201ca-dee2-4307-a605-7fef220c485d · outbound

This paper cites Multimodal foundation models: From specialists to general-purpose assistants.

LLaVA-OneVision: Easy Visual Task Transfer Multimodal foundation models: From specialists to general-purpose assistants

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.635892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:74fa102471b63b845c71102ccad99c69344c47b811c16700340cb8737ab2b379

Observation cc03fc8b-d50d-4277-a391-b08d86a5a1a4 · outbound

This paper cites Llava-next: Tackling multi-image, video, and 3d in large multimodal models, June 2024.

LLaVA-OneVision: Easy Visual Task Transfer Llava-next: Tackling multi-image, video, and 3d in large multimodal models, June 2024

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.637755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:2a9aefa080bc80e79d0ff26b00bf4454278b0fb2f0f83451f821b7dd55d6887b

Observation c7f20297-8b3c-4512-af63-9b8c8d1e3fd2 · outbound

This paper cites Fine-tuning multimodal llms to follow zero-shot demonstrative instructions.

LLaVA-OneVision: Easy Visual Task Transfer Fine-tuning multimodal llms to follow zero-shot demonstrative instructions

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.639766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:3b09d4cb49330cfbc4a259a530ce64c77ce744a8ced208403abd67108078ff0e

Observation af34bcc1-16cc-41c3-ac41-7e8054bfa2de · outbound

This paper cites Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions.

LLaVA-OneVision: Easy Visual Task Transfer Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:23:49.618766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:5d9bf767be2611e13c93064c1926b7ced664f16c86f52ebae9258c9beb03c23a

Observation b2126fea-1174-459d-bbae-e5081aef274a · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

LLaVA-OneVision: Easy Visual Task Transfer Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.641626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:03111bfd12e1f8abebb1a7c0ef736753d8cc3327d0784b4ecd6dde907aa4aec1

Observation 87379fcf-b02c-4021-ab2e-7cc62491152f · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

LLaVA-OneVision: Easy Visual Task Transfer Llama-vid: An image is worth 2 tokens in large language models

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.643354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:b67c5dd0de1df5ed52a527d1d6a443a36d27c438409bd09c2e0f4958f278a28d

Observation 2db2e25c-e584-4d85-a25a-ba409ac6c683 · outbound

This paper cites Mini-gemini: Mining the potential of multi-modality vision language models.

LLaVA-OneVision: Easy Visual Task Transfer Mini-gemini: Mining the potential of multi-modality vision language models

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.645698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:35f4e37a69de21e99f57e36260db5ea8e191073da4ec8d313214043a88462a5a

Observation 27c83f54-934c-476c-ba70-a69c5ea1e487 · outbound

This paper cites Storygan: A sequential conditional gan for story visualization.

LLaVA-OneVision: Easy Visual Task Transfer Storygan: A sequential conditional gan for story visualization

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.647707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:16b30309256e8031e096bbc0dff6d00b3ba356580dbc5cda1565de2a4d19cd93

Observation c8f15db4-19d6-4d83-a753-40dcc3c4c9ee · outbound

This paper cites Super-clevr: A virtual benchmark to diagnose domain robustness in visual reasoning.

LLaVA-OneVision: Easy Visual Task Transfer Super-clevr: A virtual benchmark to diagnose domain robustness in visual reasoning

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.649751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:28e889d83299f15336ce9837a1381602a36651778daa5ebc1917791e29f5989f

Observation f349e4c0-686b-4786-adaa-53a8edddfd56 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

LLaVA-OneVision: Easy Visual Task Transfer Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:08:01.613687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:1f1c04f105eed535beab945a841866a9433471ffd99b57669d86abb41cf887dd

Observation b219c21d-97bb-4821-80fc-731976455243 · outbound

This paper cites Vila: On pre-training for visual language models.

LLaVA-OneVision: Easy Visual Task Transfer Vila: On pre-training for visual language models

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.651418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:d9873299b40f1c65f06404174ff1bd8904add35b943a706562740dd559fb362a

Observation 4a79c2c9-be72-4508-8a9e-932f30570a87 · outbound

This paper cites Lawrence Zitnick, and Piotr Dollár.

LLaVA-OneVision: Easy Visual Task Transfer Lawrence Zitnick, and Piotr Dollár

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.653459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:916ebdaef02e64e2f377666844d00e6af90052fe33483090d8ddb01d5169bf29

Observation f21353e8-6e08-4dd1-875e-7d26ac3e8eae · outbound

This paper cites Visual spatial reasoning.

LLaVA-OneVision: Easy Visual Task Transfer Visual spatial reasoning

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.655462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:bfb67fdec0810ebf0abace1ea7001c53b09b60d5483f782e24b65c17de50d6c3

Observation 0cec71d8-58ef-48e2-81f2-a141b92d1efc · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

LLaVA-OneVision: Easy Visual Task Transfer Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-14T17:34:57.031001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:4f0693af6c744986542c473ba028fad155af85a6ada536d9d770d60cb786069c

Observation 90ed9301-76ee-44d2-942d-a7e524a229ce · outbound

This paper cites Improved baselines with visual instruction tuning.

LLaVA-OneVision: Easy Visual Task Transfer Improved baselines with visual instruction tuning

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.657619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:ef4f566670ebe723f6ae1ad0d91415c30ba4fa7a763b0ff96e04a3b861e0953d

Observation 296a0492-df03-4f34-83d4-1409f908b7af · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

LLaVA-OneVision: Easy Visual Task Transfer Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.659664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:982028573a448aa3b446be43c2f9ccac4c3d0e9cd7c41ffe30a0a1d68dcc5332

Observation df059886-a620-40c8-809c-44d9f22f6ddc · outbound

This paper cites Visual instruction tuning.

LLaVA-OneVision: Easy Visual Task Transfer Visual instruction tuning

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.661437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:bb1e6668020cb78a05871cc628a8f14872bf0410284173ce2f58fb658dc61904

Observation 3a933ec1-7f26-4d31-9304-f0cb90657c0b · outbound

This paper cites What large language models bring to text-rich vqa?.

LLaVA-OneVision: Easy Visual Task Transfer What large language models bring to text-rich vqa?

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.663255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:3580719829cdc7680a19b7ab22036b891ce194e0abb1e92d0b10bf61e1a3e6fc

Observation d7382f9a-093c-4f1b-b359-133d9c7af79c · outbound

This paper cites What Large Language Models Bring to Text-rich VQA?.

LLaVA-OneVision: Easy Visual Task Transfer What Large Language Models Bring to Text-rich VQA?

Reference 85

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:23:49.604786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:4bab4a7ea8a25d145acf65f60e016959e444b848aa3fbf75b4e0e0e59a3a6b68

Observation 679b2922-8129-4a51-a545-9f4b041ca91c · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? Technical Report.

LLaVA-OneVision: Easy Visual Task Transfer Mmbench: Is your multi-modal model an all-around player? Technical Report

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.665085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:aac7bb9518efbb9095cd70ca31a08fa53e0f63ec2cabdf48372af5320cda0ad0

Observation 6ad57d36-1cb8-4920-9b89-03599fb5c899 · outbound

This paper cites Video detail caption.

LLaVA-OneVision: Easy Visual Task Transfer Video detail caption

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.666813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:07ea5e6c096dd66331223fdf6e27c3d6f66cfbe58a505996091563ad86da4c46

Observation 15c7f4b3-7a98-457c-b8fb-1726088e38ae · outbound

This paper cites The flan collection: Designing data and methods for effective instruction tuning.

LLaVA-OneVision: Easy Visual Task Transfer The flan collection: Designing data and methods for effective instruction tuning

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.668564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:fe60d183c1d4d3c0af7ce58be8b1edb4fb26e3cabd6b28c6b815620f67235e05

Observation 8a071b03-c99c-40c0-892a-fbdcfbee4807 · outbound

This paper cites Deepseek-vl: towards real-world vision-language under- standing.

LLaVA-OneVision: Easy Visual Task Transfer Deepseek-vl: towards real-world vision-language under- standing

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.670487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:780b5894855167c3542165e581a90cd2fa10f8ef1c8cd483e9bafa9d7537bf60

Observation 46eb10de-5d99-40f8-9b91-ed5e2b195334 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

LLaVA-OneVision: Easy Visual Task Transfer MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:30:15.750849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:7f855ed9ee2e9c9a7f0df50034b4da79ce18920b3be90c0d97eaac4a781ac06b

Observation e30d940b-ceaf-4804-a7d7-704711b348f8 · outbound

This paper cites Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning.

LLaVA-OneVision: Easy Visual Task Transfer Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.672505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:17424e2e226c04cb3ac7611135f3c967a8a6dc64a28007145da32260ec9869b6

Observation 8d34a9a0-7931-44fc-ade0-5176bdd9e566 · outbound

This paper cites Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning.

LLaVA-OneVision: Easy Visual Task Transfer Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.674472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:30abaceb9b90d26bf2de0cd6cce2b137be924a65fbd135f185ec871a3cf5ac2f

Observation 0ef779fb-b910-4310-95b4-699b47a8e80c · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

LLaVA-OneVision: Easy Visual Task Transfer Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.676605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:4213fac6c0e0991d782feaf15793782f62f0a974973a910dd59e6f47750c7a38

Observation 4039bad7-b895-4c21-9fd6-4e39e408e529 · outbound

This paper cites Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning.

LLaVA-OneVision: Easy Visual Task Transfer Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.678897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:1c9146e53788c582c785b330c8275083f47572cf0f28cabc4fe9e8ccc8e38899

Observation 79454796-eea8-4dbc-b54d-4353a6dd42a8 · outbound

This paper cites Iconqa: A new benchmark for abstract diagram understanding and visual language reasoning.

LLaVA-OneVision: Easy Visual Task Transfer Iconqa: A new benchmark for abstract diagram understanding and visual language reasoning

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.680874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:8276093719b9ad92a2fba75c1ea974f30083d283f91aaad8a617fea87e9ad404

Observation ea453277-00f1-42d0-a90c-d7d711bff62d · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

LLaVA-OneVision: Easy Visual Task Transfer Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:36:18.685087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:ac9e298b57fe2643ddda7d3bf5ad0b97525347031a9d3d32df6d85930974f498

Observation 26a70b3e-c92f-4f6e-b92a-d4e222da2445 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

LLaVA-OneVision: Easy Visual Task Transfer Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.683295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:20a500f14dc839cb93d4a3c4eb0627f8c9de6235c51557a48d7cab23362cda69

Observation cbbad9f4-b95f-44e3-a807-34efa192c4c6 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.

LLaVA-OneVision: Easy Visual Task Transfer Egoschema: A diagnostic benchmark for very long-form video language understanding

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.685493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:d8fe813e24fdeb4cb59f50957ca2b6667502499cc8d0341e20fa8beb7cf8ff04

Observation 1fe2b577-ad07-40b4-a9db-b016eeaaf124 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

LLaVA-OneVision: Easy Visual Task Transfer Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.687408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:9a67c92fad3e4b406eebd725fb349b91f8de2b221bd8f8cac82cec903ad52c1b

Observation c289ac2f-f956-4e93-b0a0-4ba45a1478dd · outbound

This paper cites The iam-database: an english sentence database for offline handwriting recognition.

LLaVA-OneVision: Easy Visual Task Transfer The iam-database: an english sentence database for offline handwriting recognition

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T14:23:49.689261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:445772a8dff7347a931f29caceb94099939e6bf6f4010b51cb00fedf8bc4a8c3

Pith citing papers

Observation 0570e303-6a28-46cf-ac4e-364240baf1ae · inbound

MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems? cites this paper.

MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems? LLaVA-OneVision: Easy Visual Task Transfer

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T01:29:30.202749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T01:29:30.032408Z digest=sha256:9ee4b70fed3fe8896df22847beb9141147a2668754a3fd509ae08e87f4f93f0f

Observation 6f7651f1-27ca-4e49-83b1-cbe48b54ede8 · inbound

Hallucination of Multimodal Large Language Models: A Survey cites this paper.

Hallucination of Multimodal Large Language Models: A Survey LLaVA-OneVision: Easy Visual Task Transfer

Reference 101

Resolution
verified exact
local_arxiv, observed 2026-05-11T12:33:33.927578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T12:33:32.631346Z digest=sha256:34087d465e707aa0763cc9237fd725c1277605e593e5e582bf5244dc9875b804

Observation 2c152a20-a77a-413d-abf3-ce299d96e22f · inbound

MLVU: Benchmarking Multi-task Long Video Understanding cites this paper.

MLVU: Benchmarking Multi-task Long Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:55:26.585070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T19:55:26.333923Z digest=sha256:2da9db03ee3ea099a5b115cc754d18bdce9bf6d70d8e9c48f8b15f01160f0795

Observation eba91837-dbc0-4889-be20-f007cefdde2a · inbound

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs cites this paper.

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-11T02:44:53.537663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T02:44:53.284345Z digest=sha256:09a2d0f9a2a182a81ca2a98fc7218a274bf8ef8be70057f777c564af1df4b502

Observation ac23ea09-a587-4b6e-b33f-5b690a3ffe08 · inbound

LVBench: An Extreme Long Video Understanding Benchmark cites this paper.

LVBench: An Extreme Long Video Understanding Benchmark LLaVA-OneVision: Easy Visual Task Transfer

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:55:30.182459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T11:55:30.048525Z digest=sha256:474ac5085f3a238b8355bff47f86d3b324952284c54671f0933a2a76f606bc7b

Observation 2680417b-f22a-4d16-9b97-a82fac67b69f · inbound

LongVILA: Scaling Long-Context Visual Language Models for Long Videos cites this paper.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos LLaVA-OneVision: Easy Visual Task Transfer

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:51:25.534046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:ad99963aeb135868ae82c019cff778d894ff4d667840cab086a1804658c8a77a

Observation d3264127-4082-440d-b911-47e4aef871af · inbound

MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark cites this paper.

MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark LLaVA-OneVision: Easy Visual Task Transfer

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-14T00:51:48.352594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-14T00:51:48.163349Z digest=sha256:bda95d484c9df32744aedf1871c3a32ec1ad55259988c99d502d58b4d1ce7b90

Observation 0a1fa6ab-bb1e-4d10-8053-2d0f6c1ab2e4 · inbound

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models cites this paper.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:55:12.646452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:d0f702fced195c59513993c1f874bbc485219c57163ec7d1372ade79b81bb911

Observation e462826b-9284-47bd-a9c0-47d812daf8d6 · inbound

Emu3: Next-Token Prediction is All You Need cites this paper.

Emu3: Next-Token Prediction is All You Need LLaVA-OneVision: Easy Visual Task Transfer

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:56:08.485056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T10:56:06.418360Z digest=sha256:3a201210d563a33b62afa66762e707754b448db7fc8501e06db31f3a67b80eb9

Observation 42b2e88a-7121-4d11-826f-031c99e0148b · inbound

LLaVA-Video: Video Instruction Tuning With Synthetic Data cites this paper.

LLaVA-Video: Video Instruction Tuning With Synthetic Data LLaVA-OneVision: Easy Visual Task Transfer

Reference 158

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T23:20:33.148644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:e1e1b5e5935f29a917cdfa70c9b72c4c70f5009853728906eb68672aab237cc9

Observation 0d29a7cc-cbba-45a5-890d-5e6d3fea689f · inbound

Pixtral 12B cites this paper.

Pixtral 12B LLaVA-OneVision: Easy Visual Task Transfer

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-14T23:53:29.895469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T23:53:29.862702Z digest=sha256:5c1847b9e06c82caea75ce9574fb60ba8413f6c5a13478326c5373ac0c024e40

Observation c06c2da2-63ee-4f38-9835-722a7a36d9ad · inbound

VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents cites this paper.

VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents LLaVA-OneVision: Easy Visual Task Transfer

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:37:25.899474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T15:37:25.781240Z digest=sha256:f27ab9bfda8a98339efc41eeceb0a5aead1bd78a6ec984f340e415d6abcaf80b

Observation 0fa72a6b-fc80-4bf0-bf2b-61ff6ea989ba · inbound

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation cites this paper.

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation LLaVA-OneVision: Easy Visual Task Transfer

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:09:16.134516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T22:09:16.001309Z digest=sha256:4b88445e36261d15e25098514b37a06a1930b6cf9be201da8a8a54066e02f2e5

Observation 514efd96-2704-4416-b953-0debf9d077a6 · inbound

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding cites this paper.

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-16T13:53:33.665517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:53:33.585035Z digest=sha256:e878ebcc9aab15dff0d3954f4ecbae72871c49f8af224dc367d6f5489afdd37a

Observation 18dd4ca0-d149-44c2-b463-214f2f9c2526 · inbound

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization cites this paper.

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization LLaVA-OneVision: Easy Visual Task Transfer

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-16T09:16:17.549123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T09:16:17.150383Z digest=sha256:655957c61eb2e5e632b50cf1fe6b7f61c057c959d7a44e3b93717bd376cdec74

Observation 1ea11bc5-4b4f-455c-9579-39a22dbbaec1 · inbound

Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction cites this paper.

Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction LLaVA-OneVision: Easy Visual Task Transfer

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:09:41.686056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-18T04:09:41.494136Z digest=sha256:c81c28369ea0db1f7bab18445e36b97c84dfe83de8eedd4599543efc91636f9f

Observation de0a1d31-fe8c-4354-aea6-046e16d7fbde · inbound

NVILA: Efficient Frontier Visual Language Models cites this paper.

NVILA: Efficient Frontier Visual Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.116592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:30a39eff380eb05ae2cbdc6cf2612843390b15dce3955514dc556442cc7df9cf

Observation 151592b6-8e8e-4e49-8135-ac20532f0abd · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling LLaVA-OneVision: Easy Visual Task Transfer

Reference 124

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.058774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:a1b61787954be010ef6b7f483f31ffec7f71267b9dba15573ce6c67cfeb34eb7

Observation 2415b5c5-f4d0-4c83-9154-f9deeb18a3b5 · inbound

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding cites this paper.

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:09:23.568350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T10:09:21.542356Z digest=sha256:4e04428b7b815d42e2b411c25011d430e5e61be121fc87544bd76779a0cf7f45

Observation 136d2089-23bd-4ed7-bdfe-6060310f1c1f · inbound

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning cites this paper.

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning LLaVA-OneVision: Easy Visual Task Transfer

Reference 296

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T07:51:13.406086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-17T07:51:12.953777Z digest=sha256:59875a27abb906e3b063586f6d91ca9fe215dfb67014e56dda05853caa2a29ed

Observation 93469cd3-eb59-4e3c-8195-0cb0d796480e · inbound

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces cites this paper.

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces LLaVA-OneVision: Easy Visual Task Transfer

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:27:44.077838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T09:27:43.919941Z digest=sha256:f293b9b5bce93399a3797f1387b119e90a81b7c1f6b5f1cbb492042ef63793f3

Observation c27ac01b-8481-4ccd-a77b-7a572bde8ea8 · inbound

Progressive Multimodal Reasoning via Active Retrieval cites this paper.

Progressive Multimodal Reasoning via Active Retrieval LLaVA-OneVision: Easy Visual Task Transfer

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T11:55:10.312257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:55:10.312257Z digest=sha256:dc3e1429f1791b1f53a12faae9c1a026e78552e0be81bab1a6f9a405a1b0aadb

Observation d6ea7ed6-1071-4b89-96a8-992875141e06 · inbound

EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues cites this paper.

EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues LLaVA-OneVision: Easy Visual Task Transfer

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:35.279293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:36:35.279293Z digest=sha256:3d179faaabaa287e27b43a79e7cbf308c2ad7c323247f452ddc041681f0dbbed

Observation 9e03463c-dc7a-4dfc-9543-94655fca8ffc · inbound

Error-driven Data-efficient Large Multimodal Model Tuning cites this paper.

Error-driven Data-efficient Large Multimodal Model Tuning LLaVA-OneVision: Easy Visual Task Transfer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:32.014466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:32.014466Z digest=sha256:8a5ab9811d23b9508cab649471bc61693e8e987fbfdb66ad7e74f4c682168e17

Observation 7366d366-22a8-42ed-b489-fff6b0b1a64c · inbound

PruneVid: Visual Token Pruning for Efficient Video Large Language Models cites this paper.

PruneVid: Visual Token Pruning for Efficient Video Large Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:52.135759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:49:52.135759Z digest=sha256:a73f99fb47281f9a42045a11bc10c24cda289c967cbb047ac72568d90e728082

Observation be2a0fa4-5fcd-42d2-8c89-e5a8607e13d9 · inbound

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation cites this paper.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation LLaVA-OneVision: Easy Visual Task Transfer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.145499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.145499Z digest=sha256:cd738525f855183d9ced1ba3d9ef5e993ff64612e34098ff810443d0ea6f84fd

Observation d7f22d14-e1ba-485c-a2d6-23804eac654f · inbound

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities cites this paper.

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities LLaVA-OneVision: Easy Visual Task Transfer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:37.764851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:41:37.764851Z digest=sha256:bf9b9d2af8fa9c36d9a96bad8132fae76f3c2c2187cf9f50cd5c6289b5c83769

Observation f8bebfc0-b866-4763-aef4-58a647a6b53c · inbound

Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment cites this paper.

Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment LLaVA-OneVision: Easy Visual Task Transfer

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T00:47:12.224469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:47:12.224469Z digest=sha256:6c064e8172bde056fd1caa529e3a1bd489807d66e68b8c8d429526eea86b55c2

Observation 463dba09-af75-409e-a9ab-143d5bb24204 · inbound

Hear the Scene: Audio-Enhanced Text Spotting cites this paper.

Hear the Scene: Audio-Enhanced Text Spotting LLaVA-OneVision: Easy Visual Task Transfer

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:07.835018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:07.835018Z digest=sha256:c916666c8aaa99faa7f50626b31d1a351ce3045e299e553724009c85abb38b0d

Observation 00ee3d80-6cce-4d23-a9b5-4a1f09edcb9d · inbound

MBQ: Modality-Balanced Quantization for Large Vision-Language Models cites this paper.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.195975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.195975Z digest=sha256:1dc96feb1b8d709fd428ef0fb56bf78f9decae9b696374838519e026d728221b

Observation 94026a25-586b-472d-9aca-2226bc1caefe · inbound

From Elements to Design: A Layered Approach for Automatic Graphic Design Composition cites this paper.

From Elements to Design: A Layered Approach for Automatic Graphic Design Composition LLaVA-OneVision: Easy Visual Task Transfer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T00:06:12.278164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:06:12.278164Z digest=sha256:94a96c78aa5d2df8607f37200e9ed1028ea83437e84766bff7ce01decf002671

Observation 661e7cc0-3ca7-438a-a978-04ee1868d002 · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning LLaVA-OneVision: Easy Visual Task Transfer

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:33:26.804574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:8ea21a76c3bdc69c77e94ed736c278c3d03626429279158d107b513dbafe5686

Observation 707e2774-c404-43ff-8072-de78d6e6b17f · inbound

Probing Visual Language Priors in VLMs cites this paper.

Probing Visual Language Priors in VLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T22:55:53.996989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:55:53.996989Z digest=sha256:1a2bb8f2506d3a6bc02c1376e70ea40458ba501b178c76c2bc29275f7d5c5325

Observation 06e36acb-10d6-4385-93c0-b600a5db02d5 · inbound

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling cites this paper.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LLaVA-OneVision: Easy Visual Task Transfer

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.422956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:4bbb88724b2cfefdbe4d6b6073b3b7a354f79d9bd80c7b70d10aefda64da093a

Observation 6a1c409b-61bf-48f0-8294-2bdf2f324a5f · inbound

Online Video Understanding: OVBench and VideoChat-Online cites this paper.

Online Video Understanding: OVBench and VideoChat-Online LLaVA-OneVision: Easy Visual Task Transfer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:57:40.040203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:57:40.040203Z digest=sha256:1d34b39cdf6b8860bafb1f90758d5176a67740110fb8e94d74166194829e23a6

Observation 4a727e5f-c7b7-400c-9958-13972947ccbc · inbound

VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM cites this paper.

VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM LLaVA-OneVision: Easy Visual Task Transfer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:30.214027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:30.214027Z digest=sha256:998ed7e6b34d8f2d18083fb24759cbc96735b2f814b5cc6d0b9b9e5c905ef680

Observation ced4e771-125b-4515-b76e-0d3ba48b78d7 · inbound

CultureVLM: Characterizing and Improving Cultural Understanding of Vision-Language Models for over 100 Countries cites this paper.

CultureVLM: Characterizing and Improving Cultural Understanding of Vision-Language Models for over 100 Countries LLaVA-OneVision: Easy Visual Task Transfer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:36:03.762120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:36:03.762120Z digest=sha256:7a25e42cf4d918d89b0eb7067a08e7cbc4cd101adbca1c46b0664aa9edaae9b2

Observation 70559f8a-0c0f-4998-ae61-7b3f90067184 · inbound

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models cites this paper.

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:55.384987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:55.384987Z digest=sha256:eba5d01e0b0336d6746139a6a69d610bb69bf1efae0feb62d41dee95d9a4d93c

Observation e24928e3-7f57-46d1-a256-8ad9b5f43bf1 · inbound

HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding cites this paper.

HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:56.346349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:56.346349Z digest=sha256:10ec31baa8dea94e44de5a4337a2b08bfab118f49f419a44ceeea8c1cd9ddf1c

Observation 095f3798-dbbe-4038-ae20-9cd4c6adc80e · inbound

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction cites this paper.

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction LLaVA-OneVision: Easy Visual Task Transfer

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:08:19.699963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T21:08:19.570050Z digest=sha256:a7b67cd0fd6753a13b6e66d6631efec3162753d13f2f997649e5f281c9499cc4

Observation 245ac6f4-058b-4f12-9fe3-b532b4b7107b · inbound

FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models cites this paper.

FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T23:09:25.036482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:09:25.036482Z digest=sha256:a1f9db53a36d0564f09f4f5ce05561b59a3e463cdc4be8895993ea2a90eda1a0

Observation 7700c7d5-726f-4f78-8fbc-28b2963c9066 · inbound

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models cites this paper.

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-23T05:45:28.337895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T05:44:31.546843Z digest=sha256:1939cc843bd18281ba6e899a58cfa0a1b76347bb719ba5cce3463781fbd4fbe8

Observation 86a06654-be6c-4862-9611-717fdea28d43 · inbound

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos cites this paper.

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos LLaVA-OneVision: Easy Visual Task Transfer

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:39:22.495677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:39:22.340737Z digest=sha256:e22da14120fd86498081e0d89fcfa042cddac651a982f09d5a1a015424601970

Observation ee277cb2-d287-40fa-b6ce-2d5dba39fffc · inbound

Feedback-Driven Vision-Language Alignment with Minimal Human Supervision cites this paper.

Feedback-Driven Vision-Language Alignment with Minimal Human Supervision LLaVA-OneVision: Easy Visual Task Transfer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T21:34:35.139694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:34:35.139694Z digest=sha256:60e70f43cbb7bc0e755a44537f42b44dd06002d64db73de0f94bebd4ad20e430

Observation aa2e30a8-481b-46af-81f2-23513e7f024c · inbound

InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection cites this paper.

InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection LLaVA-OneVision: Easy Visual Task Transfer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T21:33:21.154696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:33:21.154696Z digest=sha256:d9cd06d9a491255f93fb41a587f21a5584a9ef2e19d41ec943fd8783b673afdc

Observation 19835f3a-e59f-45cf-b729-4f0f20fd8f72 · inbound

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs cites this paper.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.962372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.962372Z digest=sha256:29e164fc9ad4214e4defdd3cb85cc4a10ba8c4a8f136b3f9ced703eeac69653f

Observation 7bb1aab0-1a22-477b-876e-037b4f056aca · inbound

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios cites this paper.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios LLaVA-OneVision: Easy Visual Task Transfer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.324818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.324818Z digest=sha256:37601dd38cba53a87105493c2b70a63defeb42bbb1de9c3d6a60b5de2643791c

Observation 8af571e2-811a-4dea-b0dd-c08022b6a2b7 · inbound

LongViTU: Instruction Tuning for Long-Form Video Understanding cites this paper.

LongViTU: Instruction Tuning for Long-Form Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T21:23:57.915886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:23:57.915886Z digest=sha256:3944e2b628ea5313c8fa916db2565e1cec7dddf82a3ba82762f65fd3b9ab6115

Observation 94eb469f-ed67-42d2-87c5-e0d1e6ff5edd · inbound

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding cites this paper.

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:02:37.460440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T06:01:00.775721Z digest=sha256:08634f117cbe761ea6e7114cfaec2ab38708a2cbd95c5a7d1380e41ab8d86864

Observation 6b3e57b4-b5e3-4591-917b-d272259502de · inbound

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark cites this paper.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark LLaVA-OneVision: Easy Visual Task Transfer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.596503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.596503Z digest=sha256:4e932c2b20ed67024c1ef777d32ccc40d2da6142c176f30e8a5bb94a2db201ca

Observation a808c9b6-1b81-496f-8207-a2cd1df8271f · inbound

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models cites this paper.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.574358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.574358Z digest=sha256:b2fb0d872715a231d2e84e0871741a6847149a8b7b73051cc1b71165956c590e

Observation 7fdb3330-8930-43c6-945b-ef2c112488e4 · inbound

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design cites this paper.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design LLaVA-OneVision: Easy Visual Task Transfer

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.001752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.001752Z digest=sha256:6f98e3dd90acb8eb840e79f240ad7db6e94b9a19f519ec8ba801575d468d1efb

Observation 88d7ab99-7b94-4b1d-b854-299054cfac5a · inbound

Personalized Preference Fine-tuning of Diffusion Models cites this paper.

Personalized Preference Fine-tuning of Diffusion Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T21:01:56.804272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:01:56.804272Z digest=sha256:6476a23b4a061cdd0eac4622c57b7eaac90b240cb5281cd503fc3373b0c5218d

Observation 7cff97f2-83c7-460b-8ba4-38f3988fdfee · inbound

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness cites this paper.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness LLaVA-OneVision: Easy Visual Task Transfer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.026558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.026558Z digest=sha256:17b064a2e5bae661c3019a3d3d1e1548ab8777099d3d11704f3c6d4fdd7f3a60

Observation 138298ac-2959-49fd-afe0-2437fa49bdf8 · inbound

Embodied Scene Understanding for Vision Language Models via MetaVQA cites this paper.

Embodied Scene Understanding for Vision Language Models via MetaVQA LLaVA-OneVision: Easy Visual Task Transfer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:04.592076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:04.592076Z digest=sha256:0bcfe75087b9cdcbac1bed8eaac815f5e4180017906cc943cdcf2e30a3142d44

Observation 822b496c-ab5a-4bb3-b965-eeaea77617c3 · inbound

Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis cites this paper.

Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis LLaVA-OneVision: Easy Visual Task Transfer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T20:03:12.840613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:03:12.840613Z digest=sha256:40ca34fde8702446cd7fe137bb1eb2ac4ede7b7fa844ee9ecb29f939771222a3

Observation 5c9a4e94-fe01-4c2f-9b26-57de76e9657b · inbound

Universal Actions for Enhanced Embodied Foundation Models cites this paper.

Universal Actions for Enhanced Embodied Foundation Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T19:29:39.783689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:29:39.783689Z digest=sha256:3172bbde91b5fbe1d7ea3a9021cffd28fcd0939b5974fab0bd612d3d86def5db

Observation 7b034409-e9db-4e28-9444-bbfdec447f22 · inbound

HiMix: Reducing Computational Complexity in Large Vision-Language Models cites this paper.

HiMix: Reducing Computational Complexity in Large Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T19:18:49.272790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:18:49.272790Z digest=sha256:a4d4e799c8b7581a3c11584dac292802093220e9ade23cf387de2e005797f1cf

Observation d67c6540-880e-4403-8880-9d79099455cb · inbound

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling cites this paper.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling LLaVA-OneVision: Easy Visual Task Transfer

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:52:20.731060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:eb4e5437855130871932cc02073c5e453eb9405051c7aca427b19b60dd8e4ba4

Observation ad5bddca-0af7-47e9-b6d5-f44aa79fa8b4 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-11T01:19:59.859900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:783c6d78c0dde20212a29ebc1d08db308ea7a04cf599112705611ad89093b1f2

Observation bfd974dc-4d90-4fbd-adb7-e120ae28f162 · inbound

ReasVQA: Advancing VideoQA with Imperfect Reasoning Process cites this paper.

ReasVQA: Advancing VideoQA with Imperfect Reasoning Process LLaVA-OneVision: Easy Visual Task Transfer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T15:56:37.552702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:56:37.552702Z digest=sha256:784f4e80ef2d74ef5c87c4dd456627dbf34ccdc91f4c4a96499a26bd5a9c5273

Observation bc0b2524-7f91-4dc2-ba18-60771d57269b · inbound

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos cites this paper.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos LLaVA-OneVision: Easy Visual Task Transfer

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-14T00:32:41.237653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:7f60daf2461947a1dd659b0ad6d886c278a325065184d50f448229444cdb9992

Observation 2e3060e2-e581-4b3f-9f95-c8b9535aecbd · inbound

Pix2Cap-COCO: Advancing Visual Comprehension via Pixel-Level Captioning cites this paper.

Pix2Cap-COCO: Advancing Visual Comprehension via Pixel-Level Captioning LLaVA-OneVision: Easy Visual Task Transfer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:55.434519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:55.434519Z digest=sha256:6d05bd1da468b38f707f429ad3ba9eb8fbe6acd265610fe75d28cd5fa967ecfa

Observation 6d969358-0c24-4406-bef4-d42986ae8cd2 · inbound

Temporal Preference Optimization for Long-Form Video Understanding cites this paper.

Temporal Preference Optimization for Long-Form Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.158917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.158917Z digest=sha256:99298f54617e37fd2c181f8456c57e3bf1f56770902f1b6e86151c048de59600

Observation 1e7fa420-0d1b-4873-b353-6313d8aa430e · inbound

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models cites this paper.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.597651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.597651Z digest=sha256:310025c3eb04941a030dc71a7c31e419474e85f54082312d62a7271abe5166aa

Observation 18e7c0ae-9925-4d1b-a088-169b569ebbf9 · inbound

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step cites this paper.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step LLaVA-OneVision: Easy Visual Task Transfer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.324481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.324481Z digest=sha256:786dbbbf6d4d2add0576bf8ac8bd9e7bc6fd9c488130c40fcd54b32885942f45

Observation 792c9708-dcc9-4a30-8475-549b859867f3 · inbound

Redundancy Principles for MLLMs Benchmarks cites this paper.

Redundancy Principles for MLLMs Benchmarks LLaVA-OneVision: Easy Visual Task Transfer

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-10T18:27:55.282687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:27:55.282687Z digest=sha256:907880027514423aacc8190dfbe7aefabd66d9e7c094b6f65422d3f0e10219dd

Observation 140b4b66-190e-45ad-9a71-68564329d3f7 · inbound

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models cites this paper.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T15:18:49.965049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:18:49.965049Z digest=sha256:68466903bc9987759cebb40a94e057247c17549e7011279d185d4a5a6abcbea4

Observation 6886b653-5f38-4ed7-9357-ef3f1153e04b · inbound

HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation cites this paper.

HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation LLaVA-OneVision: Easy Visual Task Transfer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T14:58:25.810775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:58:25.810775Z digest=sha256:a4b299d3308fba3c145cfffd71b3ad57efcc232ee63b277a6916128171371947

Observation 0809cf4c-adbe-43f7-94eb-a2ecc17b0674 · inbound

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models cites this paper.

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:33.879095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:33.879095Z digest=sha256:d38eb796a6fc74ac35a25e402bb688403e0b15e540015d8ff5925838a94c6fcb

Observation 53d7b1d6-d0b3-41a9-8f8d-0ae6458b32f5 · inbound

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding cites this paper.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.229080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.229080Z digest=sha256:ab5dbf390be158e07c7cb9d8f39b1f0e5d59e3cb12b068ab097197afca103f57

Observation 8a0cea23-b7f3-4ccd-b990-e06dc059f830 · inbound

Return of the Encoder: Maximizing Parameter Efficiency for SLMs cites this paper.

Return of the Encoder: Maximizing Parameter Efficiency for SLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:27.692002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:27.692002Z digest=sha256:476d3a2df2cd2229fefd0181713d6a7fb65d107e60234b18615cac7214d8ac0f

Observation aab48173-787c-4e01-acf6-46778c48d019 · inbound

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding cites this paper.

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-10T10:48:17.837710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:48:17.837710Z digest=sha256:0c4ff9389129d950f04ba2051a95ad74b104a5c071cae945571ca1d8bb6cc61c

Observation 76bde895-cbad-43ba-996f-a18531bfe81c · inbound

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models cites this paper.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.227783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.227783Z digest=sha256:59c8199edbe835cc9ead90cd1bd4a38848c4c6eb0ab0875b707624560ed3f58a

Observation a6dbdbe5-26c2-4a7e-8986-ed95ddb263f6 · inbound

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs cites this paper.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.712002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.712002Z digest=sha256:d27f25f4708f88e615ebd93ce3f80200c05e3c2cb023aea52fd14307215789e7

Observation 3ef84a01-1a7c-4ad1-99c4-8492b79118c5 · inbound

Poison as Cure: Visual Noise for Mitigating Object Hallucinations in LVMs cites this paper.

Poison as Cure: Visual Noise for Mitigating Object Hallucinations in LVMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T21:10:05.898096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:10:05.898096Z digest=sha256:e0918607e4c3356033de01af84e87154fb1cd9e6df7726d6a07f43f17da89820

Observation c2e222d3-8f4c-4324-a005-b9586adcc476 · inbound

AIN: The Arabic INclusive Large Multimodal Model cites this paper.

AIN: The Arabic INclusive Large Multimodal Model LLaVA-OneVision: Easy Visual Task Transfer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T20:17:40.320719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:17:40.320719Z digest=sha256:9ea727bc59346e7ee46c730d4da355a6c798156f14d53a392e8a6beb0c40f82d

Observation 20b3b639-bf3c-474b-8144-d1e2523c39ce · inbound

Hypo3D: Exploring Hypothetical Reasoning in 3D cites this paper.

Hypo3D: Exploring Hypothetical Reasoning in 3D LLaVA-OneVision: Easy Visual Task Transfer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T17:12:16.442473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:12:16.442473Z digest=sha256:c0f64de908b4dcaaf8420c778e288f0d74ff6dc0758c9d43e882d3913bade53e

Observation 6db63981-1b8b-49ee-8a9a-716399238624 · inbound

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models cites this paper.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.852478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.852478Z digest=sha256:a45e40f80f4700f10f4dbadd61c435f2133103cc20b97a625e9d98165c99b9ce

Observation 9e54a25f-f822-42c6-aa14-b949d6a8f2cb · inbound

D-Attn: Decomposed Attention for Large Vision-and-Language Models cites this paper.

D-Attn: Decomposed Attention for Large Vision-and-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T14:05:14.363667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:05:14.363667Z digest=sha256:705177a56dbd13295c54480cf3b9bb0c169741c7b4c1acdf1780cd758ef9c1f1

Observation 3fd78b47-52aa-4346-917f-985d606fbd1e · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-17T05:53:26.263650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:5acbfd77c8c7760af1b04a53633d88cb51025587049cd7c9162d96886aa5bc28

Observation 2ccad04d-6854-46c4-9f16-3d646fad8786 · inbound

Ola: Pushing the Frontiers of Omni-Modal Language Model cites this paper.

Ola: Pushing the Frontiers of Omni-Modal Language Model LLaVA-OneVision: Easy Visual Task Transfer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.172390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.172390Z digest=sha256:3fd52cf30b0900cd44d44c7caeab76659da182489729cfbd2ed88f025c11ece0

Observation 4ebe73cf-d0f2-4ead-8403-41637f201321 · inbound

PerPO: Perceptual Preference Optimization via Discriminative Rewarding cites this paper.

PerPO: Perceptual Preference Optimization via Discriminative Rewarding LLaVA-OneVision: Easy Visual Task Transfer

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T06:01:13.273044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T06:01:13.273044Z digest=sha256:935547bf9f8355aaa89181923e97f51e7157562d15e9e0a6eb5b1c98c6481ff5

Observation 9d9897c6-5d50-4db0-afda-4117e9e9b6bb · inbound

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding cites this paper.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.861883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.861883Z digest=sha256:de202faefeac105e2dc3329505d96288c90e157aedf3f5db9e4bd17dfc820a78

Observation d68e8ec1-3498-43aa-bcbb-d638cf8dffe7 · inbound

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models cites this paper.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.715605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.715605Z digest=sha256:76c2b4df868047962f0e769feb7f2312d2560a647badd57162050aeb8f17db8f

Observation 66cc9f87-b8fd-41f8-a287-329117361fa8 · inbound

CTR-Driven Advertising Image Generation with Multimodal Large Language Models cites this paper.

CTR-Driven Advertising Image Generation with Multimodal Large Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:04.392994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:21:04.392994Z digest=sha256:85d672eca149640e69be7e9b17c2a3749fe4895a79ff898f7a08ebe729b05b2b

Observation 4788e605-f9db-45c5-ba01-a6e35624d98a · inbound

Space-Aware Instruction Tuning: Dataset and Benchmark for Guide Dog Robots Assisting the Visually Impaired cites this paper.

Space-Aware Instruction Tuning: Dataset and Benchmark for Guide Dog Robots Assisting the Visually Impaired LLaVA-OneVision: Easy Visual Task Transfer

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T13:36:46.002032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:36:46.002032Z digest=sha256:4f8dcf3763f0ddd94a07ba1faf9b3a5ff02f68f2948fe833beb57ae2aab57e84

Observation 20974f62-a450-491e-bb83-fa987735ae22 · inbound

Salamandra Technical Report cites this paper.

Salamandra Technical Report LLaVA-OneVision: Easy Visual Task Transfer

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-08T04:58:33.074279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:58:33.074279Z digest=sha256:3418f1a4737bc74d0a334b619032a9999f35c2cca2b131325e1b0c5a0360ef9f

Observation 58cde45a-7803-4029-8cba-4a3418c7745b · inbound

A Benchmark for Crime Surveillance Video Analysis with Large Models cites this paper.

A Benchmark for Crime Surveillance Video Analysis with Large Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T21:57:23.526197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:57:23.526197Z digest=sha256:fe8de1b46c69a95c12b101304266e7016df5b4658f0398350d1f8a3364e486cc

Observation 1f0f5200-02ce-46c2-a9b1-e48cbf77f46e · inbound

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence cites this paper.

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence LLaVA-OneVision: Easy Visual Task Transfer

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:00.804050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:06:00.804050Z digest=sha256:efd5ecc477c9b1070a6cf026ba5a7c7fba9614047ac48ffe2084fe27319ba83c

Observation 664a80c7-ddfa-489f-a597-752f7529b8da · inbound

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment cites this paper.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment LLaVA-OneVision: Easy Visual Task Transfer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.932313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.932313Z digest=sha256:507680a53ea33eec32b5d63e7e35a3f0877f730a684785f41565c4af7a79c35b

Observation 989e0b49-9604-4c3c-bfbd-75e55220e621 · inbound

MapNav: A Novel Memory Representation via Annotated Semantic Maps for Vision-and-Language Navigation cites this paper.

MapNav: A Novel Memory Representation via Annotated Semantic Maps for Vision-and-Language Navigation LLaVA-OneVision: Easy Visual Task Transfer

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-23T03:02:27.127822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-23T02:58:50.014240Z digest=sha256:23c55c3ed31a5c9c393199d55f5034ed73c92704aa9643aab2b7dd4036be713f

Observation 8ad6b221-ce0e-456f-8e74-0ce47edd838d · inbound

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations cites this paper.

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations LLaVA-OneVision: Easy Visual Task Transfer

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T19:45:18.892417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:45:18.892417Z digest=sha256:e463d3e92cbc5ebe47b35dc43822a2818d73f23de0785cdb20cab845cc238b32

Observation e19f28f0-ed57-49c5-90b5-9df326138136 · inbound

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs cites this paper.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs LLaVA-OneVision: Easy Visual Task Transfer

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:22:28.089575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:bc9fb6b999c763f4662385d97f1834a2b5af903c55c4c54174475fedd2f15e1d

Observation 330e4fcc-fab3-4e27-bc4b-3d8af310f41e · inbound

Visual-RFT: Visual Reinforcement Fine-Tuning cites this paper.

Visual-RFT: Visual Reinforcement Fine-Tuning LLaVA-OneVision: Easy Visual Task Transfer

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-13T22:16:16.551067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T22:16:16.528682Z digest=sha256:43fa23cabf2dd13f80196645c9d68e1cbcafbd221ea734539c2e4ba5e0284da8

Observation 3e32795d-619c-4477-a0c3-853d97f5530b · inbound

Seeing is Understanding: Unlocking Causal Attention into Modality-Mutual Attention for Multimodal LLMs cites this paper.

Seeing is Understanding: Unlocking Causal Attention into Modality-Mutual Attention for Multimodal LLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-23T01:25:16.427674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T01:23:01.892612Z digest=sha256:675cdadf641fde5cb34b517026cdab0da9f4f363eddd0d60f94a4f69e0810580

Observation e63186cb-5e27-413b-aa9d-a0f6e46c4dd9 · inbound

Unified Reward Model for Multimodal Understanding and Generation cites this paper.

Unified Reward Model for Multimodal Understanding and Generation LLaVA-OneVision: Easy Visual Task Transfer

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-14T00:44:30.899721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T00:44:30.558048Z digest=sha256:d6a799cf0cd91fb0f14cee789b1de8772ac6a748e83ada96490b5333b5ea4728

Observation 37fcc567-e408-4830-b985-deb0766b5266 · inbound

FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO cites this paper.

FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO LLaVA-OneVision: Easy Visual Task Transfer

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-23T00:22:18.751954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T00:21:51.621582Z digest=sha256:ebfefe741ecd891e90fa7e27f2ebaaa600e537f496006ae6893c3ebd780deea0

Observation 5f478673-d4e7-4e5c-a799-3f71b3291d25 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey LLaVA-OneVision: Easy Visual Task Transfer

Reference 275

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:18:53.611634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:e7f401f9437710e0824dc867565ea6bcbea5feac92312839a0411cb120f9f8b2

Observation 31b893a0-eb22-432f-9f81-e3f62936c929 · inbound

GuideDog: A Real-World Egocentric Multimodal Dataset for Blind and Low-Vision Accessibility-Aware Guidance cites this paper.

GuideDog: A Real-World Egocentric Multimodal Dataset for Blind and Low-Vision Accessibility-Aware Guidance LLaVA-OneVision: Easy Visual Task Transfer

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-22T23:52:17.035631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T23:49:53.830731Z digest=sha256:1dd5250bf8863c482f89000a49431475b762a6f5707384e323f2c27c1cdf10e4