Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T14:23:49.412830Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 100 of 179 outbound references and 100 inbound Pith citation observations for arXiv:2408.03326.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T14:23:49.412830Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T11:55:10.312257Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
100 of 179 outbound references displayed
External citation measurements
29
pith, observed 2026-08-05T02:28:24.338817Z
Observation 80aaa1d5-a0d1-4057-b0fb-c768eb867b2d · outbound
LLaVA-OneVision: Easy Visual Task Transfer Tallyqa: Answering complex counting questions
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8299c215-5244-4383-ad68-54ec3ce42dac · outbound
LLaVA-OneVision: Easy Visual Task Transfer Mathqa: Towards interpretable math word problem solving with operation-based formalisms
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3980e647-3419-4456-88e6-1200d3f2b3f2 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Claude-3.5
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 49780143-6ee2-445d-9afd-48c018046913 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Vqa: Visual question answering
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 939a577e-2304-43bd-a4e1-eb3fe5aba008 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Scanqa: 3d question answering for spatial scene understanding
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fff6ee15-b520-4502-bea8-745273ed97a9 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Scanqa: 3d question answering for spatial scene understanding
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 09b9f03d-3afb-4f80-b33b-2b88bcce453f · outbound
LLaVA-OneVision: Easy Visual Task Transfer Vision datasets: A benchmark for vision-based industrial inspection
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 36401b18-0905-4bb6-a6b7-7c5f545b0bb7 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4cc965fd-d81a-4718-b465-ca12e151f77f · outbound
LLaVA-OneVision: Easy Visual Task Transfer Visual question answering on image sets
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9977947e-a56d-4700-b6e6-ed7d5c078497 · outbound
LLaVA-OneVision: Easy Visual Task Transfer PaliGemma: A versatile 3B VLM for transfer
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 57fae0da-a23c-4282-8bc6-11e1d162d150 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Scene text visual question answering
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b11632f6-694c-4c6d-84ad-1067a68c4a81 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ba37d1e6-51c1-4e9d-a3ce-3614b774efa3 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Textocr-gpt4v
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a19211ea-8f19-46ac-b957-f00b74c18c02 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Mapqa: A dataset for question answering on choropleth maps
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e467d038-9853-47e7-b9ba-17e74180245d · outbound
LLaVA-OneVision: Easy Visual Task Transfer WebQA: Multihop and Multimodal QA
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 19822b22-57b7-44b1-9d38-ca6665b1f39d · outbound
LLaVA-OneVision: Easy Visual Task Transfer ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 62947ea0-eb12-401e-b3c2-43811352c8d0 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Unigeo: Unifying geometry logical reasoning via reformulating mathematical expression
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 244068af-5d7e-4930-9116-5ed59e936f3d · outbound
LLaVA-OneVision: Easy Visual Task Transfer Xing, and Liang Lin
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8babf22d-2ce6-4233-9f01-2eb45017b6cb · outbound
LLaVA-OneVision: Easy Visual Task Transfer Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 94c68049-8bc0-4ff0-94e3-41252d02cebd · outbound
LLaVA-OneVision: Easy Visual Task Transfer ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 13c3cee9-ad9c-4bf5-9fc0-a5837391ad38 · outbound
LLaVA-OneVision: Easy Visual Task Transfer ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 447de910-863d-4632-9536-226cceb1fdb0 · outbound
LLaVA-OneVision: Easy Visual Task Transfer InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 77253167-594a-4623-b5e2-128dcad032f3 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Hitab: A hierarchical table dataset for question answering and natural language generation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 04b06fc2-0556-4778-b1ca-40832d375154 · outbound
LLaVA-OneVision: Easy Visual Task Transfer PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f22179d7-39a8-4176-a85d-f58bcb331c17 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3e0c6430-5282-44e5-95a6-f0016fe55e8a · outbound
LLaVA-OneVision: Easy Visual Task Transfer Instructblip: Towards general-purpose vision-language models with instruction tuning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c22a8d7d-9323-48d2-8ec6-a634e3a95fd8 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Neural naturalist: Generating fine-grained image comparisons
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation db47c7bf-afaf-4fae-a598-b945cc5cf105 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Mme: A comprehensive evaluation benchmark for multimodal large language models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3011ee19-4410-42db-93f3-0292ec3202c3 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0e6c63e6-ab24-49f0-85c4-c660d45bb2fe · outbound
LLaVA-OneVision: Easy Visual Task Transfer Dreamsim: Learning new dimensions of human visual similarity using synthetic data
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d256730f-a2e3-4983-90d8-0043a8ba6f24 · outbound
LLaVA-OneVision: Easy Visual Task Transfer BLINK: Multimodal Large Language Models Can See but Not Perceive
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c8550678-7331-49be-a443-8f261fe114cf · outbound
LLaVA-OneVision: Easy Visual Task Transfer G-llava: Solving geometric problem with multi-modal large language model
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 38e0cf1a-67e1-4c36-947d-c0b3eb208f91 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ef769a95-7c90-4816-89cb-1101658638a1 · outbound
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ee4d18be-d8a8-4d3a-8433-57f5ae56d34d · outbound
LLaVA-OneVision: Easy Visual Task Transfer Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3c03def1-69a2-4cb7-82f1-7faa39ef1274 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Imagine this! scripts to compositions to videos
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2d47fe6a-6734-42ee-a1ba-b5c0909a299c · outbound
LLaVA-OneVision: Easy Visual Task Transfer Vizwiz grand challenge: Answering visual questions from blind people
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 877bde4c-0fbc-4d83-b9a9-3416de912c39 · outbound
LLaVA-OneVision: Easy Visual Task Transfer 3d-llm: Injecting the 3d world into large language models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9f16cc92-efff-4b33-8e86-5e857d6aca25 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Image change captioning by learning from an auxiliary task
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation aecba392-37a3-4508-94c3-ce749f3a07ed · outbound
LLaVA-OneVision: Easy Visual Task Transfer Huang, Francis Ferraro, Nasrin Mostafazadeh, Ishan Misra, Jacob Devlin, Aish- warya Agrawal, Ross Girshick, Xiaodong He, Pushmeet Kohli, Dhruv Batra, et al
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c8e012eb-f212-4ee1-b62d-4391ebfea9ec · outbound
LLaVA-OneVision: Easy Visual Task Transfer Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a8891752-b795-4f50-8f6d-f7041698b6de · outbound
LLaVA-OneVision: Easy Visual Task Transfer Hq-edit: A high-quality dataset for instruction-based image editing
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 009f1ef6-5956-4b13-bc90-5d9d2262356f · outbound
LLaVA-OneVision: Easy Visual Task Transfer Lim, and Edward H
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8a890935-c89e-4a05-a816-4319b646dab0 · outbound
LLaVA-OneVision: Easy Visual Task Transfer The amazing mysteries of the gutter: Drawing inferences between panels in comic book narratives
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation edf95c37-ef1b-485f-934b-4277e539d0f2 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Learning to Describe Differences Between Pairs of Similar Images
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b187c262-76b1-44b1-9e40-22657ef7bd7b · outbound
LLaVA-OneVision: Easy Visual Task Transfer Learning to describe differences between pairs of similar images
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7c9a2926-243f-46e2-999e-7107c13e8884 · outbound
LLaVA-OneVision: Easy Visual Task Transfer MANTIS: Interleaved Multi-Image Instruction Tuning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1a8ea4f3-11bf-4324-ad5a-3532696c0aef · outbound
LLaVA-OneVision: Easy Visual Task Transfer Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f24d1fd9-eea2-4afa-871c-acd1ac9ede1e · outbound
LLaVA-OneVision: Easy Visual Task Transfer Dvqa: Understanding data visualizations via question answering
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f52a5e4b-2ffe-4e2b-83fb-722adfdf2b32 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Figureqa: An annotated figure dataset for visual reasoning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a7731beb-ab5f-42e2-b610-b74dc1f08f26 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Prismatic vlms: Investigating the design space of visually-conditioned language models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b56f4798-8e8e-43ea-8c81-c08d90c27881 · outbound
LLaVA-OneVision: Easy Visual Task Transfer GeomVerse: A Systematic Evaluation of Large Models for Geometric Reasoning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4e69fa9d-8692-4b84-9521-c652b6d6df8a · outbound
LLaVA-OneVision: Easy Visual Task Transfer A diagram is worth a dozen images
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 910b65a7-0897-43f4-8873-c8c33340a03e · outbound
LLaVA-OneVision: Easy Visual Task Transfer A diagram is worth a dozen images
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 729ca3a5-b14d-40fa-aac9-6f9a80b095d4 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 227e1777-e9f2-4c22-a978-4b7a29c8e5b3 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e21be52d-e899-4b21-a1fd-be0231c46d5b · outbound
LLaVA-OneVision: Easy Visual Task Transfer The hateful memes challenge: Detecting hate speech in multimodal memes
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 54881f83-74b0-44a8-8905-1a6b3d11b93c · outbound
LLaVA-OneVision: Easy Visual Task Transfer Ocr-free document understanding transformer
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 78f3114f-fe43-484b-8dc0-97e79c7e5058 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Shamma, Michael S
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8bb47279-fe45-4ad2-bbcb-9caa9c469f91 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Image retrieval from contextual descriptions
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9a059159-be65-4fe1-b559-1a0c37a94e5b · outbound
LLaVA-OneVision: Easy Visual Task Transfer Sharegpt-4o: Comprehensive multimodal annotations with gpt-4o
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 00792893-240a-4162-a2bf-46c9319827f9 · outbound
LLaVA-OneVision: Easy Visual Task Transfer A dataset of clinically generated visual questions and answers about radiology images
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dfac752f-7085-48e8-9d86-61e7ad93252d · outbound
LLaVA-OneVision: Easy Visual Task Transfer What matters when building vision-language models? Technical Report
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d02dd82a-a468-4742-8b8d-ce036e2b3e9a · outbound
LLaVA-OneVision: Easy Visual Task Transfer Llava-next: What else influences visual instruction tuning beyond data?, May 2024
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 386fdd86-e05c-4908-8f51-5ee8d0da447e · outbound
LLaVA-OneVision: Easy Visual Task Transfer Llava-next: Stronger llms supercharge multimodal capabilities in the wild, May 2024
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b6d57fb6-77cc-422a-add0-a0574805be89 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Seed-bench: Benchmarking multimodal llms with generative comprehension
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 50d201ca-dee2-4307-a605-7fef220c485d · outbound
LLaVA-OneVision: Easy Visual Task Transfer Multimodal foundation models: From specialists to general-purpose assistants
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cc03fc8b-d50d-4277-a391-b08d86a5a1a4 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Llava-next: Tackling multi-image, video, and 3d in large multimodal models, June 2024
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c7f20297-8b3c-4512-af63-9b8c8d1e3fd2 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Fine-tuning multimodal llms to follow zero-shot demonstrative instructions
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation af34bcc1-16cc-41c3-ac41-7e8054bfa2de · outbound
LLaVA-OneVision: Easy Visual Task Transfer Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b2126fea-1174-459d-bbae-e5081aef274a · outbound
LLaVA-OneVision: Easy Visual Task Transfer Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 87379fcf-b02c-4021-ab2e-7cc62491152f · outbound
LLaVA-OneVision: Easy Visual Task Transfer Llama-vid: An image is worth 2 tokens in large language models
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2db2e25c-e584-4d85-a25a-ba409ac6c683 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Mini-gemini: Mining the potential of multi-modality vision language models
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 27c83f54-934c-476c-ba70-a69c5ea1e487 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Storygan: A sequential conditional gan for story visualization
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c8f15db4-19d6-4d83-a753-40dcc3c4c9ee · outbound
LLaVA-OneVision: Easy Visual Task Transfer Super-clevr: A virtual benchmark to diagnose domain robustness in visual reasoning
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f349e4c0-686b-4786-adaa-53a8edddfd56 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b219c21d-97bb-4821-80fc-731976455243 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Vila: On pre-training for visual language models
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4a79c2c9-be72-4508-8a9e-932f30570a87 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Lawrence Zitnick, and Piotr Dollár
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f21353e8-6e08-4dd1-875e-7d26ac3e8eae · outbound
LLaVA-OneVision: Easy Visual Task Transfer Visual spatial reasoning
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0cec71d8-58ef-48e2-81f2-a141b92d1efc · outbound
LLaVA-OneVision: Easy Visual Task Transfer Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 90ed9301-76ee-44d2-942d-a7e524a229ce · outbound
LLaVA-OneVision: Easy Visual Task Transfer Improved baselines with visual instruction tuning
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 296a0492-df03-4f34-83d4-1409f908b7af · outbound
LLaVA-OneVision: Easy Visual Task Transfer Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation df059886-a620-40c8-809c-44d9f22f6ddc · outbound
LLaVA-OneVision: Easy Visual Task Transfer Visual instruction tuning
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3a933ec1-7f26-4d31-9304-f0cb90657c0b · outbound
LLaVA-OneVision: Easy Visual Task Transfer What large language models bring to text-rich vqa?
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d7382f9a-093c-4f1b-b359-133d9c7af79c · outbound
LLaVA-OneVision: Easy Visual Task Transfer What Large Language Models Bring to Text-rich VQA?
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 679b2922-8129-4a51-a545-9f4b041ca91c · outbound
LLaVA-OneVision: Easy Visual Task Transfer Mmbench: Is your multi-modal model an all-around player? Technical Report
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6ad57d36-1cb8-4920-9b89-03599fb5c899 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Video detail caption
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 15c7f4b3-7a98-457c-b8fb-1726088e38ae · outbound
LLaVA-OneVision: Easy Visual Task Transfer The flan collection: Designing data and methods for effective instruction tuning
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8a071b03-c99c-40c0-892a-fbdcfbee4807 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Deepseek-vl: towards real-world vision-language under- standing
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 46eb10de-5d99-40f8-9b91-ed5e2b195334 · outbound
LLaVA-OneVision: Easy Visual Task Transfer MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e30d940b-ceaf-4804-a7d7-704711b348f8 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8d34a9a0-7931-44fc-ade0-5176bdd9e566 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0ef779fb-b910-4310-95b4-699b47a8e80c · outbound
LLaVA-OneVision: Easy Visual Task Transfer Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4039bad7-b895-4c21-9fd6-4e39e408e529 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 79454796-eea8-4dbc-b54d-4353a6dd42a8 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Iconqa: A new benchmark for abstract diagram understanding and visual language reasoning
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ea453277-00f1-42d0-a90c-d7d711bff62d · outbound
LLaVA-OneVision: Easy Visual Task Transfer Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 26a70b3e-c92f-4f6e-b92a-d4e222da2445 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Video-chatgpt: Towards detailed video understanding via large vision and language models
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cbbad9f4-b95f-44e3-a807-34efa192c4c6 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Egoschema: A diagnostic benchmark for very long-form video language understanding
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1fe2b577-ad07-40b4-a9db-b016eeaaf124 · outbound
LLaVA-OneVision: Easy Visual Task Transfer Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c289ac2f-f956-4e93-b0a0-4ba45a1478dd · outbound
LLaVA-OneVision: Easy Visual Task Transfer The iam-database: an english sentence database for offline handwriting recognition
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0570e303-6a28-46cf-ac4e-364240baf1ae · inbound
MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems? LLaVA-OneVision: Easy Visual Task Transfer
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6f7651f1-27ca-4e49-83b1-cbe48b54ede8 · inbound
Hallucination of Multimodal Large Language Models: A Survey LLaVA-OneVision: Easy Visual Task Transfer
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2c152a20-a77a-413d-abf3-ce299d96e22f · inbound
MLVU: Benchmarking Multi-task Long Video Understanding LLaVA-OneVision: Easy Visual Task Transfer
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation eba91837-dbc0-4889-be20-f007cefdde2a · inbound
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs LLaVA-OneVision: Easy Visual Task Transfer
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ac23ea09-a587-4b6e-b33f-5b690a3ffe08 · inbound
LVBench: An Extreme Long Video Understanding Benchmark LLaVA-OneVision: Easy Visual Task Transfer
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2680417b-f22a-4d16-9b97-a82fac67b69f · inbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos LLaVA-OneVision: Easy Visual Task Transfer
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d3264127-4082-440d-b911-47e4aef871af · inbound
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark LLaVA-OneVision: Easy Visual Task Transfer
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0a1fa6ab-bb1e-4d10-8053-2d0f6c1ab2e4 · inbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e462826b-9284-47bd-a9c0-47d812daf8d6 · inbound
Emu3: Next-Token Prediction is All You Need LLaVA-OneVision: Easy Visual Task Transfer
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 42b2e88a-7121-4d11-826f-031c99e0148b · inbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data LLaVA-OneVision: Easy Visual Task Transfer
Reference 158
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0d29a7cc-cbba-45a5-890d-5e6d3fea689f · inbound
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c06c2da2-63ee-4f38-9835-722a7a36d9ad · inbound
VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents LLaVA-OneVision: Easy Visual Task Transfer
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0fa72a6b-fc80-4bf0-bf2b-61ff6ea989ba · inbound
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation LLaVA-OneVision: Easy Visual Task Transfer
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 514efd96-2704-4416-b953-0debf9d077a6 · inbound
LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding LLaVA-OneVision: Easy Visual Task Transfer
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 18dd4ca0-d149-44c2-b463-214f2f9c2526 · inbound
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization LLaVA-OneVision: Easy Visual Task Transfer
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1ea11bc5-4b4f-455c-9579-39a22dbbaec1 · inbound
Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction LLaVA-OneVision: Easy Visual Task Transfer
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation de0a1d31-fe8c-4354-aea6-046e16d7fbde · inbound
NVILA: Efficient Frontier Visual Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 151592b6-8e8e-4e49-8135-ac20532f0abd · inbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling LLaVA-OneVision: Easy Visual Task Transfer
Reference 124
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2415b5c5-f4d0-4c83-9154-f9deeb18a3b5 · inbound
DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding LLaVA-OneVision: Easy Visual Task Transfer
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 136d2089-23bd-4ed7-bdfe-6060310f1c1f · inbound
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning LLaVA-OneVision: Easy Visual Task Transfer
Reference 296
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 93469cd3-eb59-4e3c-8195-0cb0d796480e · inbound
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces LLaVA-OneVision: Easy Visual Task Transfer
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c27ac01b-8481-4ccd-a77b-7a572bde8ea8 · inbound
Progressive Multimodal Reasoning via Active Retrieval LLaVA-OneVision: Easy Visual Task Transfer
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6ea7ed6-1071-4b89-96a8-992875141e06 · inbound
EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues LLaVA-OneVision: Easy Visual Task Transfer
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e03463c-dc7a-4dfc-9543-94655fca8ffc · inbound
Error-driven Data-efficient Large Multimodal Model Tuning LLaVA-OneVision: Easy Visual Task Transfer
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7366d366-22a8-42ed-b489-fff6b0b1a64c · inbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be2a0fa4-5fcd-42d2-8c89-e5a8607e13d9 · inbound
A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation LLaVA-OneVision: Easy Visual Task Transfer
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7f22d14-e1ba-485c-a2d6-23804eac654f · inbound
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities LLaVA-OneVision: Easy Visual Task Transfer
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8bebfc0-b866-4763-aef4-58a647a6b53c · inbound
Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment LLaVA-OneVision: Easy Visual Task Transfer
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 463dba09-af75-409e-a9ab-143d5bb24204 · inbound
Hear the Scene: Audio-Enhanced Text Spotting LLaVA-OneVision: Easy Visual Task Transfer
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00ee3d80-6cce-4d23-a9b5-4a1f09edcb9d · inbound
MBQ: Modality-Balanced Quantization for Large Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94026a25-586b-472d-9aca-2226bc1caefe · inbound
From Elements to Design: A Layered Approach for Automatic Graphic Design Composition LLaVA-OneVision: Easy Visual Task Transfer
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 661e7cc0-3ca7-438a-a978-04ee1868d002 · inbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning LLaVA-OneVision: Easy Visual Task Transfer
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 707e2774-c404-43ff-8072-de78d6e6b17f · inbound
Probing Visual Language Priors in VLMs LLaVA-OneVision: Easy Visual Task Transfer
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06e36acb-10d6-4385-93c0-b600a5db02d5 · inbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LLaVA-OneVision: Easy Visual Task Transfer
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6a1c409b-61bf-48f0-8294-2bdf2f324a5f · inbound
Online Video Understanding: OVBench and VideoChat-Online LLaVA-OneVision: Easy Visual Task Transfer
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a727e5f-c7b7-400c-9958-13972947ccbc · inbound
VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM LLaVA-OneVision: Easy Visual Task Transfer
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ced4e771-125b-4515-b76e-0d3ba48b78d7 · inbound
CultureVLM: Characterizing and Improving Cultural Understanding of Vision-Language Models for over 100 Countries LLaVA-OneVision: Easy Visual Task Transfer
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70559f8a-0c0f-4998-ae61-7b3f90067184 · inbound
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e24928e3-7f57-46d1-a256-8ad9b5f43bf1 · inbound
HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding LLaVA-OneVision: Easy Visual Task Transfer
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 095f3798-dbbe-4038-ae20-9cd4c6adc80e · inbound
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction LLaVA-OneVision: Easy Visual Task Transfer
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 245ac6f4-058b-4f12-9fe3-b532b4b7107b · inbound
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7700c7d5-726f-4f78-8fbc-28b2963c9066 · inbound
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 86a06654-be6c-4862-9611-717fdea28d43 · inbound
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos LLaVA-OneVision: Easy Visual Task Transfer
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ee277cb2-d287-40fa-b6ce-2d5dba39fffc · inbound
Feedback-Driven Vision-Language Alignment with Minimal Human Supervision LLaVA-OneVision: Easy Visual Task Transfer
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa2e30a8-481b-46af-81f2-23513e7f024c · inbound
InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection LLaVA-OneVision: Easy Visual Task Transfer
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19835f3a-e59f-45cf-b729-4f0f20fd8f72 · inbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs LLaVA-OneVision: Easy Visual Task Transfer
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bb1aab0-1a22-477b-876e-037b4f056aca · inbound
Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios LLaVA-OneVision: Easy Visual Task Transfer
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8af571e2-811a-4dea-b0dd-c08022b6a2b7 · inbound
LongViTU: Instruction Tuning for Long-Form Video Understanding LLaVA-OneVision: Easy Visual Task Transfer
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94eb469f-ed67-42d2-87c5-e0d1e6ff5edd · inbound
LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding LLaVA-OneVision: Easy Visual Task Transfer
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6b3e57b4-b5e3-4591-917b-d272259502de · inbound
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark LLaVA-OneVision: Easy Visual Task Transfer
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a808c9b6-1b81-496f-8207-a2cd1df8271f · inbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fdb3330-8930-43c6-945b-ef2c112488e4 · inbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design LLaVA-OneVision: Easy Visual Task Transfer
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88d7ab99-7b94-4b1d-b854-299054cfac5a · inbound
Personalized Preference Fine-tuning of Diffusion Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cff97f2-83c7-460b-8ba4-38f3988fdfee · inbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness LLaVA-OneVision: Easy Visual Task Transfer
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 138298ac-2959-49fd-afe0-2437fa49bdf8 · inbound
Embodied Scene Understanding for Vision Language Models via MetaVQA LLaVA-OneVision: Easy Visual Task Transfer
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 822b496c-ab5a-4bb3-b965-eeaea77617c3 · inbound
Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis LLaVA-OneVision: Easy Visual Task Transfer
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c9a4e94-fe01-4c2f-9b26-57de76e9657b · inbound
Universal Actions for Enhanced Embodied Foundation Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b034409-e9db-4e28-9444-bbfdec447f22 · inbound
HiMix: Reducing Computational Complexity in Large Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d67c6540-880e-4403-8880-9d79099455cb · inbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling LLaVA-OneVision: Easy Visual Task Transfer
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ad5bddca-0af7-47e9-b6d5-f44aa79fa8b4 · inbound
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding LLaVA-OneVision: Easy Visual Task Transfer
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bfd974dc-4d90-4fbd-adb7-e120ae28f162 · inbound
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process LLaVA-OneVision: Easy Visual Task Transfer
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc0b2524-7f91-4dc2-ba18-60771d57269b · inbound
Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos LLaVA-OneVision: Easy Visual Task Transfer
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2e3060e2-e581-4b3f-9f95-c8b9535aecbd · inbound
Pix2Cap-COCO: Advancing Visual Comprehension via Pixel-Level Captioning LLaVA-OneVision: Easy Visual Task Transfer
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d969358-0c24-4406-bef4-d42986ae8cd2 · inbound
Temporal Preference Optimization for Long-Form Video Understanding LLaVA-OneVision: Easy Visual Task Transfer
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e7fa420-0d1b-4873-b353-6313d8aa430e · inbound
IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18e7c0ae-9925-4d1b-a088-169b569ebbf9 · inbound
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step LLaVA-OneVision: Easy Visual Task Transfer
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 792c9708-dcc9-4a30-8475-549b859867f3 · inbound
Redundancy Principles for MLLMs Benchmarks LLaVA-OneVision: Easy Visual Task Transfer
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 140b4b66-190e-45ad-9a71-68564329d3f7 · inbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6886b653-5f38-4ed7-9357-ef3f1153e04b · inbound
HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation LLaVA-OneVision: Easy Visual Task Transfer
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0809cf4c-adbe-43f7-94eb-a2ecc17b0674 · inbound
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53d7b1d6-d0b3-41a9-8f8d-0ae6458b32f5 · inbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding LLaVA-OneVision: Easy Visual Task Transfer
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a0cea23-b7f3-4ccd-b990-e06dc059f830 · inbound
Return of the Encoder: Maximizing Parameter Efficiency for SLMs LLaVA-OneVision: Easy Visual Task Transfer
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aab48173-787c-4e01-acf6-46778c48d019 · inbound
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding LLaVA-OneVision: Easy Visual Task Transfer
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76bde895-cbad-43ba-996f-a18531bfe81c · inbound
LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6dbdbe5-26c2-4a7e-8986-ed95ddb263f6 · inbound
RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs LLaVA-OneVision: Easy Visual Task Transfer
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ef84a01-1a7c-4ad1-99c4-8492b79118c5 · inbound
Poison as Cure: Visual Noise for Mitigating Object Hallucinations in LVMs LLaVA-OneVision: Easy Visual Task Transfer
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2e222d3-8f4c-4324-a005-b9586adcc476 · inbound
AIN: The Arabic INclusive Large Multimodal Model LLaVA-OneVision: Easy Visual Task Transfer
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20b3b639-bf3c-474b-8144-d1e2523c39ce · inbound
Hypo3D: Exploring Hypothetical Reasoning in 3D LLaVA-OneVision: Easy Visual Task Transfer
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6db63981-1b8b-49ee-8a9a-716399238624 · inbound
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e54a25f-f822-42c6-aa14-b949d6a8f2cb · inbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fd78b47-52aa-4346-917f-985d606fbd1e · inbound
WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs LLaVA-OneVision: Easy Visual Task Transfer
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2ccad04d-6854-46c4-9f16-3d646fad8786 · inbound
Ola: Pushing the Frontiers of Omni-Modal Language Model LLaVA-OneVision: Easy Visual Task Transfer
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ebe73cf-d0f2-4ead-8403-41637f201321 · inbound
PerPO: Perceptual Preference Optimization via Discriminative Rewarding LLaVA-OneVision: Easy Visual Task Transfer
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d9897c6-5d50-4db0-afda-4117e9e9b6bb · inbound
UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding LLaVA-OneVision: Easy Visual Task Transfer
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d68e8ec1-3498-43aa-bcbb-d638cf8dffe7 · inbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66cc9f87-b8fd-41f8-a287-329117361fa8 · inbound
CTR-Driven Advertising Image Generation with Multimodal Large Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4788e605-f9db-45c5-ba01-a6e35624d98a · inbound
Space-Aware Instruction Tuning: Dataset and Benchmark for Guide Dog Robots Assisting the Visually Impaired LLaVA-OneVision: Easy Visual Task Transfer
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20974f62-a450-491e-bb83-fa987735ae22 · inbound
Salamandra Technical Report LLaVA-OneVision: Easy Visual Task Transfer
Reference 111
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58cde45a-7803-4029-8cba-4a3418c7745b · inbound
A Benchmark for Crime Surveillance Video Analysis with Large Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f0f5200-02ce-46c2-a9b1-e48cbf77f46e · inbound
Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence LLaVA-OneVision: Easy Visual Task Transfer
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 664a80c7-ddfa-489f-a597-752f7529b8da · inbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment LLaVA-OneVision: Easy Visual Task Transfer
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 989e0b49-9604-4c3c-bfbd-75e55220e621 · inbound
MapNav: A Novel Memory Representation via Annotated Semantic Maps for Vision-and-Language Navigation LLaVA-OneVision: Easy Visual Task Transfer
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8ad6b221-ce0e-456f-8e74-0ce47edd838d · inbound
A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations LLaVA-OneVision: Easy Visual Task Transfer
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e19f28f0-ed57-49c5-90b5-9df326138136 · inbound
Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs LLaVA-OneVision: Easy Visual Task Transfer
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 330e4fcc-fab3-4e27-bc4b-3d8af310f41e · inbound
Visual-RFT: Visual Reinforcement Fine-Tuning LLaVA-OneVision: Easy Visual Task Transfer
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3e32795d-619c-4477-a0c3-853d97f5530b · inbound
Seeing is Understanding: Unlocking Causal Attention into Modality-Mutual Attention for Multimodal LLMs LLaVA-OneVision: Easy Visual Task Transfer
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e63186cb-5e27-413b-aa9d-a0f6e46c4dd9 · inbound
Unified Reward Model for Multimodal Understanding and Generation LLaVA-OneVision: Easy Visual Task Transfer
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 37fcc567-e408-4830-b985-deb0766b5266 · inbound
FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO LLaVA-OneVision: Easy Visual Task Transfer
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5f478673-d4e7-4e5c-a799-3f71b3291d25 · inbound
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey LLaVA-OneVision: Easy Visual Task Transfer
Reference 275
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 31b893a0-eb22-432f-9f81-e3f62936c929 · inbound
GuideDog: A Real-World Egocentric Multimodal Dataset for Blind and Low-Vision Accessibility-Aware Guidance LLaVA-OneVision: Easy Visual Task Transfer
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.