Pith. sign in

Paper Citation Record · LEDGER

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

As of 11 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 100 inbound Pith citation observations for arXiv:2509.23661.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.23661 v3

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T10:53:26.574350Z

measured 112 of 112 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 100 of 129 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T18:15:04.579056Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact10
  • verified fuzzy2
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation cbc9b1a2-80e2-4ec9-9357-be6e956ad599 · outbound

This paper cites Qwen2.5-VL Technical Report.

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training Qwen2.5-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:53:26.605927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T10:53:26.574350Z digest=sha256:e67c688397fd4c56f2c8cc7df84a05ddc1eb5a2f3175e2f09dd083ca99710fb3

Observation b62095e3-85b8-401e-96ee-90db625f3260 · outbound

This paper cites Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models.

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:53:26.615927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T10:53:26.574350Z digest=sha256:626b82d45761377c84258aee28595cfa8c34dfadf8eb5ba11406c3f0754c6fef

Observation 6be7f074-8350-42ff-ae4e-2dfaa8007e4d · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:53:26.625638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T10:53:26.574350Z digest=sha256:28f609453a12e8a38597a4ab26dc7c0ca91b00869c54857302a6eb7e4bd996e5

Observation 0ccb065e-fc53-4243-8d35-1912f4ee181b · outbound

This paper cites Seed1.5-VL Technical Report.

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training Seed1.5-VL Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:53:26.633881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T10:53:26.574350Z digest=sha256:bc5a533a318e1146b60f64ad98ee7f65d8146c17c726a4e4880b77f9474137cd

Observation 0e9fd3db-5621-447e-a107-d9caf1fe1d66 · outbound

This paper cites Improved baselines with visual instruction tuning.

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training Improved baselines with visual instruction tuning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:53:26.690243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T10:53:26.574350Z digest=sha256:f987c0be9465f734e3d1763c57600ecdb5ab9ae3f82a2121c29caa445104b7de

Observation 6e7f1377-5f59-4b6e-a128-dc2373558a68 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:53:26.659111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T10:53:26.574350Z digest=sha256:4db29de011c3776a711a72a793a5b887801c573c82807e751e8ad68d894ffbb1

Observation d4721ba2-2861-46cd-984a-826ed8ccff76 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:53:26.665304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T10:53:26.574350Z digest=sha256:5d0ae48f246393d995ad8d01b91416516e1261745ec27acb36677a7ab8eb4b05

Observation cf0b588f-8e48-4056-aa29-6a671db1bf53 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:53:26.674948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T10:53:26.574350Z digest=sha256:11c19f58b7a0306b0f628c231b9918ccba329c122701be1f5c6a9c003cd8c700

Observation 3b1c8059-5c00-4693-8df0-5593e2087e77 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:53:26.683427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T10:53:26.574350Z digest=sha256:8ccff19d130e47c2d8858a5d51b72a6656f774460293c0714c6aeaaed546e7dd

Observation 9afd1751-b2fb-4b5d-a371-f26283a6438e · outbound

This paper cites Fengshenbang 1.0: Being the Foundation of Chinese Cognitive Intelligence.

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training Fengshenbang 1.0: Being the Foundation of Chinese Cognitive Intelligence

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:53:26.642822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T10:53:26.574350Z digest=sha256:a566f82a744f1073403a84ecc22b9ce0c5d0acc13f4cd735cd9cafeca813afb9

Observation c7c9e0a2-5806-4133-bc6d-1ce5d4549367 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:53:26.651501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T10:53:26.574350Z digest=sha256:2ce1af8d4be07fb16427a5b7a584abec404384bd5df3af91389db79844ce3146

Observation c92bab24-d994-458e-8c1d-ba831700dccb · outbound

This paper cites Qwen2.5-VL with Same LLM ToenableafaircomparisonwithQwen2.5-VL,wetrainLLaVA-Onevision-1.5-3BbasedonQwen2.5- 3B-Instruct.

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training Qwen2.5-VL with Same LLM ToenableafaircomparisonwithQwen2.5-VL,wetrainLLaVA-Onevision-1.5-3BbasedonQwen2.5- 3B-Instruct

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T10:53:26.696416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T10:53:26.574350Z digest=sha256:6b2915ffb1ea1dbecf5b8fd1cd750cf9dfe57f1608a952ee6a5c13f4f95a6d47

Pith citing papers

Observation b917bfce-3735-45c4-80af-82bfe0998b31 · inbound

Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs cites this paper.

Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-17T05:35:13.215458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:35:13.118221Z digest=sha256:f6748e2b1dbf5c271f74efbb4300872b18f9c711cca1cb64f4094883de6193af

Observation 1ac081a8-7e28-4f8c-bf34-9c7875622ae9 · inbound

RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension cites this paper.

RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T18:15:04.579056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:15:04.579056Z digest=sha256:592f1914570166326f8ec35b464890adedfae09f49a5c30b55efd0aebd97e516

Observation 5356b479-aeb8-46c3-bb09-b613b6c54de4 · inbound

$M^3-Verse$: A "Spot the Difference" Challenge for Large Multimodal Models cites this paper.

$M^3-Verse$: A "Spot the Difference" Challenge for Large Multimodal Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T14:57:29.090391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:57:29.090391Z digest=sha256:9f8dfd3773da556440b7e49b2d45566d201b23d35e836a84290448c8e396b66c

Observation a56e1fd8-4266-4845-a1f8-4e3861645a2d · inbound

Generative Semantic Multi-Object Tracking: A Large-Scale Benchmark and an MLLM-Driven Reasoning Framework cites this paper.

Generative Semantic Multi-Object Tracking: A Large-Scale Benchmark and an MLLM-Driven Reasoning Framework LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T11:28:04.980968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:28:04.980968Z digest=sha256:c6f48deeda945a9dcdcdb693a55bfece8660a677957a2612dbedafb92a4b26c1

Observation 5716265c-27fa-4e24-a474-7d6ee936cdbc · inbound

PersonalAlign: Hierarchical Implicit Intent Alignment for Personalized GUI Agent with Long-Term User-Centric Records cites this paper.

PersonalAlign: Hierarchical Implicit Intent Alignment for Personalized GUI Agent with Long-Term User-Centric Records LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-16T14:32:59.922873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T14:32:36.257645Z digest=sha256:5d3702f627658a442e3fa1048c1fa1da4ed92a936361adc3ebd2f59628f57f44

Observation 5a194843-e605-457b-8ed1-d8df4eb8ec76 · inbound

Dual Latent Memory for Visual Multi-agent System cites this paper.

Dual Latent Memory for Visual Multi-agent System LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T06:04:50.083863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:04:50.083863Z digest=sha256:3b478b572f04bb753b1ead8592bed8023b5d11874d17d2f25b4d0fefeb1ed9fb

Observation 349738ae-b01c-42f3-91c0-9874107e2033 · inbound

R1-SyntheticVL: Is Synthetic Data from Generative Models Ready for Multimodal Large Language Model? cites this paper.

R1-SyntheticVL: Is Synthetic Data from Generative Models Ready for Multimodal Large Language Model? LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T05:06:43.872810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:06:43.872810Z digest=sha256:86604b5f0a3505557d3ed7263394fddd1ffd5b490f8196a460323211e69a4052

Observation 73b4d90e-ff62-4713-ab13-0b5c884139de · inbound

From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models cites this paper.

From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T03:16:49.965608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:16:49.965608Z digest=sha256:bba70bad8cb467f3ed24e7d3118b8b3c6d812ff80eda001862779373ae36613c

Observation 85f3831f-f314-4016-b1fa-e21ad07c8f9b · inbound

How to Take a Memorable Picture? Empowering Users with Actionable Feedback cites this paper.

How to Take a Memorable Picture? Empowering Users with Actionable Feedback LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T20:57:10.768968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:57:10.768968Z digest=sha256:a9439e90afca011c862f2cd9222a16768edd7caa6edb653247d3e7b6e27fcf93

Observation 1055f925-b328-461c-b4ae-ae9e7c4202ec · inbound

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models cites this paper.

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:26:26.935702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T18:25:21.621268Z digest=sha256:0d104590d4ebc35f4d6db664bdd05287124386282857170c39b7a80d98755334

Observation ed61e2d9-dd4d-4810-bf4c-937f630c789a · inbound

SCP: Spatial Causal Prediction in Video cites this paper.

SCP: Spatial Causal Prediction in Video LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-15T16:50:11.662033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T16:47:44.523606Z digest=sha256:0b738a0461e0cb27f89554495d7e8e1634d877214ad938b8d09e72688170520e

Observation 127ea1cc-24a0-4cb0-bdb4-b7814eccc74c · inbound

What if? Emulative Simulation with World Models for Situated Reasoning cites this paper.

What if? Emulative Simulation with World Models for Situated Reasoning LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-15T13:51:30.008232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:51:30.008232Z digest=sha256:1bca44ceee774115efdc9e9fcafe5d33daba9e2da5179ee000019daed4000482

Observation aba40587-7c77-41d3-ba35-b127dbd705b6 · inbound

SlowBA: An efficiency backdoor attack towards VLM-based GUI agents cites this paper.

SlowBA: An efficiency backdoor attack towards VLM-based GUI agents LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-15T12:42:28.672749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T12:42:28.672749Z digest=sha256:41ccf5fd3c27a3b683634f6e59019b7569a49574165986e7f0a58e14d89a237e

Observation 1e00bf86-2b95-41c9-bcc9-462e83cb5da0 · inbound

HiMemVLN: Enhancing Reliability of Open-Source Zero-Shot Vision-and-Language Navigation with Hierarchical Memory System cites this paper.

HiMemVLN: Enhancing Reliability of Open-Source Zero-Shot Vision-and-Language Navigation with Hierarchical Memory System LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T18:10:17.054079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:10:17.054079Z digest=sha256:cc056d028c41d9b2d3d424558fb6160551e03d570e39e6714f94f460babbecb1

Observation 1947a29a-bdd4-4bc0-a9d5-acd628fd37d9 · inbound

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering cites this paper.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:871fd10c3c52e7550ead9b17ad278a4ea2550be3abf69c1fb588a9bde1b7060c

Observation 3142730b-f9d2-408f-9a3f-7ecc346f62ca · inbound

DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline Optimization cites this paper.

DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline Optimization LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-21T11:10:02.348735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T11:05:07.336827Z digest=sha256:108657fba9f16c7a6647268fabfa45971857d4212d10d54b3ca78ba4eea9b7f0

Observation 5bdf1458-ab84-49fb-bfbe-bae8e742e76b · inbound

Peel neighborhoods cites this paper.

Peel neighborhoods LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T20:08:04.785209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T20:08:04.785209Z digest=sha256:f1841c2a6ebb9e26e2f12b165fbb9a3864640deafeec1894db4e0165da2b009f

Observation 2e90120b-beb8-4475-95d0-04bff2add184 · inbound

ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding cites this paper.

ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-14T22:39:32.904073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T22:39:06.113655Z digest=sha256:0c3e3a08f297020b549b468a94ab572b2ed829c71153794091f8c5195904ba00

Observation cd311869-f0bd-4e8a-b264-8c4536981707 · inbound

Token Warping Helps MLLMs Look from Nearby Viewpoints cites this paper.

Token Warping Helps MLLMs Look from Nearby Viewpoints LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-13T21:08:17.284561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T21:07:55.062113Z digest=sha256:58660833445eefd2bbd62c18cf22bb03e00c86e9aba62750261a053f70348e7b

Observation 01e30d6b-ae9a-4cde-9452-8ffd76e762b7 · inbound

BoxComm: Benchmarking Category-Aware Commentary Generation and Narration Rhythm in Boxing cites this paper.

BoxComm: Benchmarking Category-Aware Commentary Generation and Narration Rhythm in Boxing LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:53:26.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T19:10:07.699225Z digest=sha256:6987558e6eddcfa485e216e5de273120d969447513524cf7011a15e7c90ef2ac

Observation cfbfdf4f-b70c-4886-bfa7-560635bcffe1 · inbound

Steering the Verifiability of Multimodal AI Hallucinations cites this paper.

Steering the Verifiability of Multimodal AI Hallucinations LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:53:26.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:19:45.405698Z digest=sha256:44ce358f1335d664fc3ca273fa3813f65a8c9a854d274dc3b2c1408fae1362a8

Observation 85be7987-200e-4b19-a14a-0db85f1bb29d · inbound

Enhancing MLLM Spatial Understanding via Active 3D Scene Exploration for Multi-Perspective Reasoning cites this paper.

Enhancing MLLM Spatial Understanding via Active 3D Scene Exploration for Multi-Perspective Reasoning LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:53:26.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T19:16:46.753641Z digest=sha256:1d960a5712113b1d865e44e7a6ccfdf388744caa4d4bd29f2416cdc6070735c2

Observation 787b359f-03c1-4185-a051-bdb401ced8dc · inbound

Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models cites this paper.

Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:53:26.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:46:26.869644Z digest=sha256:9db3ed8cf0f9c84ab21085234fa54c25b65c2260745b7a5e92d61978bb8f15be

Observation 5016fc86-0c22-4adc-9087-09269cfb3ac9 · inbound

Small Vision-Language Models are Smart Compressors for Long Video Understanding cites this paper.

Small Vision-Language Models are Smart Compressors for Long Video Understanding LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:53:26.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:38:49.654073Z digest=sha256:e1c4f8e5f7fac887c83aeccb4b3856b2c281c296a89d9ea3c436c0ddf7bc85b5

Observation 7e4d366d-480f-497a-a5ed-b996f6d4fb22 · inbound

Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models cites this paper.

Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:53:26.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:14:11.941977Z digest=sha256:678cefda82505a60313dfda034debb35dc01cc3f3a9633fea187f0481072d2e2

Observation 108cb2b1-22d1-4f94-adda-44c29233d8dc · inbound

Targeted Interpretable Safety Neuron Enhancement for Multilingual Vision-Language Large Models cites this paper.

Targeted Interpretable Safety Neuron Enhancement for Multilingual Vision-Language Large Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:53:26.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:12:34.909229Z digest=sha256:349a30c6974eb69a8997755d702c9791d27ec43efec5280092cbddacf2ab457b

Observation 703b41d5-9773-42e5-91b4-932b3c967f8a · inbound

Targeted Interpretable Safety Neuron Enhancement for Multilingual Vision-Language Large Models cites this paper.

Targeted Interpretable Safety Neuron Enhancement for Multilingual Vision-Language Large Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T16:34:02.688999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:34:02.688999Z digest=sha256:ae90ea0cd5d72dd66d01721c4b96a742d1a68a029928ff9cc3c94312a8b7001e

Observation 5feff337-67bd-4dd0-b6f4-d500cb2c9ad0 · inbound

PinpointQA: A Dataset and Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos cites this paper.

PinpointQA: A Dataset and Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:53:26.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:40:39.180502Z digest=sha256:40e52157599f6824bef1d41a84fa4053a8940a259318b5c8ae1fad2e017462c7

Observation 8c4c3159-2089-4c1f-8f42-2019eabc6b59 · inbound

PinpointQA: A Dataset and Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos cites this paper.

PinpointQA: A Dataset and Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:39:48.790068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T06:36:57.544235Z digest=sha256:b61765621e192057855f9efafc8f093f76755c7f8cd34adadaf66922664384d4

Observation a4065dd2-6766-4ff8-8f39-da886683f0ab · inbound

Boosting Visual Instruction Tuning with Self-Supervised Guidance cites this paper.

Boosting Visual Instruction Tuning with Self-Supervised Guidance LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:53:26.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:27:52.208827Z digest=sha256:d6ee3e9909a5eeb926df408af3490b3c7ec1ed20dd936541cd82f65c056bbb79

Observation d3f719bc-52d5-461c-a042-fceb0656903c · inbound

PersonaVLM: Long-Term Personalized Multimodal LLMs cites this paper.

PersonaVLM: Long-Term Personalized Multimodal LLMs LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-15T08:05:15.453050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T08:02:10.327523Z digest=sha256:af16726f05704ad76cf959e42e709674d652cee4ce1bdc3f47ac676ee2fa50e4

Observation 63e716c8-fc00-47bc-a5af-85a4da609d5f · inbound

SLQ: Bridging Modalities via Shared Latent Queries for Retrieval with Frozen MLLMs cites this paper.

SLQ: Bridging Modalities via Shared Latent Queries for Retrieval with Frozen MLLMs LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:53:26.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T14:03:23.807710Z digest=sha256:294ff0317e546c21f3a03b7e92436196b96e151e335e9c83a908037b4f0ee057

Observation ea5dc257-7758-4616-aec3-430dfd6e0574 · inbound

SLQ: Bridging Modalities via Shared Latent Queries for Retrieval with Frozen MLLMs cites this paper.

SLQ: Bridging Modalities via Shared Latent Queries for Retrieval with Frozen MLLMs LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:53:26.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:25.022963Z digest=sha256:114fe60d261c7bd68e7c0e3bf6f8d40c7e6f1f7bda358779a3ab56b57b86b7ed

Observation ba838f1f-db2c-41f3-9f83-335e075b4a11 · inbound

Seek-and-Solve: Benchmarking MLLMs for Visual Clue-Driven Reasoning in Daily Scenarios cites this paper.

Seek-and-Solve: Benchmarking MLLMs for Visual Clue-Driven Reasoning in Daily Scenarios LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:53:26.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:04:34.328439Z digest=sha256:1a04e8ca1e76bdd1b545457f40387d958ddb1ed2dbc4074dfa66709432adead0

Observation 99182cbd-e134-48a7-92f6-eb31c28a6a36 · inbound

VisPCO: Visual Token Pruning Configuration Optimization via Budget-Aware Pareto-Frontier Learning for Vision-Language Models cites this paper.

VisPCO: Visual Token Pruning Configuration Optimization via Budget-Aware Pareto-Frontier Learning for Vision-Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:53:26.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T11:28:44.565497Z digest=sha256:b43c5703ef5f49e33e09deca5b585c87d4f10dcaab4aa01c7b65b9b716a5565e

Observation d1212e51-d4f0-4391-be2d-9181e458135d · inbound

BARD: Bridging AutoRegressive and Diffusion Vision-Language Models Via Highly Efficient Progressive Block Merging and Stage-Wise Distillation cites this paper.

BARD: Bridging AutoRegressive and Diffusion Vision-Language Models Via Highly Efficient Progressive Block Merging and Stage-Wise Distillation LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:53:26.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T14:28:47.505755Z digest=sha256:63e2d60c1ad76911c7f3ab6074f9142eda27b474a47381971546f24e88616d4e

Observation 1d8947d4-f66f-4005-ada7-56ffeb0eaded · inbound

Towards Temporal Compositional Reasoning in Long-Form Sports Videos cites this paper.

Towards Temporal Compositional Reasoning in Long-Form Sports Videos LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T19:27:29.843866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:27:29.843866Z digest=sha256:e8bdf3468062efc53ba2c3bda24804b6d14d871286d8ef8318f34ec026390298

Observation 14176f09-46b7-4be3-8289-cc7cf2de2db3 · inbound

SpikingBrain2.0: Brain-Inspired Foundation Models for Efficient Long-Context and Cross-Platform Inference cites this paper.

SpikingBrain2.0: Brain-Inspired Foundation Models for Efficient Long-Context and Cross-Platform Inference LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:53:26.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T12:18:23.898779Z digest=sha256:e5f78470d956fdbe5d780e6b9fd3c1726244e630ed02f10bb70b014b8d25aab3

Observation 8da5f080-67dc-441f-965e-cb07b768a296 · inbound

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation cites this paper.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:53:26.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:31:26.325118Z digest=sha256:d25da606e81d3e32bdd811d7c89903f8e5720fdaaf8a4e8a0cf15cf0c7040dd4

Observation aa21ed8b-d792-4f37-8981-a01aba58df48 · inbound

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation cites this paper.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.182896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:9c243f150cc7a4e3fbbeae50f4328fe77852dda113ba028fcce69b06f3c20e7a

Observation c4a7cbcb-c5c9-4b51-b8f7-42da4fcf56c7 · inbound

Scaling Video Understanding via Compact Latent Multi-Agent Collaboration cites this paper.

Scaling Video Understanding via Compact Latent Multi-Agent Collaboration LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:53:26.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T20:11:11.410051Z digest=sha256:4dece4c22fb9f54a5036515a57eb8cb63bd514a1bb4ef7a6f0133a5effd6ae91

Observation d57b3385-acb7-4ee0-a523-0827cb857b0a · inbound

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs cites this paper.

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:53:26.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:53:06.494640Z digest=sha256:483844bf9d37b2dd8f9c7bd1c4f9532ffc42a4fbef753ded6c15d8ac436d1a74

Observation bcc3743c-f4d9-4dd8-b78e-ea3c3e13b86f · inbound

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs cites this paper.

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:53:26.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:49:15.136031Z digest=sha256:d333c78b3ba998e23c860f42a892b97c0d382f590516711f8ccedd1799163f52

Observation 3f00c458-fd10-4d3a-8806-2a235c9b2cb5 · inbound

GameScope: A Multi-Attribute, Multi-Codec Benchmark Dataset for Gaming Video Quality Assessment cites this paper.

GameScope: A Multi-Attribute, Multi-Codec Benchmark Dataset for Gaming Video Quality Assessment LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:53:26.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T15:02:17.115842Z digest=sha256:4dd59b12f8a1801e57959aff6808628a676fa103f6077b05855f9969c87a466f

Observation 6d95d961-fc09-4be7-b773-9da99db45b7f · inbound

SurgCheck: Do Vision-Language Models Really Look at Images in Surgical VQA? cites this paper.

SurgCheck: Do Vision-Language Models Really Look at Images in Surgical VQA? LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:53:26.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T17:27:57.384336Z digest=sha256:62648debb7de312f611d95798aa4931fccfe1e3b7fd3235f7a9629527a42314f

Observation 9aa22368-ea91-4392-ab56-2a7fdc87daa5 · inbound

Causal Probing for Internal Visual Representations in Multimodal Large Language Models cites this paper.

Causal Probing for Internal Visual Representations in Multimodal Large Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:53:26.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T11:58:08.853836Z digest=sha256:e1610cf2a06666fb0eaa8e9ad29b5768eefab8859a322d83ecc45bd6e6e43c38

Observation 8597090d-1cd3-4f16-bf27-eea09536c982 · inbound

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding cites this paper.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:53:26.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T02:30:55.939351Z digest=sha256:533acc1f447fa53f0df5f8753b802dad13c1806a7d8651e1d0a1fdd71ee4d698

Observation 54b888d4-6da8-46bb-85a6-2b2ad46d75a3 · inbound

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding cites this paper.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:53:26.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:0ddff5c263ed90dc503a8cd9cac0a7553d45713e8147b6f8d9c9b19c1bfe6b61

Observation d1c355a4-e3f6-43a2-a814-e2cc36b37582 · inbound

GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning cites this paper.

GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:53:26.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T02:03:52.566413Z digest=sha256:430d9f359bc884800bd01646ef670a05504a389f20d29ef08b5bc722ab4cd648

Observation 37d812f7-ce55-421c-85e6-4b13a27bb05e · inbound

ZAYA1-VL-8B Technical Report cites this paper.

ZAYA1-VL-8B Technical Report LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:53:26.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:15:16.607346Z digest=sha256:e39fc12fba8399b7c4baabe99cbf2362c78a73a81b53e48eea066e723a1b8f80

Observation eb7ac702-33ad-4a53-b580-b7dce4cd9dbb · inbound

Evading Visual Aphasia: Contrastive Adaptive Semantic Token Pruning for Vision-Language Models cites this paper.

Evading Visual Aphasia: Contrastive Adaptive Semantic Token Pruning for Vision-Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:53:26.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T02:29:22.152053Z digest=sha256:c9905dfdaf02725f1d3acc64401d4ccc3d1b9c5a49d056dfb09eae92167844a1

Observation 73b8cc84-5552-42c3-b5cb-f0bae577c1d0 · inbound

Sens-VisualNews: A Benchmark Dataset for Sensational Image Detection cites this paper.

Sens-VisualNews: A Benchmark Dataset for Sensational Image Detection LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:53:26.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:14:13.461412Z digest=sha256:a68734aa24cf670f321ad088d1ba611d0476481263af003787a3e2831e355443

Observation e948a9e3-9bbc-4c6f-a842-538d8c5569fe · inbound

LDDR: Linear-DPP-Based Dynamic-Resolution Frame Sampling for Video MLLMs cites this paper.

LDDR: Linear-DPP-Based Dynamic-Resolution Frame Sampling for Video MLLMs LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-13T03:22:10.102663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T03:20:42.684286Z digest=sha256:ad8c64cd3c672dbe8c8f363d8fe2e26e482f095744a87d566611a138970462f0

Observation 5383bb70-7c6e-4e24-b80a-001e835af3c2 · inbound

Logit-Attention Divergence: Mitigating Position Bias in Multi-Image Retrieval via Attention-Guided Calibration cites this paper.

Logit-Attention Divergence: Mitigating Position Bias in Multi-Image Retrieval via Attention-Guided Calibration LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:27:01.711374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T01:26:45.182597Z digest=sha256:acab25ad29e75990c67b7b64587ba40d7cced24236c7925a967339a430f74116

Observation 45c912ff-38df-4f70-831e-63ef8f6060cd · inbound

Instruction Lens Score: Your Instruction Contributes a Powerful Object Hallucination Detector for Multimodal Large Language Models cites this paper.

Instruction Lens Score: Your Instruction Contributes a Powerful Object Hallucination Detector for Multimodal Large Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T05:52:22.127221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T05:50:47.441551Z digest=sha256:9a8e878ace574fad17dc7d969ccc4921f1b58bb849ff3b38fff87872e135dc69

Observation 5d5bfda9-6dea-41ac-bdf6-41c6312e63a4 · inbound

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models cites this paper.

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 130

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T20:22:56.380945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-14T20:13:18.813131Z digest=sha256:0f21cdb122a290595f2948e1489e919b421f0b9012f12674603847fcd9ccdc06

Observation a476c70a-76b1-4475-aff1-ec1cdfa79336 · inbound

ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both cites this paper.

ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-15T03:14:53.478527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T03:09:58.411261Z digest=sha256:38318b10440038ee56e435c6febe8313ba341f8ffdc97310346f854f8187af11

Observation 0259e5b1-bb7b-419c-9120-6bdd82715438 · inbound

GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions cites this paper.

GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-20T18:53:38.987244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T18:49:18.815456Z digest=sha256:76be7a10ed04997aa3d2c7781f67d96d4825fdb33ac6e8ab5b74b5c40789735a

Observation 5933b99b-2729-44a8-a312-490afc1eac6a · inbound

EgoInteract: Synthetic Egocentric Videos Generation for Interaction Understanding and Anticipation cites this paper.

EgoInteract: Synthetic Egocentric Videos Generation for Interaction Understanding and Anticipation LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:03:13.579091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T11:01:17.353843Z digest=sha256:2ecc42e3a6de3437e678157b7ebcbf863213fa076a6e52ac87385ef3e566af4c

Observation 6872e68a-820a-4737-9509-552ffa257b09 · inbound

EgoInteract: Synthetic Egocentric Videos Generation for Interaction Understanding and Anticipation cites this paper.

EgoInteract: Synthetic Egocentric Videos Generation for Interaction Understanding and Anticipation LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:25:23.734356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:22:17.582854Z digest=sha256:ae643ea8fbcf96eec159e801a4792a015ee233c7b1f7fff978647930d06f43c8

Observation 28594ffb-d7c4-4999-b4ed-626fb5da731e · inbound

INAR-VL: Input-Aware Routing for Edge-Cloud Vision-Language Inference cites this paper.

INAR-VL: Input-Aware Routing for Edge-Cloud Vision-Language Inference LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:38:59.940512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T20:37:06.957970Z digest=sha256:80053473ffd92bf8076d6c4e2d02d4d3eb0c9b28fadc9205b8f553fb099b789c

Observation dd4052d0-c5d6-47fe-86fa-3139ad344cc1 · inbound

CP-MoE: Consistency-Preserving Mixture-of-Experts for Continual Learning cites this paper.

CP-MoE: Consistency-Preserving Mixture-of-Experts for Continual Learning LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:39:53.506540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T08:37:10.313340Z digest=sha256:7867038b15840df38350122f701aeac93a1709348b0b9318ef8134cb4e10dc7e

Observation 1081bbc3-cc39-4d01-bf51-4457b6d742a4 · inbound

RISE: Reliable Improvement in Self-Evolving Vision-Language Models cites this paper.

RISE: Reliable Improvement in Self-Evolving Vision-Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:39:40.607807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T05:38:26.590720Z digest=sha256:5cffdece769eb6b200813067dfede34aabcb259b6fd1dd0690344a0170a9b0f6

Observation d72db277-76e2-45dd-9276-881537fc991a · inbound

RISE: Reliable Improvement in Self-Evolving Vision-Language Models cites this paper.

RISE: Reliable Improvement in Self-Evolving Vision-Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:24:57.596388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:19:09.596950Z digest=sha256:1477893c49269bfcbb8f431db635c9d397bea4417f4971b95da9cd1628af6d76

Observation cfe8a004-722e-446c-8021-a22c89f909e7 · inbound

Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning cites this paper.

Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:33:57.700008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:33:51.535737Z digest=sha256:1a3295e95d89381b427063378ed84c96001259837b63b9d47986dbd765e7cd58

Observation 3cb2db30-8ff8-449d-a5ed-8a34bfb4b918 · inbound

Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning cites this paper.

Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:45:24.148505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:43:38.328972Z digest=sha256:47f5c11417dcc9fa86cfafc62297353b505f4bcfa776bead6cd3a2d315ac3f2a

Observation 922854f8-c61e-4047-b37f-f0bbb1a97a98 · inbound

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding cites this paper.

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:55:25.030147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:51:49.390597Z digest=sha256:2b369491a5ab540ede7713760d8930b3f6a07dd42ec76ee75b9b6b8fcf72360d

Observation 8b00eee8-c5f7-44fb-8807-5c5d731ff68e · inbound

MetaphorVU: Towards Metaphorical Video Understanding cites this paper.

MetaphorVU: Towards Metaphorical Video Understanding LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:44:01.337297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:43:20.101830Z digest=sha256:32c65542083c8c26841e8c0f78f0d67181c2702fa51c3435faf2a72ff0531b71

Observation ea9195d4-8be6-4066-ac1e-c19e6997d5b4 · inbound

Towards Open-World Referring Expression Comprehension: A Benchmark with Training-free Multi-task Consistency Checker cites this paper.

Towards Open-World Referring Expression Comprehension: A Benchmark with Training-free Multi-task Consistency Checker LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T23:24:02.152246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T23:15:02.455554Z digest=sha256:02c9bed71bbf75fd702764c176c96c2fd519e24179641ed5033929abc4ce9776

Observation 581aa147-7652-47d4-b5e1-341d2bfcc9db · inbound

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence cites this paper.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.610224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:5153491846e6e1e9544a5e0a81e9ef8416b4533bb210b1a326d0a1a3e37d48e4

Observation 22c7d3f0-56c9-4a51-b678-7d0ec76415b3 · inbound

HiKEY: Hierarchical Multimodal Retrieval for Open-Domain Document Question Answering cites this paper.

HiKEY: Hierarchical Multimodal Retrieval for Open-Domain Document Question Answering LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T07:43:13.820514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T07:39:14.379668Z digest=sha256:b8465f1270e706671321ede51d80e69f41cb21a5d50a755a8d744bf9e3e200a4

Observation da34d152-e821-4a29-8496-35baa38ce090 · inbound

PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding cites this paper.

PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:23:15.910626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T08:13:42.526597Z digest=sha256:2b1bb2beb1de0f02fa1938717a8e31795e376878b88ebdbb5428f2dc14d97ca7

Observation e6283773-defb-4958-89f7-07a6107e90ee · inbound

LoMo: Local Modality Substitution for Deeper Vision-Language Fusion cites this paper.

LoMo: Local Modality Substitution for Deeper Vision-Language Fusion LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:23:15.124540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T08:20:05.081980Z digest=sha256:8893358bcb8d17a61ace3032e54e41f997254ddc8a6994c8e19dffd43f39fa28

Observation df923450-20fe-43a9-a62d-1c3b979d3bdb · inbound

Seeing Before Agreeing: Aligning Multi-Agent Consensus with Visual Evidence cites this paper.

Seeing Before Agreeing: Aligning Multi-Agent Consensus with Visual Evidence LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T23:42:50.112005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T23:25:06.490634Z digest=sha256:bf4f7d08cb2718651bfb87cc5b91046de6c2c4152ad9d069dd923791c21ab4b4

Observation 9616b357-95b2-49d2-a8fe-551a5dff4033 · inbound

Zamba2-VL Technical Report cites this paper.

Zamba2-VL Technical Report LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:26:00.679535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T22:34:20.970856Z digest=sha256:ec0432590794b7431e0d8db20cae8e040495a558822066ae8bbf2f5914ec844c

Observation d9117f3e-77ae-447c-b39f-e4aa0fa140cc · inbound

InsightVQA: High-Dimensional Emotion-Cognitive Visual Question Answering Benchmark cites this paper.

InsightVQA: High-Dimensional Emotion-Cognitive Visual Question Answering Benchmark LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-28T15:02:18.800254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T14:57:12.827012Z digest=sha256:17bf239a9e7eedf2ae49b6d878c6b19b53db6e0a5b661e55afd863e8df85257c

Observation 9e8a6090-5e48-4629-a5d8-ae8f01b80629 · inbound

Visual Instruction Tuning Aligns Modalities through Abstraction cites this paper.

Visual Instruction Tuning Aligns Modalities through Abstraction LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:56:29.785632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:25:20.950792Z digest=sha256:48450f288dfcbfa4c41be4b8f016e790d92e856743a9b5ca704f68d8bd688210

Observation b35417bc-f33f-49ef-a6d3-0349718ddb53 · inbound

Benchmark Everything Everywhere All at Once cites this paper.

Benchmark Everything Everywhere All at Once LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:46:59.041205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:5f059067fa31ebb829984e998f3bd179b4b226a289aeafa21b5dcf6ab9c584c8

Observation 67bf9386-ea11-4b0f-99dd-e994e8c4f025 · inbound

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention cites this paper.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.877757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:78fb3aae5f0550fcfe02119ed2f2de845e6f75d0ddfaca7325c629924eb431a7

Observation a41ade6a-6a63-4175-b91c-519e332121d4 · inbound

TVI-CoT: Text-Visual Interleaved Chain-of-Thought Reasoning for Multimodal Understanding cites this paper.

TVI-CoT: Text-Visual Interleaved Chain-of-Thought Reasoning for Multimodal Understanding LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:27:25.908380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T18:54:34.353940Z digest=sha256:4aa4de8ff7d2612bdbb9a3e980f6753240811ce9b1c3575890a77bff12d122e4

Observation b6231428-e782-496f-be68-21e46aa0112f · inbound

Look Less, Reason More: Block-wise Attention Skipping for Efficient Multimodal LLMs cites this paper.

Look Less, Reason More: Block-wise Attention Skipping for Efficient Multimodal LLMs LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:17:26.122906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T19:02:59.403688Z digest=sha256:06b610f25a589d928e6a6048e5425e610f004ea5fefe6d0fe06aeba85ef2e014

Observation 9c7231c6-d17f-4431-801e-68e93c8b7e2e · inbound

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs cites this paper.

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:57:32.394178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T16:12:53.387567Z digest=sha256:111531a32c24d9b844b250072bb1c051922d569a224fd6b6707b9560e756cf72

Observation e1095336-8700-4016-ad38-021264c5b1fd · inbound

Q-Fold: Query-Aware Focus-Context Spatio-Temporal Folding for Long Video Understanding cites this paper.

Q-Fold: Query-Aware Focus-Context Spatio-Temporal Folding for Long Video Understanding LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:27:56.048823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T10:04:29.739632Z digest=sha256:8acba932bbfab52df4b0c6e6530d56e8ba3e6096ab7fff8e28514872a3f27585

Observation 9fb5d31f-d23d-43a2-885a-1a023423e044 · inbound

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers cites this paper.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 215

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:28:29.702432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:f3e993f189e31119b77f130a97ef59d271bdf255fab902c6e27e22b77c5e9915

Observation b8dc02e2-c7d2-4d29-b1fa-b908df80c43a · inbound

EventDrive: Event Cameras for Vision-Language Driving Intelligence cites this paper.

EventDrive: Event Cameras for Vision-Language Driving Intelligence LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:57.091374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:26:43.752335Z digest=sha256:ff62e05dd10112b6b55a8dd755ff2d6dafde65c364e9ab17912d21f843f91fae

Observation 4b6bb2d1-56cb-45d7-8363-9822f964601b · inbound

PorTEXTO: A European Portuguese Benchmark for Visual Text Extraction cites this paper.

PorTEXTO: A European Portuguese Benchmark for Visual Text Extraction LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T00:09:14.505087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T21:28:31.600252Z digest=sha256:19c828fd4957da6cf42f781ba9f96326414864befdefea90b32a8890f266f96a

Observation ba74f4e7-5a76-40a1-8d61-6fa971326ead · inbound

PorTEXTO: A European Portuguese Benchmark for Visual Text Extraction cites this paper.

PorTEXTO: A European Portuguese Benchmark for Visual Text Extraction LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T22:27:25.279901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-02T22:18:03.991161Z digest=sha256:b8893478f7e84f0d5abb56742162a33156f5afa01e38f215c1729c7627a9f2c7

Observation 2ae1fd90-c27c-4c8d-9da8-c5abf5a1c386 · inbound

AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model cites this paper.

AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:19:12.958668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T21:22:53.792778Z digest=sha256:295041819718d2dc739d4f67548ab56769208dd9436d95e9a8449fa85da45a34

Observation 60e186ea-3219-4e3d-9880-16036e32a177 · inbound

AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model cites this paper.

AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-01T07:35:28.806321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T07:34:37.974692Z digest=sha256:6071cd2d159abb9ea78ebe8fc6a1dee91b8a94a8e55507b53938e826ea36ded3

Observation b6c3952d-1f89-4698-ae5e-b28ecc9cb72f · inbound

AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model cites this paper.

AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:17:24.978246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-02T22:15:14.561596Z digest=sha256:78ea425b84044578945c459d7c0971ecc345ffa4ac74fd4853b74b5d79cd530d

Observation 69465d34-1e91-4d0f-8be4-fcbe807e7c05 · inbound

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models cites this paper.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:17.966899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:83f14fc908f039f449fd5fe0bcf6feb40d516a9488c4584c69fe7a77a4ddefcb

Observation 0dcb5745-6d38-4870-8e3d-d553d0c16a24 · inbound

Occ-VLM: Occupancy Grounded Vision Language Model for Indoor Scene Understanding cites this paper.

Occ-VLM: Occupancy Grounded Vision Language Model for Indoor Scene Understanding LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T02:59:25.847804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T18:40:20.588652Z digest=sha256:fad2f94281ba09f9f0b66f16976510e255ae8a34cd6ce904639eb9b8eb9876e0

Observation 3bb3aa63-3f91-48ce-a715-d44acc1b30da · inbound

The Hidden Evolution of Disguised Visual Context inside the VLM cites this paper.

The Hidden Evolution of Disguised Visual Context inside the VLM LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T03:19:31.060994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T18:08:56.044278Z digest=sha256:cae6efe39666e7642f5a5a555d8629ffa90817c17bc690e4b737e244f95c5d7d

Observation 32b7ad4e-1322-4a12-922c-a467a5910ea3 · inbound

MotionHalluc: Diagnosing Kinematic Hallucinations in Fine-Grained Motion Reasoning cites this paper.

MotionHalluc: Diagnosing Kinematic Hallucinations in Fine-Grained Motion Reasoning LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:19:47.497930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T08:58:37.673268Z digest=sha256:aa6f88020945ce62c752de97d8cc9c55da5586f06755173fc4ec62c256da6a7f

Observation 2380fd62-8a0a-4ffd-96a2-a0a8c2e04b39 · inbound

PointVG-R: Internalizing Geometric Reasoning in MLLMs for Precise Pointing Localization via Visual Chain of Thought cites this paper.

PointVG-R: Internalizing Geometric Reasoning in MLLMs for Precise Pointing Localization via Visual Chain of Thought LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:19:57.673272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T00:46:17.094339Z digest=sha256:f1590e06b3a4063bb0c0a23b8653ea25c2e5ae88fe326fc3ae93cdfaaa0d4815

Observation 7a49d3f0-710f-4c52-bec9-feceaa80b68e · inbound

SSMNBench: Diagnosing Image-based Cross-View Human-Object Understanding via Single-View Sufficiency and Multi-View Necessity cites this paper.

SSMNBench: Diagnosing Image-based Cross-View Human-Object Understanding via Single-View Sufficiency and Multi-View Necessity LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T19:50:10.245866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-25T21:02:44.441202Z digest=sha256:347e2cd939db3f375be83d679c67db301beab5a933604280ee3165f58656cad3

Observation d5d0d744-97d9-4ced-831d-bf813336622c · inbound

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP cites this paper.

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T13:09:50.241764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T05:31:17.461916Z digest=sha256:256e3380b46cc4431ad61cbe1035dd32d9fe4248fc93857c43caf067b70a48e0

Observation f250ae5c-09ae-45ab-89a6-26bf685e844d · inbound

DiCoBench: Benchmarking Multi-Image Fine-Grained Perception via Differential and Commonality Visual Cues cites this paper.

DiCoBench: Benchmarking Multi-Image Fine-Grained Perception via Differential and Commonality Visual Cues LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-04T13:19:50.940528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T05:18:01.931929Z digest=sha256:0a55c100b3312bd2f49ff67fc0747831a365347e274527ff0290803353ddff69

Observation 561bd4ba-7118-41f8-9697-05d8be3b0324 · inbound

Position Rebinding Cache Reuse: Replay-Free Visual Revisiting for Interleaved Multimodal Reasoning cites this paper.

Position Rebinding Cache Reuse: Replay-Free Visual Revisiting for Interleaved Multimodal Reasoning LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-04T13:29:52.025953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T05:08:09.884231Z digest=sha256:8d2c6bf62356df3237ee08c0cb90da983a67f3fe6a87897bc9dfcfbd4c2b0df0

Observation 67e520c4-024a-4f67-a238-b39f9699f33a · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-01T15:35:48.920115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T01:16:16.834861Z digest=sha256:8bc9eccd3788b245354b4ca11b55a0a092db229810a8909811711b9cbf3a9ec5