Pith. sign in

Paper Citation Record · LEDGER

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

As of 11 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 9 inbound Pith citation observations for arXiv:2605.25979.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.25979 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T22:12:05.365596Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:22:31.939898Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-02T02:46:28.010713Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact42
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f1b74493-07db-4f49-a4fe-3cec07367994 · outbound

This paper cites SPARROW: Learning spatial precision and temporal referential consistency in pixel-grounded video MLLMs.arXiv:2603.12382,.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence SPARROW: Learning spatial precision and temporal referential consistency in pixel-grounded video MLLMs.arXiv:2603.12382,

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.593890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:03708df0ee82dc751af063724fd192b889f6d6d786779d93dcd74146482ee05a

Observation 581aa147-7652-47d4-b5e1-341d2bfcc9db · outbound

This paper cites LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.610224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:5153491846e6e1e9544a5e0a81e9ef8416b4533bb210b1a326d0a1a3e37d48e4

Observation c72ecbfe-6a27-46ce-95c5-c64057214303 · outbound

This paper cites Qwen3-VL Technical Report.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Qwen3-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.602205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:b55032c4fdfc6cf764301e82a7771d0457983e862a1b560e1a8ab6cab6e71351

Observation 1cfba95e-4081-4c37-b912-7bb7ee65101e · outbound

This paper cites Reasoning over Video: Evaluating How MLLMs Extract, Integrate, and Reconstruct Spatiotemporal Evidence.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Reasoning over Video: Evaluating How MLLMs Extract, Integrate, and Reconstruct Spatiotemporal Evidence

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.534505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:e5dbce7449fbb7704fc1f02474f985c37790edb129572fc63ffbaf62aea9f69f

Observation 7ab2eb34-103c-4f07-8f40-33ff95e5f780 · outbound

This paper cites Token Merging: Your ViT But Faster.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Token Merging: Your ViT But Faster

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.539418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:451b35a59e57876919924cceb519ecaf725de623ddf6a461f3a1ddfd9963e25d

Observation 016b2ce0-c946-4469-add7-ec1b9b10e22c · outbound

This paper cites SpatialBot: Precise Spatial Understanding with Vision Language Models.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.560592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:d64321e0d89a18a82506f2d884c65aba9a8591472c39d5a8334a36d658a1287e

Observation 22ddb1ee-6978-4b5b-b3a2-2accac6591c3 · outbound

This paper cites Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.585208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:5f06dfb290589c6dc5efe3cb901ba955b419d236d80b459fa590644bd4a3c308

Observation ca9de3d6-bab6-4863-899b-95acdc4c3104 · outbound

This paper cites GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.634880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:147cdceb963b36cc144c2e2a0d2f674c6ab0a671468a97574b0b4f2a544e480b

Observation f0b9d4ca-36d0-4cd7-9035-4a3688a5c578 · outbound

This paper cites Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.651251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:def9e41b18a0307cf5a4b0f3b329fd2cb31f8bec71d3267183c77fac5d0e188d

Observation 09a2392a-73ee-43b9-86d8-e8694f9d8ab5 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T22:13:59.623529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:2a730d8690415e64d8da039192bbf5932519a7a2f1fae7f0d2927fcce049545a

Observation 11f6c408-8efc-44f9-9362-086b0115b8c7 · outbound

This paper cites Small Vision-Language Models are Smart Compressors for Long Video Understanding.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Small Vision-Language Models are Smart Compressors for Long Video Understanding

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.629109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:26c3aa594665143c8e438a4d8ce44686e8e0c2e1c9ebcc929a439bfd4937a7cc

Observation ccd5db0d-2e2c-4f72-8418-558f3dae9183 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.632058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:257c68df8860b0ca89327f5c9965c0944614115f7808e35f4dbbbd9c4f7b3215

Observation 170f529f-5c12-4941-b8a8-b291450df40b · outbound

This paper cites Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.637794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:162bfeb21c54b3178c3234c6832c4c7c016023feb20015cdfcfd837043332b32

Observation 393a6d2c-8ec2-4bae-a1c3-04800cbabd82 · outbound

This paper cites Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-20T02:18:27.127261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:f22c297bcfec238b5899b8b40adffc1156403b08b02d27eeb54417225673e38b

Observation fe821ea5-a2a0-4274-9b99-441da40cc261 · outbound

This paper cites TRACE: Temporal Grounding Video LLM via Causal Event Modeling.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence TRACE: Temporal Grounding Video LLM via Causal Event Modeling

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.619875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:be631897c088ff072adfa0ddb53bdfc8014e3f80ab710716eb7933e5e81f333e

Observation 4803f7e1-22e8-4c1a-bb14-5600f3539762 · outbound

This paper cites Spa3R: Predictive spatial field modeling for 3D visual reasoning.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Spa3R: Predictive spatial field modeling for 3D visual reasoning

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:13:59.613532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:f2f0ce90e546f19e9129a30d14569c09bb622e72afe370f69c0008a044767676

Observation e20e8960-f15a-4301-9536-38f5411cd05d · outbound

This paper cites Token-efficient long video understanding for multimodal llms.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Token-efficient long video understanding for multimodal llms

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.617077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:69160cbd9430b1482fa8b27620f4936c704d2b625158c32f1842757ec112943a

Observation 6165672e-831e-4e3c-80a7-753e4d78d22c · outbound

This paper cites Agentrvos: Reasoning over object tracks for zero-shot referring video object segmentation.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Agentrvos: Reasoning over object tracks for zero-shot referring video object segmentation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.627023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:a06bc19244f050064071202b0c2e1220497ace1572b1c71b8b83404c92b8a85c

Observation f31884c4-788a-4c9a-a17b-e330fee1ded4 · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.644704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:c3854e2f08aab50143361f57ba41eab7c6aa32fffdf1d5a65f7f5fc521d0f4d8

Observation ae5d3360-d2fa-4279-b130-01683db20b90 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence VideoChat: Chat-Centric Video Understanding

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.604775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:378da1a367019eeb09d6a64fda4c8a5a06eea5d0ce19f9ab153b9d990725b1e9

Observation fb4fd70d-585c-48b0-9d6e-22d13ada3ef0 · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.641616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:59e7d2db3385192bd421394cb44b9503e3f446d3bd2cbd97624961a1ab874f88

Observation 764b775d-4ec5-44ac-8fa6-1402452ebb2e · outbound

This paper cites Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.625729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:fa2614fab745488abf6a86d53b2304277a367733c4b9eba553c433d99ec11c17

Observation 9fc1f3ba-b6a6-45d6-87c6-c59c5d4c378d · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.596700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:c0ff173e3b83bed56170778334882050034ff330ba587fb056aba3441c7ab981

Observation fe30cdd5-5070-40a4-b718-f28b247069ae · outbound

This paper cites SpaceR: Reinforcing MLLMs in Video Spatial Reasoning.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence SpaceR: Reinforcing MLLMs in Video Spatial Reasoning

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.550815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:155ccb178dcafbedd6627a68ea5af98c28d483eef33fd34c34e68a1ff670544f

Observation d36b3184-77a9-48be-9423-6b0cd57d2bdd · outbound

This paper cites VideoMolmo: Spatio-Temporal Grounding Meets Pointing.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence VideoMolmo: Spatio-Temporal Grounding Meets Pointing

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.563621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:a95002f9cfaa51590cb0544ea4fc00216a3a51b70d94899d16a02c394e36c576

Observation 62853dd5-b767-49ba-a9a9-4a13f37bcfab · outbound

This paper cites TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.539308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:faf3dd151d23f09292daabbd7e3e6384f05515353096ab31b810eae653a8dad7

Observation 3910d08d-800d-4c97-afac-42aeba8eb419 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.621151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:90134b22c61f87523b8bdca2f8febcaacb6ac72a3d470f8e0951c09dbc388596

Observation 51c68fed-389e-4425-9ecc-5cf38efdc3c0 · outbound

This paper cites A Simple Baseline for Streaming Video Understanding.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence A Simple Baseline for Streaming Video Understanding

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.549967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:974bed6cabb8a886b7e8977bf94d2b704862bb0cfa303debfe9563bcad342589

Observation b87dd21e-68ef-4313-9a82-2ef1a7024db7 · outbound

This paper cites Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.566701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:928b795c1e7cefde4952d9160a684ac75a2172ca9c0aafb106b0d0dfd05f8cb2

Observation aedd76ea-039d-426a-83ce-833ec7a451ed · outbound

This paper cites EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.599233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:b0f594fec18f39e63dd3a09088b5eb752b6cc5c99c0c01f4789aacab508e0345

Observation e28f33d9-354b-49a1-b24f-34aeae931da4 · outbound

This paper cites Onevision-encoder: Codec-aligned sparsity as a foundational principle for multimodal intelligence.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Onevision-encoder: Codec-aligned sparsity as a foundational principle for multimodal intelligence

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.524140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:673867eab7ca006d1e5820019cf2ad2a497221fc432e04f40d65b689d24790f1

Observation 7ecee9cc-9e28-48e1-bfca-af410ad6f48d · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.634667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:8482cf44101d8dd1bc38c3e789e15994e1c90e83d64aa660bb02d1c8b456897d

Observation e8e66c53-1aa6-462f-8eae-33e353ef4e39 · outbound

This paper cites Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.520753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:699ad428e97ffef01d9e207972b748f638d06629ecb28d166f82b39964ec8a2d

Observation 1289afdf-8daa-48c5-81c9-110436f9cd37 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T22:13:59.631862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T17:38:12.771444+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:8a7a6cbffde43406fddd78687caf50260bf74b769822b0636aff9039367af5df

Observation 9502d1c9-2655-4300-9f10-c4a927bcd63f · outbound

This paper cites Slow-Fast Architecture for Video Multi-Modal Large Language Models.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Slow-Fast Architecture for Video Multi-Modal Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.576656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:2a6ea14cba9fd00c3877275cb7e0defe828d9026c80ca56f524a4aec6abd8baf

Observation d5465249-d1cd-46c3-bc73-74d1dd4d8379 · outbound

This paper cites S-GRPO: Unified Post-Training for Large Vision-Language Models.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence S-GRPO: Unified Post-Training for Large Vision-Language Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.573692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:969b10a56c51c285a32b600ad00d424fd225749e628aca1d208e4c347271cb60

Observation 5b48ea6b-4a0b-444a-a5ff-4f6ef7bdb83c · outbound

This paper cites Kwai Keye-VL 1.5 Technical Report.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Kwai Keye-VL 1.5 Technical Report

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.514435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:4d22da4646162a535a9a62831407773795252d864c1b36100acf381505a2fab2

Observation d07ce9e9-344a-4537-b425-749d832bfd2c · outbound

This paper cites MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.579379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:08acfdb0fba4db889bfd5a2ef6dbffaf4afb85dc228dab667b610dadd885aa17

Observation d4292cc4-d4b9-41b1-b0f5-c9f077a36c83 · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.587708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:68ddbff0865bee6c8124dbee1f12394d9f0239a9f6f1d7b23bde928ff74e0fa4

Observation e825296b-f6c7-4712-8e04-fc4d80848624 · outbound

This paper cites Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.509917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:ccde455da6054965a5227c91ef37a0b90d66ace7fb20fd7c3965c26421cbe51d

Observation f09a84f8-cd03-4435-b572-40820617678a · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.517193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:edc5114f3e21788d7c01d35afdd05bfb975703d378356fc9948a39fc47b84ce2

Observation 9299af09-cb1b-48c1-9478-4ee03cd0ebf6 · outbound

This paper cites ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.618491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:07fcaa5971f2186e0ac012acee87ec68e35bf1f410269edb6983cdca06638102

Observation c8914bf5-4cef-41c4-9da6-e2f06246096e · outbound

This paper cites RoboRefer: Towards spatial referring with rea- soning in vision-language models for robotics.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence RoboRefer: Towards spatial referring with rea- soning in vision-language models for robotics

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.570597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:75a730d0ec62bf6aa3d619285854f3a92d9b476473d7e0933e8e8f9ff560d275

Observation a8fe1062-99fb-4d1e-8d07-2ddd18337f18 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.553695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:d059b30fdfad54c45174f67399438de50ecf29564f11d189e48024d7e216462c

Observation 0a37124e-a73b-4485-b966-0c8d1371122d · outbound

This paper cites Apollo: An Exploration of Video Understanding in Large Multimodal Models.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.596352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:91477aa8c21dcbbf9c40639ad21016ba25589fd7145adddfda404866b6597a11

Pith citing papers

Observation 840434c4-03d3-4765-814a-a88dc39941f2 · inbound

Benchmarking Visual State Tracking in Multimodal Video Understanding cites this paper.

Benchmarking Visual State Tracking in Multimodal Video Understanding LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:46:28.012629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:43:06.811228Z digest=sha256:2cc2b84bc90edf862712fa01eb018d535f6a6101a916a524088d6b90a69f6ea2

Observation a584e931-e40a-462b-bbde-022ef2d05f03 · inbound

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention cites this paper.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.892593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:940028e418481f3a091ee50329971361283eab5cb21852e51a9f7234e2058111

Observation 1fc63c2c-7128-4903-beb1-6f086b693d84 · inbound

SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe cites this paper.

SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-12T02:24:07.555850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:24:07.555850Z digest=sha256:a7fae37c760eaa8e0a0da4d8d744e77d9d9d556498ce2bc8b2cc5636c3475585

Observation 916cc133-aeb9-4a69-85be-5014ca2f787c · inbound

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding cites this paper.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:49.651285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:49.651285Z digest=sha256:63d122260f400df8ab21d90b6c95e130b903ae4407384bf950c49e7eccf5daf6

Observation 4fbf2bb5-3bfd-4746-9479-38aa8d899787 · inbound

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs cites this paper.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:11.814953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:11.814953Z digest=sha256:2bdab26b8d906dfb3fa30a2592ab0ac5f7b44add9178b519ef67642033de1a8d

Observation a1c55d83-6932-4859-8cb1-43d89de6ded3 · inbound

MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement cites this paper.

MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-01T08:37:44.391888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:37:44.391888Z digest=sha256:18f350d81d4432249f0298ce1848e5ab0f0cb8e03c6d17bbd07c31b8eb34d45c

Observation 3bde2716-bd37-44f1-a6b5-078edd0816e2 · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:13.650085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:13.650085Z digest=sha256:2661267bb2ec98d5dc259a2c8447b28b242e245c8ced1940ffbef6bc6436d58d

Observation 9e5ef122-d5b2-4e1c-ac12-b90a3088b2d5 · inbound

RefCaptioner: Multi-Reference Image-Grounded Video Captioning cites this paper.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:19.737340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:19.737340Z digest=sha256:4eb8d5f7ee2b470feaccd7305a5eb30de8f13ac0632a8f66884e464d4e728c2a

Observation e5460fb8-f207-4ee5-bc9f-a1a267c7fb62 · inbound

Trace, Verify, and Correct: A Training-Free Framework for Spatial Reasoning in Multimodal LLMs cites this paper.

Trace, Verify, and Correct: A Training-Free Framework for Spatial Reasoning in Multimodal LLMs LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:31.939898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:31.939898Z digest=sha256:7e6c4c4e4f1e43c13c21db0722ca36465d795393c13249bf6cc641e89cabcd38