Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T22:12:05.365596Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 9 inbound Pith citation observations for arXiv:2605.25979.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T22:12:05.365596Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:22:31.939898Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-02T02:46:28.010713Z
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f1b74493-07db-4f49-a4fe-3cec07367994 · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence SPARROW: Learning spatial precision and temporal referential consistency in pixel-grounded video MLLMs.arXiv:2603.12382,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 581aa147-7652-47d4-b5e1-341d2bfcc9db · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c72ecbfe-6a27-46ce-95c5-c64057214303 · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Qwen3-VL Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1cfba95e-4081-4c37-b912-7bb7ee65101e · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Reasoning over Video: Evaluating How MLLMs Extract, Integrate, and Reconstruct Spatiotemporal Evidence
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7ab2eb34-103c-4f07-8f40-33ff95e5f780 · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Token Merging: Your ViT But Faster
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 016b2ce0-c946-4469-add7-ec1b9b10e22c · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence SpatialBot: Precise Spatial Understanding with Vision Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 22ddb1ee-6978-4b5b-b3a2-2accac6591c3 · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ca9de3d6-bab6-4863-899b-95acdc4c3104 · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f0b9d4ca-36d0-4cd7-9035-4a3688a5c578 · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 09a2392a-73ee-43b9-86d8-e8694f9d8ab5 · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 11f6c408-8efc-44f9-9362-086b0115b8c7 · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Small Vision-Language Models are Smart Compressors for Long Video Understanding
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ccd5db0d-2e2c-4f72-8418-558f3dae9183 · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Video-R1: Reinforcing Video Reasoning in MLLMs
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 170f529f-5c12-4941-b8a8-b291450df40b · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 393a6d2c-8ec2-4bae-a1c3-04800cbabd82 · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fe821ea5-a2a0-4274-9b99-441da40cc261 · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence TRACE: Temporal Grounding Video LLM via Causal Event Modeling
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4803f7e1-22e8-4c1a-bb14-5600f3539762 · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Spa3R: Predictive spatial field modeling for 3D visual reasoning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e20e8960-f15a-4301-9536-38f5411cd05d · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Token-efficient long video understanding for multimodal llms
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6165672e-831e-4e3c-80a7-753e4d78d22c · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Agentrvos: Reasoning over object tracks for zero-shot referring video object segmentation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f31884c4-788a-4c9a-a17b-e330fee1ded4 · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ae5d3360-d2fa-4279-b130-01683db20b90 · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence VideoChat: Chat-Centric Video Understanding
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fb4fd70d-585c-48b0-9d6e-22d13ada3ef0 · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 764b775d-4ec5-44ac-8fa6-1402452ebb2e · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9fc1f3ba-b6a6-45d6-87c6-c59c5d4c378d · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fe30cdd5-5070-40a4-b718-f28b247069ae · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d36b3184-77a9-48be-9423-6b0cd57d2bdd · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence VideoMolmo: Spatio-Temporal Grounding Meets Pointing
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 62853dd5-b767-49ba-a9a9-4a13f37bcfab · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3910d08d-800d-4c97-afac-42aeba8eb419 · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 51c68fed-389e-4425-9ecc-5cf38efdc3c0 · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence A Simple Baseline for Streaming Video Understanding
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b87dd21e-68ef-4313-9a82-2ef1a7024db7 · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation aedd76ea-039d-426a-83ce-833ec7a451ed · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e28f33d9-354b-49a1-b24f-34aeae931da4 · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Onevision-encoder: Codec-aligned sparsity as a foundational principle for multimodal intelligence
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7ecee9cc-9e28-48e1-bfca-af410ad6f48d · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e8e66c53-1aa6-462f-8eae-33e353ef4e39 · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1289afdf-8daa-48c5-81c9-110436f9cd37 · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T17:38:12.771444+00:00.
Observation 9502d1c9-2655-4300-9f10-c4a927bcd63f · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Slow-Fast Architecture for Video Multi-Modal Large Language Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d5465249-d1cd-46c3-bc73-74d1dd4d8379 · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence S-GRPO: Unified Post-Training for Large Vision-Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5b48ea6b-4a0b-444a-a5ff-4f6ef7bdb83c · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Kwai Keye-VL 1.5 Technical Report
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d07ce9e9-344a-4537-b425-749d832bfd2c · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d4292cc4-d4b9-41b1-b0f5-c9f077a36c83 · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e825296b-f6c7-4712-8e04-fc4d80848624 · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f09a84f8-cd03-4435-b572-40820617678a · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9299af09-cb1b-48c1-9478-4ee03cd0ebf6 · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c8914bf5-4cef-41c4-9da6-e2f06246096e · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence RoboRefer: Towards spatial referring with rea- soning in vision-language models for robotics
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a8fe1062-99fb-4d1e-8d07-2ddd18337f18 · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0a37124e-a73b-4485-b966-0c8d1371122d · outbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Apollo: An Exploration of Video Understanding in Large Multimodal Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 840434c4-03d3-4765-814a-a88dc39941f2 · inbound
Benchmarking Visual State Tracking in Multimodal Video Understanding LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a584e931-e40a-462b-bbde-022ef2d05f03 · inbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1fc63c2c-7128-4903-beb1-6f086b693d84 · inbound
SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 916cc133-aeb9-4a69-85be-5014ca2f787c · inbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fbf2bb5-3bfd-4746-9479-38aa8d899787 · inbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1c55d83-6932-4859-8cb1-43d89de6ded3 · inbound
MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bde2716-bd37-44f1-a6b5-078edd0816e2 · inbound
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e5ef122-d5b2-4e1c-ac12-b90a3088b2d5 · inbound
RefCaptioner: Multi-Reference Image-Grounded Video Captioning LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5460fb8-f207-4ee5-bc9f-a1a267c7fb62 · inbound
Trace, Verify, and Correct: A Training-Free Framework for Spatial Reasoning in Multimodal LLMs LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.