Pith. sign in

Paper Citation Record · LEDGER

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

As of 22 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 12 inbound Pith citation observations for arXiv:2605.25979.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.25979 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T22:12:05.365596Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:20:20.740974Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-02T02:46:28.010713Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact42
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f1b74493-07db-4f49-a4fe-3cec07367994 · outbound

This paper cites SPARROW: Learning spatial precision and temporal referential consistency in pixel-grounded video MLLMs.arXiv:2603.12382,.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence SPARROW: Learning spatial precision and temporal referential consistency in pixel-grounded video MLLMs.arXiv:2603.12382,

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.593890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:cb6bc76d96ddbbb6bc29a644d88b5ac5b85b78737c6aea018624997c7f4e2fc4

Observation 581aa147-7652-47d4-b5e1-341d2bfcc9db · outbound

This paper cites LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.610224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:610d91d3cc3b9c9d67b7a3ce8d732327a1d846ebe9ba912737c8074633872bda

Observation c72ecbfe-6a27-46ce-95c5-c64057214303 · outbound

This paper cites Qwen3-VL Technical Report.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Qwen3-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.602205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:34169e4ddde99e048a7e832d738cc7ba3c35eb71263e068b2e3270088bc86b65

Observation 1cfba95e-4081-4c37-b912-7bb7ee65101e · outbound

This paper cites Reasoning over Video: Evaluating How MLLMs Extract, Integrate, and Reconstruct Spatiotemporal Evidence.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Reasoning over Video: Evaluating How MLLMs Extract, Integrate, and Reconstruct Spatiotemporal Evidence

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.534505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:84b35a377c247f97e7ad14f9bda6a75336f704ee65e976ad71030eac21510f80

Observation 7ab2eb34-103c-4f07-8f40-33ff95e5f780 · outbound

This paper cites Token Merging: Your ViT But Faster.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Token Merging: Your ViT But Faster

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.539418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:5363ac64159463b7059882e605d7cf394cb3264fb597341c140950d5b140e48d

Observation 016b2ce0-c946-4469-add7-ec1b9b10e22c · outbound

This paper cites SpatialBot: Precise Spatial Understanding with Vision Language Models.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.560592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:efb0f137f600eb96302c8774c6cc76b7f922aeb0f72fe16a23537fe0392f9ade

Observation 22ddb1ee-6978-4b5b-b3a2-2accac6591c3 · outbound

This paper cites Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.585208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:9b1282f9023a761f7d53b42545f3ab1bf49de46ab6052e1a65b378881e5b55fd

Observation ca9de3d6-bab6-4863-899b-95acdc4c3104 · outbound

This paper cites GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.634880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:73de943e8f5531b50155d9caf2b6138cd2b8a6d3d25dde82a4c01ed587d67055

Observation f0b9d4ca-36d0-4cd7-9035-4a3688a5c578 · outbound

This paper cites Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.651251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:72c29b43593aa1ac2887ced24cb3e0d792299ed12ee019468eefef51afb435d6

Observation 09a2392a-73ee-43b9-86d8-e8694f9d8ab5 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T22:13:59.623529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:a18b0a84c320a2410a898f9856303152d7e504a45ba360d7213f108e45c22409

Observation 11f6c408-8efc-44f9-9362-086b0115b8c7 · outbound

This paper cites Small Vision-Language Models are Smart Compressors for Long Video Understanding.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Small Vision-Language Models are Smart Compressors for Long Video Understanding

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.629109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:ac003518c1211ac49b08189586ee7399e796a959b2c88660f85b7c7051936100

Observation ccd5db0d-2e2c-4f72-8418-558f3dae9183 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.632058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:473b7694d46f7733fb9450fbad70b52393cbfb4016d4e58be7654a8385153e87

Observation 170f529f-5c12-4941-b8a8-b291450df40b · outbound

This paper cites Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.637794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:e4949770ee7dfc8d11aff8166050734e3014d3db7e43c81c19913139d9080fc0

Observation 393a6d2c-8ec2-4bae-a1c3-04800cbabd82 · outbound

This paper cites Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-20T02:18:27.127261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:8704b537870409f7e5b6588ba766e4aa8c368078188d166d2df3fab393b62286

Observation fe821ea5-a2a0-4274-9b99-441da40cc261 · outbound

This paper cites TRACE: Temporal Grounding Video LLM via Causal Event Modeling.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence TRACE: Temporal Grounding Video LLM via Causal Event Modeling

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.619875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:d7e62bd3921018193fc126beeaac6d7f5469a95eae7aa9556e092ab9cde607a4

Observation 4803f7e1-22e8-4c1a-bb14-5600f3539762 · outbound

This paper cites Spa3R: Predictive Spatial Field Modeling for 3D Visual Reasoning.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Spa3R: Predictive Spatial Field Modeling for 3D Visual Reasoning

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-08-14T01:00:35.997207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:d2f0062fe243fbfec2105753b255be67c340acf6b10a4f130c521c163098dbf6

Observation e20e8960-f15a-4301-9536-38f5411cd05d · outbound

This paper cites Token-efficient long video understanding for multimodal llms.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Token-efficient long video understanding for multimodal llms

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.617077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:78b6258168d845e92317960e4dd73ffde3a8e1d9628ecd4dc67c7b41c69e9611

Observation 6165672e-831e-4e3c-80a7-753e4d78d22c · outbound

This paper cites Agentrvos: Reasoning over object tracks for zero-shot referring video object segmentation.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Agentrvos: Reasoning over object tracks for zero-shot referring video object segmentation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.627023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:20ae01747da2c6341115c7918c3bac286dd91f93e88e56df58dc1284d6d1e2c4

Observation f31884c4-788a-4c9a-a17b-e330fee1ded4 · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.644704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:f00bea0ae783678b5e848ae750460530fce3973f3597a8cd0593a701b9e83776

Observation ae5d3360-d2fa-4279-b130-01683db20b90 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence VideoChat: Chat-Centric Video Understanding

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.604775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:f19463246aa525b76d5c1b662cc22909c89407832c065996b28708a0e4590713

Observation fb4fd70d-585c-48b0-9d6e-22d13ada3ef0 · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.641616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:11237c54b86c3b24983b7e9fa117dd20760174d8f101198dfe8954bda74311f8

Observation 764b775d-4ec5-44ac-8fa6-1402452ebb2e · outbound

This paper cites Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.625729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:171ed24dc279fa84113ae5fdcf2224c77890c76896dc31f2dcb8a638701d189f

Observation 9fc1f3ba-b6a6-45d6-87c6-c59c5d4c378d · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.596700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:0d465f9bec0abc4de13abbcb8294809bf0f1406f87baff9517e8e647ea98fbab

Observation fe30cdd5-5070-40a4-b718-f28b247069ae · outbound

This paper cites SpaceR: Reinforcing MLLMs in Video Spatial Reasoning.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence SpaceR: Reinforcing MLLMs in Video Spatial Reasoning

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.550815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:d9e1da8ccadf0a71905e598a610458480f2764b886907c2c615a7d421f5be32d

Observation d36b3184-77a9-48be-9423-6b0cd57d2bdd · outbound

This paper cites VideoMolmo: Spatio-Temporal Grounding Meets Pointing.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence VideoMolmo: Spatio-Temporal Grounding Meets Pointing

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.563621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:4c4664127602ad362218b78e0d4666826480e79365936dcdf06505f968a4dba1

Observation 62853dd5-b767-49ba-a9a9-4a13f37bcfab · outbound

This paper cites TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.539308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:b320b39b52ea2bd22ed5e3d5cebd5a3f9ad3ef10d4ec65b75d68c22f1175c4f6

Observation 3910d08d-800d-4c97-afac-42aeba8eb419 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.621151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:c590064bd3faf70799fc2a6a9b000460806be285fd6b2b06d4ea22df86661662

Observation 51c68fed-389e-4425-9ecc-5cf38efdc3c0 · outbound

This paper cites A Simple Baseline for Streaming Video Understanding.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence A Simple Baseline for Streaming Video Understanding

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.549967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:179ad14f9e4443fe801822ffaad3ca1f84a3e2a8b4a2bd1b21bce6007bdbf6b9

Observation b87dd21e-68ef-4313-9a82-2ef1a7024db7 · outbound

This paper cites Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.566701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:1b1fd68b2236f97f51a6de31197ffbf32845a949572604d03513009d2310bdb6

Observation aedd76ea-039d-426a-83ce-833ec7a451ed · outbound

This paper cites EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.599233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:29b8f13b4452dda75a14974f87b7ca30c5b850a01ca55a22d217e23258bac19a

Observation e28f33d9-354b-49a1-b24f-34aeae931da4 · outbound

This paper cites Onevision-encoder: Codec-aligned sparsity as a foundational principle for multimodal intelligence.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Onevision-encoder: Codec-aligned sparsity as a foundational principle for multimodal intelligence

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.524140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:021ca8bbca7b6e01b63a681126c6d79182691902be11015a8b7aab961101a447

Observation 7ecee9cc-9e28-48e1-bfca-af410ad6f48d · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.634667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:58c05451564e1fd63f3d1ceb80cdbb67ef4738e24a6566603bc4d0868be3498f

Observation e8e66c53-1aa6-462f-8eae-33e353ef4e39 · outbound

This paper cites Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.520753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:b538f0c90cebc7ca8df89ca6da1a011b0c0dc115230a8512920e48079f33f4c8

Observation 1289afdf-8daa-48c5-81c9-110436f9cd37 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T22:13:59.631862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:61d91da6f87ba25bc34860009f06444e9e01a81eac0b0aa74714a4d9c2b2de48

Observation 9502d1c9-2655-4300-9f10-c4a927bcd63f · outbound

This paper cites Slow-Fast Architecture for Video Multi-Modal Large Language Models.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Slow-Fast Architecture for Video Multi-Modal Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.576656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:cb7ff1c4077debf13b9931fa6d0e6bf6e69f90e9c48a02096926252eb1e84417

Observation d5465249-d1cd-46c3-bc73-74d1dd4d8379 · outbound

This paper cites S-GRPO: Unified Post-Training for Large Vision-Language Models.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence S-GRPO: Unified Post-Training for Large Vision-Language Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.573692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:ad267025bf8f94b718687b71cb83b602d414742a48a6c02ee9ca13f4404d2fd9

Observation 5b48ea6b-4a0b-444a-a5ff-4f6ef7bdb83c · outbound

This paper cites Kwai Keye-VL 1.5 Technical Report.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Kwai Keye-VL 1.5 Technical Report

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.514435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:0404784814774d492046ce2c64187b86975841e6e492791ebf577c64ab7a66ff

Observation d07ce9e9-344a-4537-b425-749d832bfd2c · outbound

This paper cites MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.579379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:d9a5a83430aaaba847bbbdc279c82112cd792d9ae6285f024efed60cab566d5b

Observation d4292cc4-d4b9-41b1-b0f5-c9f077a36c83 · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.587708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:ba46a4fae02c60f9b93d5daead2e09270559b81162bd1631b6e82d46f7ade3a9

Observation e825296b-f6c7-4712-8e04-fc4d80848624 · outbound

This paper cites Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.509917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:0c6bffc7280510ec179067f3ca8f3be0367179b37741f3815eb44b1a4c3bca1e

Observation f09a84f8-cd03-4435-b572-40820617678a · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.517193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:ad913545a667b306b858a265249816afe7661246247709357e65c7b85f525f57

Observation 9299af09-cb1b-48c1-9478-4ee03cd0ebf6 · outbound

This paper cites ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.618491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:0be5023192d005032fba3869ceee6d32450c60b9120713a6ba85966a98486159

Observation c8914bf5-4cef-41c4-9da6-e2f06246096e · outbound

This paper cites RoboRefer: Towards spatial referring with rea- soning in vision-language models for robotics.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence RoboRefer: Towards spatial referring with rea- soning in vision-language models for robotics

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.570597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:ac0bad87327cfb1483dc0cf6337a85095fe74431332755d2738bcf814af11996

Observation a8fe1062-99fb-4d1e-8d07-2ddd18337f18 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.553695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:32be9413c87d4d8ade7142892a4dcfd0e3690230254e9fe1eab5b43a4d0b34fa

Observation 0a37124e-a73b-4485-b966-0c8d1371122d · outbound

This paper cites Apollo: An Exploration of Video Understanding in Large Multimodal Models.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.596352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:b3829a326398b55eb499fdceaae7bb9ddb90598df2fdf540a8368825920ec601

Pith citing papers

Observation 840434c4-03d3-4765-814a-a88dc39941f2 · inbound

Benchmarking Visual State Tracking in Multimodal Video Understanding cites this paper.

Benchmarking Visual State Tracking in Multimodal Video Understanding LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:46:28.012629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T10:43:06.811228Z digest=sha256:83a701b03140567e39eeb259bddaf478c6fae141c25dbf1a2be192ce154e72a9

Observation a584e931-e40a-462b-bbde-022ef2d05f03 · inbound

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention cites this paper.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.892593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:4f3a84a9acde15f52f701625bc2c00499e3d0a05fc026a16904a52476c5f00dc

Observation 1fc63c2c-7128-4903-beb1-6f086b693d84 · inbound

SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe cites this paper.

SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-12T02:24:07.555850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:24:07.555850Z digest=sha256:bc4674bdeb2802816bbd8a54231fb971e05733f1fe4e301e1c511c3ccda40d59

Observation 916cc133-aeb9-4a69-85be-5014ca2f787c · inbound

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding cites this paper.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:49.651285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:49.651285Z digest=sha256:680626a2412f5c9dd285399b2af5997d54428be0144435f4ef780003c4dc2f5a

Observation 4fbf2bb5-3bfd-4746-9479-38aa8d899787 · inbound

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs cites this paper.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:11.814953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:11.814953Z digest=sha256:65086c011c242cf4436398012c4249688368dde811308368bc167bb245dfa3f6

Observation a1c55d83-6932-4859-8cb1-43d89de6ded3 · inbound

MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement cites this paper.

MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-01T08:37:44.391888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:37:44.391888Z digest=sha256:a79834fa6757a7b14412fefdfe669ab5053010f0c7eb5b47af13bca9f78b086f

Observation 3bde2716-bd37-44f1-a6b5-078edd0816e2 · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:13.650085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:13.650085Z digest=sha256:b229854baa519ba2a3ea00ae272534849fc69750933c3b425ef753048b6bed31

Observation 9e5ef122-d5b2-4e1c-ac12-b90a3088b2d5 · inbound

RefCaptioner: Multi-Reference Image-Grounded Video Captioning cites this paper.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:19.737340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:19.737340Z digest=sha256:402530f14f4d0b8b56bc039993879c97ea912f5c0104070687096b8a8b3d80ef

Observation e5460fb8-f207-4ee5-bc9f-a1a267c7fb62 · inbound

Trace, Verify, and Correct: A Training-Free Framework for Spatial Reasoning in Multimodal LLMs cites this paper.

Trace, Verify, and Correct: A Training-Free Framework for Spatial Reasoning in Multimodal LLMs LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:31.939898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:31.939898Z digest=sha256:7a4f57a095c829354fb92e7c096a652210c8c960f7cf9b3a375ff96821f77e42

Observation fee22884-ee9d-4572-a6e7-eb3430ad59c6 · inbound

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding cites this paper.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.655128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.655128Z digest=sha256:0fb04675aeea01baeae83c9d0ef5719da15487bea72f088dc427716b366b0cca

Observation 27f0edc8-5a5e-47f1-96f6-283ff6edecb4 · inbound

Aero Realtime: Fully Aligned Input-Output Streams for Low-Latency Streaming Multimodal Generation cites this paper.

Aero Realtime: Fully Aligned Input-Output Streams for Low-Latency Streaming Multimodal Generation LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-14T04:39:27.707257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:39:27.707257Z digest=sha256:7358fa1820e78d47fcde2e10d3b950b9608e53e2f8741f47ff19152f051b2c6a

Observation 6ec6af18-8fb0-47d7-a452-2e51a8aeba86 · inbound

QV-PIC: Query-Aware Visual Position-Independent Caching for Efficient RAG Serving cites this paper.

QV-PIC: Query-Aware Visual Position-Independent Caching for Efficient RAG Serving LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-16T00:20:20.740974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:20:20.740974Z digest=sha256:7837b4fe3452f219ed50f60c99a1da93f4ae9be47e3c261342b13e99294f316d