Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2305.13292.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:11:40.910714Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T20:18:57.839456Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation c541cc23-02e0-4193-9ac5-4076649b7a59 · inbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation VideoLLM: Modeling Video Sequence with Large Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 81c23fb6-b03e-401e-be30-7c26780bc04f · inbound
A Survey on Deep Learning Techniques for Action Anticipation VideoLLM: Modeling Video Sequence with Large Language Models
Reference 203
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d701eab2-8540-42d2-8903-ccc673024f4b · inbound
SALMONN: Towards Generic Hearing Abilities for Large Language Models VideoLLM: Modeling Video Sequence with Large Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dfe2ba98-9681-486e-9835-546287e6b040 · inbound
MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems? VideoLLM: Modeling Video Sequence with Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8f243c46-e496-4bbc-a599-432d3c31c16a · inbound
How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites VideoLLM: Modeling Video Sequence with Large Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ab5faf8c-5e75-43c2-9170-1fa85c9b3b3e · inbound
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning VideoLLM: Modeling Video Sequence with Large Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 010363d5-1c73-4422-9399-4cab4a8f1a34 · inbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models VideoLLM: Modeling Video Sequence with Large Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d481ca96-e0bd-4cfd-809f-c21b49e89785 · inbound
Polymath: A Challenging Multi-modal Mathematical Reasoning Benchmark VideoLLM: Modeling Video Sequence with Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5e2d9d77-a5ee-42a6-8377-2f5e94b38832 · inbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs VideoLLM: Modeling Video Sequence with Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0b2b3f3-62bc-4850-aaa8-558c9c09418e · inbound
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning VideoLLM: Modeling Video Sequence with Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a03daa1-2d5c-431d-b20d-263c1ca55de5 · inbound
TriPSS: A Tri-Modal Keyframe Extraction Framework Using Perceptual, Structural, and Semantic Representations VideoLLM: Modeling Video Sequence with Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a83717f8-8b63-476c-98ca-bf04eebc46aa · inbound
Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision VideoLLM: Modeling Video Sequence with Large Language Models
Reference 215
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbb85bc4-a452-4b8e-8689-a4428efd383a · inbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning VideoLLM: Modeling Video Sequence with Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73726ca9-4546-48a2-8657-49909d9a5b56 · inbound
NeuroVoxel-LM: Language-Aligned 3D Perception via Dynamic Voxelization and Meta-Embedding VideoLLM: Modeling Video Sequence with Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dfa47f3-9922-4b81-8194-46f26f8dda1d · inbound
Bidirectional Action Sequence Learning for Long-term Action Anticipation with Large Language Models VideoLLM: Modeling Video Sequence with Large Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 367ec431-a646-46d0-a3c7-c31da349848e · inbound
Training-Free Multimodal Large Language Model Orchestration VideoLLM: Modeling Video Sequence with Large Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 25faba18-f547-4682-8c3b-082d20a552dd · inbound
Training-Free Multimodal Large Language Model Orchestration VideoLLM: Modeling Video Sequence with Large Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a96d9ce7-a49e-49c9-a47c-9a59b83bc134 · inbound
Time-Scaling State-Space Models for Dense Video Captioning VideoLLM: Modeling Video Sequence with Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 505e4779-9191-4b37-96e9-7e41a182180c · inbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey VideoLLM: Modeling Video Sequence with Large Language Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ca07bbee-5094-45c2-b6bd-53e77596f81a · inbound
Scaling Video Understanding via Compact Latent Multi-Agent Collaboration VideoLLM: Modeling Video Sequence with Large Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 86af773a-8a1e-4a66-8a69-88ca542db661 · inbound
Reasoning-Guided Grounding: Elevating Video Anomaly Detection through Multimodal Large Language Models VideoLLM: Modeling Video Sequence with Large Language Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation df02794c-9aaf-4d32-856d-c0aa80f3e09f · inbound
An Attribute-Based Measure of Video Complexity VideoLLM: Modeling Video Sequence with Large Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ec893098-6043-4d53-9a47-88dc7f1c0ffb · inbound
InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning VideoLLM: Modeling Video Sequence with Large Language Models
Reference 149
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a924f3ed-4503-416d-899e-a1de50b91659 · inbound
MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning VideoLLM: Modeling Video Sequence with Large Language Models
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 204689fb-d777-4c79-b01f-1969c1657c9f · inbound
GMoT: Gated Motion-Aware Tokenization for Fine-Grained Micro-Gesture Video Reasoning with Multimodal LLMs VideoLLM: Modeling Video Sequence with Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.