Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2403.15377.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:02:29.351108Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T06:39:37.518508Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 0ea28090-2a9b-430f-a743-afc02aa2e0ea · inbound
How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 120
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6482212f-2fea-4ee8-af46-15f76d04a960 · inbound
Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 67648cd9-4a98-42e9-9df2-3427496945b5 · inbound
LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation be600ca6-544c-4333-806e-450eb4add321 · inbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 257
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9b3a2c62-41ad-4638-849f-bff192cadec6 · inbound
DOLLAR: Few-Step Video Generation via Distillation and Latent Reward Optimization InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d4ddbc8d-88bb-4485-b83d-33783b4f2547 · inbound
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e76c4e79-5a2b-4e07-83b2-a35350833ea4 · inbound
LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fe182bdf-a1eb-452c-ab36-07d2a8315a47 · inbound
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b85f535a-1e35-4405-ac04-20df1c5043b4 · inbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbc88a1a-335d-41cc-9183-f0aab1822dcd · inbound
RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f82f8a4-75f2-4b0b-bd9f-12cdf4e124f5 · inbound
HCQA-1.5 @ Ego4D EgoSchema Challenge 2025 InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c51851bd-3b04-4f8f-9f48-14366c20f7cd · inbound
HuMoCon: Concept Discovery for Human Motion Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c096c6d-ac30-465c-a225-b79619168202 · inbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34c1518b-6a96-4e58-ad3a-81998148f13c · inbound
VideoMolmo: Spatio-Temporal Grounding Meets Pointing InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a206936a-a89b-4446-b0ca-5da22eb7bbec · inbound
An Empirical study on LLM-based Log Retrieval for Software Engineering Metadata Management InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30471e61-9c6b-4718-a4bb-197fbf6e0fd0 · inbound
DejaVid: Encoder-Agnostic Learned Temporal Matching for Video Classification InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c6c2b82-1304-4e98-8301-7975bacc3446 · inbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26244bf5-0094-4a2f-a175-3efb65d83ac2 · inbound
How Far Can Off-the-Shelf Multimodal Large Language Models Go in Online Episodic Memory Question Answering? InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2293dc2-fba5-4555-8a7b-352c33d71e45 · inbound
LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e683b406-530d-4883-8e6f-29049fa30f24 · inbound
VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3f95cab1-e444-4ba3-b563-c49f4618a8e4 · inbound
Whom to Respond To? A Transformer-Based Model for Multi-Party Social Robot Interaction InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49664a06-07a5-49c3-a551-aeaea6f87f3f · inbound
HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63b3f5bf-aa89-4242-8139-ecca80d15c9e · inbound
AdsQA: Towards Advertisement Video Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4d6e6dc-1985-459c-a7f7-9e3438c90669 · inbound
StreamingVLM: Real-Time Understanding for Infinite Video Streams InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b4991dee-1341-458f-860d-3b3bdb084ac4 · inbound
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9d8ecc9e-0f2c-4a83-b438-04eb4b7f16cc · inbound
VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c27af26f-fadf-4167-a8b4-78c00b2f34a4 · inbound
AdaFocus: Adaptive Relevance-Diversity Sampling with Zero-Cache Look-back for Efficient Long Video Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3d550601-08e8-4fef-bd0a-d450e29c01e8 · inbound
When Vision Speaks for Sound InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 825cb563-f5c1-4de2-8225-143668a3cd13 · inbound
GIRL-DETR: Gradient-Isolated Reinforcement Learning for Video Moment Retrieval InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d949fff9-4f47-4bd9-9828-14cf0d8d1e52 · inbound
VidMsg: A Benchmark for Implicit Message Inference in Short Videos InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d76c4a66-6383-49dc-b854-8df5bd1deb6e · inbound
GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f705bd2b-15de-4e9e-a1c3-f31dd6f8147c · inbound
HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 32b0cf74-2c11-4ecd-993b-6de5dc48818f · inbound
HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f73a5272-7603-4603-9b33-02aa4aa36cda · inbound
MentalThink: Shaping Thoughts in Mental SVG World InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 141
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5be6167-3f15-4bf3-8fd8-5b7bbee536d4 · inbound
Reinforcement Learning: From Algorithms To Foundation Models InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.