Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T02:52:20.643070Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 84 inbound Pith citation observations for arXiv:2501.12386.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T02:52:20.643070Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:41:08.446932Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
37 of 37 outbound references displayed
External citation measurements
1
pith, observed 2026-08-05T02:28:24.338817Z
Observation 07d9e05a-eb3a-484f-b201-1502b89bf29a · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Cosmos World Foundation Model Platform for Physical AI
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4e2e81ef-2ec9-40a3-be2b-9b2916aabacf · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Qwen Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 939ad889-be6d-4612-ac51-676f8947dfee · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 879cafa7-3662-4c5c-a00f-474d3eb50d28 · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Token Merging: Your ViT But Faster
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1e93dfae-f8d2-48cb-bd97-4d65717864ef · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling InternLM2 Technical Report
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c04387ad-e5a2-4f71-b3a1-cb412a049b00 · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling HourVideo: 1-Hour Video-Language Understanding
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation be7b2f50-db56-46cf-a2b5-3f3ee2c0b42b · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5e544394-3655-4db0-a292-74efbb8a418d · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling USP: A Unified Sequence Parallelism Approach for Long Context Generative AI
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bcd3f53a-9953-4561-8386-c85bc699baec · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a1e3ce7b-6cf8-4833-9414-91be823de78e · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 40ea47f5-5c90-42a4-879d-4c6e346af6a6 · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 11b0d50c-dfca-4844-9b5b-a3276475e058 · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 18c528c3-c6ae-451a-aef5-04f7daf09b2f · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling The Kinetics Human Action Video Dataset
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d67c6540-880e-4403-8880-9d79099455cb · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling LLaVA-OneVision: Easy Visual Task Transfer
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f1b88b23-ddb9-408e-bb7e-bfcb08664a1d · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c17e7587-fead-4bb2-a55f-bc474b366077 · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Ring Attention with Blockwise Transformers for Near-Infinite Context
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e17bb074-8708-42ef-85d6-ad5c89fc49cc · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3ed63893-240d-40cb-8740-2cf5cc93a0af · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling GPT-4 Technical Report
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7757c5cc-869c-40b8-871e-e694fe99aeb6 · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling SAM 2: Segment Anything in Images and Videos
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e8d91143-00b3-4879-899c-184ca49812c0 · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 74733113-350f-4079-852c-d6a93c738232 · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 790a7fd1-3f3b-4abd-829e-7bf975876128 · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8467a30b-0658-47f3-8e17-081184e9bdde · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dcaf2cd8-1f50-401c-83f5-de3f9366e583 · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2cd6e254-0fd9-4e7c-bce5-86945bd97a24 · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Longllava: Scaling multi-modal llms to 1000 images efficiently via hybrid architecture
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e8d8b3c2-c456-4eff-af15-513248c8a146 · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 716f8125-2ca2-4693-9d43-eb94caa2de63 · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling HawkEye: Training Video-Text LLMs for Grounding Text in Videos
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 56b398d2-a2a2-4efb-9779-f18fdb255c2f · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d5166116-14a6-4c38-b05d-f5764bbb5c74 · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e2170d37-63ab-44c0-b67a-6e242c29aca4 · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d9d4332a-5f1e-483f-a493-4b5634d2952e · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 50d8304b-f818-4452-a6c0-04aa645e7471 · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Vript: A Video Is Worth Thousands of Words
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9710c9e3-ac62-487f-8521-303ae1e425ac · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 64f6bba2-b9e5-4be4-a078-cbdb17883279 · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling NExT-Chat: An LMM for Chat, Detection and Segmentation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 957433df-c785-46cb-a36e-7d1aaf9c4f52 · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c4101e25-5707-4f19-9c40-cc2457a97e13 · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling MLVU: Benchmarking Multi-task Long Video Understanding
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 51649a7d-f4f2-4ec0-8e87-21456cbedbd9 · outbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Apollo: An Exploration of Video Understanding in Large Multimodal Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a0ccc2c1-0ac2-406c-8c64-f23499a54bad · inbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 33c74869-a7ea-414b-91da-4bde70336a44 · inbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a97207d3-7a08-4b04-b7be-0cd06df75772 · inbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ffb3ee0-50c0-4ea5-affe-48f2ee8e725d · inbound
ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b4f82a9-424b-4afe-922d-3d11b78ce6a3 · inbound
QuickVideo: Real-Time Long Video Understanding with System Algorithm Co-Design InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40cb993f-0961-4de2-82ad-085c8318420d · inbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0d7da37b-986e-4383-9fe7-4e56a55a6701 · inbound
Temporal Consistency Constrained Transferable Adversarial Attacks with Background Mixup for Action Recognition InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f54c8e76-bdc0-480b-9fc8-a497a0ec08de · inbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01a233a0-30b5-42cc-a9b7-5a273bf9102d · inbound
Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee4104bc-0d5b-416d-9007-440cd0f92730 · inbound
DisTime: Distribution-based Time Representation for Video Large Language Models InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1ad3c69-2093-45df-9890-bc1709f523ce · inbound
SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c8498e34-385e-46d5-b571-f9cbbe09859c · inbound
Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e4c774c-c2f9-4a2b-8269-0bdcd2437833 · inbound
UNIC: Unified In-Context Video Editing InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cf9e3ab-02f8-44a7-a346-1c46941f6bbd · inbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 580bc404-e098-47ea-8eca-2769ff68229a · inbound
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc628e12-ea92-4ca2-8b7b-a2916806fa5f · inbound
Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61606932-3186-4d92-91fb-d4541afcc799 · inbound
CyberV: Cybernetics for Test-time Scaling in Video Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc595d87-8263-4001-ad95-c20fdf1fbb47 · inbound
Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8c1bebf-3b71-4951-8ed2-20c1839bf3c5 · inbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdd91320-29dd-4a1d-9682-0486a7bc0404 · inbound
Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b26fa6d3-8e11-4c3e-9537-a920ae3bceae · inbound
NavComposer: Composing Language Instructions for Navigation Trajectories through Action-Scene-Object Modularization InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e461aa4-bd38-4e8b-a4bb-e5313aa4cb95 · inbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 537ae9d2-8b8b-4ee2-87fc-d2868ca37de7 · inbound
"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e018751-258d-4358-8242-05dba5a50bca · inbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ac3d27f-7c44-4e0a-be78-6d8ba40ee92e · inbound
VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5076d94-07c6-4475-8fb1-31568095960e · inbound
TennisTV: Do Multimodal Large Language Models Understand Tennis Rallies? InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6c2e6875-f1a1-4088-bc1c-21e4d053fcd5 · inbound
Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation af50635d-a1d9-4977-9180-3439cbf037d9 · inbound
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b064333-15cc-4f7f-9284-4c40a6317838 · inbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 125e64e3-7bd1-4b90-98ff-66ee8ca3c617 · inbound
OneThinker: All-in-one Reasoning Model for Image and Video InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f969c771-3294-4789-b06b-0369962a0a19 · inbound
Adapting MLLMs for Nuanced Video Retrieval InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 511079d9-318b-46a3-8e6e-da53b0d1ba2b · inbound
$M^3-Verse$: A "Spot the Difference" Challenge for Large Multimodal Models InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8eca61ed-2e08-4c71-9a6e-711a3e782706 · inbound
Streaming Video Instruction Tuning InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b5c2d4b6-449e-48e8-ba9b-73b57c33bf8e · inbound
CoVR-R:Reason-Aware Composed Video Retrieval InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0eb35f2e-c345-452b-ac95-bdcc3ba2d825 · inbound
Diagnosing Long-Video Quantitative Reasoning in Multimodal LLMs via Enumeration and Counting InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a8b1fb9-ff9f-42df-a6b5-7d0c193524d5 · inbound
Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9927fd3a-fbcf-4b98-8340-fe8b16e5cc32 · inbound
InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d75ee4db-3bfd-4077-ada7-08f5479605c9 · inbound
How Should Video LLMs Output Time? An Analysis of Efficient Temporal Grounding Paradigms InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ba20b7fd-ddc6-422c-b8ae-4b189c29cccf · inbound
Decoupled Similarity for Task-Aware Token Pruning in Large Vision-Language Models InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bc29fe2b-741c-42aa-b69d-71c896e9466f · inbound
All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bc44a8a6-a9d4-4297-a4aa-4a1f6216f892 · inbound
Grounding Video Reasoning in Physical Signals InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fbce2764-2f30-44e6-9315-2647701224ee · inbound
High-Speed Vision Improves Zero-Shot Semantic Understanding of Human Actions InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c37adac5-757e-451d-8091-6c3842fc4730 · inbound
From Priors to Perception: Grounding Video-LLMs in Physical Reality InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 923b11be-4673-4dcc-ab52-11d504621cea · inbound
VISD: Enhancing Video Reasoning via Structured Self-Distillation InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cf8b4a35-bc2a-4db0-b47f-2e8b5783d3fa · inbound
VISD: Enhancing Video Reasoning via Structured Self-Distillation InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 212bc8f4-2e67-4968-b456-18f3da2e09b4 · inbound
VISD: Enhancing Video Reasoning via Structured Self-Distillation InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 35b43260-bc1a-4853-8630-7f139cb4eca7 · inbound
VISD: Enhancing Video Reasoning via Structured Self-Distillation InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f258556c-d103-4168-afd5-85ef3680fb2f · inbound
Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b04166d4-610d-4ffc-9b28-986a6891da3d · inbound
EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b81a58f8-962f-4907-9aa2-117c800457fc · inbound
TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0c3eca2b-ff6e-4ab1-810b-149a51feefe4 · inbound
TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ede6c788-8b9f-49a7-8cbd-e05aa70ac9c4 · inbound
CoRDS: Coreset-based Representative and Diverse Selection for Streaming Video Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 58ae6413-6598-4b95-937b-02f8a737aa6d · inbound
Video-Zero: Self-Evolution Video Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fef7ad78-f74b-42a8-9b45-2ff7f068697c · inbound
LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 83ae1637-3a12-4761-a6c8-4f0e0de9c6c0 · inbound
LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3face163-f13a-4359-8295-62fd8ad61842 · inbound
OProver: A Unified Framework for Agentic Formal Theorem Proving InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bf36e788-09c0-43e7-8528-2b51fd61e1ed · inbound
MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation eb83c77e-b09d-4845-80da-8151f417c158 · inbound
IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bb48f86c-314b-441a-9a64-3c18019f3738 · inbound
VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 670c8e33-42ac-4af3-9923-6d3663bc5fc1 · inbound
Masked Diffusion Vision-Language Models for Temporal Action Localization InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6e7de1fb-0a2f-4b9c-a6a3-039326139e56 · inbound
ViCuR: Visual Cues as Recoverable Privilege for Multimodal On-Policy Distillation InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0f97e9b7-e167-413e-886c-b5181e0c233a · inbound
GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b62462ba-6f47-43e3-948c-df820a424819 · inbound
Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 76fafa33-ac4f-443b-845c-5b6137108df7 · inbound
CoCoSI: Collaborative Cognitive Map Construction for Spatial Intelligence InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8ae19305-c3cc-420b-abc8-6aa4de49f69c · inbound
MultiToP: Learning to Patch Visual Tokens to Mitigate Hallucinations in Video Large Multimodal Models InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 067fdded-660c-4c46-95ad-f2f2b80fb3ae · inbound
InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 40caaa14-d8ec-4074-a426-ddab0dadf133 · inbound
LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c15dff4b-100b-4359-9947-c668c84245f7 · inbound
Latent Visual Diffusion Reasoning with Monte Carlo Tree Search InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 74cbdf07-f726-49e6-bbac-7ad254b47189 · inbound
MotionAtlas: Detailed Region Captioning for Motion-Centric Videos InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 739ba1f6-1d06-49be-a72b-87a3066511ab · inbound
Learning to Deny: Action Denial in Multimodal Large Language Models InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f29c3843-0f5c-4a33-83c4-b1b4a5e870e8 · inbound
Bridging Video Understanding and Generation in a Unified Framework InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5add77d4-1b8a-497d-908a-e85aa98bd5ce · inbound
LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5e8d3556-65e6-4e26-8f18-241baa4085a2 · inbound
Natural Language Camera Movement Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e882bf37-90d4-499d-b6eb-528ea4968e99 · inbound
Probing Identity-Specific Motion Signatures: A Controlled Diagnostic Study InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0b6ef65-6f2c-4d51-b3ef-f0f043eb7e6f · inbound
DynTrace: Tracking Dynamic Object Evidence for 4D Spatio-Temporal Reasoning in MLLMs InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbcfc400-3eab-4288-add5-b74a12a64c15 · inbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa7e5358-e336-43c6-930b-f843dab8d307 · inbound
MAC 2026: Advancing Micro-Action Analysis Towards Fine-Grained Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a282df78-dfdb-4c02-a1fa-68bff1bba73f · inbound
Continual Video-MLLM Adaptation over Evolving Domains InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff3267ed-70a9-484c-bfc9-b1e7b5a9304d · inbound
V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f64894f-968d-4087-8589-5ad2eec00747 · inbound
TimePLE: Rethinking Temporal Representation for Video Temporal Grounding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2e2663e-7d84-4879-84ce-54ed7852962f · inbound
FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0c17133-5f7a-4704-b073-b60832060a27 · inbound
FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32732981-9dac-4e5e-b682-39fa7b19ab3a · inbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac560be2-623d-42f1-83e5-80d01c5921d6 · inbound
Think in Sets for Streaming Video Token Compression InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.