Pith. sign in

Paper Citation Record · LEDGER

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

As of 11 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 84 inbound Pith citation observations for arXiv:2501.12386.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12386 v3

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T02:52:20.643070Z

measured 121 of 121 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 84 of 84 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:41:08.446932Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact32
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 07d9e05a-eb3a-484f-b201-1502b89bf29a · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Cosmos World Foundation Model Platform for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:52:20.670876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:efbecdcdfeba46af9cf89df4ce21b109993e83ec4fc40c990c1068bb929f929e

Observation 4e2e81ef-2ec9-40a3-be2b-9b2916aabacf · outbound

This paper cites Qwen Technical Report.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Qwen Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:52:20.675691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:efcc0564701fa0091df405938c7557620613152f8651e87b5d6d75ae27cb139c

Observation 939ad889-be6d-4612-ac51-676f8947dfee · outbound

This paper cites One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.680766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:c15dccae5b3f916f70f44af375757ccdb30bda55b27c5ebfb09f1bc0316d4644

Observation 879cafa7-3662-4c5c-a00f-474d3eb50d28 · outbound

This paper cites Token Merging: Your ViT But Faster.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Token Merging: Your ViT But Faster

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:52:20.685487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:8f42f5aa24df0f1e0006f33e555523da8f14d76c3110daff61aaa6dd9741bee1

Observation 1e93dfae-f8d2-48cb-bd97-4d65717864ef · outbound

This paper cites InternLM2 Technical Report.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling InternLM2 Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:52:20.689675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:46487386a15bc036b1f5bd5ae9e21517d93e6777cfb7f3ae1d181306776db80a

Observation c04387ad-e5a2-4f71-b3a1-cb412a049b00 · outbound

This paper cites HourVideo: 1-Hour Video-Language Understanding.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling HourVideo: 1-Hour Video-Language Understanding

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.694510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:0616fd9e5129aa1e3949e741239e66c8392712a1688d1722620ceda1cc73a23a

Observation be7b2f50-db56-46cf-a2b5-3f3ee2c0b42b · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:83f20761796c715a45cb2ee275a431253e5b8a66eaba3c6a2e89e0a0a342404f

Observation 5e544394-3655-4db0-a292-74efbb8a418d · outbound

This paper cites USP: A Unified Sequence Parallelism Approach for Long Context Generative AI.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling USP: A Unified Sequence Parallelism Approach for Long Context Generative AI

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.705108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:3cba234c846aa0f719dd16d239c9461d77390efa5a780784a2b8cf23cd72de77

Observation bcd3f53a-9953-4561-8386-c85bc699baec · outbound

This paper cites MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.710199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:8716f191366900dc85bfdad8e049ae0b39637b07627715e146597b851fc546ff

Observation a1e3ce7b-6cf8-4833-9414-91be823de78e · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.714773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:87c4ecc36e07cef04c70acc15da798dbbffa825d530128f7601430ae15e355f9

Observation 40ea47f5-5c90-42a4-879d-4c6e346af6a6 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:52:20.718503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:2615aaaba954e610c8db0a3b5f35fc0e3360ea541d649347c1db6b08752f3466

Observation 11b0d50c-dfca-4844-9b5b-a3276475e058 · outbound

This paper cites DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T02:52:20.722917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:8cbdd7e751edb21207b2dbf835bad31f790b5906981ffb7f6f95aedb0c2340d1

Observation 18c528c3-c6ae-451a-aef5-04f7daf09b2f · outbound

This paper cites The Kinetics Human Action Video Dataset.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling The Kinetics Human Action Video Dataset

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:52:20.727050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:1467f972311e6a9799521ffc3f74777c99f94e1cb36ad07abb37859cdbcc4c35

Observation d67c6540-880e-4403-8880-9d79099455cb · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling LLaVA-OneVision: Easy Visual Task Transfer

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:52:20.731060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:eb4e5437855130871932cc02073c5e453eb9405051c7aca427b19b60dd8e4ba4

Observation f1b88b23-ddb9-408e-bb7e-bfcb08664a1d · outbound

This paper cites OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.735829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:68402f47c040ca386e57541f4cd1a45504483b5900ed4f2a5947e133bb65839a

Observation c17e7587-fead-4bb2-a55f-bc474b366077 · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:52:20.740404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:74614e757e5b25a16ab7e02c06f7d7bbc4e99d333547bf6a186b742740a0e2de

Observation e17bb074-8708-42ef-85d6-ad5c89fc49cc · outbound

This paper cites VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.745049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:6f2af2c059bc6d5b371dd3fea02adb1faacf877ec88bdaf75f5bfea3b50d59ea

Observation 3ed63893-240d-40cb-8740-2cf5cc93a0af · outbound

This paper cites GPT-4 Technical Report.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling GPT-4 Technical Report

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:52:20.749277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:e7b5bd0607352c458a587fefc781c703922d8f12346025d6f17c2c69abd0226a

Observation 7757c5cc-869c-40b8-871e-e694fe99aeb6 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling SAM 2: Segment Anything in Images and Videos

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T02:52:20.752983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:1e2f9b74d30b126426661992727d71b3766de65d9ff1a30e526c9d2fb776e727

Observation e8d91143-00b3-4879-899c-184ca49812c0 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T02:52:20.757255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:37a2c40b452616dcedb0625aff10a111d8808618f959000be1ebd31eca9fb3c0

Observation 74733113-350f-4079-852c-d6a93c738232 · outbound

This paper cites TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.762202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:53b284ff8c29195882e0c98c18d6d702c7f2fe1a43589f4e531765a621a70c2d

Observation 790a7fd1-3f3b-4abd-829e-7bf975876128 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T02:52:20.765897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:9bb7d123bc08c21366cd29a48a80894f18b72d6e918769a042c1eb72a6057a1a

Observation 8467a30b-0658-47f3-8e17-081184e9bdde · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.770114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:5c403722f1063bbac1d4a068b880ba82b37bf39afee9575c99887a68566e3f2c

Observation dcaf2cd8-1f50-401c-83f5-de3f9366e583 · outbound

This paper cites Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.775068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:bfb06c5bb3e3eb6e766fed7d7b148465ae93aa0bd3b2b69fdbe9dd9468cb1200

Observation 2cd6e254-0fd9-4e7c-bce5-86945bd97a24 · outbound

This paper cites Longllava: Scaling multi-modal llms to 1000 images efficiently via hybrid architecture.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Longllava: Scaling multi-modal llms to 1000 images efficiently via hybrid architecture

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.779477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:8aef549eff13b7660348790a76d330ad532375578e3446a322cdfd5d9068cf55

Observation e8d8b3c2-c456-4eff-af15-513248c8a146 · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:52:20.783565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:fe1dda7bc6dfe3e7dd4633099b927d4b14955b3468747ed43771ae557d525fd5

Observation 716f8125-2ca2-4693-9d43-eb94caa2de63 · outbound

This paper cites HawkEye: Training Video-Text LLMs for Grounding Text in Videos.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T02:52:20.788080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:230f7fd3005aaaed72ba32c3d72d7a5d6f56a026cb0355762e775fb49c5a6cd2

Observation 56b398d2-a2a2-4efb-9779-f18fdb255c2f · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:30:12.519784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:a24d7fa807b262646c10badef5cb4686c62f38fa8777c15222b0090f63ebf24a

Observation d5166116-14a6-4c38-b05d-f5764bbb5c74 · outbound

This paper cites VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.796481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:6fc73722e25e84bf767b5178f2c49dd250b3ed0e1460670156865639339263ea

Observation e2170d37-63ab-44c0-b67a-6e242c29aca4 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.572615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:19871d900dfa787c3024fea9f3b49063fc39a1b9ce2582921935966dfa86423e

Observation d9d4332a-5f1e-483f-a493-4b5634d2952e · outbound

This paper cites Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.806469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:96c605e08670aa6aea4686f994212bcfe97e9d83a1effac2335e5e496ec4d71f

Observation 50d8304b-f818-4452-a6c0-04aa645e7471 · outbound

This paper cites Vript: A Video Is Worth Thousands of Words.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Vript: A Video Is Worth Thousands of Words

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.810354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:01ca2eb3e171ce3b072e31d058cac52c79604002e626f827da7270f1323f6d71

Observation 9710c9e3-ac62-487f-8521-303ae1e425ac · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.894768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:16145e5c98eef6ad5f18477fae786d1df1c2e759cbad55614f6977534bb94bee

Observation 64f6bba2-b9e5-4be4-a078-cbdb17883279 · outbound

This paper cites NExT-Chat: An LMM for Chat, Detection and Segmentation.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling NExT-Chat: An LMM for Chat, Detection and Segmentation

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.818808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:c5d052c7c7e762257e891a73cff9b3c52dc0d3b0d076d806f3e0893ad99aea29

Observation 957433df-c785-46cb-a36e-7d1aaf9c4f52 · outbound

This paper cites LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.823715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:e3055d3c5dc60db05c81d4bc99387aef57b27ed748dc6cb350dd804dcd715657

Observation c4101e25-5707-4f19-9c40-cc2457a97e13 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling MLVU: Benchmarking Multi-task Long Video Understanding

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:52:20.827990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:4f4b4dadbf41d10310d3d023f2692a018e76b4ce3db7439d4298d5a81e95aa6f

Observation 51649a7d-f4f2-4ec0-8e87-21456cbedbd9 · outbound

This paper cites Apollo: An Exploration of Video Understanding in Large Multimodal Models.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.833571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:a5779ac6a39a8b843fbbff6c9ebbbf78ad632ec3bed189aa663579dfef91d613

Pith citing papers

Observation a0ccc2c1-0ac2-406c-8c64-f23499a54bad · inbound

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning cites this paper.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.835573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:7120769de62b77e256e3fde5f39d40671f3b5d84efd5200308a274575d075d44

Observation 33c74869-a7ea-414b-91da-4bde70336a44 · inbound

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? cites this paper.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:08.446932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:08.446932Z digest=sha256:216d15575a5a09422e1ca7d3c4a665bd3f91e2d0de2de229cda55e5356eab0fc

Observation a97207d3-7a08-4b04-b7be-0cd06df75772 · inbound

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation cites this paper.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:41.079838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:41.079838Z digest=sha256:53d497be4d44a6d0d535dc4493af845cb24a7c7a2a737f31ae1b88583870d311

Observation 8ffb3ee0-50c0-4ea5-affe-48f2ee8e725d · inbound

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning cites this paper.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:59.072257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:59.072257Z digest=sha256:de63e113d8ba957c9300fc9c2c3c47346dfc101ff2a3855a58f4789f2065fcfa

Observation 3b4f82a9-424b-4afe-922d-3d11b78ce6a3 · inbound

QuickVideo: Real-Time Long Video Understanding with System Algorithm Co-Design cites this paper.

QuickVideo: Real-Time Long Video Understanding with System Algorithm Co-Design InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:01.009435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:01.009435Z digest=sha256:12f7365eedd2aaae2b1ffed31e07cfa0ebf0dd516d35b42cc9d44261b1d28898

Observation 40cb993f-0961-4de2-82ad-085c8318420d · inbound

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning cites this paper.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.399421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:ceb805d3ab83184e7b66ef465c79e8721d8e09c934a0c26ebcf8d4c45e5efe9b

Observation 0d7da37b-986e-4383-9fe7-4e56a55a6701 · inbound

Temporal Consistency Constrained Transferable Adversarial Attacks with Background Mixup for Action Recognition cites this paper.

Temporal Consistency Constrained Transferable Adversarial Attacks with Background Mixup for Action Recognition InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:30.557505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:30.557505Z digest=sha256:979c85dd2b873e8a7be1683a5dba618d1d1dcf9410d809dfa58304999c71a4bc

Observation f54c8e76-bdc0-480b-9fc8-a497a0ec08de · inbound

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval cites this paper.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.104101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.104101Z digest=sha256:6974429b478ac58da9bae386630c3e3ca1de6fef06206b7b27d9e19db1126928

Observation 01a233a0-30b5-42cc-a9b7-5a273bf9102d · inbound

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought cites this paper.

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:13.309276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:13.309276Z digest=sha256:21e3b625bfec548f3e863706ca227240482e4b3557c74a7884e8c0ca8d370918

Observation ee4104bc-0d5b-416d-9007-440cd0f92730 · inbound

DisTime: Distribution-based Time Representation for Video Large Language Models cites this paper.

DisTime: Distribution-based Time Representation for Video Large Language Models InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:54.252300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:54.252300Z digest=sha256:6fb9e14621bd26a123152cf599d6987d82419f4167457baf76eb4a16b5a81393

Observation f1ad3c69-2093-45df-9890-bc1709f523ce · inbound

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics cites this paper.

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.835573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T21:22:36.902119Z digest=sha256:d5979f6ee1269b1c26d86b33d879329d7c4bdbf07acd48ed2cde6646418b5c06

Observation c8498e34-385e-46d5-b571-f9cbbe09859c · inbound

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency cites this paper.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:47.149960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:47.149960Z digest=sha256:b7c23b15934186ec666a6cc9dfe8a99300895ecdb1c8a4446ac1533515722083

Observation 9e4c774c-c2f9-4a2b-8269-0bdcd2437833 · inbound

UNIC: Unified In-Context Video Editing cites this paper.

UNIC: Unified In-Context Video Editing InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:51:42.973201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:51:42.973201Z digest=sha256:001a160b32d6633c440d572b5a3d645d850de2c5c6b049933c90fa9bd7af925b

Observation 8cf9e3ab-02f8-44a7-a346-1c46941f6bbd · inbound

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs cites this paper.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.618482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.618482Z digest=sha256:ab3dee21f25732a7d0a630ceed1ce6d73a02661e076fc31c62c9fac120968fad

Observation 580bc404-e098-47ea-8eca-2769ff68229a · inbound

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos cites this paper.

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:25:36.311280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:25:36.311280Z digest=sha256:dbaf8f75e71ef6194bf088ba716c04ecefd0d49965101a528cc472f04223eaa4

Observation fc628e12-ea92-4ca2-8b7b-a2916806fa5f · inbound

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation cites this paper.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:43.067246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:43.067246Z digest=sha256:5da831521d97ac8d879ce74b1c70ea6860ca911e90d7b556688686b51bd8bc70

Observation 61606932-3186-4d92-91fb-d4541afcc799 · inbound

CyberV: Cybernetics for Test-time Scaling in Video Understanding cites this paper.

CyberV: Cybernetics for Test-time Scaling in Video Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:47.266447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:47.266447Z digest=sha256:932ae87549989218750726c3ad763139406561f6646e2a1fc62f3470236ee409

Observation dc595d87-8263-4001-ad95-c20fdf1fbb47 · inbound

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought cites this paper.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.389244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.389244Z digest=sha256:4c52862b41fb90e5a7c807b2b8270458aeac94d047709527949e397dbfb4c78a

Observation a8c1bebf-3b71-4951-8ed2-20c1839bf3c5 · inbound

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos cites this paper.

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T04:22:56.040389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:22:56.040389Z digest=sha256:bb6efd75e8067fbdac6e0235059417bd1128e6e95e99cd261dac8edfa9ab5cc4

Observation bdd91320-29dd-4a1d-9682-0486a7bc0404 · inbound

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning cites this paper.

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:36.092629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:36.092629Z digest=sha256:f81bf4871b2bb56aa0ccb6149452eb7c6c92f33da94fb7da5e5ebe2d8b2003e2

Observation b26fa6d3-8e11-4c3e-9537-a920ae3bceae · inbound

NavComposer: Composing Language Instructions for Navigation Trajectories through Action-Scene-Object Modularization cites this paper.

NavComposer: Composing Language Instructions for Navigation Trajectories through Action-Scene-Object Modularization InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:26:40.342191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:26:40.342191Z digest=sha256:b065eaf621b43fbabc5e5bb2a809cd581be1fdec2857bbc1db0721d3a3d349ec

Observation 9e461aa4-bd38-4e8b-a4bb-e5313aa4cb95 · inbound

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning cites this paper.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.583224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.583224Z digest=sha256:0c5225ad47ec1f00abed70a7781db00e6fbea3760acd473bb8d09d25a299f346

Observation 537ae9d2-8b8b-4ee2-87fc-d2868ca37de7 · inbound

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth cites this paper.

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T05:15:43.265563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:15:43.265563Z digest=sha256:5c4c09a141280ab794b4f5472c2097f6206b679c700fd24154666073720a2ecc

Observation 0e018751-258d-4358-8242-05dba5a50bca · inbound

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models cites this paper.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.444811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.444811Z digest=sha256:aa4cb4401ede8c381d050afd5ae58f67cc51a7d6b1c87d4372cfd253841ff32d

Observation 2ac3d27f-7c44-4e0a-be78-6d8ba40ee92e · inbound

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering cites this paper.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:41.523576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:41.523576Z digest=sha256:53749a30057fedefa875ef50a5bc1b646530e1317eccaa5172c179d82726bccc

Observation b5076d94-07c6-4475-8fb1-31568095960e · inbound

TennisTV: Do Multimodal Large Language Models Understand Tennis Rallies? cites this paper.

TennisTV: Do Multimodal Large Language Models Understand Tennis Rallies? InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-18T16:41:37.834324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T16:40:16.630602Z digest=sha256:d4cb093fdddb3de313de3b5229973e4c840baa051182896eeb0dc5289093acd8

Observation 6c2e6875-f1a1-4088-bc1c-21e4d053fcd5 · inbound

Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents cites this paper.

Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-18T12:41:22.584480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-18T12:40:19.544260Z digest=sha256:68cc5579aa36cbb81febd056e2f78c233ff89e1d1831ec2c1ee3cc5da3d87206

Observation af50635d-a1d9-4977-9180-3439cbf037d9 · inbound

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs cites this paper.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:58.011777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:58.011777Z digest=sha256:9980364e2f69be3ab14f591cb46187036b054b58765f28534bf6758871410788

Observation 5b064333-15cc-4f7f-9284-4c40a6317838 · inbound

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding cites this paper.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:07.235364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:07.235364Z digest=sha256:e79bfb08f31062b6c817ad13d21022b2f6d348b26bd5e62c57d62e8b30e1adce

Observation 125e64e3-7bd1-4b90-98ff-66ee8ca3c617 · inbound

OneThinker: All-in-one Reasoning Model for Image and Video cites this paper.

OneThinker: All-in-one Reasoning Model for Image and Video InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.835573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:09:39.820651Z digest=sha256:3e37e514fc121fb4e515ec944cf93203b475568882d83bfb4906b701666485a8

Observation f969c771-3294-4789-b06b-0369962a0a19 · inbound

Adapting MLLMs for Nuanced Video Retrieval cites this paper.

Adapting MLLMs for Nuanced Video Retrieval InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.835573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T22:20:09.051957Z digest=sha256:ef584aabb57b756f4d544200b83d3bec14707e1429a9478806a49455903f9016

Observation 511079d9-318b-46a3-8e6e-da53b0d1ba2b · inbound

$M^3-Verse$: A "Spot the Difference" Challenge for Large Multimodal Models cites this paper.

$M^3-Verse$: A "Spot the Difference" Challenge for Large Multimodal Models InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T14:57:29.163851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:57:29.163851Z digest=sha256:bcf6ded02fd1b11e87ad2cd71c6362e1c8c63562817b797fd16724be610e898f

Observation 8eca61ed-2e08-4c71-9a6e-711a3e782706 · inbound

Streaming Video Instruction Tuning cites this paper.

Streaming Video Instruction Tuning InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.835573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T19:44:11.032898Z digest=sha256:f906c170737884764ddc244247fde469e5eef32eee45d2f941eb2ef326b5c022

Observation b5c2d4b6-449e-48e8-ba9b-73b57c33bf8e · inbound

CoVR-R:Reason-Aware Composed Video Retrieval cites this paper.

CoVR-R:Reason-Aware Composed Video Retrieval InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-13T21:37:55.887477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T21:37:55.887477Z digest=sha256:c27d175ee5854510bb680dd8175fb316c1dfe4015969d9fe58da2d54d8abbe0b

Observation 0eb35f2e-c345-452b-ac95-bdcc3ba2d825 · inbound

Diagnosing Long-Video Quantitative Reasoning in Multimodal LLMs via Enumeration and Counting cites this paper.

Diagnosing Long-Video Quantitative Reasoning in Multimodal LLMs via Enumeration and Counting InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-13T15:29:31.834566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:29:31.834566Z digest=sha256:111bbcdb487c0bcef01a9f95b984bc3dec561c9dd4a1da7845b671529f402070

Observation 3a8b1fb9-ff9f-42df-a6b5-7d0c193524d5 · inbound

Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning cites this paper.

Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.835573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T19:46:16.975267Z digest=sha256:15a2a53a78fc232a9784ec138d6743fec30cb46f5778e8f7fc415a46b7a8c0c1

Observation 9927fd3a-fbcf-4b98-8340-fe8b16e5cc32 · inbound

InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding cites this paper.

InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.835573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T18:06:33.139310Z digest=sha256:1d61367ca7402eba54fadf9607779fe28b84702cb3ea50dbc86ee17188a64967

Observation d75ee4db-3bfd-4077-ada7-08f5479605c9 · inbound

How Should Video LLMs Output Time? An Analysis of Efficient Temporal Grounding Paradigms cites this paper.

How Should Video LLMs Output Time? An Analysis of Efficient Temporal Grounding Paradigms InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.835573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T16:50:48.559703Z digest=sha256:1816a287e6d139a696b2e1c901ca91c50953221959a5a8ef0726070db9daf933

Observation ba20b7fd-ddc6-422c-b8ae-4b189c29cccf · inbound

Decoupled Similarity for Task-Aware Token Pruning in Large Vision-Language Models cites this paper.

Decoupled Similarity for Task-Aware Token Pruning in Large Vision-Language Models InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.835573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T15:15:04.261855Z digest=sha256:e9809915426831df84780f5451b2ffdad808194ac34e11cb7567b35de625c8c9

Observation bc29fe2b-741c-42aa-b69d-71c896e9466f · inbound

All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding cites this paper.

All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.835573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T15:26:55.369840Z digest=sha256:58440c7a9a7c3f46f17005e6dfbe19e174cd028c42dc90867aacfc59162a3e27

Observation bc44a8a6-a9d4-4297-a4aa-4a1f6216f892 · inbound

Grounding Video Reasoning in Physical Signals cites this paper.

Grounding Video Reasoning in Physical Signals InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.835573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-09T22:29:21.240010Z digest=sha256:9e8cea1373914b62e8a6cdb1fbbd0bd7def8141eaf1a21b03f3a14db836177b7

Observation fbce2764-2f30-44e6-9315-2647701224ee · inbound

High-Speed Vision Improves Zero-Shot Semantic Understanding of Human Actions cites this paper.

High-Speed Vision Improves Zero-Shot Semantic Understanding of Human Actions InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.835573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-09T19:49:57.016399Z digest=sha256:c929a96214ca7aebc81bb93fd1bebbc9aa95e5600d836456a9e4f7d52cb592a0

Observation c37adac5-757e-451d-8091-6c3842fc4730 · inbound

From Priors to Perception: Grounding Video-LLMs in Physical Reality cites this paper.

From Priors to Perception: Grounding Video-LLMs in Physical Reality InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.835573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T17:41:23.233366Z digest=sha256:92e776091e7de2779df684fb29d1cbc048ccc68f54b3e7f962fab5cc42055b1c

Observation 923b11be-4673-4dcc-ab52-11d504621cea · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.835573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T14:06:27.953376Z digest=sha256:9811d92096c57993e52d96b285275abc774fd4c5c28555c6b17f395e2dd8ec67

Observation cf8b4a35-bc2a-4db0-b47f-2e8b5783d3fa · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.835573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T01:49:41.654207Z digest=sha256:27e9105c300abccf7dc7828e9b490186fc102d68d0644b6de6f62ef4c0c2f83d

Observation 212bc8f4-2e67-4968-b456-18f3da2e09b4 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.835573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:35:59.553683Z digest=sha256:70fc12357803a62cc0f831356b4c4558ad716ec94b17ba88bd6d93807f613a82

Observation 35b43260-bc1a-4853-8630-7f139cb4eca7 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:10:23.843719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-25T06:08:19.956833Z digest=sha256:0843dfff18e53170a5df647f0616bd53b96be80a2b75b4ebea56152e44e42a60

Observation f258556c-d103-4168-afd5-85ef3680fb2f · inbound

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding cites this paper.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.835573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:2e1cfe96ad9bdedc46f2bc6e88230299800ec9a12b4df403caf44f0465f4738a

Observation b04166d4-610d-4ffc-9b28-986a6891da3d · inbound

EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding cites this paper.

EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 105

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T02:52:20.835573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T04:37:46.090777Z digest=sha256:66bc96cee8fc094625bd951036c2b82549bc73dfd0b3443e50d6d365c1a90ccf

Observation b81a58f8-962f-4907-9aa2-117c800457fc · inbound

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models cites this paper.

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.835573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:13:21.487431Z digest=sha256:3ee2805ccbc227be851d4ffb3ec56aa6e59ebffb0955856a79243b51d24971f7

Observation 0c3eca2b-ff6e-4ab1-810b-149a51feefe4 · inbound

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models cites this paper.

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.835573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T06:53:42.726350Z digest=sha256:4834b912653f5b0719c8e12064ebf3641617e4607c33464796dadd7d064637b4

Observation ede6c788-8b9f-49a7-8cbd-e05aa70ac9c4 · inbound

CoRDS: Coreset-based Representative and Diverse Selection for Streaming Video Understanding cites this paper.

CoRDS: Coreset-based Representative and Diverse Selection for Streaming Video Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.835573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T02:31:19.662372Z digest=sha256:9bd8823908e550c96154b35d0a6c1d937ea9527aa9d5f66116ddb5e827a90d48

Observation 58ae6413-6598-4b95-937b-02f8a737aa6d · inbound

Video-Zero: Self-Evolution Video Understanding cites this paper.

Video-Zero: Self-Evolution Video Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:35:04.405551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T21:32:16.939563Z digest=sha256:fff036cffa60e604ac791e8711524a2083db00d2465f9dbc1fcb2dd01c44397b

Observation fef7ad78-f74b-42a8-9b45-2ff7f068697c · inbound

LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs cites this paper.

LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-20T14:58:24.740744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T14:54:12.264474Z digest=sha256:958cc9251688cdfb44ba79822678692c04dfbebd6e17fe8a1bb5df15d8b4e43b

Observation 83ae1637-3a12-4761-a6c8-4f0e0de9c6c0 · inbound

LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs cites this paper.

LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:55:47.286739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T19:45:01.311800Z digest=sha256:2864110f870312bf34472204865bfd739049d8b16abc35287f76e09b9cc6cc3d

Observation 3face163-f13a-4359-8295-62fd8ad61842 · inbound

OProver: A Unified Framework for Agentic Formal Theorem Proving cites this paper.

OProver: A Unified Framework for Agentic Formal Theorem Proving InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T14:48:23.444378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-20T14:43:46.517807Z digest=sha256:3c85b3de92bd126b8c973b6d6b5440dd3fc35f4301b513eb8d7deff56cd9630f

Observation bf36e788-09c0-43e7-8528-2b51fd61e1ed · inbound

MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering cites this paper.

MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-22T07:21:13.185135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T07:16:33.817183Z digest=sha256:a68b7ccfeb02b50e234894ab099888d0f46a1ada412f5f994ebb2ce2050c6bf2

Observation eb83c77e-b09d-4845-80da-8151f417c158 · inbound

IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams cites this paper.

IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:33:51.057493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T18:24:57.881644Z digest=sha256:4efd757c6c2288b51e7632ad98e9fcda7f02b384501a9299fc3ca9984db70c9a

Observation bb48f86c-314b-441a-9a64-3c18019f3738 · inbound

VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer cites this paper.

VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:03:25.898452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T13:01:42.880738Z digest=sha256:16731056ef459b35c65fcb05c4d42bf1f7d5c37e15e8a003f5381f572a6c88ca

Observation 670c8e33-42ac-4af3-9923-6d3663bc5fc1 · inbound

Masked Diffusion Vision-Language Models for Temporal Action Localization cites this paper.

Masked Diffusion Vision-Language Models for Temporal Action Localization InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:03:13.693211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T07:56:52.524815Z digest=sha256:8ee3ea44148ab38bc1e69a98c5a4e1f5976479a0ed26ec1a4ccd6e15200066f0

Observation 6e7de1fb-0a2f-4b9c-a6a3-039326139e56 · inbound

ViCuR: Visual Cues as Recoverable Privilege for Multimodal On-Policy Distillation cites this paper.

ViCuR: Visual Cues as Recoverable Privilege for Multimodal On-Policy Distillation InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:36:56.986840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T01:59:15.154873Z digest=sha256:41f85ad63c14d0885237dee13eeae529f455fa02866411bb567cb7477cfceb39

Observation 0f97e9b7-e167-413e-886c-b5181e0c233a · inbound

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning cites this paper.

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-07-02T07:56:47.318780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T06:33:32.090913Z digest=sha256:21f10be42f0b7b163b6feb5671f3c6f514f9be5a1c80f9e240f2fc3d160e85b3

Observation b62462ba-6f47-43e3-948c-df820a424819 · inbound

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding cites this paper.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 97

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:07:12.785043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:2487e564d6c2c8b0905e09cd5cb575eb133d9a257fb4a788c7e5adebbf27ad03

Observation 76fafa33-ac4f-443b-845c-5b6137108df7 · inbound

CoCoSI: Collaborative Cognitive Map Construction for Spatial Intelligence cites this paper.

CoCoSI: Collaborative Cognitive Map Construction for Spatial Intelligence InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-03T03:57:38.936566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T14:17:08.713863Z digest=sha256:917ed91904ed2267d169614df836ad7dd2aeeb610c2aaf43a419a74a9f82a405

Observation 8ae19305-c3cc-420b-abc8-6aa4de49f69c · inbound

MultiToP: Learning to Patch Visual Tokens to Mitigate Hallucinations in Video Large Multimodal Models cites this paper.

MultiToP: Learning to Patch Visual Tokens to Mitigate Hallucinations in Video Large Multimodal Models InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:27:56.120244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T10:02:58.341050Z digest=sha256:a5675ea008486faeb387875df2494f4c99bbb28879f4b4c6bcdffebea8587f60

Observation 067fdded-660c-4c46-95ad-f2f2b80fb3ae · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:48:03.182059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:d324eb245359d960589a082dc54494001d4a5c57412308efae1d771966022732

Observation 40caaa14-d8ec-4074-a426-ddab0dadf133 · inbound

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams cites this paper.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 93

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:38:56.157736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:119e20c9bc9c525f4b5c86d62a06f78eed762cf4bcb79c5730f4ffd1e2fb7500

Observation c15dff4b-100b-4359-9947-c668c84245f7 · inbound

Latent Visual Diffusion Reasoning with Monte Carlo Tree Search cites this paper.

Latent Visual Diffusion Reasoning with Monte Carlo Tree Search InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:03:52.403983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T04:55:37.555809Z digest=sha256:cd9d96fe2c45763abd9f9e652a344cff862bfcbadf9e98a13e1dc7f04bbc5fde

Observation 74cbdf07-f726-49e6-bbac-7ad254b47189 · inbound

MotionAtlas: Detailed Region Captioning for Motion-Centric Videos cites this paper.

MotionAtlas: Detailed Region Captioning for Motion-Centric Videos InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:24:21.122131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T07:21:46.783970Z digest=sha256:4aa8b7bc48b3cead74a60f45581af063fd8e591ad11064de183d7950c614162a

Observation 739ba1f6-1d06-49be-a72b-87a3066511ab · inbound

Learning to Deny: Action Denial in Multimodal Large Language Models cites this paper.

Learning to Deny: Action Denial in Multimodal Large Language Models InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 72

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T09:45:39.577296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-01T06:21:09.996386Z digest=sha256:843cc9135330ae09c7136d7b35d4a3bfc8eafef500c58ddd933e2a01d058a56d

Observation f29c3843-0f5c-4a33-83c4-b1b4a5e870e8 · inbound

Bridging Video Understanding and Generation in a Unified Framework cites this paper.

Bridging Video Understanding and Generation in a Unified Framework InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:05:40.298995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-01T05:57:54.653504Z digest=sha256:e86ddfd6775607f39033329fabb708eda840b0ebf47454cbe638304f011b8066

Observation 5add77d4-1b8a-497d-908a-e85aa98bd5ce · inbound

LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension cites this paper.

LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T15:48:35.012694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-03T15:41:03.548799Z digest=sha256:4a983460d0cd917bef48655d55bcec84e3f44b655213408b7db0639711252835

Observation 5e8d3556-65e6-4e26-8f18-241baa4085a2 · inbound

Natural Language Camera Movement Understanding cites this paper.

Natural Language Camera Movement Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-12T05:16:19.750976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:16:19.750976Z digest=sha256:5f85255dfb331b7014edde361b2e00f44fee98bd4f09ae9ae21a87eb4f1c2850

Observation e882bf37-90d4-499d-b6eb-528ea4968e99 · inbound

Probing Identity-Specific Motion Signatures: A Controlled Diagnostic Study cites this paper.

Probing Identity-Specific Motion Signatures: A Controlled Diagnostic Study InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T01:03:30.708082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:03:30.708082Z digest=sha256:2bcedd922a30bd0b16cf8d5f3a9a643224fda239cac86f32509be2acf8c7cba6

Observation d0b6ef65-6f2c-4d51-b3ef-f0f043eb7e6f · inbound

DynTrace: Tracking Dynamic Object Evidence for 4D Spatio-Temporal Reasoning in MLLMs cites this paper.

DynTrace: Tracking Dynamic Object Evidence for 4D Spatio-Temporal Reasoning in MLLMs InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T06:34:30.007079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:34:30.007079Z digest=sha256:2d1b4aa75cb14d445010cb932235fc9b2c3b7c3b88565c062ac6c2a4e3315c13

Observation dbcfc400-3eab-4288-add5-b74a12a64c15 · inbound

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding cites this paper.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:40.385436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:40.385436Z digest=sha256:d60bd7dfd11e04c24ff8b9dc41c34cab1b0f461365a0b0d972e4817ab1c5933e

Observation aa7e5358-e336-43c6-930b-f843dab8d307 · inbound

MAC 2026: Advancing Micro-Action Analysis Towards Fine-Grained Understanding cites this paper.

MAC 2026: Advancing Micro-Action Analysis Towards Fine-Grained Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T07:40:00.692260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:40:00.692260Z digest=sha256:40207e39b58f770a9f1ec936bbf9d51f7734bf64d2133c7088ded90b6435afac

Observation a282df78-dfdb-4c02-a1fa-68bff1bba73f · inbound

Continual Video-MLLM Adaptation over Evolving Domains cites this paper.

Continual Video-MLLM Adaptation over Evolving Domains InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T14:37:47.442834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:37:47.442834Z digest=sha256:7b2577eddfaf422a2d905e448e618c8e4a0a2780b6a84fb7a6c8cef1ab3041a9

Observation ff3267ed-70a9-484c-bfc9-b1e7b5a9304d · inbound

V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure cites this paper.

V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T08:23:24.495566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:23:24.495566Z digest=sha256:b8090a8dcfcb49f4238c1a8fcfe84d53ffd01a15c1d751c40e236e210e6a6cd1

Observation 2f64894f-968d-4087-8589-5ad2eec00747 · inbound

TimePLE: Rethinking Temporal Representation for Video Temporal Grounding cites this paper.

TimePLE: Rethinking Temporal Representation for Video Temporal Grounding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-31T23:32:54.425149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:32:54.425149Z digest=sha256:1e92673047ec3a973d07657d9776c65dbf3f56caa80c4b977fa0eaa382e8530a

Observation f2e2663e-7d84-4879-84ce-54ed7852962f · inbound

FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding cites this paper.

FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T02:59:25.397951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:59:25.397951Z digest=sha256:1237ea2ac1148d4069b0bccc0026105bc8a6cc9b8c83c179e6cf5f8970e5be92

Observation f0c17133-5f7a-4704-b073-b60832060a27 · inbound

FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding cites this paper.

FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T01:46:11.788344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:46:11.788344Z digest=sha256:ad1b41ed82564ede9a5f40bb1c8edc0f3282466c608ca4cb3d4d4d489af9af9b

Observation 32732981-9dac-4e5e-b682-39fa7b19ab3a · inbound

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding cites this paper.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:32.975156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:32.975156Z digest=sha256:8b9f18b5ff9e812b7c67ad08bc4649471b058ac08e3425942d45b79938ca6fc1

Observation ac560be2-623d-42f1-83e5-80d01c5921d6 · inbound

Think in Sets for Streaming Video Token Compression cites this paper.

Think in Sets for Streaming Video Token Compression InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T00:31:40.530514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:31:40.530514Z digest=sha256:188be750be97ccb01fcfa26cf8149850a1bce9afcd3857e5817ae129675fafec