Pith. sign in

Paper Citation Record · LEDGER

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

As of 22 August 2026, this Paper Citation Record lists 100 of 156 outbound references and 18 inbound Pith citation observations for arXiv:2412.09596.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.09596 v1

Coverage vector

measured 100 of 156 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:56:06.396297Z

measured 118 of 118 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:31:56.590340Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:19:29.961741Z

Reference resolution

100 of 156 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c744f3c7-613e-4012-927b-bda6284c14bc · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:05.967600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:05.967600Z digest=sha256:f0adad0df543032c8b24c527c6a8901d34ec75df47aefabb80e5d328498aad18

Observation 433c683d-8d92-4f14-85ee-792f70a05c4a · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Common Voice: A Massively-Multilingual Speech Corpus

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:05.972313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:05.972313Z digest=sha256:57091163bf6ee94b5695c7707a6d84be757b7d3a5877b9922c66b027f71947aa

Observation 0f4e8061-edb1-45d5-97d7-4f46078fbfbe · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:05.976903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:05.976903Z digest=sha256:036f42976be880f4bcd8d205c91f9e1507fa20f50fd0bb8c5015302604cf1069

Observation 98fc5988-08af-47c5-ac23-1052fa794e5f · outbound

This paper cites Qwen-VL: A frontier large vision-language model with versatile abilities.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Qwen-VL: A frontier large vision-language model with versatile abilities

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:05.981367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:05.981367Z digest=sha256:8aaaad5d12e096b871b42cdea22e98ab93b753f7706c47f8a0f4c643a0068e2d

Observation 5c8e06a1-69d8-44c5-92c7-4c59c54d1563 · outbound

This paper cites Baichuan 2: Open large-scale language models.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Baichuan 2: Open large-scale language models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:05.985702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:05.985702Z digest=sha256:6ac9c6b6765034633fc2eec8c25be7176aae4fb20aa45a6a2c708584a2a8799e

Observation 07517be1-07f1-4da5-adc8-f83d54209323 · outbound

This paper cites Introducing our multimodal models, 2023.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Introducing our multimodal models, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:05.989916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:05.989916Z digest=sha256:bed998c8d7dbca4847744f87e30130f645ca4afaad696cc2f9c2fd01d276a56c

Observation 6221e4a6-15d4-485a-bbda-504726b387a0 · outbound

This paper cites Language models are few-shot learners.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Language models are few-shot learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:05.994464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:05.994464Z digest=sha256:67ad5fd2a662c023820e9a6931c5c1f8511e491acb46e4fd932537e193d81a57

Observation 33144b44-ea54-4a49-9432-fc3d8d4130e5 · outbound

This paper cites Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:05.999221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:05.999221Z digest=sha256:fb030950e394da2f8fdf870e4cf141cf74523e3f5ef1affbe4786e3842a9bc25

Observation b98253d3-693c-4aca-84f2-925a25f511c0 · outbound

This paper cites InternLM2 Technical Report.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions InternLM2 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.003409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.003409Z digest=sha256:e719eb7cf72e6d500e17a361e1f4ec5e85f187a12d8f68a57f4eaab68a45a819

Observation 3a64ae17-e955-46bb-8cc6-57455852f026 · outbound

This paper cites DualFocus: Integrating Macro and Micro Perspectives in Multi-modal Large Language Models.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions DualFocus: Integrating Macro and Micro Perspectives in Multi-modal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.007679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.007679Z digest=sha256:bbe44c288bbc7f7e08d9361b1747a5fe6d0e8cd924190df1db193fa3f5b2b315

Observation 30b85e59-2d2d-4431-a9e5-654e150edb9e · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.012180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.012180Z digest=sha256:45cf003812eb12050e9c5a349e5e6ad7f368303fd42ec46a6c9913cfc9497467

Observation 7bc294ea-2e32-44f6-88da-4378cbd2281f · outbound

This paper cites Videollm-online: Online video large language model for streaming video, 2024.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Videollm-online: Online video large language model for streaming video, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.016693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.016693Z digest=sha256:3550175291df9178e80b1645c9c58c650f71c6c0aa65a1d01c80a94d8a926b47

Observation 780a2c6b-aba3-47fd-a748-1eb4aa5c67bf · outbound

This paper cites EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.020705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.020705Z digest=sha256:7e57a238f573564446caa348626db5ebdf76733d3fd9b8509db85ece025b4de3

Observation 090ce6d6-c9d9-4002-951d-3a79a94709ed · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.025445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.025445Z digest=sha256:364a545c35705fdafd26633641e23f06686fc09d405343700bdfa528846d9e5f

Observation 6f00bb27-f81e-463f-84e7-8ec8086f4590 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.030029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.030029Z digest=sha256:9896d70e4d1608851b17d36fb31dfd388a8c83a2f1eaf2a77f1694fda10e47fc

Observation cb2ee4e5-48e0-4649-9b6f-76ad751b43f8 · outbound

This paper cites Timemarker: A versatile video-llm for long and short video understanding with superior temporal localiza- tion ability, 2024.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Timemarker: A versatile video-llm for long and short video understanding with superior temporal localiza- tion ability, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.034260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.034260Z digest=sha256:3e1d26d415963bf3aa178a86db258af472ed73a5bf98fc1b143ff9e85c2e1e34

Observation 61ee2236-00d0-4f6b-9844-27ba13cc289e · outbound

This paper cites PaLI-X: On Scaling up a Multilingual Vision and Language Model.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions PaLI-X: On Scaling up a Multilingual Vision and Language Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.038559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.038559Z digest=sha256:97e7c5cdb1555affd9d013a3e20f2962e60e3c9d3cac3b0cba55d38a21721a2a

Observation acedf2d3-392f-40b6-bcee-b0d4bfa8d97b · outbound

This paper cites Pali-3 vision language models: Smaller, faster, stronger, 2023.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Pali-3 vision language models: Smaller, faster, stronger, 2023

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.042911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.042911Z digest=sha256:d8e9b3b11831aac9752532c19783295665170dd8ab1c0e6fb873522ddf9f5886

Observation 0c298f8e-bdf7-4563-876b-5029fb33866e · outbound

This paper cites Pali: A jointly-scaled multilingual language-image model, 2023.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Pali: A jointly-scaled multilingual language-image model, 2023

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.046874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.046874Z digest=sha256:feb5254d8ede7207debd306d3afc75e09e9a180c0db60f9fd95becc0fc441e99

Observation 2f652eb5-bfd8-4aed-b378-af6ac8ff9e63 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.051011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.051011Z digest=sha256:bb0a391e99f6574487137ab9b772ce055703a195b97c7769454dd0ed40706987

Observation 6212b70a-fe1a-46ad-9e61-0815d76825a7 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.055542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.055542Z digest=sha256:f5470b69dda23fc59837ba520f5376fd27568ff171ca18bbf6e09bb2bc8f7c7d

Observation 034adb00-2a66-4d98-9cc3-66c647eb824b · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites,.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.059946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.059946Z digest=sha256:d8bdfea0ee692e1cba8740ddf3f4a3989b470d33536a2c9ab75ce13473f4c585

Observation f569de39-f504-4371-9747-4f3991c3c63b · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.064527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.064527Z digest=sha256:4f628cbf63a25d20c1962e14b3bdce6d261cb9a6e4c390ac546686fd56e3c148

Observation 6c2c3b17-0ef8-48d2-ad46-163a6c64e327 · outbound

This paper cites Palm: Scaling language modeling with pathways.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Palm: Scaling language modeling with pathways

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.068926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.068926Z digest=sha256:cb3e78dfada4be52fabc94c2a32405ad9e042fb38b5274b834a500dac2756988

Observation 4e82ad13-742a-48d1-9434-5ba479ec6862 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.073206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.073206Z digest=sha256:1d392094501d2f44652e2e14c2cfe5dd1a800eb3ba24d920135bbeab3275c6bf

Observation f61d49ce-c6f5-4c8c-bc14-5fdd891ac026 · outbound

This paper cites Qwen2-Audio Technical Report.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Qwen2-Audio Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.077568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.077568Z digest=sha256:b01c5929baef3ffd3d02efe18f69aa37b07a04f8a5c770fe8ed94592a53a5350

Observation afbb0b0e-74d8-4c0e-8c6b-41523560362e · outbound

This paper cites The CHiME-8 DASR Challenge for Generalizable and Array Agnostic Distant Automatic Speech Recognition and Diarization.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions The CHiME-8 DASR Challenge for Generalizable and Array Agnostic Distant Automatic Speech Recognition and Diarization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.081913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.081913Z digest=sha256:a71cfeb7bf057ee74afcb68963669cb55e9bf6258b401555bfb5da63f699758f

Observation 1cb52534-6eeb-4a53-a6b5-b7daa14d7426 · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.086250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.086250Z digest=sha256:ce69e98e578e412f24d3e2d2217f6b22a3c8b03d085823d6276f8222801372a7

Observation f5074f80-6bff-4f65-b3da-415a02bedbfe · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.090532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.090532Z digest=sha256:db22c85d2ce0a79444fc5625f89d53ac24186c147ad9ff2ee8a4dfe19156f7b2

Observation ee520e93-d1f3-4aee-9f7c-4a245cd46a58 · outbound

This paper cites InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.095069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.095069Z digest=sha256:10edd4ee355a71d7909418767949070e86d6455a2affb327edf3919fd95f09c4

Observation 84b67cb4-3541-4f43-85c7-683cdc86f827 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions PaLM-E: An Embodied Multimodal Language Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.099727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.099727Z digest=sha256:30b00f51e83b96a697e12bb685fa527ba9614caceb74b38279726d000701b16a

Observation 05591220-12b2-4031-8c5c-8a4b5c0294c8 · outbound

This paper cites ActivityNet: A large-scale video benchmark for human activity understanding.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions ActivityNet: A large-scale video benchmark for human activity understanding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.104175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.104175Z digest=sha256:da9e97c3cb87c1ff44309aad057599e20efcb61d01d3cf5f8c7402e264dfeea4

Observation 54b48044-aca8-490e-834f-1607bc646cd8 · outbound

This paper cites Videoagent: A memory-augmented mul- timodal agent for video understanding, 2024.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Videoagent: A memory-augmented mul- timodal agent for video understanding, 2024

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.108328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.108328Z digest=sha256:3e586a763133d3cb12c2d32734bc082b2fa325b93e89c0806dba666f7b71c053

Observation c530d3e2-5ced-4dde-9fcb-81f40a943c60 · outbound

This paper cites MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.112524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.112524Z digest=sha256:6fd1c5242550e3d24e1e193ef921009a1394ccd8ddf77d57b79c0dcb1a5080ee

Observation 7af27ba0-f27c-4d72-8557-73e62e5183fb · outbound

This paper cites FSD50K: An Open Dataset of Human-Labeled Sound Events.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions FSD50K: An Open Dataset of Human-Labeled Sound Events

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.117111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.117111Z digest=sha256:87901320aa0a9a515bbf7ffb94eadc7478f6c48d977368318ac37e8b6f26aca6

Observation b40d57bb-ec1e-4364-b7a9-13a64f2269e7 · outbound

This paper cites A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.121595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.121595Z digest=sha256:79e47b0ba5c44ec72294efe2160d9232fbac19b9c4c2398519e1b63fd7c07f56

Observation 7a76eeb1-e83b-453d-84f6-2df57cff5257 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.125920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.125920Z digest=sha256:5b42cceabbc3adbc54e2ee959e381be1096358876cc1359ef47a8439dfd10211

Observation df22b326-c263-45d5-bb1e-49cd100fcf04 · outbound

This paper cites Vita: Towards open-source interactive omni multimodal llm, 2024.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Vita: Towards open-source interactive omni multimodal llm, 2024

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.130244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.130244Z digest=sha256:e3a275842730a328a54dab31f6c1051cfef1ccc81ac96a495802ecfe36238390

Observation 00424ac6-d2ba-43b3-97f1-b527e0d1f771 · outbound

This paper cites AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.134482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.134482Z digest=sha256:678c89912b4df95929c56b95ea83f4e12066d8ee62923ba640b1e44665e5f2f5

Observation 485d39c7-7c65-4899-8652-babf6d6eec75 · outbound

This paper cites Funasr: A fundamental end-to-end speech recognition toolkit.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Funasr: A fundamental end-to-end speech recognition toolkit

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.138923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.138923Z digest=sha256:a16d43fef1ac13d856998f893cdc7e563f1a9cc0459ce379cff55cce931a42d8

Observation 47cbe73f-7057-44cc-ba56-b0d1aa1327fc · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Ego4d: Around the world in 3,000 hours of egocentric video

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.143412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.143412Z digest=sha256:faa6c1d1426e9a495ca3339c0ba20c39661e3557329400094b4cba9cfe50964c

Observation be6177a0-c5b6-44bd-aebf-05d551f1470c · outbound

This paper cites Onellm: One framework to align all modalities with language, 2023.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Onellm: One framework to align all modalities with language, 2023

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.147720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.147720Z digest=sha256:fb86cf9f78ba2c23ce14a2773820b0bca3b1ebbf978c3df3316035af236219aa

Observation 938bcdb0-1865-49dc-8ce8-f2956f94d0db · outbound

This paper cites MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.151966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.151966Z digest=sha256:b9dc060046be4c29bbf78c744e3cfbce14bda450d9ccd9ab2cad0ba6421a843c

Observation e38e553c-8288-480d-a3ff-1a0cb72be693 · outbound

This paper cites From Image to Video, what do we need in multimodal LLMs?.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions From Image to Video, what do we need in multimodal LLMs?

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.156584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.156584Z digest=sha256:5c00f5c593f461f21d0d97eed5d1a9602eb281e284659aa1bc599109d05892a8

Observation 6f106160-c0f9-4cd2-b1de-a61fa3a8525c · outbound

This paper cites Video ReCap: Recursive Captioning of Hour-Long Videos.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Video ReCap: Recursive Captioning of Hour-Long Videos

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.160859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.160859Z digest=sha256:7fd3ed3c82fe41128adbd5bbc668b22d8c2e72996ce38136b76b7d9713859cde

Observation 59568157-554a-406f-9da7-05a407233827 · outbound

This paper cites Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, et al.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, et al

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.165340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.165340Z digest=sha256:01154ae77b88e08ddf501f778fab2239e76f1cd1e3245a22813eee298a8fb9d0

Observation dac03a0f-6b47-4df2-a572-f1a68f5e6f17 · outbound

This paper cites Mixtral of Experts.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Mixtral of Experts

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.169616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.169616Z digest=sha256:0aee4fdf267c00eb4ea96c90723712a3972440f28bdab82f8c48b968471ee902

Observation 71ce98f5-8e6d-426d-80ab-599c88a6d78b · outbound

This paper cites Mantis: Interleaved multi- image instruction tuning, 2024.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Mantis: Interleaved multi- image instruction tuning, 2024

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.173760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.173760Z digest=sha256:33bb031323a54d632bfcac07e9b9e583f7692bf8269e58ff8fd30cafdfc55f59

Observation 23f79b5f-948f-49b4-be34-ed9cf150219f · outbound

This paper cites Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.177770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.177770Z digest=sha256:a3229beac30f69efc7c1563a18795b1ba10ac2b513cc9e16290fc5d4ac5bfc10

Observation 776b7648-a119-449b-8d8f-2d49fa3dc781 · outbound

This paper cites Language Repository for Long Video Understanding.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Language Repository for Long Video Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.181876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.181876Z digest=sha256:7529d9138935f9a15611c82e01992f93ddfd0f49c9b7e1429e7d48491879606c

Observation 39c479a1-4dd8-4cdc-95c0-8885ea7b8a1f · outbound

This paper cites Scaling Laws for Neural Language Models.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Scaling Laws for Neural Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.186924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.186924Z digest=sha256:b2cb9d982e3023fee6ae9c789b384185a10939022c69538204781dc37d1d9516

Observation af49cee5-8323-4dd3-8938-380d20bbf0d2 · outbound

This paper cites An image grid can be worth a video: Zero-shot video question answering using a vlm, 2024.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions An image grid can be worth a video: Zero-shot video question answering using a vlm, 2024

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.190904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.190904Z digest=sha256:42504fa03edfd7ea47b6507935a2ec906a912620bfa755267673d63f97774a5f

Observation b6dd74f8-a085-45d3-8b99-8296c79fc8d3 · outbound

This paper cites Audio set classification with attention model: A probabilistic perspective.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Audio set classification with attention model: A probabilistic perspective

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.195228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.195228Z digest=sha256:dc29a96d1ef21762942d1b96214d1393946c90d69421f983aa640fac02c066de

Observation 22e4558d-670b-4e36-82a7-6124f9f2c2cd · outbound

This paper cites The open images dataset v4.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions The open images dataset v4

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.199222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.199222Z digest=sha256:9c5362f0ed93888a9575b9e98021d4622f3cff0d0963a36efb03ecef92cb76d6

Observation 7da56cc2-2180-4ecd-9f8d-bbcaf71cc812 · outbound

This paper cites A path towards autonomous machine intelli- gence version 0.9.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions A path towards autonomous machine intelli- gence version 0.9

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.203499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.203499Z digest=sha256:6c2224c71ed8a2037ab4561f41aa75c93142134fb9479e1d31582b93879bd6ef

Observation c34e72a2-4864-4edd-bf00-21930523295b · outbound

This paper cites Otter: A multi-modal model with in-context instruction tuning.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Otter: A multi-modal model with in-context instruction tuning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.207820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.207820Z digest=sha256:9f6e952357ef1d4b491c740242c620a56b69887543e8a52ca60a245d9014a2c4

Observation b1afa105-ceb0-4888-9091-51d4b00482d6 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions LLaVA-OneVision: Easy Visual Task Transfer

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.211962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.211962Z digest=sha256:1a8d092ab6a822da4da17e269fbc619de5ab8d6f8ca418f8f87427183c31f1d6

Observation 38b506b7-86ed-4a00-8a13-ac6cc65c444e · outbound

This paper cites Aria: An open multimodal native mixture- of-experts model, 2024.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Aria: An open multimodal native mixture- of-experts model, 2024

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.216516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.216516Z digest=sha256:52a2dda8554742feb0ea4b9a749f796e8c3afbe5030638160b766809e33ab3aa

Observation 8ad76a10-5bb5-430a-a9fb-20d3772ac61f · outbound

This paper cites SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.220677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.220677Z digest=sha256:2d05b534cc364fdf234721c312493fb233bbee5ff31da6a1b62573f256319bf1

Observation 31dcb5ac-6e71-45d9-a861-40bf71b16d7c · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions VideoChat: Chat-Centric Video Understanding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.224978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.224978Z digest=sha256:6dcda720e3d253bb8c564a3ba4eb842a08db1ba1829b92a79b9d2f08cb4a5f95

Observation 790c9991-459e-4d10-85d6-6d44328e20d7 · outbound

This paper cites Mvbench: A comprehensive multi- modal video understanding benchmark, 2023.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Mvbench: A comprehensive multi- modal video understanding benchmark, 2023

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.229323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.229323Z digest=sha256:4109461049d8febeb0dc9ba9af7a46c7e649fb2f2fe17de214a84d8cb195e11e

Observation e49de970-4143-4a15-8f23-a5d8efdc613e · outbound

This paper cites Mvbench: A comprehensive multi-modal video under- standing benchmark.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Mvbench: A comprehensive multi-modal video under- standing benchmark

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.233554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.233554Z digest=sha256:0e5e49f19c10ff83866dbfb7d46bd4708c1fe65a5533fe8585223a17e0fc838b

Observation 964cd314-0fc9-4b2f-84cc-0ddab4b10069 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.238363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.238363Z digest=sha256:56485bfd104152cc0feb8617481313be40fd19547072c9508083b470065ff42e

Observation 75e74ca5-1747-484b-b908-478480df5e54 · outbound

This paper cites Ocean-omni: To understand the world with omni-modality,.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Ocean-omni: To understand the world with omni-modality,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.242596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.242596Z digest=sha256:ab9d9b496750905db361b16652d718cb432360317c88118d5c9d97d4a2db9b48

Observation bc4509b7-6fec-4873-bcc1-fda0f4c35220 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Llama-vid: An image is worth 2 tokens in large language models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.246617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.246617Z digest=sha256:90755f053d0b3355368a6566425ada78156db69dd2529fded44b5572edaaca8f

Observation e913b9c3-9757-449d-9d54-55e3442ae1c2 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.250684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.250684Z digest=sha256:3ea3edcd6be14f8977e5882f82971ff5ca3ab562f8aff355a654656dcb3940e7

Observation 9bd4898b-b214-4606-8419-3699604de887 · outbound

This paper cites StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.254787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.254787Z digest=sha256:21b1c781b563a23c1c4410bb73241658444278d64f27ac36e022cd4f81f277f9

Observation f81de65e-2723-438b-990e-d782a044e5b1 · outbound

This paper cites Vila: On pre-training for visual language models, 2024.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Vila: On pre-training for visual language models, 2024

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.259320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.259320Z digest=sha256:6e422d36b319b8affc7386a6acc16a4ac97483fc7702ec2e4dbd21b2020c2cfa

Observation 908aaa5a-ee8e-4526-8777-f8e2b3f0ccd6 · outbound

This paper cites Microsoft coco: Common objects in context.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Microsoft coco: Common objects in context

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.263520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.263520Z digest=sha256:a736b8cf66deef7ddfac98f5f1a38e02d324e2affad2d569bc6617577215241e

Observation 3e9a1db8-0bde-43fc-a99b-502e8b213b33 · outbound

This paper cites SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.267557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.267557Z digest=sha256:0f5bfabad13691fb1f2f84e359fd667984f2bb0b1f6cbfeb2c961e5d5a7b7326

Observation b91696b1-eac5-4dd0-bce6-de6e3dfe1063 · outbound

This paper cites Visual instruction tuning.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Visual instruction tuning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.271959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.271959Z digest=sha256:7a0a85629e55b28a61d9fb25555d4564bc365e7d1338fa01a41b973a0e272588

Observation 9e9f8983-2816-491e-a802-47cb02a7a7e5 · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.276119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.276119Z digest=sha256:eb325f2d5d4273adb93d9bc324b786b182dac56f0c19b1b744e0ac2e367622d3

Observation 471d60f4-d974-49da-b27b-db695c1966d6 · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions TempCompass: Do Video LLMs Really Understand Videos?

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.280632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.280632Z digest=sha256:e36877f8ef51ccb7cf0ca5575129c639e0e01aecaa96ff4f0c9e8072cf7555a2

Observation efa31c04-731a-4b27-809e-0af8506c2415 · outbound

This paper cites A convnet for the 2020s, 2022.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions A convnet for the 2020s, 2022

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.285412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.285412Z digest=sha256:b90f2f195bbe65d3a413037028c66982f0634d3a136b07d4663dabb71126319d

Observation bbe6ba7c-1f0a-4542-9429-521a7d73908e · outbound

This paper cites RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.289494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.289494Z digest=sha256:ea5d8aa946c3df04f28f8bcf0ce4dccdaeabb2c576ddf1b5d34ce472e0d32480

Observation b3052a53-f009-4e02-8f81-069b7ade084c · outbound

This paper cites ASCEND: A Spontaneous Chinese-English Dataset for Code-switching in Multi-turn Conversation.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions ASCEND: A Spontaneous Chinese-English Dataset for Code-switching in Multi-turn Conversation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.293867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.293867Z digest=sha256:a41fa8c9435214884877a925cb89a58aaa9d2b6076392e0549fad2845c3a1f52

Observation ddb41289-f379-443c-a354-ad0be15de7cb · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.298328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.298328Z digest=sha256:b7ca44bc50a235fa196f7c25e4c4de98f174fd3070a8e9ef624a5f393cad2910

Observation c8a474e0-d0b3-410d-8da3-75f94ae18b0f · outbound

This paper cites Language Model Can Listen While Speaking.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Language Model Can Listen While Speaking

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.302551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.302551Z digest=sha256:c794bf5f807e0eee73a49c20bfafc990785d87477581b51ec9b2fbe6b890d4e5

Observation 979a71f5-dd93-49fe-8169-d3df4766ba75 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.306978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.306978Z digest=sha256:b8d0a2786b3201a01ecf4495b75bf8b85557260caaa98e59bd258d6fefcb7f55

Observation facc0219-2a24-4ea2-a415-bd5a42daf1f1 · outbound

This paper cites Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.311193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.311193Z digest=sha256:91a0588cd3c30c14422039b6a8b332bc5da0890989083166ff7722f115197307

Observation 7ad79243-ad23-40e8-a037-e19ba7fe2db6 · outbound

This paper cites an unresolved cited work.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Unresolved cited work

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.315666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.315666Z digest=sha256:703abfc4ec8470298e03225c571cc69e93a34d773a68600017c346246f9c3127

Observation 851f6ad7-4674-4a86-8985-68275265cac7 · outbound

This paper cites Gpt-4 technical report, 2023.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Gpt-4 technical report, 2023

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.319843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.319843Z digest=sha256:203b52820b25774911d5f006575b74082e7d459e6780f93fca6d7a93110c3d4a

Observation 884dbce1-66dc-423d-9e37-64ada99cc10c · outbound

This paper cites Gpt-4v(ision) system card, 2023.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Gpt-4v(ision) system card, 2023

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.323991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.323991Z digest=sha256:99be7fbb7000f118ec4ffa8f2f170f89c22d69d44746850396237f00f32b7f5f

Observation 8e63f14e-b146-4b1b-b86e-a68550d75380 · outbound

This paper cites GPT-4o System Card.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions GPT-4o System Card

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.328088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.328088Z digest=sha256:33330bacc3b69f318cb1d7192d0218da718b502b48819bb08dd2f5641625f84b

Observation 374ef120-9fa9-4ffb-94f7-c06721450ed3 · outbound

This paper cites Dinov2: Learning robust vi- sual features without supervision, 2024.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Dinov2: Learning robust vi- sual features without supervision, 2024

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.332491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.332491Z digest=sha256:9f66112be3da0da9bb6dea0a7b4d5b5667ebeb8de80d53aab64dcc017694e70c

Observation 16806f12-98b9-4869-ab7f-99cd7d1281cc · outbound

This paper cites Training language models to follow instructions with human feed- back.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Training language models to follow instructions with human feed- back

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.336689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.336689Z digest=sha256:fa260ef8a618e694cf904701f2868dda196c9adede818d35a87139ade7d4010a

Observation 239884da-b5b2-4be2-8af8-86b08bd9c539 · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Librispeech: an asr corpus based on public domain audio books

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.341083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.341083Z digest=sha256:46f66d348d041b08268949d1736646c5e5b446bdd2210d8ecce650ec7fdaf1c1

Observation 3eb47b21-5ca2-43e1-9c23-2c984cb84dba · outbound

This paper cites Kosmos-2: Grounding multimodal large language models to the world.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Kosmos-2: Grounding multimodal large language models to the world

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.345145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.345145Z digest=sha256:a509f1c66d23911799cde6ef61e970bd95fd1256fdb0ed03195fc80b70d6cfdc

Observation d41ee99a-730f-42d7-9d9c-b86cb7b87b1f · outbound

This paper cites Streaming long video understanding with large language models, 2024.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Streaming long video understanding with large language models, 2024

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.349190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.349190Z digest=sha256:2915f9520ab37afdbc0036fccfb1dc04a18b5823b69ba63bd2b2d86ddf9274d6

Observation 7d13233f-fab6-41d3-920a-7da4e789612b · outbound

This paper cites Introducing Qwen-7B: Open foundation and human-aligned models (of the state-of-the-arts), 2023.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Introducing Qwen-7B: Open foundation and human-aligned models (of the state-of-the-arts), 2023

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.353385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.353385Z digest=sha256:796d7e13ca475510c03d194f1fa3c05e3c8b99891d849cccf873bbfb6e4ebd79

Observation b8ce8c6d-f0bf-4265-b191-1e90aff6faab · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Learn- ing transferable visual models from natural language super- vision

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.357461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.357461Z digest=sha256:9a8f1f133912a6ed79595a13f7d35cf8c0e0af85ee811bca5abdf1710b77ef52

Observation 0faa449f-d6a8-4cce-b2d3-2f19a1927e6f · outbound

This paper cites Robust speech recognition via large-scale weak supervision, 2022.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Robust speech recognition via large-scale weak supervision, 2022

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.361633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.361633Z digest=sha256:395e73a77b742925f63a6eec89f1a7b251f3dc3b2e21b2be2b15f0bc30234377

Observation c2724e02-687e-444a-a2a1-7b815eecfcad · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Robust speech recognition via large-scale weak supervision

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.365615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.365615Z digest=sha256:96d02b3e65f93506abef85662d3da55630120d00ebb06bf1e2d26c64e12c7beb

Observation f11389cc-a2fc-427d-b7c0-2be58b05813b · outbound

This paper cites Ryoo, Honglu Zhou, Shrikant Kendre, Can Qin, Le Xue, Manli Shu, Silvio Savarese, Ran Xu, Caiming Xiong, and Juan Carlos Niebles.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Ryoo, Honglu Zhou, Shrikant Kendre, Can Qin, Le Xue, Manli Shu, Silvio Savarese, Ran Xu, Caiming Xiong, and Juan Carlos Niebles

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.369775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.369775Z digest=sha256:667ade9cd8d32df93c5a8183d424bd7df5f977f765cfeccda51013807dacee5e

Observation 7afc7d95-6afd-43d9-a757-5e118e823574 · outbound

This paper cites LAION-5B: An open large-scale dataset for training next generation image-text models.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions LAION-5B: An open large-scale dataset for training next generation image-text models

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.374215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.374215Z digest=sha256:d13922f86404731a9f3792d1046515dfcabdd772b0074862a19c296c2f33f2bc

Observation e40cbb6c-00dd-44ce-b0d7-197e02296213 · outbound

This paper cites Audio-Visual LLM for Video Understanding.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Audio-Visual LLM for Video Understanding

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.378321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.378321Z digest=sha256:cb558f88340af5ef4243e73eb900b274703d657c9ee5051a106f129004161058

Observation d8dddf8f-ccea-48c0-8946-f4dff78ec74b · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.382504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.382504Z digest=sha256:48915a773faac8c836f844e6ff875e84cb55170656951ab95ce223286c429db3

Observation 834f2aaa-61d9-4671-8b90-1a1b2562d80a · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.386904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.386904Z digest=sha256:b6605be72838740265e487951ad60ecc6cd8dbad6a0a638aaeb38721857fc49b

Observation 347597a9-b7f1-4867-963b-3f4764f64f39 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Moviechat: From dense token to sparse memory for long video understanding

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.391276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.391276Z digest=sha256:6fd03adee31081a94de6635377537097df67d6337a44376977ea5bcdca57dc64

Observation 47c3e7f5-047f-441f-9cfd-9fe183439f7b · outbound

This paper cites MovieChat+: Question-aware Sparse Memory for Long Video Question Answering.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions MovieChat+: Question-aware Sparse Memory for Long Video Question Answering

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.396297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.396297Z digest=sha256:f28ec071f08a77ff9e3b05c404ab3849ecc7708150e7b036aefc3902bde241eb

Pith citing papers

Observation 62bdc033-4193-46d2-9b07-93013461b598 · inbound

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding cites this paper.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.372307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.372307Z digest=sha256:be051c8834e29bf7b271f7a5d0f7d12f7a769e9ee73b833d7f7d516b5d6291b0

Observation c4b67586-0fda-4dfd-a13e-849b0c11f5c8 · inbound

Ola: Pushing the Frontiers of Omni-Modal Language Model cites this paper.

Ola: Pushing the Frontiers of Omni-Modal Language Model InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.376946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.376946Z digest=sha256:4cfa4fcdb2211e4918079fbeaea843dcd1c4aed6d306f62e9736b489250de4e4

Observation 75b57934-9208-4775-af26-56708e043fd2 · inbound

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs cites this paper.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:22:27.783266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:516b99c3409f74dcaf269aa14d3f1a26a7f67d102ca7716a8cb5a6dd422d8d2c

Observation 61a438b0-bfb6-49f1-9949-94f0d8549a88 · inbound

Towards Understanding Camera Motions in Any Video cites this paper.

Towards Understanding Camera Motions in Any Video InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.590340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.590340Z digest=sha256:714ff4fcaeb1059db4501581416f9e087e0ae41ad4bded8fb10acb9540e69fd1

Observation e4f1732d-a2a3-497c-926c-d717722e1696 · inbound

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos cites this paper.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.458880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.458880Z digest=sha256:a21f8863b1eefac951c4674c1be1a84027810e19833b6a1fbcd046f016339199

Observation e7381476-3277-4121-a530-b252841a80b1 · inbound

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence cites this paper.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:43.022715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:43.022715Z digest=sha256:0a38405f0712d9a6dea2485164caa5fcb639ca7efb60025b72d8731aa604a62e

Observation 0adb5e62-b5c0-4719-9158-042dcd1b604f · inbound

Know-MRI: A Knowledge Mechanisms Revealer&Interpreter for Large Language Models cites this paper.

Know-MRI: A Knowledge Mechanisms Revealer&Interpreter for Large Language Models InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:35.710503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:35.710503Z digest=sha256:78ac16bea35fc7d062bea9ec52bf92f61939e54ce8f00820f0aa3ce84b6a428e

Observation 4cef6c13-6340-4263-8359-99ce2afbf1c6 · inbound

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models cites this paper.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:34.167998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:34.167998Z digest=sha256:b17749a0249e02ab41aaddb8c42f36b2ebad9cbc1e77b53ea8c39e66c8fce6bb

Observation 3ec31ea1-a90e-4493-a56b-c09be0b7b9bb · inbound

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context cites this paper.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:14.198672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:14.198672Z digest=sha256:cee77c95ed17b8f75e2abea14ea04b851233b51b8cc452dcc6da2e256fd397fd

Observation 6b46ae8c-9509-456b-aa26-197f2265ce28 · inbound

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes cites this paper.

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:11.657893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:11.657893Z digest=sha256:c206410cd9c9f243a5dcb50e1f848bd70df03fedc14844de8f30d591c7e742d5

Observation c6146844-f8ee-49d0-9297-b77bace798b1 · inbound

Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning cites this paper.

Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-21T16:44:16.021335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T16:43:11.995960Z digest=sha256:adb6f367ec390a83b4cdc9c7585948b60261a17f1b6823a39aae707c7716b5e1

Observation 12da96cb-a329-43fa-9797-2526cd9f7f49 · inbound

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning cites this paper.

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:25:30.155575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T14:22:25.660785Z digest=sha256:73e841a5a6990f2868a9aa6670c6758c1773d34300b477d5c9873680d7f4abc2

Observation e404bb10-364e-4aa7-977c-1807aaf36f69 · inbound

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games cites this paper.

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:19:13.658047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T21:17:02.332687Z digest=sha256:4a52d563da8ac33f1bc5127d12011db3e9d94966a66be895f7eb9d6e83f998f9

Observation d8002c6b-4f07-46a7-b86c-10a96f7a00a3 · inbound

ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference cites this paper.

ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:19:29.964055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-26T18:17:53.013043Z digest=sha256:aefb0656c912b1dbc54f94dc352d7e6cd1b2e9bec09735c6453e894780385195

Observation a25e7ef2-4c1f-492c-8e1e-295760706f59 · inbound

Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory cites this paper.

Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-11T06:35:35.951554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T06:35:35.951554Z digest=sha256:2b04e13f8c2df67a9ae96cbec555aa2ce376ceab422586bb5470f7b844617c98

Observation c7153a65-386a-4160-b939-63be0ae1703d · inbound

FOLIO: Focused Semantic Memory for Streaming Video Understanding cites this paper.

FOLIO: Focused Semantic Memory for Streaming Video Understanding InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:49.748388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:49.748388Z digest=sha256:c020f771c4d39d100c1f0b103812c572a4cdf3d908450735913e949972f444e3

Observation 8a354390-4034-47f5-b432-5c9212b803d4 · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 134

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.654651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.654651Z digest=sha256:fe037757b82ac271965829e6643eb8c288c9e6aefd9bc1291559e8e0c13c4fac

Observation 23380163-05c6-4472-aa1b-ba3150efe153 · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:13.894956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:13.894956Z digest=sha256:7ef14f5151bc3a068f5cf829030ece5e350b7b1c7e1c4cf8cadc6fe04144d594