Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T16:56:06.396297Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 100 of 156 outbound references and 18 inbound Pith citation observations for arXiv:2412.09596.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T16:56:06.396297Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:31:56.590340Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T03:19:29.961741Z
100 of 156 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c744f3c7-613e-4012-927b-bda6284c14bc · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 433c683d-8d92-4f14-85ee-792f70a05c4a · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Common Voice: A Massively-Multilingual Speech Corpus
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f4e8061-edb1-45d5-97d7-4f46078fbfbe · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98fc5988-08af-47c5-ac23-1052fa794e5f · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Qwen-VL: A frontier large vision-language model with versatile abilities
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c8e06a1-69d8-44c5-92c7-4c59c54d1563 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Baichuan 2: Open large-scale language models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07517be1-07f1-4da5-adc8-f83d54209323 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Introducing our multimodal models, 2023
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6221e4a6-15d4-485a-bbda-504726b387a0 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Language models are few-shot learners
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33144b44-ea54-4a49-9432-fc3d8d4130e5 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b98253d3-693c-4aca-84f2-925a25f511c0 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions InternLM2 Technical Report
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a64ae17-e955-46bb-8cc6-57455852f026 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions DualFocus: Integrating Macro and Micro Perspectives in Multi-modal Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30b85e59-2d2d-4431-a9e5-654e150edb9e · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bc294ea-2e32-44f6-88da-4378cbd2281f · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Videollm-online: Online video large language model for streaming video, 2024
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 780a2c6b-aba3-47fd-a748-1eb4aa5c67bf · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 090ce6d6-c9d9-4002-951d-3a79a94709ed · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f00bb27-f81e-463f-84e7-8ec8086f4590 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb2ee4e5-48e0-4649-9b6f-76ad751b43f8 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Timemarker: A versatile video-llm for long and short video understanding with superior temporal localiza- tion ability, 2024
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61ee2236-00d0-4f6b-9844-27ba13cc289e · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions PaLI-X: On Scaling up a Multilingual Vision and Language Model
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acedf2d3-392f-40b6-bcee-b0d4bfa8d97b · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Pali-3 vision language models: Smaller, faster, stronger, 2023
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c298f8e-bdf7-4563-876b-5029fb33866e · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Pali: A jointly-scaled multilingual language-image model, 2023
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f652eb5-bfd8-4aed-b378-af6ac8ff9e63 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6212b70a-fe1a-46ad-9e61-0815d76825a7 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 034adb00-2a66-4d98-9cc3-66c647eb824b · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f569de39-f504-4371-9747-4f3991c3c63b · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c2c3b17-0ef8-48d2-ad46-163a6c64e327 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Palm: Scaling language modeling with pathways
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e82ad13-742a-48d1-9434-5ba479ec6862 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f61d49ce-c6f5-4c8c-bc14-5fdd891ac026 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Qwen2-Audio Technical Report
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afbb0b0e-74d8-4c0e-8c6b-41523560362e · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions The CHiME-8 DASR Challenge for Generalizable and Array Agnostic Distant Automatic Speech Recognition and Diarization
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cb52534-6eeb-4a53-a6b5-b7daa14d7426 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Instructblip: Towards general- purpose vision-language models with instruction tuning,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5074f80-6bff-4f65-b3da-415a02bedbfe · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee520e93-d1f3-4aee-9f7c-4a245cd46a58 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84b67cb4-3541-4f43-85c7-683cdc86f827 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions PaLM-E: An Embodied Multimodal Language Model
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05591220-12b2-4031-8c5c-8a4b5c0294c8 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions ActivityNet: A large-scale video benchmark for human activity understanding
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54b48044-aca8-490e-834f-1607bc646cd8 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Videoagent: A memory-augmented mul- timodal agent for video understanding, 2024
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c530d3e2-5ced-4dde-9fcb-81f40a943c60 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7af27ba0-f27c-4d72-8557-73e62e5183fb · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions FSD50K: An Open Dataset of Human-Labeled Sound Events
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b40d57bb-ec1e-4364-b7a9-13a64f2269e7 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a76eeb1-e83b-453d-84f6-2df57cff5257 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df22b326-c263-45d5-bb1e-49cd100fcf04 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Vita: Towards open-source interactive omni multimodal llm, 2024
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00424ac6-d2ba-43b3-97f1-b527e0d1f771 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 485d39c7-7c65-4899-8652-babf6d6eec75 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Funasr: A fundamental end-to-end speech recognition toolkit
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47cbe73f-7057-44cc-ba56-b0d1aa1327fc · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Ego4d: Around the world in 3,000 hours of egocentric video
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be6177a0-c5b6-44bd-aebf-05d551f1470c · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Onellm: One framework to align all modalities with language, 2023
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 938bcdb0-1865-49dc-8ce8-f2956f94d0db · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e38e553c-8288-480d-a3ff-1a0cb72be693 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions From Image to Video, what do we need in multimodal LLMs?
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f106160-c0f9-4cd2-b1de-a61fa3a8525c · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Video ReCap: Recursive Captioning of Hour-Long Videos
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59568157-554a-406f-9da7-05a407233827 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, et al
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dac03a0f-6b47-4df2-a572-f1a68f5e6f17 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Mixtral of Experts
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71ce98f5-8e6d-426d-80ab-599c88a6d78b · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Mantis: Interleaved multi- image instruction tuning, 2024
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23f79b5f-948f-49b4-be34-ed9cf150219f · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 776b7648-a119-449b-8d8f-2d49fa3dc781 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Language Repository for Long Video Understanding
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39c479a1-4dd8-4cdc-95c0-8885ea7b8a1f · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Scaling Laws for Neural Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af49cee5-8323-4dd3-8938-380d20bbf0d2 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions An image grid can be worth a video: Zero-shot video question answering using a vlm, 2024
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6dd74f8-a085-45d3-8b99-8296c79fc8d3 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Audio set classification with attention model: A probabilistic perspective
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22e4558d-670b-4e36-82a7-6124f9f2c2cd · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions The open images dataset v4
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7da56cc2-2180-4ecd-9f8d-bbcaf71cc812 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions A path towards autonomous machine intelli- gence version 0.9
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c34e72a2-4864-4edd-bf00-21930523295b · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Otter: A multi-modal model with in-context instruction tuning
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1afa105-ceb0-4888-9091-51d4b00482d6 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions LLaVA-OneVision: Easy Visual Task Transfer
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38b506b7-86ed-4a00-8a13-ac6cc65c444e · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Aria: An open multimodal native mixture- of-experts model, 2024
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ad76a10-5bb5-430a-a9fb-20d3772ac61f · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31dcb5ac-6e71-45d9-a861-40bf71b16d7c · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions VideoChat: Chat-Centric Video Understanding
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 790c9991-459e-4d10-85d6-6d44328e20d7 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Mvbench: A comprehensive multi- modal video understanding benchmark, 2023
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e49de970-4143-4a15-8f23-a5d8efdc613e · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Mvbench: A comprehensive multi-modal video under- standing benchmark
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 964cd314-0fc9-4b2f-84cc-0ddab4b10069 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75e74ca5-1747-484b-b908-478480df5e54 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Ocean-omni: To understand the world with omni-modality,
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc4509b7-6fec-4873-bcc1-fda0f4c35220 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Llama-vid: An image is worth 2 tokens in large language models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e913b9c3-9757-449d-9d54-55e3442ae1c2 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bd4898b-b214-4606-8419-3699604de887 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f81de65e-2723-438b-990e-d782a044e5b1 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Vila: On pre-training for visual language models, 2024
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 908aaa5a-ee8e-4526-8777-f8e2b3f0ccd6 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Microsoft coco: Common objects in context
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e9a1db8-0bde-43fc-a99b-502e8b213b33 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b91696b1-eac5-4dd0-bce6-de6e3dfe1063 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Visual instruction tuning
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e9f8983-2816-491e-a802-47cb02a7a7e5 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 471d60f4-d974-49da-b27b-db695c1966d6 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions TempCompass: Do Video LLMs Really Understand Videos?
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efa31c04-731a-4b27-809e-0af8506c2415 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions A convnet for the 2020s, 2022
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbe6ba7c-1f0a-4542-9429-521a7d73908e · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3052a53-f009-4e02-8f81-069b7ade084c · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions ASCEND: A Spontaneous Chinese-English Dataset for Code-switching in Multi-turn Conversation
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddb41289-f379-443c-a354-ad0be15de7cb · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8a474e0-d0b3-410d-8da3-75f94ae18b0f · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Language Model Can Listen While Speaking
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 979a71f5-dd93-49fe-8169-d3df4766ba75 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation facc0219-2a24-4ea2-a415-bd5a42daf1f1 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ad79243-ad23-40e8-a037-e19ba7fe2db6 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Unresolved cited work
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 851f6ad7-4674-4a86-8985-68275265cac7 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Gpt-4 technical report, 2023
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 884dbce1-66dc-423d-9e37-64ada99cc10c · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Gpt-4v(ision) system card, 2023
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e63f14e-b146-4b1b-b86e-a68550d75380 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions GPT-4o System Card
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 374ef120-9fa9-4ffb-94f7-c06721450ed3 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Dinov2: Learning robust vi- sual features without supervision, 2024
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16806f12-98b9-4869-ab7f-99cd7d1281cc · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Training language models to follow instructions with human feed- back
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 239884da-b5b2-4be2-8af8-86b08bd9c539 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Librispeech: an asr corpus based on public domain audio books
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3eb47b21-5ca2-43e1-9c23-2c984cb84dba · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Kosmos-2: Grounding multimodal large language models to the world
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d41ee99a-730f-42d7-9d9c-b86cb7b87b1f · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Streaming long video understanding with large language models, 2024
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d13233f-fab6-41d3-920a-7da4e789612b · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Introducing Qwen-7B: Open foundation and human-aligned models (of the state-of-the-arts), 2023
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8ce8c6d-f0bf-4265-b191-1e90aff6faab · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Learn- ing transferable visual models from natural language super- vision
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0faa449f-d6a8-4cce-b2d3-2f19a1927e6f · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Robust speech recognition via large-scale weak supervision, 2022
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2724e02-687e-444a-a2a1-7b815eecfcad · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Robust speech recognition via large-scale weak supervision
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f11389cc-a2fc-427d-b7c0-2be58b05813b · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Ryoo, Honglu Zhou, Shrikant Kendre, Can Qin, Le Xue, Manli Shu, Silvio Savarese, Ran Xu, Caiming Xiong, and Juan Carlos Niebles
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7afc7d95-6afd-43d9-a757-5e118e823574 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions LAION-5B: An open large-scale dataset for training next generation image-text models
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e40cbb6c-00dd-44ce-b0d7-197e02296213 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Audio-Visual LLM for Video Understanding
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8dddf8f-ccea-48c0-8946-f4dff78ec74b · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 834f2aaa-61d9-4671-8b90-1a1b2562d80a · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 347597a9-b7f1-4867-963b-3f4764f64f39 · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Moviechat: From dense token to sparse memory for long video understanding
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47c3e7f5-047f-441f-9cfd-9fe183439f7b · outbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions MovieChat+: Question-aware Sparse Memory for Long Video Question Answering
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62bdc033-4193-46d2-9b07-93013461b598 · inbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4b67586-0fda-4dfd-a13e-849b0c11f5c8 · inbound
Ola: Pushing the Frontiers of Omni-Modal Language Model InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75b57934-9208-4775-af26-56708e043fd2 · inbound
Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 61a438b0-bfb6-49f1-9949-94f0d8549a88 · inbound
Towards Understanding Camera Motions in Any Video InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4f1732d-a2a3-497c-926c-d717722e1696 · inbound
TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7381476-3277-4121-a530-b252841a80b1 · inbound
SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0adb5e62-b5c0-4719-9158-042dcd1b604f · inbound
Know-MRI: A Knowledge Mechanisms Revealer&Interpreter for Large Language Models InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cef6c13-6340-4263-8359-99ce2afbf1c6 · inbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ec31ea1-a90e-4493-a56b-c09be0b7b9bb · inbound
HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b46ae8c-9509-456b-aa26-197f2265ce28 · inbound
HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
Reference 103
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6146844-f8ee-49d0-9297-b77bace798b1 · inbound
Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 12da96cb-a329-43fa-9797-2526cd9f7f49 · inbound
Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e404bb10-364e-4aa7-977c-1807aaf36f69 · inbound
Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d8002c6b-4f07-46a7-b86c-10a96f7a00a3 · inbound
ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a25e7ef2-4c1f-492c-8e1e-295760706f59 · inbound
Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7153a65-386a-4160-b939-63be0ae1703d · inbound
FOLIO: Focused Semantic Memory for Streaming Video Understanding InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a354390-4034-47f5-b432-5c9212b803d4 · inbound
X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
Reference 134
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23380163-05c6-4472-aa1b-ba3150efe153 · inbound
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.