Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T00:44:49.747631Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 100 of 122 outbound references and 0 inbound Pith citation observations for arXiv:2607.14935.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T00:44:49.747631Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 122 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4a5ccee9-7d80-4eb4-980c-aedb59733c49 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbcfc400-3eab-4288-add5-b74a12a64c15 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7e61b8a-07a7-4a88-98f4-a78be078cb89 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e667046c-a725-4642-b587-c2328b08047c · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89442b5b-01c1-4281-b165-e4ae54133cd4 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef734fdc-7083-4879-a6ff-04ef4bbef4a0 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Qwen2 Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3002662-927d-4f41-84b7-8cd3dfe506a1 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Qwen3 Technical Report
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48755b6a-f903-4acc-923a-ccfe3364939c · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Videochat-r1
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e95a7ad-809f-46f5-8825-6b314a59a3ec · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7eb20e68-e4ab-472d-83e9-bc2fcebe7712 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a14c559-b0a9-4999-89b1-aa7072fa0657 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3971ffea-3617-48e5-a5ad-d0b190556c65 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Timechat: A time-sensitive multimodal large language model for long video understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 896580a9-b2dd-4e5f-928b-d905b606de14 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Timesuite: Improving mllms for long video understanding via grounded tuning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ee61b8d-42f9-49b3-bdf3-4d68bfebf8cd · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Timelens: Rethinking video temporal grounding with multimodal llms
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a97708be-9b01-4dac-b427-131bd70fac0e · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Videollm-online: Online video large language model for streaming video
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e5440df-9915-40ca-8717-78f8973438bf · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a369fec-b7b9-4308-81b4-7288778b1d35 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Online video understanding: Ovbench and videochat-online
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd5440ab-3792-4d09-8638-349f095ca67a · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Dispider: Enabling video llms with active real-time interaction via disentangled perception, decision, and reac- tion
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98d24d20-5989-423a-807d-fd5b57368f76 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Streamforest: Efficient online video understanding with persistent event memory
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e546c26-dd84-4100-afa7-2c56326fc8b1 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Streaming Video Instruction Tuning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 993fd87c-d1d1-4a9c-9cb8-d355b9fc2d39 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Streambridge: Turning your offline video large language model into a proactive streaming assistant
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84c47551-8d13-4eb5-afe6-0ea131a643b8 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Gemini 3: News and announcements
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdfce8b2-b132-4d59-be38-3009f0e0f2bf · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Kimi K2.5: Visual Agentic Intelligence
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f292ffb8-0f86-4e74-879e-f18ccb366a37 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Seed1.8 Model Card: Towards Generalized Real-World Agency
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e387c711-e657-4b7c-89db-b3d318a62123 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68902810-0686-4fd3-a7f9-7b0fbb7bfd67 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Llava-video: Video instruction tuning with synthetic data
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f310e039-88e8-4489-a3f5-4b83df3475d5 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Spoken moments: Learning joint audio-visual representations from video descriptions
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 696b1e40-7be9-4046-870a-3400740071da · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Vript: A Video Is Worth Thousands of Words
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 348f95ff-4641-4035-9846-921b4270211e · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Kimi-VL Technical Report
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ad1ffdd-369f-409c-b585-466a4e9b6ca0 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Visual instruction tuning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fbbbd9f-5d8a-48f9-8681-349e8f704aba · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Caprl: Stimulating dense image caption capabilities via reinforcement learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc39791f-29a1-412a-8774-f3fba5cbe5e9 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adf665ae-a5d2-46df-a260-54c49dd84a5f · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Densefusion-1m: Merging vision experts for comprehensive multimodal perception, 2024
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e086231-e4f4-4f51-b110-bb8fb79de085 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Conceptual captions: A cleaned, hyper- nymed, image alt-text dataset for automatic image captioning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04b79d1e-0edf-457b-bda4-415a13500e9f · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Frozen in time: A joint video and image encoder for end-to-end retrieval
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f1b7b9b-5867-4fee-917d-998c90074fef · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding The Kinetics Human Action Video Dataset
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9070223b-b3a6-407a-930f-bb3cadac97ad · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Perception Encoder: The best visual embeddings are not at the output of the network
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc8bdbe5-db5f-4970-bc1b-09f5f5777cae · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97b1d4bd-f720-4045-b127-9119a204ce99 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Movie Description
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2487f5b2-24d7-41e8-9d0e-57d5915cece2 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Vatex: A large-scale, high-quality multilingual dataset for video-and-language research, 2020
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d1616b3-84c4-4329-b3fc-b43ea5e67a88 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Activitynet: A large-scale video benchmark for human activity understanding
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3a2e685-6b4c-491d-acd5-1ad19717cea8 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Hollywood in homes: Crowdsourcing data collection for activity understanding
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2722b1ba-4233-4f03-834d-c6f0f5029497 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding The” something something” video database for learning and evaluating visual common sense
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb7ea558-f755-4e43-98a0-275f5ac5a278 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Ego4d: Around the world in 3,000 hours of egocentric video
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5334e404-e128-4616-891d-a8d35781854c · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Sharegemini: Scaling up video caption data for multimodal large language models, June 2024
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 434eb6fd-db63-4997-8139-4e4e8e4601b0 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Sharegpt-4o: Comprehensive multimodal annotations with gpt-4o, 2024
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c212413-8de1-446d-97f7-3302e944f610 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Bee: A high-quality corpus and full-stack suite to unlock advanced fully open mllms, 2026
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73f7c1a8-e954-4cd3-b94e-18816196ddd5 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Unresolved cited work
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05740471-e970-4649-8eae-b3c488fdc495 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding TVQA: Localized, Compositional Video Question Answering
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfe6d3b7-0d2d-478d-aad9-a23b8f8f4043 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Finevideo
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63fe83c7-4305-4236-8c2d-31689e10f41c · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding CinePile: A Long Video Question Answering Dataset and Benchmark
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44f6e26a-3390-417a-a7dc-68dcecdda5ec · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7aad7ba-4663-404f-86ec-3a64470c839f · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding TGIF-QA: Toward Spatio-Temporal Reasoning in Visual Question Answering
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 933de6f8-6ff6-4e45-937e-4ccaa01c61c8 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Tenenbaum
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddffa66c-6b18-4507-8f6a-462f594e4986 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Exo2ego: Exocentric knowledge guided mllm for egocentric video understanding, 2025
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29679ae7-b005-4d36-ab11-651511c0719a · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Motionbench: Benchmarking and improving fine-grained video motion understanding for vision language models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12757dcf-7036-41ce-b4ac-9ceee6d0bd8b · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Lasot: A high-quality benchmark for large-scale single object tracking
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92921bda-1059-4880-aed9-db3eb7094c78 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Got-10k: A large high-diversity benchmark for generic object tracking in the wild
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68d5b2d2-ca47-477e-9453-4f25bba290d2 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijaya- narasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, Cordelia Schmid, and Jitendra Malik
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 797fe49c-7856-46dc-a804-f06fab2fc2cd · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Egovqa-an egocentric video question answering benchmark dataset
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28367e16-b65d-4c7e-84da-1042e0956299 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Learning transferable temporal primitives for video reasoning via synthetic videos, 2026
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e21f2666-0810-4545-b42d-eb15d92e4cfc · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Tomato: Assessing visual temporal reasoning capabilities in multimodal foundation models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e3ba013-2d83-496e-9ee5-82869176d604 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Lost in Time: A New Temporal Benchmark for VideoLLMs
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc81c15b-8707-4181-a9b9-f212d9ff8779 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Tempcompass: Do video llms really understand videos? In Findings of the Association for Computational Linguistics: ACL 2024 , pages 8731–8772
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57d2e2c4-13c4-4e1d-a7e6-9879a171b493 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1af6863-d03b-4ec8-8da1-c04869640ed1 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding LVBench: An Extreme Long Video Understanding Benchmark
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84ee4e48-c29d-4129-aa2e-0182b72b4e4e · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a0e508e-08bc-41c5-a574-c1a893454cb8 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c9669d4-2c06-4e4e-bba6-df62de1d7772 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Mmvu: Measuring expert-level multi-discipline video understanding
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cf8ffcd-ab76-4593-9b85-c9087da6b0f0 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Minerva: Evaluating complex video reasoning
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f09832a2-f0ac-4d92-aefe-e83c8dbc9759 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb923698-d405-4226-9789-6d1f8bdb5cc4 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Tall: Temporal activity localization via language query
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2685caf0-64d8-4b74-8a34-57bfd9f32434 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Dense-captioning events in videos
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ec15f71-9897-4104-ab87-3af549c81da9 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Detecting moments and highlights in videos via natural language queries
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8536bac6-8a77-49a4-be09-1afb22f38f08 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Vidi: Large Multimodal Models for Video Understanding and Editing
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40348e98-5b43-4527-8e75-a38ed0aeb684 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Vidi2: Large multimodal models for video understanding and creation
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54903a52-59db-465b-91bf-80471d354aa0 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Momentseeker: A task-oriented benchmark for long-video moment retrieval
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adf89ab0-e959-4a7b-84ad-7c967815939f · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding GPT-5 system card
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e079e2a-41aa-4261-81bf-284bf742459e · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 527104fb-143b-4554-aacb-595ec68ac764 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Introducing claude sonnet 4.5
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29f11d55-cb24-4841-99a3-6ac23f236f6a · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bb62b33-3fc7-4574-9664-dd78ac3c8c3b · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Eagle 2.5: Boosting long-context post-training for frontier vision-language models
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f6503ae-9746-4fe7-918f-cbccd66be913 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34055a81-c4a9-4c9c-957e-3431b35f15f0 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb5c00dc-8b5c-4208-a414-9e766e2d3148 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Unresolved cited work
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45268e19-3c2e-44e5-8398-aa21ee8d91ec · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Streamingbench: Assessing the gap for mllms to achieve streaming video understanding
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7a17479-bc58-446d-8e99-0bea900400df · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding River: A real-time interaction benchmark for video llms
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8ee7c1c-267d-4f8c-b05d-d9d00aca60ac · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Em-Garde: A Propose-Match Framework for Proactive Streaming Video Understanding
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7523c331-160b-4a39-b74a-911205ccb356 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a95b4f6-12f8-44f1-b33f-20cffae494f8 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Livecc: Learning video llm with streaming speech transcription at scale
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eef673cb-ab0d-41eb-96c0-09b8a853e2c3 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Timechat-online: 80% visual tokens are naturally redundant in streaming videos
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7dbb6e1-fc38-43aa-a75a-89864475aee9 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding StreamingVLM: Real-Time Understanding for Infinite Video Streams
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a57edb3-07ad-4d15-83ec-8fb67b370d25 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Mm- duet2: Enhancing proactive interaction of video mllms with multi-turn reinforcement learning
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43fc2d0a-4e5a-49ff-8900-0cd58958772e · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34cdd0cd-7ef9-42b9-8cc2-6c5f5665e243 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding VideoChat: Chat-Centric Video Understanding
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52d95c55-6cca-42a4-8338-c243094e56f0 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Video-chatgpt: Towards detailed video understanding via large vision and language models
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a1eb399-e080-47c6-b0d6-af99858ce804 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2fcb43b-74dd-4bf0-bd01-d9d6382cb93f · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding LLaVA-OneVision: Easy Visual Task Transfer
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 916cc133-aeb9-4a69-85be-5014ca2f787c · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5939ccc1-f8be-42e1-91cf-a1dc4dec07c9 · outbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding AWT: Transferring Vision-Language Models via Augmentation, Weighting, and Transportation
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.