Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T23:30:00.457974Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 100 inbound Pith citation observations for arXiv:2305.06355.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T23:30:00.457974Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T05:41:47.548723Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
59 of 59 outbound references displayed
External citation measurements
90
pith, observed 2026-08-05T02:28:24.338817Z
Observation 4f1d7275-bcfb-42f1-8244-508f8918f414 · outbound
VideoChat: Chat-Centric Video Understanding Openflamingo, March 2023
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 243c31df-d978-462a-b328-eee78fd535be · outbound
VideoChat: Chat-Centric Video Understanding Frozen in time: A joint video and image encoder for end-to-end retrieval
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0b04bd30-56ac-4b1e-97a0-6c38aacacf6b · outbound
VideoChat: Chat-Centric Video Understanding Language models are few-shot learners
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 38cbb19d-6359-4c6a-8cbd-94ac5dabec98 · outbound
VideoChat: Chat-Centric Video Understanding Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d6438f78-fafb-4231-869f-24ece27e1510 · outbound
VideoChat: Chat-Centric Video Understanding InternVideo-Ego4D: A Pack of Champion Solutions to Ego4D Challenges
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e3d6e35d-5704-4e6a-80c4-1fd9563f1deb · outbound
VideoChat: Chat-Centric Video Understanding Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f8fcdbca-4181-4598-b423-7a3f47e98c71 · outbound
VideoChat: Chat-Centric Video Understanding Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5a610f2f-ed84-41c7-ad5b-5ff2f14d57fa · outbound
VideoChat: Chat-Centric Video Understanding Scaling Instruction-Finetuned Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 06c4edb0-be53-4a0a-a1e8-1d2d72fec63e · outbound
VideoChat: Chat-Centric Video Understanding Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 448d3caa-667f-4136-bec8-b84ba513e11e · outbound
VideoChat: Chat-Centric Video Understanding Stablelm: Stability ai language models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation df5dde4e-4fda-4135-b59e-247aacefc6ec · outbound
VideoChat: Chat-Centric Video Understanding An empirical study of training end-to-end vision-and- language transformers
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ab787156-95d0-45d3-be18-1de421898834 · outbound
VideoChat: Chat-Centric Video Understanding VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 16a193ea-9225-410a-a0c7-982b383dc647 · outbound
VideoChat: Chat-Centric Video Understanding Scaling up vision-language pre-training for image captioning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4a99b478-b6f0-497d-913d-cb6931eee182 · outbound
VideoChat: Chat-Centric Video Understanding Language Is Not All You Need: Aligning Perception with Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5d6d4010-0a8e-4b89-a0fd-e76c14b4cc0b · outbound
VideoChat: Chat-Centric Video Understanding Tag2Text: Guiding Vision-Language Model via Image Tagging
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 90bcb349-2d2f-40f7-8277-c3e4861f24d1 · outbound
VideoChat: Chat-Centric Video Understanding Dolphin: General video interaction platform based on llms
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6fc4f452-f14b-4c0a-8de7-91bd62e3ea64 · outbound
VideoChat: Chat-Centric Video Understanding Visual genome: Connecting language and vision using crowdsourced dense image annotations
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 072ee880-f327-46cd-82a0-9d9f90a051b9 · outbound
VideoChat: Chat-Centric Video Understanding BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dc230461-5565-47ef-92c1-357270b0a820 · outbound
VideoChat: Chat-Centric Video Understanding Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a5b874da-8fef-43d5-8329-ae9d9f30a5a9 · outbound
VideoChat: Chat-Centric Video Understanding UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ab5c4fc4-b242-4bda-93f8-b7df99450714 · outbound
VideoChat: Chat-Centric Video Understanding Unmasked Teacher: Towards Training-Efficient Video Foundation Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 95867805-c515-4620-b73c-0fb3e2bf0c0a · outbound
VideoChat: Chat-Centric Video Understanding LAVENDER: Unifying Video-Language Understanding as Masked Language Modeling
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 78b1790f-89ba-42ad-9701-6e8853b2cb3b · outbound
VideoChat: Chat-Centric Video Understanding Learning Spatiotemporal Features via Video and Text Pair Discrimination
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 93260ef7-5c24-485d-93a9-050fee0c7104 · outbound
VideoChat: Chat-Centric Video Understanding TaskMatrix.AI: Completing Tasks by Connecting Foundation Models with Millions of APIs
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d5d1a375-3258-4c0a-8cce-5d2ae4dcc3a5 · outbound
VideoChat: Chat-Centric Video Understanding Visual instruction tuning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dd337dbf-d56d-4135-a28d-da0da5248840 · outbound
VideoChat: Chat-Centric Video Understanding InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8fa12758-6454-4ca9-93b2-12ee0c0b55f7 · outbound
VideoChat: Chat-Centric Video Understanding End-to-end learning of visual representations from uncurated instructional videos
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b3b60760-69fe-4464-a775-e10d6d0dfea3 · outbound
VideoChat: Chat-Centric Video Understanding Cross-Task Generalization via Natural Language Crowdsourcing Instructions
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fb55e015-b5ab-44a4-8637-d8ecb7375671 · outbound
VideoChat: Chat-Centric Video Understanding Gpt-4 technical report
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3fdb47b0-ef44-4665-a138-0980a99c4e32 · outbound
VideoChat: Chat-Centric Video Understanding Chatgpt: Optimizing language models for dialogue
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c8762e9d-e69c-4a4d-a13b-fc8b84b1dbec · outbound
VideoChat: Chat-Centric Video Understanding Im2text: Describing images using 1 million captioned photographs
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 37268e7b-c1b3-4473-ac6f-e599e20b9bed · outbound
VideoChat: Chat-Centric Video Understanding Training language models to follow instructions with human feedback
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1f5592a9-5b6c-4a81-bcce-e25bde195f02 · outbound
VideoChat: Chat-Centric Video Understanding Robust Speech Recognition via Large-Scale Weak Supervision
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 837ee617-26ac-4b68-91ef-73d73f32acd8 · outbound
VideoChat: Chat-Centric Video Understanding Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b18076b8-4fd7-4dcd-90da-eb9297708857 · outbound
VideoChat: Chat-Centric Video Understanding Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7825d19f-eb53-4f0d-bfd3-6f58d4a5bcf9 · outbound
VideoChat: Chat-Centric Video Understanding How Much Can CLIP Benefit Vision-and-Language Tasks?
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 74912950-f8ee-4edc-b5f1-af1b1f977abf · outbound
VideoChat: Chat-Centric Video Understanding HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1a772ef6-6cb4-483e-925b-dac2e4448560 · outbound
VideoChat: Chat-Centric Video Understanding Murphy, and Cordelia Schmid
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e12117b1-6c89-4e1c-a853-cc27b8436508 · outbound
VideoChat: Chat-Centric Video Understanding EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 34774cd3-7749-47dd-b53b-e80404f37492 · outbound
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8837ab19-0e7f-4e4c-9102-0676edff31b7 · outbound
VideoChat: Chat-Centric Video Understanding Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c1539f79-b1e3-49b0-a27e-acefe3128c87 · outbound
VideoChat: Chat-Centric Video Understanding LLaMA: Open and Efficient Foundation Language Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3a02762d-64c5-4419-932f-1de2d9e32d0a · outbound
VideoChat: Chat-Centric Video Understanding All in One: Exploring Unified Video-Language Pre-training
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8289f058-a571-428a-8d0a-0d3825f37c8d · outbound
VideoChat: Chat-Centric Video Understanding Videomae v2: Scaling video masked autoencoders with dual masking
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5187eb9a-34b1-494a-9c95-6e71d691d44c · outbound
VideoChat: Chat-Centric Video Understanding Internimage: Exploring large-scale vision foundation models with deformable convolutions
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 01890349-770e-4c2c-9b11-0c15604e7c9d · outbound
VideoChat: Chat-Centric Video Understanding InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 71cca06b-0916-4d52-a063-8629de8e0e6c · outbound
VideoChat: Chat-Centric Video Understanding Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 52ccaa0b-309f-4256-bf89-57f0f0adba60 · outbound
VideoChat: Chat-Centric Video Understanding GRiT: A Generative Region-to-text Transformer for Object Understanding
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b9b1d4cd-aa91-48eb-b583-b569e1e5952e · outbound
VideoChat: Chat-Centric Video Understanding VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 013852b1-7a80-4f7e-b826-b5e1872d8a59 · outbound
VideoChat: Chat-Centric Video Understanding MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 25df6251-3b0c-4dbc-8297-b07522936587 · outbound
VideoChat: Chat-Centric Video Understanding FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 90037464-168e-4e64-8291-c62f5b392ac3 · outbound
VideoChat: Chat-Centric Video Understanding mplug-owl: Modularization empowers large language models with multimodality
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 33860558-ca6c-4577-a7c3-470807f3a221 · outbound
VideoChat: Chat-Centric Video Understanding Florence: A New Foundation Model for Computer Vision
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5345611c-17fc-423d-9288-17a8db0c6650 · outbound
VideoChat: Chat-Centric Video Understanding Merlot reserve: Neural script knowledge through vision and language and sound
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e4788a76-ddfa-4a4e-b0ca-ca9660cf8c48 · outbound
VideoChat: Chat-Centric Video Understanding Merlot: Multimodal neural script knowledge models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4ad12f5e-f804-442c-856d-7776cf171457 · outbound
VideoChat: Chat-Centric Video Understanding GLM-130B: An Open Bilingual Pre-trained Model
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 00743eca-d43d-4f66-8b69-91eb324e821b · outbound
VideoChat: Chat-Centric Video Understanding OPT: Open Pre-trained Transformer Language Models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8a858f91-89d1-4bed-a889-c4f5b6260a7f · outbound
VideoChat: Chat-Centric Video Understanding MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation acf393d8-c28a-4406-8a0d-58d01bf290c3 · outbound
VideoChat: Chat-Centric Video Understanding Describe the following image concisely
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ae4c2164-6c10-4cf3-84a5-8080ac845dcc · inbound
Otter: A Multi-Modal Model with In-Context Instruction Tuning VideoChat: Chat-Centric Video Understanding
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 15850b87-1bb4-4c0d-8e85-e048d6ba2f49 · inbound
A Survey on Multimodal Large Language Models VideoChat: Chat-Centric Video Understanding
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a365cf16-1b32-4160-b406-b9cbe99f8809 · inbound
A Comprehensive Overview of Large Language Models VideoChat: Chat-Centric Video Understanding
Reference 272
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 14ee9171-1cfa-45b2-95a4-f627cf2fd942 · inbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation VideoChat: Chat-Centric Video Understanding
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6abe26b8-437b-4edb-af3c-8c6a28a9c8bc · inbound
SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension VideoChat: Chat-Centric Video Understanding
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a811ed80-2b38-4a2e-9e15-e364ded8ccdb · inbound
A Survey of Hallucination in Large Foundation Models VideoChat: Chat-Centric Video Understanding
Reference 130
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 48b1f484-3639-4ead-a4b8-01f3bb01bddc · inbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration VideoChat: Chat-Centric Video Understanding
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 59472c27-18c4-4b95-b968-f6b1b8535212 · inbound
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection VideoChat: Chat-Centric Video Understanding
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 884e9606-2235-40c7-b760-7076483c38d2 · inbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark VideoChat: Chat-Centric Video Understanding
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 42f7cdf9-0992-4f98-b24e-edcbcea51670 · inbound
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks VideoChat: Chat-Centric Video Understanding
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 239ac708-0def-4865-a4f0-ecb9cda9ba29 · inbound
Agent AI: Surveying the Horizons of Multimodal Interaction VideoChat: Chat-Centric Video Understanding
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1b876105-a280-431c-ad71-0b69f8f9fd21 · inbound
TempCompass: Do Video LLMs Really Understand Videos? VideoChat: Chat-Centric Video Understanding
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a1e21a52-fc73-4886-9710-2f2d0bce3ea3 · inbound
How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites VideoChat: Chat-Centric Video Understanding
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fac87720-fb17-40cd-bd06-1a749f8a1017 · inbound
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning VideoChat: Chat-Centric Video Understanding
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a73ef674-3145-4c45-a6a8-0b1191e21e94 · inbound
MLVU: Benchmarking Multi-task Long Video Understanding VideoChat: Chat-Centric Video Understanding
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bd69b2c4-4d4e-44ed-91b8-551ffb3f7747 · inbound
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives VideoChat: Chat-Centric Video Understanding
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4015d4dd-a620-425c-920c-7210c1415dc2 · inbound
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output VideoChat: Chat-Centric Video Understanding
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 00f85e25-659b-4493-a6d5-69c64df88929 · inbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models VideoChat: Chat-Centric Video Understanding
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a6cc0f65-2612-433c-aac3-aaeb76eeb852 · inbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models VideoChat: Chat-Centric Video Understanding
Reference 222
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0652cd27-6b42-445f-8066-cb5879896a28 · inbound
CogVLM2: Visual Language Models for Image and Video Understanding VideoChat: Chat-Centric Video Understanding
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 945090e3-98eb-48ff-a09f-1ec94a9f1394 · inbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data VideoChat: Chat-Centric Video Understanding
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 79b5e6c1-457f-4a0e-a0ab-f3722dbca13d · inbound
VidHal: Benchmarking Temporal Hallucinations in Vision LLMs VideoChat: Chat-Centric Video Understanding
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ac1221c8-b61b-4081-8dd9-81e78e4a7685 · inbound
TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos VideoChat: Chat-Centric Video Understanding
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 56d1e175-a9ee-43b8-95c6-8e79068f73e4 · inbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling VideoChat: Chat-Centric Video Understanding
Reference 130
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 419570b1-1595-48dd-9558-1b75b0c809b7 · inbound
Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks VideoChat: Chat-Centric Video Understanding
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 729ca95e-ca18-4b1a-b029-2b6b82e69251 · inbound
Friends-MMC: A Dataset for Multi-modal Multi-party Conversation Understanding VideoChat: Chat-Centric Video Understanding
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d0bae63-f0bf-4c94-845d-8e3eb0568304 · inbound
HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks VideoChat: Chat-Centric Video Understanding
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bc44888b-8097-4936-ade0-db1a18d98b24 · inbound
Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment VideoChat: Chat-Centric Video Understanding
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab47944e-5706-4a95-8f8d-40b98bedb0f2 · inbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model VideoChat: Chat-Centric Video Understanding
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 349855bf-91db-4b2c-b76a-3e2ed6b978dd · inbound
CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval VideoChat: Chat-Centric Video Understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c3bf99c-f3e6-405f-a78d-f220407e7788 · inbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling VideoChat: Chat-Centric Video Understanding
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 225d188f-b04d-49e7-b30b-e3790e154e73 · inbound
Online Video Understanding: OVBench and VideoChat-Online VideoChat: Chat-Centric Video Understanding
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ce8e1bb-2c42-4e4a-9e91-a9918787bddf · inbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs VideoChat: Chat-Centric Video Understanding
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9447c53-292f-4a83-a700-71de2682f56d · inbound
Visual Large Language Models for Generalized and Specialized Applications VideoChat: Chat-Centric Video Understanding
Reference 146
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52152dbd-8747-428f-ad2a-f626394b2450 · inbound
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models VideoChat: Chat-Centric Video Understanding
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0218fc29-b73d-4169-9e62-bce6235bc050 · inbound
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos VideoChat: Chat-Centric Video Understanding
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 725677b8-7fb1-41e6-9608-fd9506f47da7 · inbound
H-MBA: Hierarchical MamBa Adaptation for Multi-Modal Video Understanding in Autonomous Driving VideoChat: Chat-Centric Video Understanding
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1730d6a1-7a2d-4750-9cbe-f79c1e5e14d0 · inbound
LongViTU: Instruction Tuning for Long-Form Video Understanding VideoChat: Chat-Centric Video Understanding
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08b776a3-e1f5-41b6-88f4-15a3c3c6c2a2 · inbound
LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding VideoChat: Chat-Centric Video Understanding
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2c5330c3-8e67-4e3d-a3e0-c4252774ad4f · inbound
VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video Captioning VideoChat: Chat-Centric Video Understanding
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e33504e-cdb5-4ea3-9099-cc2a3bfd6e84 · inbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness VideoChat: Chat-Centric Video Understanding
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 567c6c6b-9aa7-4829-8019-bd2f95cc1b56 · inbound
LeapVAD: A Leap in Autonomous Driving via Cognitive Perception and Dual-Process Thinking VideoChat: Chat-Centric Video Understanding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cbe9747-b6f5-453c-befc-6c2420fc9e32 · inbound
Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis VideoChat: Chat-Centric Video Understanding
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34fe1b59-3cca-4a37-b96a-1ea327675c95 · inbound
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding VideoChat: Chat-Centric Video Understanding
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 74613560-8f9d-42b4-8ac0-78b3b5004867 · inbound
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process VideoChat: Chat-Centric Video Understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4755349-b04c-4fda-8bd6-7c526c9e1545 · inbound
Temporal Preference Optimization for Long-Form Video Understanding VideoChat: Chat-Centric Video Understanding
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b1dd5db-994e-46ee-a79a-44ee7d9636a0 · inbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding VideoChat: Chat-Centric Video Understanding
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39898d34-2355-4ed0-82d8-51ca383af8ed · inbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler VideoChat: Chat-Centric Video Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7813284f-12ab-4ccb-bca4-624e0e947a5e · inbound
Understanding Long Videos via LLM-Powered Entity Relation Graphs VideoChat: Chat-Centric Video Understanding
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a29aa8fa-f608-47fe-bd84-ffa7172425ed · inbound
$\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation VideoChat: Chat-Centric Video Understanding
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58b706f8-00fe-41a4-8b9f-decae703afd1 · inbound
Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective VideoChat: Chat-Centric Video Understanding
Reference 139
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9573671c-4511-4233-b5bb-cc2db817550c · inbound
Survey on AI-Generated Media Detection: From Non-MLLM to MLLM VideoChat: Chat-Centric Video Understanding
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ff34654-d8bb-444b-98b8-ecc56fb6f257 · inbound
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey VideoChat: Chat-Centric Video Understanding
Reference 198
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 61c4b2fd-55ba-44d4-aff9-e9fe527c97fd · inbound
MusicInfuser: Making Video Diffusion Listen and Dance VideoChat: Chat-Centric Video Understanding
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 660889a1-6a46-4982-a916-8f1ef2e247ea · inbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning VideoChat: Chat-Centric Video Understanding
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 96dde0f3-38d0-406b-b44d-5595cf56c437 · inbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models VideoChat: Chat-Centric Video Understanding
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3f9b5974-c0ef-464f-93fd-7d8253ab8b18 · inbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation VideoChat: Chat-Centric Video Understanding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98ed9a7a-80ee-42db-9d19-2b2be5e63aed · inbound
RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language VideoChat: Chat-Centric Video Understanding
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faaad85f-8620-4213-b570-1c8ea8c4ffc9 · inbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion VideoChat: Chat-Centric Video Understanding
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8de95b1e-1fd5-40ea-b53d-1fe02be8f7fa · inbound
RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models VideoChat: Chat-Centric Video Understanding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 982636d1-66c4-4956-ba0f-a16b661028ac · inbound
Aggregated Structural Representation with Large Language Models for Human-Centric Layout Generation VideoChat: Chat-Centric Video Understanding
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 055477c3-d512-49c2-84aa-4834b945b15b · inbound
Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought VideoChat: Chat-Centric Video Understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57916962-adad-4b7f-9fd7-09b260c88146 · inbound
HuMoCon: Concept Discovery for Human Motion Understanding VideoChat: Chat-Centric Video Understanding
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bf2a2c0-36b0-4b24-abcb-c95a924b766b · inbound
Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? VideoChat: Chat-Centric Video Understanding
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 23b3972f-bd30-42bf-b519-bc4af4101549 · inbound
Adversarial Attacks against Closed-Source MLLMs via Feature Optimal Alignment VideoChat: Chat-Centric Video Understanding
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfb42be7-2600-43d1-bc3f-806577ce11ec · inbound
VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos VideoChat: Chat-Centric Video Understanding
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a7c88d5-139f-40ae-9010-2a4162c88189 · inbound
Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence VideoChat: Chat-Centric Video Understanding
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e02ec366-64f2-4304-8fe3-353a6c33a29a · inbound
Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence VideoChat: Chat-Centric Video Understanding
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 60b66c9c-95ff-4c99-aa4f-806c7132fb93 · inbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders VideoChat: Chat-Centric Video Understanding
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21d7e694-12db-4d82-a923-8fecfdde71d9 · inbound
DisTime: Distribution-based Time Representation for Video Large Language Models VideoChat: Chat-Centric Video Understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ddc9ed9-fd2d-4530-ba5f-289c8eae72c8 · inbound
VUDG: A Dataset for Video Understanding Domain Generalization VideoChat: Chat-Centric Video Understanding
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac08881b-48a1-4659-a175-f4a53b64c0d8 · inbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering VideoChat: Chat-Centric Video Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f13594f2-885c-4c3f-9901-677686589a63 · inbound
Period-LLM: Extending the Periodic Capability of Multimodal Large Language Model VideoChat: Chat-Centric Video Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db66c58f-dd38-4da6-85ac-4bbd1d2bfdf4 · inbound
Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues VideoChat: Chat-Centric Video Understanding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 096c6815-5245-45ce-baca-b8e6e4d6cf63 · inbound
ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding VideoChat: Chat-Centric Video Understanding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc6bae13-b3ab-420e-981c-09a913f748a7 · inbound
Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models VideoChat: Chat-Centric Video Understanding
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaedcbe4-f741-478f-a9e9-13c0ef7ab61c · inbound
Unraveling Spatio-Temporal Foundation Models via the Pipeline Lens: A Comprehensive Review VideoChat: Chat-Centric Video Understanding
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 360fda69-e337-43f1-8982-249ddb5c9bf0 · inbound
FormFactory: An Interactive Benchmarking Suite for Multimodal Form-Filling Agents VideoChat: Chat-Centric Video Understanding
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4f70289-08d2-46c8-ae70-8bcb2c2b402d · inbound
Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency VideoChat: Chat-Centric Video Understanding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5affa91c-1097-4287-9971-fc2a790710a4 · inbound
Vid-SME: Membership Inference Attacks against Large Video Understanding Models VideoChat: Chat-Centric Video Understanding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb72792c-376e-40b7-80b9-4ed91c6044b2 · inbound
DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding VideoChat: Chat-Centric Video Understanding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5bc3b82-eb86-4917-8094-c69ed13df073 · inbound
LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs VideoChat: Chat-Centric Video Understanding
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fdc5dd3-0d57-469d-bf4a-92a2918ac617 · inbound
EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? VideoChat: Chat-Centric Video Understanding
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9dc851b-57e1-45f1-978b-2206b840f82e · inbound
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning VideoChat: Chat-Centric Video Understanding
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7621811-7c51-449c-ae3c-e6cb18deddc6 · inbound
VideoMolmo: Spatio-Temporal Grounding Meets Pointing VideoChat: Chat-Centric Video Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7856dab-999d-4b83-a02d-48a4ba13866f · inbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing VideoChat: Chat-Centric Video Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1796d316-0616-4672-8205-6ed41d4acc80 · inbound
Pts3D-LLM: Studying the Impact of Token Structure for 3D Scene Understanding With Large Language Models VideoChat: Chat-Centric Video Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f837d44-eabf-4ebd-b8c0-9ed41b93dcd8 · inbound
Proactive Assistant Dialogue Generation from Streaming Egocentric Videos VideoChat: Chat-Centric Video Understanding
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0161f92-a0c5-4a8c-a35f-243b89c77a80 · inbound
Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding VideoChat: Chat-Centric Video Understanding
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20fd0780-27b5-4a90-872a-f5ccf4089a5d · inbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VideoChat: Chat-Centric Video Understanding
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c19432be-2e3c-44ab-82f5-9ad99d1453ae · inbound
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding VideoChat: Chat-Centric Video Understanding
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4bd1b8c-c516-46d8-a784-fb266f7e1a4f · inbound
EgoM2P: Egocentric Multimodal Multitask Pretraining VideoChat: Chat-Centric Video Understanding
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43752052-9380-4f73-bef0-41ec3dc11c36 · inbound
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning VideoChat: Chat-Centric Video Understanding
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dafab7fb-eddf-4b09-a8f5-610c60bdfd5c · inbound
TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision VideoChat: Chat-Centric Video Understanding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bca8184-4516-4ef8-9cd2-32287af706f8 · inbound
Outside Knowledge Conversational Video (OKCV) Dataset -- Dialoguing over Videos VideoChat: Chat-Centric Video Understanding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 382a5726-de98-4256-b7eb-0cd9cc1efef9 · inbound
SmartHome-Bench: A Comprehensive Benchmark for Video Anomaly Detection in Smart Homes Using Multi-Modal Large Language Models VideoChat: Chat-Centric Video Understanding
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd4ad7ce-5871-44c9-9b2f-0bc008ee470e · inbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model VideoChat: Chat-Centric Video Understanding
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0abc37a-0e0c-421d-b03e-a92d76fd7b81 · inbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning VideoChat: Chat-Centric Video Understanding
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ff9199b-effe-478f-8564-e0160dc4a7d3 · inbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations VideoChat: Chat-Centric Video Understanding
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc5d28cb-ac01-4903-b3b0-4cb656dca795 · inbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering VideoChat: Chat-Centric Video Understanding
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.