Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T06:30:22.431538Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 100 inbound Pith citation observations for arXiv:2307.06942.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T06:30:22.431538Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T04:37:15.144767Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T06:39:37.513098Z
82 of 82 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 95d1c26f-548a-4136-8067-c6cfcd01561d · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Language models are few-shot learners
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5bb18b84-cbd8-4449-be72-c518267fe6af · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dac10ac8-654b-4bfb-9d5d-f3000bf2472d · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Advancing high-resolution video-language representation with large-scale video transcriptions
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e5cafb74-dfd6-4c80-9962-45d18d7e4bb0 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Merlot: Multimodal neural script knowledge models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cdaf469d-3f9c-43f5-8234-d16b7499359d · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Merlot reserve: Neural script knowledge through vision and language and sound
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e40856e8-dc68-43e7-9f66-700ad2458cd8 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Frozen in time: A joint video and image encoder for end-to-end retrieval
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c1f02b5c-6bb4-47f6-9628-33cb6967eb42 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Flamingo: a Visual Language Model for Few-Shot Learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 88cba3f0-c804-4be0-8290-0d029175e7fe · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Openflamingo
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8af34b3e-4c11-4f9e-afe1-2eefefb5c0bc · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 14ee9171-1cfa-45b2-95a4-f627cf2fd942 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation VideoChat: Chat-Centric Video Understanding
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation feddb060-ac08-41fd-85b3-b92b79efb313 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0c548432-2503-4b28-aa4c-2660944622c1 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Yfcc100m: The new data in multimedia research
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a8b754b8-ef3f-4433-a0f2-1bb520f06616 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fd66b88d-e31f-4c93-a908-a2f9208d0830 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 848c0e15-0457-42db-b585-9264930a2f00 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Scaling up vision-language pre-training for image captioning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3febff2d-9d60-4b7d-80a6-5b27b63ba616 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation RedCaps: web-curated image-text data created by the people, for the people
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 11e2c600-e6e8-4053-bd5c-c219331e5efe · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 282f5ea3-eb52-430f-88ec-76a2c54de8b7 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Opendatalab: Empowering general artificial intelligence with open datasets
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fb00fe86-e1e5-4de8-baff-04656ad4b9e5 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation LAION-5B: An open large-scale dataset for training next generation image-text models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 940d0324-343e-4aa6-97eb-fc6be49eac82 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Wit: Wikipedia- based image text dataset for multimodal multilingual machine learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c0869013-6a87-41ef-a2fa-73804324fa1b · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Unmasked Teacher: Towards Training-Efficient Video Foundation Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 54d264cc-8c0c-47c5-aab0-0a567bbbf34b · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Learning audio-video modalities from image captions
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation eb79fada-a208-49ce-8d62-902b259f8fea · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation End-to-end learning of visual representations from uncurated instructional videos
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 48d86566-db18-4702-8d62-034dec52029b · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Learning Spatiotemporal Features via Video and Text Pair Discrimination
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation eed293b8-bc7f-4447-a1a7-672174934fd9 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 82c9ad0b-52d7-4917-8e8f-d6e77fa45beb · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b880b3eb-92b2-4175-a4e2-503d4694ffc4 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation An empirical study of training end-to-end vision-and- language transformers
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c46e81a1-d775-4997-a9d0-c1b7f6d1e4f4 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation How Much Can CLIP Benefit Vision-and-Language Tasks?
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0674e91f-b242-4d8b-aa76-cbcaa4141965 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a58f1f40-4620-4dae-b76e-87b7f3c29a00 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Murphy, and Cordelia Schmid
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation eeeb16d2-56d4-4b15-815b-ea6df4f5bb95 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Actbert: Learning global-local video-text representations
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dbc082f8-2ef1-41da-babd-92fd7526efe7 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e1b1d153-e310-46b5-bb3a-e651e945d2eb · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation InternVideo-Ego4D: A Pack of Champion Solutions to Ego4D Challenges
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 39c9893c-d8db-41f2-9ecb-70439aca86fe · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Learning transferable spatiotemporal representations from natural script knowledge
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3d770f5c-d9a9-4c27-9943-6e07d6fa0a35 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation TVTSv2: Learning Out-of-the-box Spatiotemporal Visual Representations at Scale
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c541cc23-02e0-4193-9ac5-4076649b7a59 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation VideoLLM: Modeling Video Sequence with Large Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 84278e2b-ac46-43a1-ad71-5666f994b8ae · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6b61d34b-dd18-43e6-8397-b55c889d3622 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Videomae v2: Scaling video masked autoencoders with dual masking
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 10fb84ee-7504-4aef-be19-cc3458e42a42 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation LAVENDER: Unifying Video-Language Understanding as Masked Language Modeling
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 44773781-818c-4857-ae18-e6dd84bce805 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation All in One: Exploring Unified Video-Language Pre-training
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 62b232a8-918e-44f4-bd7a-45d429e693fb · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 875e6e67-3a47-4203-818f-5958c8b8cda0 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e8ea9a48-177e-4db5-a889-4efe78cf0f9f · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7e688be5-ca48-41f7-aeaf-ac5f652c5230 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation VLAB: Enhancing Video Language Pre-training by Feature Adapting and Blending
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 099be21f-a0d2-4b88-af76-dc4d29561050 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Msr-vtt: A large video description dataset for bridging video and language
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bc37413a-8254-47bf-8784-37cdede5650c · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Localizing moments in video with natural language
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 61bc369c-f8b8-4d8b-9123-506988804bf2 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Movie description
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 92ea5ba5-243d-4604-9f6e-dd9bf7afb8ce · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Towards automatic learning of procedures from web instructional videos
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f9f03ebb-e0fe-41b3-9073-342c21be33aa · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation How2: A Large-scale Dataset for Multimodal Language Understanding
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1ae40008-4d4d-4eb7-a171-4dc497b4871d · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Dense-captioning events in videos
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 38a704ec-a40f-4f74-a192-1bc11d1d17ec · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Learning Video Representations from Textual Web Supervision
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 29d284a9-ffaa-4af5-8fc1-59f24e8d5a05 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Activitynet: A large- scale video benchmark for human activity understanding
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 75e68bc9-c824-4e4c-a217-b95e102d6610 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Quo vadis, action recognition? a new model and the kinetics dataset
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0ca669ae-6ce7-4b9d-8604-772cf490f5a9 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation something something
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d960859f-8771-43df-95e7-9f5cbc1efae0 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation On the effectiveness of task granularity for transfer learning
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 08f9679f-6c53-41b8-a7b7-83a71c8920e7 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d907cba7-fb7b-4aaf-bdfe-ebb1a6ffc708 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Co-grounding networks with semantic attention for referring expression comprehension in videos
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 242784b2-d58d-40a1-a6b8-dcf40b6c91d0 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Tubedetr: Spatio-temporal video grounding with transformers
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation aa9cd3df-61c6-42f3-8196-26777eaa1f70 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3b609533-6cb3-4a1a-95c7-23944a314510 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Revealing Single Frame Bias for Video-and-Language Learning
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7b59a083-4cb4-41c2-a172-3ac2e76c5913 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Tag2Text: Guiding Vision-Language Model via Image Tagging
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c01a9896-eef6-4f6c-87f9-d29f73179f02 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 74c720e5-13bc-458e-a066-05815af0ca20 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Gonzalez, Ion Stoica, and Eric P
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b2cfd191-bb24-47c8-acd9-cf79f305d177 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1cf8049d-d493-42f7-992f-bc9383f4ce99 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Language Is Not All You Need: Aligning Perception with Language Models
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 83655539-7804-4fc7-a7dd-c87582ab9c46 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 82431a86-ef52-4e7c-8c81-b99a31413f23 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Learning transferable visual models from natural language supervision
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 78b9f14c-977b-4333-9ef4-b4c9ba44604c · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation An image is worth 16x16 words: Transformers for image recognition at scale
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 332a12fc-907a-4aaa-b2a8-aeb1f0a7b9d8 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Representation Learning with Contrastive Predictive Coding
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f1159579-42ed-42a8-bc45-88216509d28e · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Flashattention: Fast and memory- efficient exact attention with io-awareness
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5b0763c5-fe42-473d-ae4b-031e297c053a · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation DataComp: In search of the next generation of multimodal datasets
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d75a59a5-7492-4341-8443-e0ed414b8f02 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c1537e4c-6d45-4028-9d4f-637e32f8178c · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4bef6936-55f5-457b-97e4-cc6a42a8ea06 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation A dataset for movie description
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 709b1691-9f44-4736-b18e-c5a8a860039c · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Collecting highly parallel data for paraphrase evaluation
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2731fb06-4ed2-4576-a267-d91340cf1486 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation High-resolution image synthesis with latent diffusion models
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 178147fe-575a-4c32-ab64-1d9112c9e7fe · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Align your latents: High-resolution video synthesis with latent diffusion models
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 84576d2a-14ed-4846-b78e-e9537e428f95 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3703360c-de3f-477f-a19a-db1a914ff5e0 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation MIMIC-IT: Multi-Modal In-Context Instruction Tuning
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e4eabf86-1565-4c35-b383-74844550e4c2 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video Generation
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6b5ba0e3-75ec-414d-9f00-f2c22bd3c324 · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Videofusion: Decomposed diffusion models for high-quality video generation
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9491e551-bfef-487e-a4b7-907886e0de7c · outbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Long video generation with time-agnostic vqgan and time-sensitive transformer
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 30cf6ee8-b393-4a98-a875-f59fc3f89e8e · inbound
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 153
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2cef23cf-dd78-421f-8fb5-c06e9eed97b9 · inbound
World Model on Million-Length Video And Language With Blockwise RingAttention InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 22cc7e27-cc01-4b20-8d2a-bcb0ed010feb · inbound
NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9915c743-fac1-41b1-8c82-2680e3f2a07e · inbound
VideoPhy: Evaluating Physical Commonsense for Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 106
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 254ea295-dc89-47aa-88d6-a4c06d781fba · inbound
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b1f9af37-9ea6-41d4-97e8-84f9661fd234 · inbound
Emu3: Next-Token Prediction is All You Need InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation adfe7ec5-250f-4a63-a73e-c224ec7f4af0 · inbound
DOLLAR: Few-Step Video Generation via Distillation and Latent Reward Optimization InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 310a8a70-87d0-4e22-a6fe-f963c762f638 · inbound
Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 180
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9fe933a-7450-4abe-84e2-758c87633611 · inbound
Multilevel Semantic-Aware Model for AI-Generated Video Quality Assessment InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a2033c9-7e3d-468f-8566-2ddd04618277 · inbound
STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation addab3fb-7910-4c90-b151-5e3a3adc0d1d · inbound
Tuning-Free Long Video Generation via Global-Local Collaborative Diffusion InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2db45f1e-d9ec-4042-84cf-291503112f36 · inbound
Training-Free Motion-Guided Video Generation with Enhanced Temporal Consistency Using Motion Consistency Loss InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8d8b3c2-c456-4eff-af15-513248c8a146 · inbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4b653dc5-c7c2-4fc4-b10e-e1a3f159239a · inbound
Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9559f07b-405e-443b-a987-5ea42ef568f3 · inbound
Video Latent Flow Matching: Optimal Polynomial Projections for Video Interpolation and Extrapolation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18044e88-1e2a-46ac-bbdf-2208eaf7b3f0 · inbound
Goku: Flow Based Video Generative Foundation Models InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8652f3e-296d-4467-9c6a-28b536d0992c · inbound
TripoSG: High-Fidelity 3D Shape Synthesis using Large-Scale Rectified Flow Models InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 127
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation af0bf21b-6be7-4169-b4fe-d7b0c07e597d · inbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8a76d122-1a03-4998-9597-ac78e3d31d83 · inbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42352660-1d21-46ab-8edc-e8126a2874be · inbound
RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fde07f22-2d32-4440-ba1c-660bc9eb2124 · inbound
InfLVG: Reinforce Inference-Time Consistent Long Video Generation with GRPO InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8daffc84-86b1-4a83-8ec4-10a2186913ec · inbound
OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d3d1fc9-29a7-43ef-ba22-2208dff8b6e8 · inbound
Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion Transformers InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 213c9c4c-bd20-43cf-8c8d-d90308fb481e · inbound
DisTime: Distribution-based Time Representation for Video Large Language Models InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9559444-4464-46cf-b125-2bc3146cfffa · inbound
AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e871c747-e2c3-4241-9229-d221f6bd513f · inbound
Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 260a8e1d-40b1-41b9-8a9a-644b5b27499b · inbound
ContentV: Efficient Training of Video Generation Models with Limited Compute InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dea69658-9404-4c83-ad71-b20e3413cc57 · inbound
Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b53b07d8-dadc-49c1-b29b-1c4f81601bb7 · inbound
A Watermark for Auto-Regressive Image Generation Models InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a27074b-2994-4a4c-a2fd-60636dec451c · inbound
Fake it till You Make it: Reward Modeling as Discriminative Prediction InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2aabb4b-3746-44aa-a81c-b58cff382e30 · inbound
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f0879d8-0e8e-4004-86de-263b2740616c · inbound
CI-VID: A Coherent Interleaved Text-Video Dataset InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27f8fedf-7160-4197-ae13-899b34717564 · inbound
Semantic Frame Interpolation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4545fb88-e229-402d-a9d1-46d9a3e1f23b · inbound
LoViC: Efficient Long Video Generation with Context Compression InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59763ee7-6e67-4e6d-b335-f5a1fe341371 · inbound
Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d0b065b-b671-4a08-90de-556ab4257c7a · inbound
VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85d6e32a-0eb4-48bd-b14e-ec1ec0be758d · inbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e61e9f2e-3702-4292-8b84-2a4633fad0de · inbound
Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a967589a-3573-4762-9131-6d72587aaac4 · inbound
B4DL: A Benchmark for 4D LiDAR LLM in Spatio-Temporal Understanding InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 17bc81e9-9c7c-408d-a652-5150043f612e · inbound
TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b12fb950-5c87-4313-9cd8-23b23ac671e4 · inbound
Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1fe35c4-46c6-4983-9916-c3608a9d0417 · inbound
From Black Box to Transparency: Enhancing Automated Interpreting Assessment with Explainable AI in College Classrooms InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bf63ee5-3352-4211-a8ea-1e88f6c7b3f0 · inbound
A Survey on Video Temporal Grounding with Multimodal Large Language Model InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 140
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a89c0a1-e112-4afe-8afb-42f563ea1a3a · inbound
Effectively obtaining acoustic, visual and textual data from videos InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 116
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91e45fc0-417d-4d4c-96fd-6edd7b3dcf78 · inbound
Video Understanding by Design: How Datasets Shape Video Models InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 176
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4474b6c3-8f6c-4003-add1-d6376784c4be · inbound
Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e939595a-26c5-4c43-b8ed-3a0572510185 · inbound
VC-Inspector: Advancing Reference-free Evaluation of Video Captions with Factual Analysis InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation aee86242-127c-4615-a89d-8336e02dae48 · inbound
MVAD: A Benchmark Dataset for Multimodal AI-Generated Video-Audio Detection InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4d45bc8b-0946-43aa-a9d9-501499f26280 · inbound
VABench: A Comprehensive Benchmark for Audio-Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation da4facba-bf19-4448-a034-e86ebbe38fb5 · inbound
Adapting MLLMs for Nuanced Video Retrieval InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7ffc17f1-87cf-46b9-890f-4755ca5b67c0 · inbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 865c0116-97e4-4f43-95e6-330b4bc095c6 · inbound
Efficient Scaling of LLM Training with Flexible Context Parallelism InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8244475b-1246-4855-a1dd-e3fa90962bbc · inbound
TrajTok: Learning Trajectory Tokens enables better Video Understanding InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e60ffcc2-33f3-4934-8738-43266ad76682 · inbound
TrajTok: Learning Trajectory Tokens enables better Video Understanding InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0ff4ce5-78ba-4e71-b9a6-eff494c887c4 · inbound
Insights from Visual Cognition: Understanding Human Action Dynamics with Overall Glance and Refined Gaze Transformer InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b849b742-36a3-4513-883f-3057e2185568 · inbound
MotionScape: A Large-Scale Real-World Highly Dynamic UAV Video Dataset for World Models InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0c625605-0c7a-4735-af4e-b9ec9d873b1f · inbound
InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0f891e84-a2a2-42fb-a7b7-fb9efba14814 · inbound
InstrAct: Towards Action-Centric Understanding in Instructional Videos InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5b7f40ef-af3c-4425-bdb8-671a05b9ca51 · inbound
How Should Video LLMs Output Time? An Analysis of Efficient Temporal Grounding Paradigms InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2c0d6f8c-470e-411d-9c68-42a0f99148c6 · inbound
LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 179
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 742815fa-753f-4906-89a3-f07e4c404c68 · inbound
UniMesh: Unifying 3D Mesh Understanding and Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7dcbb625-6023-46af-a3da-cc78a202cc6d · inbound
Seeing Fast and Slow: Learning the Flow of Time in Videos InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 447ca9d5-c470-4332-a75a-34cd3b686a02 · inbound
MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c2bdacff-4003-4031-b5e5-0c12abde9be9 · inbound
MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3f6d8964-8e9a-444a-af15-d1d4f59abbff · inbound
Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dc920f33-c18d-4b0c-9ff0-e691a0cafe1e · inbound
DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7f9d89da-9e81-46ff-81cc-d196e3df9427 · inbound
TMD-Bench: A Multi-Level Evaluation Paradigm for Music-Dance Co-Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8ff7a753-c4b2-42bf-8bb2-7175ff015597 · inbound
OZ-TAL: Online Zero-Shot Temporal Action Localization InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0994a83c-63ec-4923-a22f-f68a8437b17c · inbound
CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8dc8fe02-ae02-4b84-9d97-9f945aa585c6 · inbound
GTA: Advancing Image-to-3D World Generation via Geometry Then Appearance Video Diffusion InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 724a98c9-12e5-45fb-8535-f0d10359438c · inbound
TeDiO: Temporal Diagonal Optimization for Training-Free Coherent Video Diffusion InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 31da3f52-8af4-4f63-8e02-7696c037be8f · inbound
Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video Reconstruction InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e3c78d92-5a10-434d-986d-f9a4f6fce380 · inbound
HAVEN: Hierarchically Aligned Multimodal Benchmark for Unified Video Understanding InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8fcd00f2-b360-4a31-8a50-1836b20ec72d · inbound
Aero-World: Action-Conditioned Aerial Video Generation from Inertial Controls InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 983cb006-f158-4bb8-8bb6-e9c1df261456 · inbound
EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a34e84b6-6b7c-415e-a663-a463f083e867 · inbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b9a67c7d-e0c5-4340-b268-c14f00cfca16 · inbound
PhyGenHOI: Physically-Aware 4D Generation of Dynamic Human-Object Interactions InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4b773ed1-f8c1-4eca-a18f-b0b2eb3177c3 · inbound
Collaborative Few-Step Distillation and Low-Bit Quantization for Wan2.2 Dual-Expert Video Diffusion Models InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d5cb3a71-66a6-4d92-be6d-bd5d0ad139e2 · inbound
VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 69aac290-bb62-4c8b-bb0b-55ea4a9876e1 · inbound
ViMax: Agentic Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 47973b93-7910-4b31-8771-bccf6678b9ec · inbound
ViMax: Agentic Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23b586ec-97ca-436d-a597-73fc5c754578 · inbound
CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f614416f-0818-47b8-8dae-1f0979eeb8a7 · inbound
SpecLoR: Spectral Lookahead Rectification for Motion-Coherent Text-to-Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dd62c66e-c7f5-4afc-8aaf-da23ca716ddc · inbound
MLT-Dedup: Efficient Large-Scale Online Video Deduplication via Multi-Level Representations and Spatial-Temporal Matching InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6c608ea4-a649-44f6-bf37-4a5cb9fbff96 · inbound
GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ffce0e3c-5d26-47a0-a99c-3e810f98f6c5 · inbound
GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c94e94a9-48f4-4772-9768-39ac77418ff8 · inbound
HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 13c6f3a5-74aa-4e65-b246-c1be83f2fecb · inbound
MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c611cf4c-bfe3-45e4-9a0a-522ce4031e0e · inbound
MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99b8ea70-9a6f-41fc-a2e7-9b7ae6699945 · inbound
Incentivizing Vision Language Models to Search for Long Video Question Answering InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b5a1d80-68c3-46b4-9643-65678c44f226 · inbound
MentalThink: Shaping Thoughts in Mental SVG World InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 154
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 885445ab-eacf-483f-a953-47177b84ab8c · inbound
Reinforcement Learning: From Algorithms To Foundation Models InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03137473-91ef-4984-9fa1-c623d1cf536b · inbound
ShotPlan: Cinematic Video Generation with Learnable Planning Token InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c439224-b816-4c81-a181-933ffe2909de · inbound
Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0f8e8f5-cf37-4d83-9e26-a66c6e860f4a · inbound
MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9450afd1-f176-4fce-86ea-0b9877e2b8d1 · inbound
EgoPlay: Event-Triggered Video Editing for Egocentric Streams InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90122d2e-2650-4361-8e45-90c25df1d96c · inbound
Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb1f23e9-3732-4a02-a01e-8c4497451a29 · inbound
CineWeaver: Training-Free Reference-Controllable Multi-Shot Long Video Generation for Cinematic Storytelling InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2af81da1-5a35-4ec6-8f31-078701ba5f68 · inbound
Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a083e30-7eb0-436f-820f-6f0b30085e5d · inbound
4DHumanDiff: Direct Text-to-4DGS Generation for Consistent 360-Degree Dynamic Humans InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 147
Source-reported events for the cited work
Unavailable: canonical work link unavailable.