Pith. sign in

Paper Citation Record · LEDGER

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

As of 11 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 100 inbound Pith citation observations for arXiv:2307.06942.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.06942 v2

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T06:30:22.431538Z

measured 182 of 182 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 100 of 102 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:37:15.144767Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T06:39:37.513098Z

Reference resolution

82 of 82 outbound references displayed

  • verified exact36
  • verified fuzzy43
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 95d1c26f-548a-4136-8067-c6cfcd01561d · outbound

This paper cites Language models are few-shot learners.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Language models are few-shot learners

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.847582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:662fdaeb49c9c34b74883c56898716f09927f85486543fe0aabc42992b768ff2

Observation 5bb18b84-cbd8-4449-be72-c518267fe6af · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.767085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:397c33301c28cf1a914a985ea7d91931da9dcd8f4ece7c635648397b10e814c7

Observation dac10ac8-654b-4bfb-9d5d-f3000bf2472d · outbound

This paper cites Advancing high-resolution video-language representation with large-scale video transcriptions.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Advancing high-resolution video-language representation with large-scale video transcriptions

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.771097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:f3c86253184244292a5fd78ddaf168d90f5005a4738d15de1e3a681489eb49b4

Observation e5cafb74-dfd6-4c80-9962-45d18d7e4bb0 · outbound

This paper cites Merlot: Multimodal neural script knowledge models.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Merlot: Multimodal neural script knowledge models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.774776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:520ec47521f02e64b79842151ed072151ee03f059fc9983331b8312bcdd6a3fc

Observation cdaf469d-3f9c-43f5-8234-d16b7499359d · outbound

This paper cites Merlot reserve: Neural script knowledge through vision and language and sound.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Merlot reserve: Neural script knowledge through vision and language and sound

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.779043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:ae29535d9ee896e0634a14fcb86ac89073cd7ca63fa03b8b34a553bf6b75623a

Observation e40856e8-dc68-43e7-9f66-700ad2458cd8 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.783527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:598a24fb2c8a599b109761c4837afbd8b736f3f868c0191ac4dda024d27e711a

Observation c1f02b5c-6bb4-47f6-9628-33cb6967eb42 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Flamingo: a Visual Language Model for Few-Shot Learning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:30:22.494018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:45bb0c66aa2320cc531874ebb1162043ea79ae3df4c641e8c4f9ba1832bdba43

Observation 88cba3f0-c804-4be0-8290-0d029175e7fe · outbound

This paper cites Openflamingo.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Openflamingo

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.791306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:242e9d33914e3f905a4663c9e027323fe862c38b064c4c61f90a5a6544428436

Observation 8af34b3e-4c11-4f9e-afe1-2eefefb5c0bc · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:30:22.502036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:753b5b5d38f35f906d186ea768425b3e3123eaa5595448821e3056c1086f9b59

Observation 14ee9171-1cfa-45b2-95a4-f627cf2fd942 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation VideoChat: Chat-Centric Video Understanding

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:30:22.509068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:121f03635b774379e7d77ab25b010e96e5c36025c8426dde0e83935c869260e2

Observation feddb060-ac08-41fd-85b3-b92b79efb313 · outbound

This paper cites InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.516674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:df1d451cb451d058255545c5666b99ed5a1e69612286f92f9bb5fe211607444e

Observation 0c548432-2503-4b28-aa4c-2660944622c1 · outbound

This paper cites Yfcc100m: The new data in multimedia research.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Yfcc100m: The new data in multimedia research

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.806807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:3edb5efff6d987af1e0e9260bc4010c3af526b83783e5911a1dc661e74bafc1e

Observation a8b754b8-ef3f-4433-a0f2-1bb520f06616 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.811639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:4f4cef10d44daf00e7006aac6e036d4f908b82c249f0a85c96a0ea14432637ad

Observation fd66b88d-e31f-4c93-a908-a2f9208d0830 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.815798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:874c5b1cb790ca975e8b400085ef245a3f092a67412f6da2aa53a640b10cb209

Observation 848c0e15-0457-42db-b585-9264930a2f00 · outbound

This paper cites Scaling up vision-language pre-training for image captioning.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Scaling up vision-language pre-training for image captioning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.820372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:538047a22137b36d4a6a9728d859588f13968068a59cb53b2cfbc7b3c4285015

Observation 3febff2d-9d60-4b7d-80a6-5b27b63ba616 · outbound

This paper cites RedCaps: web-curated image-text data created by the people, for the people.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation RedCaps: web-curated image-text data created by the people, for the people

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.525115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:26312f57c8225aeecec7cf74f60ca7a840d10994fae4eeb1c0252ab6451d8341

Observation 11e2c600-e6e8-4053-bd5c-c219331e5efe · outbound

This paper cites WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.533004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:ee8656d0dbffb88a5e095deb85bdab233e2b191602bfb41a6eb2d26e3d64acba

Observation 282f5ea3-eb52-430f-88ec-76a2c54de8b7 · outbound

This paper cites Opendatalab: Empowering general artificial intelligence with open datasets.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Opendatalab: Empowering general artificial intelligence with open datasets

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.833785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:63bc7025de5bba0b5434b345a75c52030c005cecdf9b63b31f74dbffa213b6c0

Observation fb00fe86-e1e5-4de8-baff-04656ad4b9e5 · outbound

This paper cites LAION-5B: An open large-scale dataset for training next generation image-text models.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation LAION-5B: An open large-scale dataset for training next generation image-text models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:30:22.539337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:e37f358d141e2f1c94b2e963a148b0f3f138d133bb7111cabd6d3dee28fa77d0

Observation 940d0324-343e-4aa6-97eb-fc6be49eac82 · outbound

This paper cites Wit: Wikipedia- based image text dataset for multimodal multilingual machine learning.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Wit: Wikipedia- based image text dataset for multimodal multilingual machine learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.843535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:fd27929725323fbf692dce8a1fbf5caa1755f86bf20d754e8ef53851f37a5d17

Observation c0869013-6a87-41ef-a2fa-73804324fa1b · outbound

This paper cites Unmasked Teacher: Towards Training-Efficient Video Foundation Models.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Unmasked Teacher: Towards Training-Efficient Video Foundation Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:30:22.486478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:62a95786b62ee3c3397bd4c707e855ec5e57d4b0b708dc050988cc7fae265c51

Observation 54d264cc-8c0c-47c5-aab0-0a567bbbf34b · outbound

This paper cites Learning audio-video modalities from image captions.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Learning audio-video modalities from image captions

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.851443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:7a4a273fa7f4f1f54a4cb255832ef6016bf01f41fb9b537d983664da7bea09a6

Observation eb79fada-a208-49ce-8d62-902b259f8fea · outbound

This paper cites End-to-end learning of visual representations from uncurated instructional videos.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation End-to-end learning of visual representations from uncurated instructional videos

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.856294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:242d56398a5c3f8b8ca26a6df38780dd2297964aa0491f249d75a31e37833651

Observation 48d86566-db18-4702-8d62-034dec52029b · outbound

This paper cites Learning Spatiotemporal Features via Video and Text Pair Discrimination.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Learning Spatiotemporal Features via Video and Text Pair Discrimination

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.546904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:e6df326d1c61023731ca608a8e4b0e43c21b89f4e4278dc271e537f7a98cea28

Observation eed293b8-bc7f-4447-a1a7-672174934fd9 · outbound

This paper cites VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.554543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:cec0e8cedc3a003bda22f03a924e0f8386ef8fb316ba80a9605f179ecad7163d

Observation 82c9ad0b-52d7-4917-8e8f-d6e77fa45beb · outbound

This paper cites UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.561008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:5b7c4a2714cd58e6b57d1441c290ea4ae81228cffae26c9ac356a7a43f66a926

Observation b880b3eb-92b2-4175-a4e2-503d4694ffc4 · outbound

This paper cites An empirical study of training end-to-end vision-and- language transformers.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation An empirical study of training end-to-end vision-and- language transformers

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.873275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:896e838b253338adc34ce6ca7d4ce80f4360313aa1a8ff38b3d9a706fa047b1c

Observation c46e81a1-d775-4997-a9d0-c1b7f6d1e4f4 · outbound

This paper cites How Much Can CLIP Benefit Vision-and-Language Tasks?.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation How Much Can CLIP Benefit Vision-and-Language Tasks?

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.567453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:2a94f59ccfe17abd157ef60776d22a6568dd2c9462d171868ce9dc963d4070d1

Observation 0674e91f-b242-4d8b-aa76-cbcaa4141965 · outbound

This paper cites FILIP: Fine-grained Interactive Language-Image Pre-Training.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.575279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:fe04c87ee1984a0161cf2394830f1c8c91ab6a287cd4961f81f1026741b5b984

Observation a58f1f40-4620-4dae-b76e-87b7f3c29a00 · outbound

This paper cites Murphy, and Cordelia Schmid.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Murphy, and Cordelia Schmid

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.886345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:8b88f0db59dbe701ad4917c37d9eab225607a6f4826a6992e01c42a6def34b73

Observation eeeb16d2-56d4-4b15-815b-ea6df4f5bb95 · outbound

This paper cites Actbert: Learning global-local video-text representations.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Actbert: Learning global-local video-text representations

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.890487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:b74b7fac160cbbc3f49be89b063aac57f02e101f76be1d5fa8eabe62b16e4ab0

Observation dbc082f8-2ef1-41da-babd-92fd7526efe7 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.574468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:f46f2e7480578b1d1d1861df0b95aca8a553c3d157c309217fb951ef2f8e13f9

Observation e1b1d153-e310-46b5-bb3a-e651e945d2eb · outbound

This paper cites InternVideo-Ego4D: A Pack of Champion Solutions to Ego4D Challenges.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation InternVideo-Ego4D: A Pack of Champion Solutions to Ego4D Challenges

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.588303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:ca32e443bf3448768a3889c61e8d2f8febbfdd73566189a86361e7569e3d594e

Observation 39c9893c-d8db-41f2-9ecb-70439aca86fe · outbound

This paper cites Learning transferable spatiotemporal representations from natural script knowledge.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Learning transferable spatiotemporal representations from natural script knowledge

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.905098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:f37d79feb37a43b76c840daa7f02be035d0a627f9aeb2adb609ce971b395123d

Observation 3d770f5c-d9a9-4c27-9943-6e07d6fa0a35 · outbound

This paper cites TVTSv2: Learning Out-of-the-box Spatiotemporal Visual Representations at Scale.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation TVTSv2: Learning Out-of-the-box Spatiotemporal Visual Representations at Scale

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:30:22.594600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:87ec38f9e048ef01f45b608743d4409869e19592282ab8f29c2a0b158c523c1c

Observation c541cc23-02e0-4193-9ac5-4076649b7a59 · outbound

This paper cites VideoLLM: Modeling Video Sequence with Large Language Models.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation VideoLLM: Modeling Video Sequence with Large Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.601941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:bddc63964440f508ee7b8a6a7d932a40e5d604a7e9dab0e51d6cfa8089738fc9

Observation 84278e2b-ac46-43a1-ad71-5666f994b8ae · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.919086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:a8f2ca1ae73274baf986cb34495ab49802bb913814f6968ae173cfdc0ef1cf83

Observation 6b61d34b-dd18-43e6-8397-b55c889d3622 · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Videomae v2: Scaling video masked autoencoders with dual masking

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.923758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:e8f393806b410081405c53e8dcaa5e203a5908164f2d2e32adf04ef887f4f54f

Observation 10fb84ee-7504-4aef-be19-cc3458e42a42 · outbound

This paper cites LAVENDER: Unifying Video-Language Understanding as Masked Language Modeling.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation LAVENDER: Unifying Video-Language Understanding as Masked Language Modeling

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.608373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:b021ff51c491959523997632ceaec531ef1af14d48b3e6c16606816dd86058f6

Observation 44773781-818c-4857-ae18-e6dd84bce805 · outbound

This paper cites All in One: Exploring Unified Video-Language Pre-training.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation All in One: Exploring Unified Video-Language Pre-training

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.616379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:5ed316b7562c8c938e2f8dfe335caf24fdedc055bea04bb824f61009656e451e

Observation 62b232a8-918e-44f4-bd7a-45d429e693fb · outbound

This paper cites VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.623094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:1cfb111f230aad2c1d4fc5a4ad9def9b2db80856d1ef2701a6295af773bc937d

Observation 875e6e67-3a47-4203-818f-5958c8b8cda0 · outbound

This paper cites VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.630119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:71b4117bc2ecb03326ba63ee8f23fa449c70300e9632b4385feed95c24f827cc

Observation e8ea9a48-177e-4db5-a889-4efe78cf0f9f · outbound

This paper cites mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.637374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:a0d036e8a506afdfaac33b4424d9ce6b8e52d5bfdfd50f0732ff0971c3eacf9e

Observation 7e688be5-ca48-41f7-aeaf-ac5f652c5230 · outbound

This paper cites VLAB: Enhancing Video Language Pre-training by Feature Adapting and Blending.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation VLAB: Enhancing Video Language Pre-training by Feature Adapting and Blending

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.643920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:6766303cf8dd996e8601caabd3cd02fd977044af94b4b2d627cca40fb001e194

Observation 099be21f-a0d2-4b88-af76-dc4d29561050 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Msr-vtt: A large video description dataset for bridging video and language

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.763014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:b91a7e47d3ff8a3f7725ba445cc6867b0d47012cbaeb7b2b1000607da43e1eb6

Observation bc37413a-8254-47bf-8784-37cdede5650c · outbound

This paper cites Localizing moments in video with natural language.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Localizing moments in video with natural language

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.787448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:d85a10a033c24aba4d0379a0638bef0cf85fa02a1029fe3bef3d6fa9e2d173f0

Observation 61bc369c-f8b8-4d8b-9123-506988804bf2 · outbound

This paper cites Movie description.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Movie description

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.795340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:552b39182a148c2294540a1483246bbad10e4d68b6b13ce10383e512fd2a109d

Observation 92ea5ba5-243d-4604-9f6e-dd9bf7afb8ce · outbound

This paper cites Towards automatic learning of procedures from web instructional videos.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Towards automatic learning of procedures from web instructional videos

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.799298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:9cd4df60361b6ab5e17d5e14464a8f2cfc1a255901a7d1d928f12f5312040630

Observation f9f03ebb-e0fe-41b3-9073-342c21be33aa · outbound

This paper cites How2: A Large-scale Dataset for Multimodal Language Understanding.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation How2: A Large-scale Dataset for Multimodal Language Understanding

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:30:22.650019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:b4f4ae004c67674dc4fc883c66acf620cf81e87204183b36bd005b1d7d2e262b

Observation 1ae40008-4d4d-4eb7-a171-4dc497b4871d · outbound

This paper cites Dense-captioning events in videos.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Dense-captioning events in videos

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.824583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:654f0bb629fda17bcd84b5e4dd752bc09e115799a7ce78848b2a7dc10e1aaa87

Observation 38a704ec-a40f-4f74-a192-1bc11d1d17ec · outbound

This paper cites Learning Video Representations from Textual Web Supervision.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Learning Video Representations from Textual Web Supervision

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.656697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:a1f391521fb9138622f1735628317adc845f49f63901bdb82ab89da178381e0e

Observation 29d284a9-ffaa-4af5-8fc1-59f24e8d5a05 · outbound

This paper cites Activitynet: A large- scale video benchmark for human activity understanding.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Activitynet: A large- scale video benchmark for human activity understanding

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.839338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:8a9e880640f95de06dbee55235762e169713542b97ebefe02e5e7805fb722ea6

Observation 75e68bc9-c824-4e4c-a217-b95e102d6610 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Quo vadis, action recognition? a new model and the kinetics dataset

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.861011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:c0be920b1679bc993cf16812932e4b4d235f45944304e78a6d7399308d68e925

Observation 0ca669ae-6ce7-4b9d-8604-772cf490f5a9 · outbound

This paper cites something something.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation something something

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.864884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:1e68a6b792017eca2f76355e1f4a2b9bdcc3704d6a836c816ffb5b07633527a5

Observation d960859f-8771-43df-95e7-9f5cbc1efae0 · outbound

This paper cites On the effectiveness of task granularity for transfer learning.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation On the effectiveness of task granularity for transfer learning

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:30:22.662786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:6e43b6d5f3a7802ef6f2991cce3e86b9c2375db1d9cf8242d8dbaa7c3b13cef8

Observation 08f9679f-6c53-41b8-a7b7-83a71c8920e7 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:30:22.669014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:b2f72031fe6e2a553b71370b9255096d6543ed3baeb3ff642a0e9680f54a8397

Observation d907cba7-fb7b-4aaf-bdfe-ebb1a6ffc708 · outbound

This paper cites Co-grounding networks with semantic attention for referring expression comprehension in videos.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Co-grounding networks with semantic attention for referring expression comprehension in videos

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.882252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:b269735be8db3b1c3049eccdce9e405815ae0669d40c2d26e444bab528a32515

Observation 242784b2-d58d-40a1-a6b8-dcf40b6c91d0 · outbound

This paper cites Tubedetr: Spatio-temporal video grounding with transformers.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Tubedetr: Spatio-temporal video grounding with transformers

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.894940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:ecc6940587e09f227da888ff873d9d2badcb973aea289118e456672c99f33cbe

Observation aa9cd3df-61c6-42f3-8196-26777eaa1f70 · outbound

This paper cites an unresolved cited work.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-05-15T06:30:22.899592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:1bbe79d963e01e56471d1180b746d6c588a7c6b6ce40a14654cdc4a19d38d138

Observation 3b609533-6cb3-4a1a-95c7-23944a314510 · outbound

This paper cites Revealing Single Frame Bias for Video-and-Language Learning.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Revealing Single Frame Bias for Video-and-Language Learning

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.675262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:51c670e18f9122b8c65afd74bd713e02e07baf892e1aee7f4ab79ec041bc0a2a

Observation 7b59a083-4cb4-41c2-a172-3ac2e76c5913 · outbound

This paper cites Tag2Text: Guiding Vision-Language Model via Image Tagging.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Tag2Text: Guiding Vision-Language Model via Image Tagging

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.682037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:1cadac515276e80165936d965a583ad797aedcd32ca162579fcaeb16c137d438

Observation c01a9896-eef6-4f6c-87f9-d29f73179f02 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.740346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:e9da4420c2a398942320f1b7e196a5bb154c338d9ba73887263a98226d47f0e2

Observation 74c720e5-13bc-458e-a066-05815af0ca20 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Gonzalez, Ion Stoica, and Eric P

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.744175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:f109146f7b5713af14a8b7d8900fe044dea8554499085dc57059c4b90b4a710c

Observation b2cfd191-bb24-47c8-acd9-cf79f305d177 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:30:22.687587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:c90b0a37fc7f5e0f090bd3d39008acfbacd8eb425403bbbda1e91efda2d024fc

Observation 1cf8049d-d493-42f7-992f-bc9383f4ce99 · outbound

This paper cites Language Is Not All You Need: Aligning Perception with Language Models.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Language Is Not All You Need: Aligning Perception with Language Models

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:32:23.026112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:abe0ca889b2f803443b8c0d5baf0965194ccaeca09d5fd50e4cf2450c5d9e243

Observation 83655539-7804-4fc7-a7dd-c87582ab9c46 · outbound

This paper cites Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.700766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:fead1d816e1c596fb964f540ecb3debe18ff73c8e084bcbf5a2925391d31fe07

Observation 82431a86-ef52-4e7c-8c81-b99a31413f23 · outbound

This paper cites Learning transferable visual models from natural language supervision.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Learning transferable visual models from natural language supervision

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.759289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:2649a2819d00cfd9685b7a33c2f69be7c9718d31c5bf656ef975273fb9376af4

Observation 78b9f14c-977b-4333-9ef4-b4c9ba44604c · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation An image is worth 16x16 words: Transformers for image recognition at scale

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.803274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:c2f0d366ce7fd122dbd9d7d87ffb9f97a7efea26e3ae43b1c51eb48f1871cac9

Observation 332a12fc-907a-4aaa-b2a8-aeb1f0a7b9d8 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Representation Learning with Contrastive Predictive Coding

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:30:22.706671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:f9bb0b94de26bd8978171024f2bfd3938b14cebc59bc2a93538916299eca09d6

Observation f1159579-42ed-42a8-bc45-88216509d28e · outbound

This paper cites Flashattention: Fast and memory- efficient exact attention with io-awareness.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Flashattention: Fast and memory- efficient exact attention with io-awareness

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.868754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:c0c8b45df1a599a82d5f1745f3c56c219b3c523c9ef135a7348f1f60901865d5

Observation 5b0763c5-fe42-473d-ae4b-031e297c053a · outbound

This paper cites DataComp: In search of the next generation of multimodal datasets.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation DataComp: In search of the next generation of multimodal datasets

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.712246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:f803c85f784e93875cc97ef365686bf60de5b4fc6dfefe7c00ac9fc434fdeb12

Observation d75a59a5-7492-4341-8443-e0ed414b8f02 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:30:22.717922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:beecbba72c6a73d535a4826d463f8d2e23f54d74c3fb2c8278998c057f847d2f

Observation c1537e4c-6d45-4028-9d4f-637e32f8178c · outbound

This paper cites Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.914480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:c8f32071839eb5dee3096a3e07092726a0bce30e32b9a417d682070b6a94d59c

Observation 4bef6936-55f5-457b-97e4-cc6a42a8ea06 · outbound

This paper cites A dataset for movie description.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation A dataset for movie description

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.747968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:4601b5493b1d0e9d7a5b74aad5ae59d177eef7c0283de5ee5b551c94e68d3c45

Observation 709b1691-9f44-4736-b18e-c5a8a860039c · outbound

This paper cites Collecting highly parallel data for paraphrase evaluation.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Collecting highly parallel data for paraphrase evaluation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.751671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:928f0c4653a332f97a0d77be99efd4204a3bd75b594a2de8ed684c8efd3efbe8

Observation 2731fb06-4ed2-4576-a267-d91340cf1486 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation High-resolution image synthesis with latent diffusion models

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.755378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:4544df71261ed1f5bd9a4d6f729fdfbb0abcd406116965209ac2f6426766d8c6

Observation 178147fe-575a-4c32-ab64-1d9112c9e7fe · outbound

This paper cites Align your latents: High-resolution video synthesis with latent diffusion models.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Align your latents: High-resolution video synthesis with latent diffusion models

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.828792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:a2455abb17331f7e4c4d15f8d11919c6534db58c0be7f9fc01b479b8f6459bba

Observation 84576d2a-14ed-4846-b78e-e9537e428f95 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:30:22.723823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:cf1443c16df7015dc8ceab0f07b1e53e30bf2ad6303d2830f715aae466e362c1

Observation 3703360c-de3f-477f-a19a-db1a914ff5e0 · outbound

This paper cites MIMIC-IT: Multi-Modal In-Context Instruction Tuning.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.729805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:213edc82dba61fcea2eb5f5a728335941f817c4d79f0b0b93f9f3afbc0e9cc6c

Observation e4eabf86-1565-4c35-b383-74844550e4c2 · outbound

This paper cites Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video Generation.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video Generation

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.735873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:0ac20b4053d406237ba6ae566c7936099fc96bc672b4310d4fd3a12d0af37a1c

Observation 6b5ba0e3-75ec-414d-9f00-f2c22bd3c324 · outbound

This paper cites Videofusion: Decomposed diffusion models for high-quality video generation.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Videofusion: Decomposed diffusion models for high-quality video generation

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.910163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:2320afb04d33a5d352f6bac22260e15da4ee2b4d55fffa647678184d7688e21f

Observation 9491e551-bfef-487e-a4b7-907886e0de7c · outbound

This paper cites Long video generation with time-agnostic vqgan and time-sensitive transformer.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Long video generation with time-agnostic vqgan and time-sensitive transformer

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:30:22.877562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:8729116e7f92ce6f136b69198e3a815029075b67289fae6a51d873d876e5fff4

Pith citing papers

Observation 30cf6ee8-b393-4a98-a875-f59fc3f89e8e · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 153

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.925421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:bffd3672c32e1294ad07ceb96a710f5f16c3a516aabf8f4763de571b8337d4c2

Observation 2cef23cf-dd78-421f-8fb5-c06e9eed97b9 · inbound

World Model on Million-Length Video And Language With Blockwise RingAttention cites this paper.

World Model on Million-Length Video And Language With Blockwise RingAttention InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:36:57.228815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T06:36:57.165551Z digest=sha256:9c03cc45a0cc48a12b82f3fe3fad748b89e23552ce1b0b927817330bb15979f7

Observation 22cc7e27-cc01-4b20-8d2a-bcb0ed010feb · inbound

NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation cites this paper.

NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 102

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:55:20.531559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:55:20.362512Z digest=sha256:a9065845c4da55fd982b14eb682e81c9d29ed1e49050e82a19fe449d603dbd93

Observation 9915c743-fac1-41b1-8c82-2680e3f2a07e · inbound

VideoPhy: Evaluating Physical Commonsense for Video Generation cites this paper.

VideoPhy: Evaluating Physical Commonsense for Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 106

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:34:37.712578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T11:34:37.599691Z digest=sha256:e2e71e072b913004b3775319826d5a1c616adbda46e8aaf55c6159d6a8922761

Observation 254ea295-dc89-47aa-88d6-a4c06d781fba · inbound

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs cites this paper.

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.925421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T02:44:53.284345Z digest=sha256:5be5f924ecdebb6473e4e41c3ba8bf81a549ad6f656a890db23c6f7147d8ddbb

Observation b1f9af37-9ea6-41d4-97e8-84f9661fd234 · inbound

Emu3: Next-Token Prediction is All You Need cites this paper.

Emu3: Next-Token Prediction is All You Need InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.925421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T10:56:06.418360Z digest=sha256:fed146dc0c3ec54e1488399965c260405e3bba472cde9a1e481fc09984668178

Observation adfe7ec5-250f-4a63-a73e-c224ec7f4af0 · inbound

DOLLAR: Few-Step Video Generation via Distillation and Latent Reward Optimization cites this paper.

DOLLAR: Few-Step Video Generation via Distillation and Latent Reward Optimization InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:02:41.848080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T06:57:50.897865Z digest=sha256:d0edcab59764f1c81179f0d17de4e7fd6de4ff489cfde9ccba4d12a78385edcb

Observation 310a8a70-87d0-4e22-a6fe-f963c762f638 · inbound

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation cites this paper.

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 180

Resolution
unresolved
no resolver link, observed 2026-08-11T04:37:15.144767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:37:15.144767Z digest=sha256:5776a1ce426db40b0b22d444aae29c0428fea2922c740fe5207c7c145e9e0bab

Observation a9fe933a-7450-4abe-84e2-758c87633611 · inbound

Multilevel Semantic-Aware Model for AI-Generated Video Quality Assessment cites this paper.

Multilevel Semantic-Aware Model for AI-Generated Video Quality Assessment InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:54.320861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:54.320861Z digest=sha256:50b9f13213e7c7a5182c4c5ed6be44438c6d766f5cf10069b854c861e722a78d

Observation 3a2033c9-7e3d-468f-8566-2ddd04618277 · inbound

STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution cites this paper.

STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T22:04:48.803216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:04:48.803216Z digest=sha256:6e13383b032c1e224ba794ba55c9c472912a8f4f4ee4f846f4fd27f945b81843

Observation addab3fb-7910-4c90-b151-5e3a3adc0d1d · inbound

Tuning-Free Long Video Generation via Global-Local Collaborative Diffusion cites this paper.

Tuning-Free Long Video Generation via Global-Local Collaborative Diffusion InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T21:39:55.119247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:39:55.119247Z digest=sha256:cef2445b688817d8a651ee96491d963029706d7cc5a06d468f4c769064a49cfc

Observation 2db45f1e-d9ec-4042-84cf-291503112f36 · inbound

Training-Free Motion-Guided Video Generation with Enhanced Temporal Consistency Using Motion Consistency Loss cites this paper.

Training-Free Motion-Guided Video Generation with Enhanced Temporal Consistency Using Motion Consistency Loss InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T20:42:45.264582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:42:45.264582Z digest=sha256:e2a65f6d0bc735ffd828884686d59b6bee4053427fd02fd4258f54d739ea5dfb

Observation e8d8b3c2-c456-4eff-af15-513248c8a146 · inbound

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling cites this paper.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:52:20.783565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:fe1dda7bc6dfe3e7dd4633099b927d4b14955b3468747ed43771ae557d525fd5

Observation 4b653dc5-c7c2-4fc4-b10e-e1a3f159239a · inbound

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models cites this paper.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.856008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.856008Z digest=sha256:caa94cd04dfe23b7e1d56b8752a1b0dfa7bc5742dde229b8a4f4baef55216fc2

Observation 9559f07b-405e-443b-a987-5ea42ef568f3 · inbound

Video Latent Flow Matching: Optimal Polynomial Projections for Video Interpolation and Extrapolation cites this paper.

Video Latent Flow Matching: Optimal Polynomial Projections for Video Interpolation and Extrapolation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-09T18:51:12.672171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:51:12.672171Z digest=sha256:84e357cfe5261f4c9ebf458e9d4aa60caff5dea3eaba9dff11c29b4056b4f935

Observation 18044e88-1e2a-46ac-bbdf-2208eaf7b3f0 · inbound

Goku: Flow Based Video Generative Foundation Models cites this paper.

Goku: Flow Based Video Generative Foundation Models InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-08T21:07:32.451450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T21:07:32.451450Z digest=sha256:a08ba1e7b7f0d383c8913f6e1707fca19d7304ac860e9cd895f75b0068e2e92d

Observation f8652f3e-296d-4467-9c6a-28b536d0992c · inbound

TripoSG: High-Fidelity 3D Shape Synthesis using Large-Scale Rectified Flow Models cites this paper.

TripoSG: High-Fidelity 3D Shape Synthesis using Large-Scale Rectified Flow Models InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 127

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T21:51:18.394902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-16T21:51:18.323840Z digest=sha256:e402d1e28f581e67c16a969654da752bcc1920dcac17eaa73b3d946e734af53a

Observation af0bf21b-6be7-4169-b4fe-d7b0c07e597d · inbound

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding cites this paper.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:40:06.538109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:f5a80dec396a3a0b6f445871240b0b2351be8523cb0bd5a42eec696ccac5c161

Observation 8a76d122-1a03-4998-9597-ac78e3d31d83 · inbound

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks cites this paper.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:31.081961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:31.081961Z digest=sha256:b0019b1fb5cc24891b65458e77f2ae146078b492fea150330928e39adbca117f

Observation 42352660-1d21-46ab-8edc-e8126a2874be · inbound

RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language cites this paper.

RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T15:19:14.859109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:19:14.859109Z digest=sha256:290b6ef00b308c042ad54a9737c67ef5702087b295d398b30a57f84c5838e5d6

Observation fde07f22-2d32-4440-ba1c-660bc9eb2124 · inbound

InfLVG: Reinforce Inference-Time Consistent Long Video Generation with GRPO cites this paper.

InfLVG: Reinforce Inference-Time Consistent Long Video Generation with GRPO InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:48:11.064681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:48:11.064681Z digest=sha256:b21621c962b533faaa6e8ef08ca1de6d4c443b4e00b49c7ad0c718644965ab17

Observation 8daffc84-86b1-4a83-8ec4-10a2186913ec · inbound

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation cites this paper.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.298215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.298215Z digest=sha256:d57f690d828f3f8bee5aa754014ce45090e8f2c610bc4ca38bc36f60ef18761b

Observation 7d3d1fc9-29a7-43ef-ba22-2208dff8b6e8 · inbound

Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion Transformers cites this paper.

Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion Transformers InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:18.607441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:22:18.607441Z digest=sha256:9cae4c019902d3265da2a065e742c7bd028ad66f3f787344a0a4a9478831a6b3

Observation 213c9c4c-bd20-43cf-8c8d-d90308fb481e · inbound

DisTime: Distribution-based Time Representation for Video Large Language Models cites this paper.

DisTime: Distribution-based Time Representation for Video Large Language Models InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:54.042453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:54.042453Z digest=sha256:de131811f35027197ba4bf7dddb66b1b03d247454135a29b31930a59267e4a14

Observation d9559444-4464-46cf-b125-2bc3146cfffa · inbound

AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation cites this paper.

AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:10.899127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:10.899127Z digest=sha256:e47610944228afdadf297a9a3d1c8622a1e45189dbed10e2dde9834f9d3a8324

Observation e871c747-e2c3-4241-9229-d221f6bd513f · inbound

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model cites this paper.

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:44.268995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:44.268995Z digest=sha256:e2817df4d0e1db3940b86001ee1c41cdbe59023a9354e2b5d765a7346777893b

Observation 260a8e1d-40b1-41b9-8a9a-644b5b27499b · inbound

ContentV: Efficient Training of Video Generation Models with Limited Compute cites this paper.

ContentV: Efficient Training of Video Generation Models with Limited Compute InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:34.955458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:34.955458Z digest=sha256:30aca345e430b56c856913da164ea3ba66edee8305a820f480c1a81da37d8ada

Observation dea69658-9404-4c83-ad71-b20e3413cc57 · inbound

Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision cites this paper.

Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:58.005525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:58.005525Z digest=sha256:d1a4b6ba493bcad6f08c8fb19bf8f8e2dfbc2275c24a25794797bb29ad99f3da

Observation b53b07d8-dadc-49c1-b29b-1c4f81601bb7 · inbound

A Watermark for Auto-Regressive Image Generation Models cites this paper.

A Watermark for Auto-Regressive Image Generation Models InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:04.292024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:04.292024Z digest=sha256:670c62b541901086d28ecd351e99eca41b90b6ae8ce45feb3035177420cbec16

Observation 2a27074b-2994-4a4c-a2fd-60636dec451c · inbound

Fake it till You Make it: Reward Modeling as Discriminative Prediction cites this paper.

Fake it till You Make it: Reward Modeling as Discriminative Prediction InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:42.096528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:42.096528Z digest=sha256:f4b14f3e0f8969332e4fe54065b37201b6362bb62fe97d7b3c21f9e967ad25a5

Observation a2aabb4b-3746-44aa-a81c-b58cff382e30 · inbound

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes cites this paper.

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T22:37:39.102494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:37:39.102494Z digest=sha256:93fb704783deda5646dea5f4f43557cd4d31fa99cec57f7fb2469446a3b4756d

Observation 2f0879d8-0e8e-4004-86de-263b2740616c · inbound

CI-VID: A Coherent Interleaved Text-Video Dataset cites this paper.

CI-VID: A Coherent Interleaved Text-Video Dataset InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:43:55.960100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:43:55.960100Z digest=sha256:8fdd9fc10ae877eb47858d913893746201990117d0c967ad661f1535dab2e685

Observation 27f8fedf-7160-4197-ae13-899b34717564 · inbound

Semantic Frame Interpolation cites this paper.

Semantic Frame Interpolation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T19:39:51.503054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:39:51.503054Z digest=sha256:11e6f03cdc41d24e7aa77e542752f405b0d47999e78c7f9926c4b08a35c39dbc

Observation 4545fb88-e229-402d-a9d1-46d9a3e1f23b · inbound

LoViC: Efficient Long Video Generation with Context Compression cites this paper.

LoViC: Efficient Long Video Generation with Context Compression InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:30.750694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:39:30.750694Z digest=sha256:fd75abe3eeced3843764e6cc55402f3b1662874c3c43b476c24829c1b138b549

Observation 59763ee7-6e67-4e6d-b335-f5a1fe341371 · inbound

Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning cites this paper.

Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:28.044936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:44:28.044936Z digest=sha256:edc6825f33518aab036411970615b66b1e30ca37aebd040c6f6b8342979d2f09

Observation 3d0b065b-b671-4a08-90de-556ab4257c7a · inbound

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding cites this paper.

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:17.724007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:17.724007Z digest=sha256:805d3340d46f12b96394598b0010fc830781b7c7b40bf19ec11699cf0a71a746

Observation 85d6e32a-0eb4-48bd-b14e-ec1ec0be758d · inbound

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation cites this paper.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.952467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.952467Z digest=sha256:b36b126e6edc6e2793f9bfce6b68f5022a7436a35f8d1b8f78eb96b04bfced92

Observation e61e9f2e-3702-4292-8b84-2a4633fad0de · inbound

Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm cites this paper.

Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T01:05:34.727296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T01:05:34.727296Z digest=sha256:7796e734c9cbd909c7daf301f8ff7e854b2e36976307a3dd54e0167d8e06d7e3

Observation a967589a-3573-4762-9131-6d72587aaac4 · inbound

B4DL: A Benchmark for 4D LiDAR LLM in Spatio-Temporal Understanding cites this paper.

B4DL: A Benchmark for 4D LiDAR LLM in Spatio-Temporal Understanding InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-19T00:11:55.966378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T00:09:57.236162Z digest=sha256:0db5b0f5ee6189f4081a81545dc12fe5cf875d652e018615fdd21aea4a254376

Observation 17bc81e9-9c7c-408d-a652-5150043f612e · inbound

TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding cites this paper.

TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T22:01:26.008102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:01:26.008102Z digest=sha256:ea94123d44bcd6294ffef4eed3b849bee067ba5cdd3c10f5e1744ceb94feb0a8

Observation b12fb950-5c87-4313-9cd8-23b23ac671e4 · inbound

Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation cites this paper.

Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:24.611355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:20:24.611355Z digest=sha256:066e92a59f3badaf12556f2fd997697a6ab8ff289affdaa8ebaf7b48bd8cd793

Observation f1fe35c4-46c6-4983-9916-c3608a9d0417 · inbound

From Black Box to Transparency: Enhancing Automated Interpreting Assessment with Explainable AI in College Classrooms cites this paper.

From Black Box to Transparency: Enhancing Automated Interpreting Assessment with Explainable AI in College Classrooms InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:50.734115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:17:50.734115Z digest=sha256:54af7a80e7bdbc07b20f0a25aa953d4c13b266c8384d12fc33325fd0a4230851

Observation 3bf63ee5-3352-4211-a8ea-1e88f6c7b3f0 · inbound

A Survey on Video Temporal Grounding with Multimodal Large Language Model cites this paper.

A Survey on Video Temporal Grounding with Multimodal Large Language Model InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 140

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:18.105472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:18.105472Z digest=sha256:a9fe56e0b6f20f35884fdc22316ff7fdaad812601e5b9bcab90912842d917107

Observation 9a89c0a1-e112-4afe-8afb-42f563ea1a3a · inbound

Effectively obtaining acoustic, visual and textual data from videos cites this paper.

Effectively obtaining acoustic, visual and textual data from videos InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:37.063914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:37.063914Z digest=sha256:2042c7ebd48b1b19caf6c9227f6eb3be3ff627e6a50f5079771a6a05f08fa9eb

Observation 91e45fc0-417d-4d4c-96fd-6edd7b3dcf78 · inbound

Video Understanding by Design: How Datasets Shape Video Models cites this paper.

Video Understanding by Design: How Datasets Shape Video Models InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 176

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:42.065033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:42.065033Z digest=sha256:7bf61771fb3a0e29993c17365417e3270f9d8b3537baed1e09b7fdb12e111898

Observation 4474b6c3-8f6c-4003-add1-d6376784c4be · inbound

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders cites this paper.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:11.018159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:11.018159Z digest=sha256:06215be38c6160fd50af000841823e8f4e4e474560095f451addfd3e2be7f767

Observation e939595a-26c5-4c43-b8ed-3a0572510185 · inbound

VC-Inspector: Advancing Reference-free Evaluation of Video Captions with Factual Analysis cites this paper.

VC-Inspector: Advancing Reference-free Evaluation of Video Captions with Factual Analysis InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-18T15:31:33.344470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T15:30:29.933558Z digest=sha256:c59108599e2a4a818baebfb8ae0ac0af95c9b9f385502938102cef57556181f4

Observation aee86242-127c-4615-a89d-8336e02dae48 · inbound

MVAD: A Benchmark Dataset for Multimodal AI-Generated Video-Audio Detection cites this paper.

MVAD: A Benchmark Dataset for Multimodal AI-Generated Video-Audio Detection InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:48:58.512894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-17T03:48:26.807495Z digest=sha256:393016e46afee079ce1aef0b5c10d50f2e3e3eeffaf79a44f1467828e866ff87

Observation 4d45bc8b-0946-43aa-a9d9-501499f26280 · inbound

VABench: A Comprehensive Benchmark for Audio-Video Generation cites this paper.

VABench: A Comprehensive Benchmark for Audio-Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:08:43.896369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:03:45.576961Z digest=sha256:324e36cf5524b9db5d6ddafc4b28bceb024fc943f0f1d9723758a6e72d623e18

Observation da4facba-bf19-4448-a034-e86ebbe38fb5 · inbound

Adapting MLLMs for Nuanced Video Retrieval cites this paper.

Adapting MLLMs for Nuanced Video Retrieval InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-05-16T22:21:18.875720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T22:20:09.051957Z digest=sha256:8d92c1cb13151b75e2cf4a5886ce0f7534c8e3050cc5cc7085a3382916339b0b

Observation 7ffc17f1-87cf-46b9-890f-4755ca5b67c0 · inbound

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing cites this paper.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:59.060875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:59.060875Z digest=sha256:4bfc3a871b7f05396562f62d539c7980a537344b955cc55ea206ee3c9742c47b

Observation 865c0116-97e4-4f43-95e6-330b4bc095c6 · inbound

Efficient Scaling of LLM Training with Flexible Context Parallelism cites this paper.

Efficient Scaling of LLM Training with Flexible Context Parallelism InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T20:59:25.698578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:59:25.698578Z digest=sha256:1eff43611d2660edb6c6f50c8e0e4f5f4cc965e3d136e2dcdffd63b7e9541f90

Observation 8244475b-1246-4855-a1dd-e3fa90962bbc · inbound

TrajTok: Learning Trajectory Tokens enables better Video Understanding cites this paper.

TrajTok: Learning Trajectory Tokens enables better Video Understanding InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:16:31.775449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T19:11:52.694778Z digest=sha256:f981305a330a4673ba8c27712dfd0bf0962f9a42f3822812ba28c5b5ba392cac

Observation e60ffcc2-33f3-4934-8738-43266ad76682 · inbound

TrajTok: Learning Trajectory Tokens enables better Video Understanding cites this paper.

TrajTok: Learning Trajectory Tokens enables better Video Understanding InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-02T20:38:44.081929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:38:44.081929Z digest=sha256:56da33a2c74c91cd9768b45647f56d32f2dfd04a9d8d60cebc52addf904131c9

Observation e0ff4ce5-78ba-4e71-b9a6-eff494c887c4 · inbound

Insights from Visual Cognition: Understanding Human Action Dynamics with Overall Glance and Refined Gaze Transformer cites this paper.

Insights from Visual Cognition: Understanding Human Action Dynamics with Overall Glance and Refined Gaze Transformer InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.925421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T18:07:41.426142Z digest=sha256:44fcfc564d741ac9adbac20189f44fcf0bcb389246596c4d532694771f591b7b

Observation b849b742-36a3-4513-883f-3057e2185568 · inbound

MotionScape: A Large-Scale Real-World Highly Dynamic UAV Video Dataset for World Models cites this paper.

MotionScape: A Large-Scale Real-World Highly Dynamic UAV Video Dataset for World Models InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.925421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T17:11:14.242522Z digest=sha256:3308c2488b3a908a4c003c4bed64ffec138a1b88b9b788262304081cd7e6a1d5

Observation 0c625605-0c7a-4735-af4e-b9ec9d873b1f · inbound

InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding cites this paper.

InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.925421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T18:06:33.139310Z digest=sha256:d1549f01094d059022f47ee295369e3158b0d90479e6d1b19ee5615efbd44dd7

Observation 0f891e84-a2a2-42fb-a7b7-fb9efba14814 · inbound

InstrAct: Towards Action-Centric Understanding in Instructional Videos cites this paper.

InstrAct: Towards Action-Centric Understanding in Instructional Videos InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.925421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T17:07:49.334104Z digest=sha256:645b282a760631b31d6a595b7087ac1d2709941b813356851295ca1a749c01b6

Observation 5b7f40ef-af3c-4425-bdb8-671a05b9ca51 · inbound

How Should Video LLMs Output Time? An Analysis of Efficient Temporal Grounding Paradigms cites this paper.

How Should Video LLMs Output Time? An Analysis of Efficient Temporal Grounding Paradigms InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.925421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T16:50:48.559703Z digest=sha256:0875581ef9771e711b1d1dadf8d2d6c81fa9f83297193a0c63a03adb8aeff0b8

Observation 2c0d6f8c-470e-411d-9c68-42a0f99148c6 · inbound

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation cites this paper.

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 179

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.925421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T15:35:37.095627Z digest=sha256:600c31bfc8f6a3338b3bd87180b530b04b1a0d68ff658d5380d7707745450b92

Observation 742815fa-753f-4906-89a3-f07e4c404c68 · inbound

UniMesh: Unifying 3D Mesh Understanding and Generation cites this paper.

UniMesh: Unifying 3D Mesh Understanding and Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:30:22.925421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T06:55:42.679323Z digest=sha256:140a57bac26c41b01b8afcfb5519c4f688c9e2a949af9187e12fb28fe8c89167

Observation 7dcbb625-6023-46af-a3da-cc78a202cc6d · inbound

Seeing Fast and Slow: Learning the Flow of Time in Videos cites this paper.

Seeing Fast and Slow: Learning the Flow of Time in Videos InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.925421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-09T21:55:17.377887Z digest=sha256:ce3b472e63a74345c4efdb4e4b5ca684fae65d1338d96d61e1f3aaa2470a9783

Observation 447ca9d5-c470-4332-a75a-34cd3b686a02 · inbound

MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation cites this paper.

MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.925421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T06:28:42.129881Z digest=sha256:31421da2d62660dab94b4912358ea9ac1def31731c5753884b2f663df609f8cc

Observation c2bdacff-4003-4031-b5e5-0c12abde9be9 · inbound

MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation cites this paper.

MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.925421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T00:50:10.509727Z digest=sha256:1174086117b45b7718bf834fc45553e5c90655c103a3664372ae55abbbafc760

Observation 3f6d8964-8e9a-444a-af15-d1d4f59abbff · inbound

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation cites this paper.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.925421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:e1f9502e190f0779d27320e75f6fcedec369c6dfbb706e1f51ef9218ca325a14

Observation dc920f33-c18d-4b0c-9ff0-e691a0cafe1e · inbound

DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation cites this paper.

DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.925421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-07T13:50:23.835653Z digest=sha256:798c9ec3b75cf2cb5a2ff21caf48bee6596f6ef33d248b4273b174a18993b07e

Observation 7f9d89da-9e81-46ff-81cc-d196e3df9427 · inbound

TMD-Bench: A Multi-Level Evaluation Paradigm for Music-Dance Co-Generation cites this paper.

TMD-Bench: A Multi-Level Evaluation Paradigm for Music-Dance Co-Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:30:22.925421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-09T16:11:16.551910Z digest=sha256:9c67cfc7112877e651d342196488f5370d6faa912badfcc2c2eb31c2e2cb947f

Observation 8ff7a753-c4b2-42bf-8bb2-7175ff015597 · inbound

OZ-TAL: Online Zero-Shot Temporal Action Localization cites this paper.

OZ-TAL: Online Zero-Shot Temporal Action Localization InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.925421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T02:48:20.005281Z digest=sha256:118b4928574f19d0c425b44ee64cb8a36b4c9ed7061a0da1cc18d0b8ffba0ab4

Observation 0994a83c-63ec-4923-a22f-f68a8437b17c · inbound

CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives cites this paper.

CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.925421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T05:38:21.587653Z digest=sha256:ea64cb286d64e4c73f2a5d5eca08f2d09111c2de785f5c9ec45988f9c5ffafce

Observation 8dc8fe02-ae02-4b84-9d97-9f945aa585c6 · inbound

GTA: Advancing Image-to-3D World Generation via Geometry Then Appearance Video Diffusion cites this paper.

GTA: Advancing Image-to-3D World Generation via Geometry Then Appearance Video Diffusion InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:30:22.925421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T19:39:16.410139Z digest=sha256:24c2affec925cb9fe721ad462e290677a2b87d078ef044a14f5336808c929dbb

Observation 724a98c9-12e5-45fb-8535-f0d10359438c · inbound

TeDiO: Temporal Diagonal Optimization for Training-Free Coherent Video Diffusion cites this paper.

TeDiO: Temporal Diagonal Optimization for Training-Free Coherent Video Diffusion InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.925421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T04:59:34.046044Z digest=sha256:fade6ac0f2a6ff9dd782712bd00a67baac02e8867581930bac81245e7afb080b

Observation 31da3f52-8af4-4f63-8e02-7696c037be8f · inbound

Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video Reconstruction cites this paper.

Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video Reconstruction InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.925421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T01:51:57.018809Z digest=sha256:57abfd53e6c2643c744a3fe6b05a91c1b6a412bff4c92391406e547b8096586d

Observation e3c78d92-5a10-434d-986d-f9a4f6fce380 · inbound

HAVEN: Hierarchically Aligned Multimodal Benchmark for Unified Video Understanding cites this paper.

HAVEN: Hierarchically Aligned Multimodal Benchmark for Unified Video Understanding InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:43:08.500431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T07:38:08.819186Z digest=sha256:498d866f9f3036bd0e347a1d734df221c3dcbf207fc147725f9276a96b30ef5e

Observation 8fcd00f2-b360-4a31-8a50-1836b20ec72d · inbound

Aero-World: Action-Conditioned Aerial Video Generation from Inertial Controls cites this paper.

Aero-World: Action-Conditioned Aerial Video Generation from Inertial Controls InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:43:05.344245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T05:42:59.996351Z digest=sha256:ae14daab545185f8bf675ce8d99d452e8cfbcbe622d08be9f03e0a2980d4cdbb

Observation 983cb006-f158-4bb8-8bb6-e9c1df261456 · inbound

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation cites this paper.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T05:00:22.396412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:c5dfbabe3f59017bda5d0724ee0db5c974148c9f0b46c532b911941608fe7b13

Observation a34e84b6-6b7c-415e-a663-a463f083e867 · inbound

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models cites this paper.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T13:44:41.227663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:089824232167b1489c690703ca6005edd7c5243d1b55592234887438fd84185a

Observation b9a67c7d-e0c5-4340-b268-c14f00cfca16 · inbound

PhyGenHOI: Physically-Aware 4D Generation of Dynamic Human-Object Interactions cites this paper.

PhyGenHOI: Physically-Aware 4D Generation of Dynamic Human-Object Interactions InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:23:15.529824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T08:17:14.536157Z digest=sha256:4c9f154e1460c4683c338d176ddf8410576e7bb1e1e71ac4420554d952795324

Observation 4b773ed1-f8c1-4eca-a18f-b0b2eb3177c3 · inbound

Collaborative Few-Step Distillation and Low-Bit Quantization for Wan2.2 Dual-Expert Video Diffusion Models cites this paper.

Collaborative Few-Step Distillation and Low-Bit Quantization for Wan2.2 Dual-Expert Video Diffusion Models InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:02:34.304002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T18:58:13.799615Z digest=sha256:8e2ec8104a80054fbb5b4b74c6250c72ee7f29b4691655998321c87ca304e2f7

Observation d5cb3a71-66a6-4d92-be6d-bd5d0ad139e2 · inbound

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning cites this paper.

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:46:56.822372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T01:52:44.785582Z digest=sha256:891b58b6f9310c0c135153ee21f632806f68a7e6af2b9017272d797f45f75c12

Observation 69aac290-bb62-4c8b-bb0b-55ea4a9876e1 · inbound

ViMax: Agentic Video Generation cites this paper.

ViMax: Agentic Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T02:46:29.093730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T10:37:21.944514Z digest=sha256:f77e02c2ca9d9556d6a7dd6898ec0cefac6d1721bf284a45255d48e21754828c

Observation 47973b93-7910-4b31-8771-bccf6678b9ec · inbound

ViMax: Agentic Video Generation cites this paper.

ViMax: Agentic Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T12:30:53.634740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:30:53.634740Z digest=sha256:5b7b6130d0498c1db5897953426884e25f94081351fb2da2f01968af1af679ae

Observation 23b586ec-97ca-436d-a597-73fc5c754578 · inbound

CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation cites this paper.

CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 71

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T00:07:28.115521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T17:30:25.371658Z digest=sha256:4ae66454a8f7c974a865960fb63b5833e506debbd8f11b3129f545a1def67c2c

Observation f614416f-0818-47b8-8dae-1f0979eeb8a7 · inbound

SpecLoR: Spectral Lookahead Rectification for Motion-Coherent Text-to-Video Generation cites this paper.

SpecLoR: Spectral Lookahead Rectification for Motion-Coherent Text-to-Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:27:56.744669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T10:00:35.608696Z digest=sha256:3f90a1266520fc48a748a4f7c4ae459a9ecd9f15f8fd6f8f095355173ed1d10c

Observation dd62c66e-c7f5-4afc-8aaf-da23ca716ddc · inbound

MLT-Dedup: Efficient Large-Scale Online Video Deduplication via Multi-Level Representations and Spatial-Temporal Matching cites this paper.

MLT-Dedup: Efficient Large-Scale Online Video Deduplication via Multi-Level Representations and Spatial-Temporal Matching InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:58:03.533522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T09:43:22.054789Z digest=sha256:8298efc4755d6698c9cef84443e30fb0ceb9725122d6db1e74b5cb31c91f07bf

Observation 6c608ea4-a649-44f6-bf37-4a5cb9fbff96 · inbound

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling cites this paper.

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-04T03:19:30.186003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T18:15:51.862192Z digest=sha256:de48bf27220d69f7d28bbca810c55eb23a659aff51d28d61c78c568dc0ef56fd

Observation ffce0e3c-5d26-47a0-a99c-3e810f98f6c5 · inbound

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling cites this paper.

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:07.007337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:07.007337Z digest=sha256:d04abdfd4b7b2e6fc86e4cbca83f9a9f3b47b535d1f39ad6bec4d08c2160fb4c

Observation c94e94a9-48f4-4772-9768-39ac77418ff8 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:39:37.514533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:d750b99dffbf5047d10fe3755d648fff854edf4a2eeffda1956a412c0e09f850

Observation 13c6f3a5-74aa-4e65-b246-c1be83f2fecb · inbound

MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control cites this paper.

MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 68

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T07:24:21.589088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T07:18:20.501369Z digest=sha256:fc5a74d9fb4229713fcfc9e8379364213a9c607e5fbc9085cf2249a56522a1f6

Observation c611cf4c-bfe3-45e4-9a0a-522ce4031e0e · inbound

MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control cites this paper.

MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-14T17:03:01.432145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:03:01.432145Z digest=sha256:2c335867b9bf0bc84fe88ade51c56291e5f1c3a2d1d9ad0691a4964154f50f7f

Observation 99b8ea70-9a6f-41fc-a2e7-9b7ae6699945 · inbound

Incentivizing Vision Language Models to Search for Long Video Question Answering cites this paper.

Incentivizing Vision Language Models to Search for Long Video Question Answering InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-12T05:50:16.895740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:50:16.895740Z digest=sha256:92229d6c5b17b6a3a1adc62de8dde8cdbc292f9cecfb921dad67a723bfc6763c

Observation 3b5a1d80-68c3-46b4-9643-65678c44f226 · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 154

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:8761485dd9e24d1097dc17a919d732a309c0098f1ed33bafc242bc5064e9c3c3

Observation 885445ab-eacf-483f-a953-47177b84ab8c · inbound

Reinforcement Learning: From Algorithms To Foundation Models cites this paper.

Reinforcement Learning: From Algorithms To Foundation Models InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-01T17:45:04.006638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:45:04.006638Z digest=sha256:45a6f84f0c2fa7e15eed89294e84592fe3e1b8f79f4c0f97b926924273aa514e

Observation 03137473-91ef-4984-9fa1-c623d1cf536b · inbound

ShotPlan: Cinematic Video Generation with Learnable Planning Token cites this paper.

ShotPlan: Cinematic Video Generation with Learnable Planning Token InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T17:22:32.633281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:22:32.633281Z digest=sha256:42c0dd6efc4362ffb5231466a22d4151098a867d9cb772719e9dc78bf345469a

Observation 2c439224-b816-4c81-a181-933ffe2909de · inbound

Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation cites this paper.

Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T14:24:46.109969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:24:46.109969Z digest=sha256:3f145074d4c0409fbde628059b27a05de6333cec51a49557fd2977917b4356f1

Observation b0f8e8f5-cf37-4d83-9e26-a66c6e860f4a · inbound

MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning cites this paper.

MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-31T23:10:27.561730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:10:27.561730Z digest=sha256:daeecaa595526d272f9c1279dcfed149473c3c0cdb6876650fdacad0366b5ddd

Observation 9450afd1-f176-4fce-86ea-0b9877e2b8d1 · inbound

EgoPlay: Event-Triggered Video Editing for Egocentric Streams cites this paper.

EgoPlay: Event-Triggered Video Editing for Egocentric Streams InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 94

Resolution
unresolved
no resolver link, observed 2026-07-31T11:46:42.233597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T11:46:42.233597Z digest=sha256:03af5301643e4ca1d61c3b98660684dab0d69286aa72d15d78e17a285e3f8538

Observation 90122d2e-2650-4361-8e45-90c25df1d96c · inbound

Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding cites this paper.

Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T09:17:00.664674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:17:00.664674Z digest=sha256:1b1e244d554f5c47eb0f74a060bedfedd842a124ab105064d0fbacce07b1441d

Observation bb1f23e9-3732-4a02-a01e-8c4497451a29 · inbound

CineWeaver: Training-Free Reference-Controllable Multi-Shot Long Video Generation for Cinematic Storytelling cites this paper.

CineWeaver: Training-Free Reference-Controllable Multi-Shot Long Video Generation for Cinematic Storytelling InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-01T13:56:37.009703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T13:56:37.009703Z digest=sha256:5f17de6983706e10fe69ffdae44c718ea20bf2a463179a93954baf118f2507f5

Observation 2af81da1-5a35-4ec6-8f31-078701ba5f68 · inbound

Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation cites this paper.

Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-30T23:47:57.752509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T23:47:57.752509Z digest=sha256:2e4799ed2a66af601f83871198dc46978723da15512510865d5576c05eaeda5e

Observation 8a083e30-7eb0-436f-820f-6f0b30085e5d · inbound

4DHumanDiff: Direct Text-to-4DGS Generation for Consistent 360-Degree Dynamic Humans cites this paper.

4DHumanDiff: Direct Text-to-4DGS Generation for Consistent 360-Degree Dynamic Humans InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 147

Resolution
unresolved
no resolver link, observed 2026-08-01T04:24:27.351945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:24:27.351945Z digest=sha256:50053f64c9666c28c252274fde3660a1008b604937607ca0392090ebd4d2912e