Pith. sign in

Paper Citation Record · LEDGER

InternVideo: General Video Foundation Models via Generative and Discriminative Learning

As of 11 August 2026, this Paper Citation Record lists 100 of 111 outbound references and 100 inbound Pith citation observations for arXiv:2212.03191.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2212.03191 v2

Coverage vector

measured 100 of 111 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T00:36:53.235740Z

measured 200 of 200 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 100 of 101 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:59:13.473248Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 111 outbound references displayed

  • verified exact21
  • verified fuzzy76
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

93
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 3904b6a7-d161-4169-b3a8-96a507d5c664 · outbound

This paper cites Nsnet: Non-saliency suppression sampler for efficient video recognition.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Nsnet: Non-saliency suppression sampler for efficient video recognition

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.508727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:9289f499d7abb9f3de760413bbac922befdddb2311d9cc035d1b65d818cecd51

Observation cbffbc11-917c-4ad1-94dd-385cab12fc97 · outbound

This paper cites Learn to cycle: Time-consistent feature discovery for action recognition.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Learn to cycle: Time-consistent feature discovery for action recognition

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.514105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:5eefa1d5ebdd285d867e8464f1fe5f6771577ae734aefda3138bac501e354557

Observation dd35ef7d-c1ef-4b64-bb92-257903cfba94 · outbound

This paper cites Self-supervising action recognition by statistical moment and subspace descriptors.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Self-supervising action recognition by statistical moment and subspace descriptors

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.516789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:390e4e8f484573f20ec19b3fec38bb6fe4083671da6fab5f8f3442f1c0e4bbef

Observation 4551cede-8f60-46b9-b39f-c6229ae96a31 · outbound

This paper cites Actionformer: Localizing moments of actions with transformers.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Actionformer: Localizing moments of actions with transformers

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.519516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:85b09be5b7e0fb2e2e7e7d28eab440010ad684acc930388bd1a84261c2a5b751

Observation e626acfb-cab5-4f20-a700-3e5a8e8b8b37 · outbound

This paper cites Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.522177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:2b57f42b2e2a300ff6e591924c67ceb04b6b3edff5820f691b6b65540b7f546a

Observation 57446b8b-8a03-4a00-80e3-ad5434efd26c · outbound

This paper cites Masked feature prediction for self-supervised visual pre-training.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Masked feature prediction for self-supervised visual pre-training

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.524483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:319c7200fdc4f892bd4e0a18e57f020c7ad751924386458368b562ffc22d6d0a

Observation 041299c2-4743-4bdf-a2d4-5450b4b6993a · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:36:53.357463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:a0812de03ac1ec55a5b82b30e2ed01ed4601f53e5a91e0cac93e5c392e5abc77

Observation afaab448-fd03-4875-8d07-d6e643c521f0 · outbound

This paper cites Multiview transformers for video recognition.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Multiview transformers for video recognition

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.526483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:ac50418539ad3319aa4b02fb25e0b88f05d0821550d817d8d6a8bfefddcaab91

Observation 66281555-6776-43b3-9d94-1a1e57adb8a1 · outbound

This paper cites Merlot: Multimodal neural script knowledge models.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Merlot: Multimodal neural script knowledge models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.528653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:852f017c48f83fc2c544051a48b02a4f618b3b9cfd183be8638d28b77cf8d921

Observation ff7620d4-1be5-4c82-843f-88a18b392ccb · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning On the Opportunities and Risks of Foundation Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:36:53.354621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:14e087a0c598c6990be49041ca5c3aee389f6bbddc18de357ae5f43fc1f7b92a

Observation 72dfe7fc-0cf1-4782-9d65-7bbbd4f9d610 · outbound

This paper cites Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.530886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:6093be1da0559b9c205d477bda42202cff7d9d6db37b1c1fd04ac826228e0c40

Observation 1df9fb52-dc1e-4933-bc0d-c08c2db0eec2 · outbound

This paper cites INTERN: A New Learning Paradigm Towards General Vision.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning INTERN: A New Learning Paradigm Towards General Vision

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.329293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:72612106d4ccfd672d95d5a21be3546ce54a1e33f35813d50a36e0e5a551c955

Observation e4da4d3c-8d47-45b6-b5ff-48a6252475e3 · outbound

This paper cites Learning transferable visual models from natural language supervision.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Learning transferable visual models from natural language supervision

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.533740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:f9cbba43b16032fe31a650088b1bb9c8fd7b7af9b0621e2eb3f93e35eda9ca89

Observation 9b2314a5-8cee-4968-9d4b-1cc13321daf3 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Scaling up visual and vision-language representation learning with noisy text supervision

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.536310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:d12833406afc74eae4e588a31604a28714081de2aec9d3145678cf62393aecf7

Observation 109f2a8e-3b42-4a52-a230-15d8417d5650 · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Florence: A New Foundation Model for Computer Vision

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:36:53.308902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:053dcd622311cd30147aca01f0f0824aa40966d1cde2ff98a58ad28769ece153

Observation 4f1d64d9-2c27-4878-a18f-787c3dc6ce92 · outbound

This paper cites Unified contrastive learning in image-text-label space.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Unified contrastive learning in image-text-label space

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.538927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:66fe1e972719a9bfa6bd9899a468ebd01a1365a28af27a7f495445582400f71c

Observation ba08958b-c143-46ad-8cf5-da2b9005a221 · outbound

This paper cites SimVLM: Simple visual language model pretraining with weak supervision.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning SimVLM: Simple visual language model pretraining with weak supervision

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.541465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:02b0912efce2a2ce84f7ee58d80f23a64110f1c4c42c9f60e1fe7fc47042f669

Observation a06c1fe5-9bea-4dad-a1f9-59470a7f62ec · outbound

This paper cites Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.544680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:6d9d51f6a60b59fa08f8469542b3fb40c0e3a952c02f4610598437740ebb9b0e

Observation f2dbac5c-c1d7-41a8-9f03-5ea139503240 · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T00:36:53.366555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:cfd731253e9dc2cd0607189098715268f970930604ec3f3634f2661c61d86ed0

Observation 398788ad-629b-4001-b807-112e6bf10745 · outbound

This paper cites BEit: BERT pre-training of image transformers.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning BEit: BERT pre-training of image transformers

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.547292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:39a6414091dd1e6437b7fbba43028efbca1a4e87f0bf8b8eb59a52f4a2e35ef1

Observation f564919a-d241-4b27-8dc5-089f0e45c941 · outbound

This paper cites Pathways: Asynchronous distributed dataflow for ml.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Pathways: Asynchronous distributed dataflow for ml

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.549438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:0d25ddf5d56d2bf3352510b88b8fa36f13b53089a0018d9fcbfdc72a47832731

Observation 10b34b71-aa82-49fb-893b-7c79dec90fe5 · outbound

This paper cites X-clip: End-to-end multi-grained contrastive learning for video-text retrieval.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning X-clip: End-to-end multi-grained contrastive learning for video-text retrieval

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.551696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:bcb43bc69a487b88cb6dc6f5a0736b453aba21b07da974dd8531f2dc0ee5ab1b

Observation e835e721-e8bb-49eb-917e-7d8738d2da5b · outbound

This paper cites VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.554578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:46f2249090044db72db97d7b2c4f657cf34ea94ec550111a0c701aacd943fdfa

Observation d592edaf-3d7a-463e-a2e4-c647938fae95 · outbound

This paper cites All in One: Exploring Unified Video-Language Pre-training.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning All in One: Exploring Unified Video-Language Pre-training

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.363109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:3536a49a6bb29d0dc035b9fc92c5623fb0c49f03be3e84c9afc1b9279d706c82

Observation ed71fbf0-5b96-40a4-8947-b5943e26dab4 · outbound

This paper cites Masked autoencoders are scalable vision learners.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Masked autoencoders are scalable vision learners

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.556938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:553c0fa633a2445ea4feb9ef3e5855a32137d17e80e923d4cd17ec18b11c94bc

Observation f6427790-add4-4399-ba0e-9dd4f2a67756 · outbound

This paper cites Learning Spatiotemporal Features via Video and Text Pair Discrimination.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Learning Spatiotemporal Features via Video and Text Pair Discrimination

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.277470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:d514441afe5d8114dde5e72139124b2e29ba8fc9f2d47cb9fb6f1998e878001c

Observation a1d8d62b-3f27-4931-b883-56bd2c53aaa5 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Quo vadis, action recognition? a new model and the kinetics dataset

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.559327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:afc77b73f28e20155687d5997bc9b490af4f90ad48d052eba1d6bdf17cdb90ac

Observation 1e26c95d-0036-43ad-9c94-acc008c81bca · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning An image is worth 16x16 words: Transformers for image recognition at scale

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.562142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:766b9bba7d37aca65e7c1f52b1346eb45f4252f1bd48830a224bc2f8d85580ec

Observation 97522ffd-971f-482e-92f4-fa7cc729d036 · outbound

This paper cites MultiMAE: Multi-modal Multi-task Masked Autoencoders.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning MultiMAE: Multi-modal Multi-task Masked Autoencoders

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T00:36:53.332561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:f536ff77b337ff57e7104180f18afd33f9d09dd4f1f3ef226acadc8b600133c0

Observation f1b3dfd1-d7f2-4246-a5ad-9a97e7ede817 · outbound

This paper cites VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.335861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:4069f1e41ed8c8d032aa1e09095070a8124da59e36ecca70f61b32eccffb7e28

Observation 43d26c3f-fb83-487a-9878-515c56343167 · outbound

This paper cites LAVENDER: Unifying Video-Language Understanding as Masked Language Modeling.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning LAVENDER: Unifying Video-Language Understanding as Masked Language Modeling

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.351832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:cfe44249cdca5acd8d308b2066f216f7d52b03d3974e9b3eecb3aa688035a4d3

Observation 37845baf-0dba-4d3b-915d-9899e0d124a2 · outbound

This paper cites Merlot reserve: Neural script knowledge through vision and language and sound.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Merlot reserve: Neural script knowledge through vision and language and sound

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.564481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:ff95630b711e44f4b7ab3967aef7558822685d8dd1af06964d9ddd23685240f2

Observation ae004afd-d140-46f7-b35a-d7c0c2c74132 · outbound

This paper cites Unsupervised visual representation learning by context prediction.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Unsupervised visual representation learning by context prediction

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.567120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:9c5cae62ecebd85d0d497b0ae144ecb45c9f10a137bd59b144351c9d09da6c98

Observation 4723c39e-4cc3-4f7d-8350-a19b0649b9f5 · outbound

This paper cites Unsupervised learning of visual representations using videos.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Unsupervised learning of visual representations using videos

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.568990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:c624a60c74d1cc48ee538bf0a3c2d6de1b8a10aede5727e9009814318cf8cf35

Observation fb44d790-7eca-4b01-8e83-85d44d9d1a76 · outbound

This paper cites Unsupervised learning of visual representations by solving jigsaw puzzles.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Unsupervised learning of visual representations by solving jigsaw puzzles

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.571331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:c6d8585c7d41ade835cec46604628ea0667899035d1f7775214be0d205698806

Observation f4619852-2280-447c-8202-519597bcd303 · outbound

This paper cites Colorful image colorization.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Colorful image colorization

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.573515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:bf42707458ce12f2c20fd3b924d5bd0788cf9265801664df5205349388b77b8d

Observation 08324d3f-53f6-433a-b574-a1e9211bf325 · outbound

This paper cites Masked Autoencoders As Spatiotemporal Learners.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Masked Autoencoders As Spatiotemporal Learners

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.282866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:65b3334105898516b337c6d3e6375b769106c9fcdd458611b09b6be1e4c8dfd2

Observation 769a66e5-c946-46e2-abf9-94ed606d4775 · outbound

This paper cites Unsupervised feature learning via non-parametric instance discrimination.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Unsupervised feature learning via non-parametric instance discrimination

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.374533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:06c2e6ef74ba04ee8c0ee1a2b76edec2554507a0e31820ab3b3b13c0b63f0aed

Observation 2f4b1f9b-58dc-4d18-9ce9-698117bd4380 · outbound

This paper cites Momentum contrast for unsupervised visual representation learning.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Momentum contrast for unsupervised visual representation learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.376775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:ea63ce0cae2a2eb65a5c1c4de9692c70bf651fa31d05c847b506856b1f4bfd03

Observation 54313e2e-2c6d-4fa7-bfb8-31add93d391e · outbound

This paper cites A simple framework for contrastive learning of visual representations.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning A simple framework for contrastive learning of visual representations

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.379017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:4f53afc1fb8e87453f20e32bfb1d514b7a225006d5d2cdaa80bbb70270ac6434

Observation a17948b0-4a8a-49c2-a591-67fba7f18663 · outbound

This paper cites Bootstrap your own latent-a new approach to self-supervised learning.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Bootstrap your own latent-a new approach to self-supervised learning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.381372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:41d9a232f657de7ee2996d699d5989e0de8ad4027cc859b26b95405fee97a771

Observation e98d55ec-8417-4f5b-b7ca-863d1716455b · outbound

This paper cites Exploring simple siamese representation learning.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Exploring simple siamese representation learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.383305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:6777811e247c6abcadb171739bec89229a935d13844bd1a8d4b4c4107606c4fe

Observation 077e4be3-83a6-4924-8e52-f8386b83e1b4 · outbound

This paper cites Generative pretraining from pixels.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Generative pretraining from pixels

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.385547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:fc7a8fac57cbf7460ba33d288c9573cb6b1cd74415865b0f98d126c7e83d278e

Observation cadc7453-4fd4-4dc2-ad4b-8fb9d1ee65e8 · outbound

This paper cites Zero-shot text-to-image generation.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Zero-shot text-to-image generation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.387682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:697ea19e6f90795114fb5a6b063a981e90965a3614a3df2c3dbbd8c3476c64e6

Observation 125ac4f9-8d1b-4e65-b57b-9c292d5d28cf · outbound

This paper cites Bevt: Bert pretraining of video transformers.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Bevt: Bert pretraining of video transformers

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.389686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:3e763f8784a1c4abac2ea0ef9d62828da6ae44e8b06563c16fa3feef91bc29ba

Observation 25cf0191-21ab-4879-bf31-b02cacfb6501 · outbound

This paper cites End-to-end learning of visual representations from uncurated instructional videos.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning End-to-end learning of visual representations from uncurated instructional videos

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.391711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:48eea6c71e472ef9aff0f9364e18fbf1d9a383c5419ccabec5517b847e988286

Observation 1eaf6c1f-7f3f-4dbd-8fe1-793a40eb840d · outbound

This paper cites VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.372393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:64fb95ba8450aa0def4c8fa1ac4052c49181148af51af1b8472a78e4963b4297

Observation 0d455d99-3861-4248-9217-1b71bfe27ee0 · outbound

This paper cites Scaling up vision- language pre-training for image captioning.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Scaling up vision- language pre-training for image captioning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.393606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:02dca55d090c1ec5f67961fa6711d6b30b8860c90a81c4e5788a93a2d4b3b191

Observation 45bb48ee-1534-460b-a89a-e4629a7ed4aa · outbound

This paper cites An empirical study of training end-to-end vision-and-language transformers.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning An empirical study of training end-to-end vision-and-language transformers

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.395470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:102f1b4deeba817a67742594d612fd88978d37156c2a81d34298087b39225ee3

Observation cb05c3d7-b7f6-4b6a-aace-20d7ebb73939 · outbound

This paper cites How Much Can CLIP Benefit Vision-and-Language Tasks?.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning How Much Can CLIP Benefit Vision-and-Language Tasks?

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.287756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:ca50ad317f9dfe20cd94b8b47b4f8db31b3c908197b20208834f2f5970faa82c

Observation 394409b5-09ee-46ff-b7a7-49f06cd927e2 · outbound

This paper cites FILIP: Fine-grained Interactive Language-Image Pre-Training.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.293435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:6a6c9fc087c321591738743663bf309417403f185619ffb7f3b38b245337451a

Observation 94c2227e-51bf-41c5-9628-3d813875b9b9 · outbound

This paper cites Murphy, and Cordelia Schmid.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Murphy, and Cordelia Schmid

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.397390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:f0a30ac36ad56d7bc2cf06c64b011d0c5819faa672473382d3f4fd3e9a1596ed

Observation a3db7f46-f5c6-47cf-ac64-c8eacf823504 · outbound

This paper cites Actbert: Learning global-local video-text representations.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Actbert: Learning global-local video-text representations

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.399479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:29cd7ded0ebf0e34f5feed56706a7391faa35f6c8dc70c596816711488c30583

Observation ab0c1b3e-9a54-42b4-9d74-cbac6ed77c38 · outbound

This paper cites Less is more: Clipbert for video-and-language learning via sparse sampling.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Less is more: Clipbert for video-and-language learning via sparse sampling

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.401395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:14cf9677820ec0226711d4880ef4b7c03266bf517827d791f34949e4606e6564

Observation ed286fd8-33ec-4656-841a-605e03c91755 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.403232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:8243ec791f79bdc0bbd9cc4eda15acbc9986f28212c6012ea82323c8b72d675d

Observation df6acf2d-ce08-4040-9031-ec1f5337f94f · outbound

This paper cites UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.339682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:f7b2e9c509e6dd76b29d8d98a2b5d1fad413f7ab5a4631e5694ae102a66c4da1

Observation daf2cd83-44df-4d94-9636-da67314f460f · outbound

This paper cites InternVideo-Ego4D: A Pack of Champion Solutions to Ego4D Challenges.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning InternVideo-Ego4D: A Pack of Champion Solutions to Ego4D Challenges

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.342706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:aaf5c80c2fbddb168249e96b942918f68135276c4567162dccb5c27476f77c75

Observation 125f031c-56af-4fb1-b9c2-54c26841c724 · outbound

This paper cites Vivit: A video vision transformer.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Vivit: A video vision transformer

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.405084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:b9fe5829fa79d78d0c8ef7501f3cd270a9d949ac37f7760bbe2507cffa8767e9

Observation a0f9a437-972d-46b2-8949-b337cfbdb66f · outbound

This paper cites Video swin transformer.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Video swin transformer

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.407101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:4a75995d870d3dd24d00ba3136fc3f13d8d689f0a677b68fac7655c34a6bda3a

Observation 3c835898-0ae6-401d-baf2-dc100efc610b · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distillation.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Align before fuse: Vision and language representation learning with momentum distillation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.409165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:aa6714bfdcd534ae1df885576ab0f65ae8e43d9c09a67f09d8e0991c58f17e45

Observation cb220de4-fd55-4eb4-88c9-1ec6f7541301 · outbound

This paper cites Bsn: Boundary sensitive network for temporal action proposal generation.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Bsn: Boundary sensitive network for temporal action proposal generation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.410972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:806671be9d8a2c456e1f27488d03b9eb1370c2117726477330a315ba6b798f02

Observation f153cb14-1f77-4a63-9bc8-deed1fe2a383 · outbound

This paper cites Bmn: Boundary-matching network for temporal action proposal generation.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Bmn: Boundary-matching network for temporal action proposal generation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.412743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:26d2c01198333c133f9cf0a849fb06c6c726758b649f8e9f5a42251ccafe2f37

Observation cfb36e00-385d-4d5f-b291-c5053f1307ec · outbound

This paper cites Augment your batch: Improving generalization through instance repetition.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Augment your batch: Improving generalization through instance repetition

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.414525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:2331b7dd05a93556fcdcd3e067fc8027774accb71a4f46c86350b173cab36d4b

Observation a556e19c-29e9-48f8-9636-437f237f3bcf · outbound

This paper cites Uniformer: Unified transformer for efficient spatial-temporal representation learning.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Uniformer: Unified transformer for efficient spatial-temporal representation learning

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.416433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:ec231fae118bf45e94289cd094e61bcf1fdb5d4f572fd552f32e7887c0b5c5f3

Observation 553d1d0d-c341-41f6-b7f2-3cc8e6a5c298 · outbound

This paper cites Temporal segment networks: Towards good practices for deep action recognition.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Temporal segment networks: Towards good practices for deep action recognition

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.418486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:583b1d9dc82a9384c7b26a5302f00429bc1481d254b325d5ca2950e1b564809c

Observation 4c47372a-b8a2-42a4-8a05-2c885d9bebc2 · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.420442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:ab2c313f4fa65a9516c939f284777819ab06232fb8af7a88db57939697f63a66

Observation 06eee4ec-f8b8-422c-8a69-cea7ef01bac1 · outbound

This paper cites Ava: A video dataset of spatio-temporally localized atomic visual actions.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Ava: A video dataset of spatio-temporally localized atomic visual actions

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.422380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:f80d67efad306d151c57fc92b82aac1461ec6bc5848da78729b8228730d2489a

Observation f365664f-37db-4ab0-9990-5cf0a2a5d10d · outbound

This paper cites something something.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning something something

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.424590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:b608e2cba7ecc6c0af99752701a1e2fa29ff0860adf6fb11d1227bf5adbd40af

Observation a79c3217-3bd8-4419-ac62-458aabe245d0 · outbound

This paper cites A Short Note about Kinetics-600.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning A Short Note about Kinetics-600

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:36:53.312698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:7484627977791030014584b79dda1aa5aa17c722d18a38a75508e99966dcc347

Observation 25073e3d-7956-4d69-b3ea-815b437f8a5d · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning A Short Note on the Kinetics-700 Human Action Dataset

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.317677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:f400540cd84b42d807a3206329db7dbd5ba0cec2a3660554d293081fcb4cf917

Observation a7473c7a-7460-4718-ae07-d8ac09ecd542 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:36:53.321624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:7943b0f7d13cf35812b60066e4214bbe63185d8ebb3b1dc1c7f63f81d581c1c4

Observation 49c0baeb-7464-4264-972c-23a6bdcbc89b · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Flamingo: a Visual Language Model for Few-Shot Learning

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:36:53.325661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:7ee18a62fc16fe4a124b18b6dee20a7b92547efad48a8a5d33e642a33f107bfc

Observation 28f5ce3e-010a-4d07-bc33-0c600825708b · outbound

This paper cites Ean: event adaptive network for enhanced action recognition.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Ean: event adaptive network for enhanced action recognition

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.426987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:322b7825ae5cf10d8834f51ca1bb60cddd7a2e3e19c64043ef286e45bc6f3aa4

Observation e6fb0d7a-bf64-464e-b8bd-c12b02062df7 · outbound

This paper cites Slowfast networks for video recognition.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Slowfast networks for video recognition

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.429068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:6f9a615c5eefed92070bda16aaba5ce6b32e159fab4bf55720edd1d977e096f5

Observation fe0f9f6c-549b-44d7-891d-afa458f2b9ed · outbound

This paper cites Temporal context aggregation network for temporal action proposal refinement.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Temporal context aggregation network for temporal action proposal refinement

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.431590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:84846f4af33b3abf01619d3d5c0d684d6ee86dab8657228d8b97785002baa718

Observation b520d146-266c-40ca-933f-6f726ee4d751 · outbound

This paper cites Tsp: Temporally-sensitive pretraining of video encoders for localization tasks.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Tsp: Temporally-sensitive pretraining of video encoders for localization tasks

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.434466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:e383ef805dc79fa368260f68c861a697a08f5d05f27203ed2c34c7eae97acc1a

Observation 2ab8484b-ddaf-4b2d-9574-27f03a4e1f38 · outbound

This paper cites Actor-context-actor relation network for spatio-temporal action localization.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Actor-context-actor relation network for spatio-temporal action localization

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.437790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:8b85974f8f6532ee2fbba72e4ed54d225fc6226d893680c9f230d3a5c18e970a

Observation 1cda90a1-64f6-429e-9734-f7175806b97a · outbound

This paper cites Relation Modeling in Spatio-Temporal Action Localization.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Relation Modeling in Spatio-Temporal Action Localization

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.348925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:22a86f0d2bec3f193e7c96dde90006b7a9ba62359090086a3e6b4773275558fe

Observation 1bbe1b0d-3e33-46ad-94c2-9f3cbdf14c85 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Activitynet: A large-scale video benchmark for human activity understanding

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.441544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:ea6f37b9faa8bae9b746011f7f85345c426a1f4c5ed8dd292a3e2798389898d1

Observation b058bb58-df2e-4099-800b-aac3f4266817 · outbound

This paper cites Hacs: Human action clips and segments dataset for recognition and temporal localization.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Hacs: Human action clips and segments dataset for recognition and temporal localization

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.444418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:326a1d941e0a84aa8b9cdeaec47f9e08eae79846c477eb56a06b8a237d4f9577

Observation 21385d48-09e2-42cc-a87e-aa3ed4f73611 · outbound

This paper cites Hmdb: a large video database for human motion recognition.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Hmdb: a large video database for human motion recognition

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.446490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:3f83f8d1af43c24d063bcefd8a7f4d9095aef1ab656d0f6f199b3aeeb86e6011

Observation a441b6ad-3c9a-4642-9523-12edc9633dfb · outbound

This paper cites in the wild.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning in the wild

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.448820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:699a3d7cbc69b7daa4c972baef752aadf68cebca6f565a51e92e51aa165b062c

Observation e58262bb-91a9-4abe-9fe1-5ecb4a666e31 · outbound

This paper cites Fineaction: A fine-grained video dataset for temporal action localization.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Fineaction: A fine-grained video dataset for temporal action localization

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.451543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:9dd0a1984bdbdb97d5eb5cb837d0adb710af98cf64233988fea0000ba299261f

Observation 305de34c-d1fc-4b85-9f6a-546e3a9cb41d · outbound

This paper cites The AVA-Kinetics Localized Human Actions Video Dataset.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning The AVA-Kinetics Localized Human Actions Video Dataset

Reference 84

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T00:36:53.369430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:7cf2c3c1bae96c0b5a39198c0d77c3e36a1671639613cd7af42bcb394f959352

Observation 95ff95e0-14bf-4c78-96b8-15bb76c2dd87 · outbound

This paper cites Microsoft coco: Common objects in context.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Microsoft coco: Common objects in context

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.454159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:a89d49b7ba3fb4c4a0b3dae620d67b797fd6d8a95ced71d463414b8eb73b3a5f

Observation 4bca6380-1687-4bec-a8b5-85130b0ec985 · outbound

This paper cites Mask r-cnn.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Mask r-cnn

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.456500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:9308e70d00f2f6d9e54a0c5b87c1034795fe1c7e0dc70f7d2ddf7f684412890a

Observation 395d12ae-d898-47ed-ac60-b6fc9cfe6bb8 · outbound

This paper cites Asynchronous interaction aggregation for action detection.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Asynchronous interaction aggregation for action detection

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.458816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:10337101f906e9ef4ef9a44cf1dc91793949de09d6113079df4ff68da94a9c33

Observation 4dc41450-fea2-4490-b286-f4a94198cb14 · outbound

This paper cites Ts2-net: Token shift and selection transformer for text-video retrieval.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Ts2-net: Token shift and selection transformer for text-video retrieval

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.461068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:96e1438c1f7e58ffa50b9506a26265492879482c7151ddc8268647107860af10

Observation 2fab92fb-d8e8-4498-bf7c-4907c4130824 · outbound

This paper cites Bridging the gap between learning in discrete and continuous environments for vision-and-language navigation.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Bridging the gap between learning in discrete and continuous environments for vision-and-language navigation

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.463399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:15a12628ed8b56823ce94842ff29f55ec9828fadff46c86c3c8ed3a77b17f46d

Observation 6a182d2a-0e0a-4dae-af1e-d1df9e27c009 · outbound

This paper cites 1st Place Solutions for RxR-Habitat Vision-and-Language Navigation Competition (CVPR 2022).

InternVideo: General Video Foundation Models via Generative and Discriminative Learning 1st Place Solutions for RxR-Habitat Vision-and-Language Navigation Competition (CVPR 2022)

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.298941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:e294e05eb5c1f8ea6cc69380482deaf5552cb335f4fcac08e59c3c56c36e2ec0

Observation 779a5061-293a-4889-b7ae-50a2d3404ad7 · outbound

This paper cites Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.304132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:df98ecfbcb3fc8fe6943b7c336ac5da72d295e104e75841d05aae252aca30373

Observation 6b002104-5b7f-4200-8540-ec9801a7d1a6 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Msr-vtt: A large video description dataset for bridging video and language

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.465903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:d206c007bf2368dd65a8370aaa0c98ac8fc34fe3b00cc6d25b842608ef501dbc

Observation 35ac8685-0577-4375-8322-4b7df5fdb57b · outbound

This paper cites Deep learning for video classification and captioning.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Deep learning for video classification and captioning

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.468135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:b31bdc0432e3a96d33e711dbb376b30dadb4dc3d063d949d2eefe6c36c5f19a8

Observation dc73f63c-03e0-4c40-b633-dcca7f1df363 · outbound

This paper cites Localizing moments in video with natural language.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Localizing moments in video with natural language

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.470503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:1b066ad17c5d0f4542a8f08f94db9dd9834cb5c44fb1817950fb25b691dbc849

Observation 4e29b554-3fed-44f5-b3c5-7e01be66b7ac · outbound

This paper cites A dataset for movie description.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning A dataset for movie description

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.473009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:ec1646192a02639dc48fe3a2d3feeb864f64ed52280c07ef2698392af0fc8774

Observation b8dc8f73-21ef-41d4-a06e-5d4b69662908 · outbound

This paper cites Vatex: A large-scale, high-quality multilingual dataset for video-and-language research.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Vatex: A large-scale, high-quality multilingual dataset for video-and-language research

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.475759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:35de63aeebf6b97ec7ebcd4881fd2cafbac185eed40adb9eb04bb11a7c366bfd

Observation 0ab77ef9-76a9-42c5-95ec-bc1c2c80ac1b · outbound

This paper cites Collecting highly parallel data for paraphrase evaluation.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Collecting highly parallel data for paraphrase evaluation

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.478587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:0d91e2284a9b6fa8b0580f66d19c16cba31b4e48fbf28bc00bf4d28a9a69ef76

Observation dc7b93fb-238a-4ac4-9d47-c0f4619a54a2 · outbound

This paper cites Tgif: A new dataset and benchmark on animated gif description.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Tgif: A new dataset and benchmark on animated gif description

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.481235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:6b62549fcaa888cf2bb3aeccef756b333aa7ebbe32060b1ebdf5833f08d8a442

Observation afc4b203-517b-48ba-9b13-feab2632f974 · outbound

This paper cites Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.484026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:66d0d4672dc9447a04f1c1f2fcb7908ed46ab3599dd10c65435d6104ad056bb7

Observation 355c2f4b-df10-4abb-bffc-11ef4e25dae4 · outbound

This paper cites Beyond the nav-graph: Vision-and-language navigation in continuous environments.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Beyond the nav-graph: Vision-and-language navigation in continuous environments

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:36:53.486843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:9267babeb8c95c23945821c0ebfe5e7fbcad06cb702ce280c15df35527eef4af

Pith citing papers

Observation 01890349-770e-4c2c-9b11-0c15604e7c9d · inbound

VideoChat: Chat-Centric Video Understanding cites this paper.

VideoChat: Chat-Centric Video Understanding InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.574468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:fea6eeca6d2de47d6ba605c10b48e38a5e6b84c2fde11fdcd395222921aad2ce

Observation dbc082f8-2ef1-41da-babd-92fd7526efe7 · inbound

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation cites this paper.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.574468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:f46f2e7480578b1d1d1861df0b95aca8a553c3d157c309217fb951ef2f8e13f9

Observation d0657f60-4e87-4833-8595-c8a29637bd6c · inbound

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment cites this paper.

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 195

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T03:27:59.069687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-17T03:27:58.952076Z digest=sha256:0958a57f48c9ba42773c7ff100a763f20e8a89181dfca5e88c84ac927e102599

Observation 5508df41-5c95-4f58-b5a3-c3927c243577 · inbound

LRM: Large Reconstruction Model for Single Image to 3D cites this paper.

LRM: Large Reconstruction Model for Single Image to 3D InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.574468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T10:11:00.777193Z digest=sha256:3976057cc39b5bbfa501321186032943bfca830062e921563aaa97563f7ce50e

Observation 107448f9-83d7-444d-b0ef-af0b68ae0758 · inbound

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark cites this paper.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.048363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:dc09adf421a127b178cd6ed9be941ed924ab3c53244faba5e38d23fcb47130b6

Observation cf9a3d9f-a4be-44ef-8b03-2752a1e04304 · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 152

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.574468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:6adbe395c0057f0f3a74d22c21db719fd81decf3c8b89eb3ef2d5cd9e19ea548

Observation 26568750-392f-4f45-b9e1-0dc79ca3b760 · inbound

Revisiting Feature Prediction for Learning Visual Representations from Video cites this paper.

Revisiting Feature Prediction for Learning Visual Representations from Video InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 287

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.574468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T12:40:23.709098Z digest=sha256:de30f9c6c50a2ad370aa7b99eb9a01a76252935be66f597eca45b4aef88bb55c

Observation 69b91bbb-8d7a-4fc6-9e55-b0d424f983e9 · inbound

Towards Long Video Understanding via Fine-detailed Video Story Generation cites this paper.

Towards Long Video Understanding via Fine-detailed Video Story Generation InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.473248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.473248Z digest=sha256:e8b6c2f3e74c8d0a0d86fa8a00ea04ea1b80114b98f8dcab7e3e652dde493dad

Observation 1f529530-4988-48b5-a831-8c247a5de911 · inbound

Multimodal Contextualized Support for Enhancing Video Retrieval System cites this paper.

Multimodal Contextualized Support for Enhancing Video Retrieval System InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:15:28.445628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T07:14:34.843867Z digest=sha256:8d7be27404089d77b68d713ce8048a7037abfb7ab07f5ce653931979f1e1178d

Observation 516ba0a7-066f-4d67-ba4b-452eaab05ab0 · inbound

Gramian Multimodal Representation Learning and Alignment cites this paper.

Gramian Multimodal Representation Learning and Alignment InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-11T14:31:23.951045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:31:23.951045Z digest=sha256:b226c49945051a7bc94a75bee531471ea6aba354d8ae1c4d07ce6ca1669706f2

Observation 31e0d036-f164-4b30-bb86-d74b0d0d4d93 · inbound

ShotVL: Human-Centric Highlight Frame Retrieval via Language Queries cites this paper.

ShotVL: Human-Centric Highlight Frame Retrieval via Language Queries InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:48.343022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:48.343022Z digest=sha256:7ca651cfa508cfb7ae39307fc468a8ae55f5567c8123755f872f464367e0cd7a

Observation a417401b-6c6e-49ee-a758-1f05bf83348f · inbound

Do Language Models Understand Time? cites this paper.

Do Language Models Understand Time? InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 192

Resolution
unresolved
no resolver link, observed 2026-08-11T12:47:17.963301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:47:17.963301Z digest=sha256:da3227bec925929e952d7e706909f9645fd6bb50cc60c89f9b8602346ccad825

Observation 0ce6ac81-ac32-4b5c-ac18-878f9ed0616d · inbound

Movie2Story: A framework for understanding videos and telling stories in the form of novel text cites this paper.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.222604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.222604Z digest=sha256:aa6dd04b340f90b3060877c8b4e126973ed9daea1b73432976eb6033be69e0ed

Observation 527a5034-1726-4fac-984e-db07468774e6 · inbound

VidCtx: Context-aware Video Question Answering with Image Models cites this paper.

VidCtx: Context-aware Video Question Answering with Image Models InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.387257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.387257Z digest=sha256:1d06201223c107d3c58371dd92cf8889b7bf819d8c931db52463484d4099034d

Observation 47ee8998-e4aa-4b7f-befe-9a345cd6ab89 · inbound

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries cites this paper.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T00:48:42.803699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:48:42.803699Z digest=sha256:e8e373bb020685ce174c9d561337960925f2bb768dbcb8456ab830a27b295209

Observation 9489c8d4-380e-4e33-aee0-489b66a22283 · inbound

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling cites this paper.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.544903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:43d116a78c6f4794a279cdfbee8d41aa2696a31a39d850aeb4a15ecb5577d9fd

Observation 8645d375-12b2-447d-bb3b-d448f37c19fa · inbound

LongViTU: Instruction Tuning for Long-Form Video Understanding cites this paper.

LongViTU: Instruction Tuning for Long-Form Video Understanding InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T21:23:58.007577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:23:58.007577Z digest=sha256:d96446c37c0cb182c6a950821854e7f06bd61048134f32fa9816d8edf8872297

Observation c63646bf-f532-4508-ba8d-8831ca267467 · inbound

An Empirical Study of Autoregressive Pre-training from Videos cites this paper.

An Empirical Study of Autoregressive Pre-training from Videos InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:08.522588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:08.522588Z digest=sha256:257927c2d2245bd7c59bfb7a5a40b36103b413d62df689be78642bd432111785

Observation 3a3acdbf-0840-45e7-9e6a-ed3fd685817b · inbound

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models cites this paper.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:12.294304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:12.294304Z digest=sha256:b77c11fe94e320cf480cb41fe679c3003875dce671866f61e134c7d8920d6164

Observation a1bf4bc0-51da-4e06-92df-b148982abc98 · inbound

Admitting Ignorance Helps the Video Question Answering Models to Answer cites this paper.

Admitting Ignorance Helps the Video Question Answering Models to Answer InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:28.840944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:28.840944Z digest=sha256:2a9228fca31ca473745c986ae76f4b5896712ea3dfa8b811d7abb2df9facfe6f

Observation 0c5358da-12bd-4211-8196-7cbcf3e17854 · inbound

SMART-Vision: Survey of Modern Action Recognition Techniques in Vision cites this paper.

SMART-Vision: Survey of Modern Action Recognition Techniques in Vision InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 173

Resolution
unresolved
no resolver link, observed 2026-08-10T16:32:15.523468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:32:15.523468Z digest=sha256:5e72ea47eab50f8eaf70f26dc18621fdf5ff6834eef4446a66be9347840f8b7e

Observation d55281f1-d88e-46cc-a9fb-13cecc4bdc98 · inbound

ReasVQA: Advancing VideoQA with Imperfect Reasoning Process cites this paper.

ReasVQA: Advancing VideoQA with Imperfect Reasoning Process InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T15:56:37.656464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:56:37.656464Z digest=sha256:9c647065309592b51aa79261da5e57340b4334f1ad8db2b9b761b3187d6b7f5e

Observation 3dc38dd8-50f9-41c9-8692-9e850f9b14ca · inbound

SpatioTemporal Learning for Human Pose Estimation in Sparsely-Labeled Videos cites this paper.

SpatioTemporal Learning for Human Pose Estimation in Sparsely-Labeled Videos InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T14:43:38.979512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:43:38.979512Z digest=sha256:f09c0ce7d66ae71434ae12fb386b589e180f37a41113e56ae3f7f31f48b17c01

Observation 6ead0cae-eefb-4090-bfb7-abaf41a42baa · inbound

Understanding Long Videos via LLM-Powered Entity Relation Graphs cites this paper.

Understanding Long Videos via LLM-Powered Entity Relation Graphs InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T13:53:27.070979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:53:27.070979Z digest=sha256:5965c8ca4fec59e20bfdcdf0b0e829a38931258235d8db55cad16bacd10d69bf

Observation 9e1eaf42-8b49-4753-b98c-53e90cf689e3 · inbound

XRF V2: A Dataset for Action Summarization with Wi-Fi Signals, and IMUs in Phones, Watches, Earbuds, and Glasses cites this paper.

XRF V2: A Dataset for Action Summarization with Wi-Fi Signals, and IMUs in Phones, Watches, Earbuds, and Glasses InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-09T21:37:01.129396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:37:01.129396Z digest=sha256:ce95856f11e801ef26c88b6e76a2eb56acabc98eab676546e184811ce65651cc

Observation 18a5d58a-d8dc-486a-bcbf-96dcf678f6e4 · inbound

Expertized Caption Auto-Enhancement for Video-Text Retrieval cites this paper.

Expertized Caption Auto-Enhancement for Video-Text Retrieval InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T10:51:29.201639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:51:29.201639Z digest=sha256:2b298306ea588707ea642d5ec87e867b573288adf575803500e320c52d06fa7a

Observation dc9bac64-023a-4caa-b918-1bec63136a6a · inbound

VideoRoPE: What Makes for Good Video Rotary Position Embedding? cites this paper.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.190284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.190284Z digest=sha256:13a655fb4d8d700efcca5af361a6efeecf093f974441236419efd6ed56836f92

Observation 6d980b77-377e-436e-b647-a44a6c4e628a · inbound

Generative Ghost: Investigating Ranking Bias Hidden in AI-Generated Videos cites this paper.

Generative Ghost: Investigating Ranking Bias Hidden in AI-Generated Videos InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T13:12:05.071047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:12:05.071047Z digest=sha256:0855e5e0a75f32192a642426bcf63f44264773d57a87e154aed139ed9114000c

Observation aa1a084a-51b0-4e57-94fe-f26759cc1cfd · inbound

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks cites this paper.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.260811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.260811Z digest=sha256:4545c58beb6f6791f20a05145fc5a133d615faae7f703d81282c5d53731a3e0d

Observation c92fe79b-13d1-4bc5-a59a-c97df7637a2d · inbound

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval cites this paper.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.013607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.013607Z digest=sha256:d1392c381a3fb20c0380d4aa707487913ac029a4a0b9cb518e16d07719e6e614

Observation 8401b309-0b16-4b9d-ac1b-30f5586f3016 · inbound

HuMoCon: Concept Discovery for Human Motion Understanding cites this paper.

HuMoCon: Concept Discovery for Human Motion Understanding InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:15.328435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:15.328435Z digest=sha256:86c0d9f6b3dc8bad1e9c8bb2c5744d8d65148dcef7329c87350628a53a4ebf8c

Observation a19f1bcc-2c7e-4e11-9527-3b1c0426aa55 · inbound

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering cites this paper.

Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:33.191576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:33.191576Z digest=sha256:46d9878fc2c69a321d16a55d0300375470a57a07922c071ac3a39042d2b7c1f3

Observation 348e432f-cd03-416b-8af0-23bf91de4165 · inbound

Improving Keystep Recognition in Ego-Video via Dexterous Focus cites this paper.

Improving Keystep Recognition in Ego-Video via Dexterous Focus InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:06.099409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:06.099409Z digest=sha256:45ea00454ca89ed762974a54ab44e1017cde8d98283dfdf9f98ad5fff671c52c

Observation 00fe6806-3f7c-4d4b-9759-d075e8f50405 · inbound

Large-scale Self-supervised Video Foundation Model for Intelligent Surgery cites this paper.

Large-scale Self-supervised Video Foundation Model for Intelligent Surgery InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:22:59.564896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:22:59.564896Z digest=sha256:69d4466436cd4fd92a566a1993d43675faf6bdc2d0decdeb94af4f3050ccc87c

Observation 0ba4e3ee-4df0-48c8-b7f6-ca213a40175a · inbound

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs cites this paper.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.495697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.495697Z digest=sha256:da78d4a5be18228ab929466aa9a6d38755041e55164966e95edee3bb27763943

Observation e067d0b9-cc70-41ab-840e-dcdd53d4e93d · inbound

EgoM2P: Egocentric Multimodal Multitask Pretraining cites this paper.

EgoM2P: Egocentric Multimodal Multitask Pretraining InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 109

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.620432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.620432Z digest=sha256:ccbb8422b945983800073015fd3546c6311b27a8daeffba0c7b3a833bb21d92a

Observation 300d6d22-2ed4-4a14-89bc-5e3f9fe69607 · inbound

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval cites this paper.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.612750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.612750Z digest=sha256:b58b029e3fc14d95f0c1863b0e1bd1f240d9d7e944ee81f72dbe7b5c4f9cdefd

Observation 1afc1f04-1945-446e-8181-a52b3e8865f1 · inbound

Vision Generalist Model: A Survey cites this paper.

Vision Generalist Model: A Survey InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 177

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:02.280523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:02.280523Z digest=sha256:97575e39eb589a27696843338ba25e2cdb4c35c70d39fbad9c9b41fb34c01095

Observation 90ff68e9-ab1c-4fff-90a3-784606db2304 · inbound

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model cites this paper.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.311995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.311995Z digest=sha256:ed9ff4a136875cdb0ad17ec75b6e410c84cd71cc4fe29fe3661cb4d689fe862e

Observation c5d6b164-6f58-4f03-a66e-b3565b0bde8d · inbound

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization cites this paper.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:25:59.990074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:25:59.990074Z digest=sha256:97b5d90898e11f19d94442cfd05f983ebe6e2ed665978debcd5d3776811918f2

Observation b084f60e-1e9e-47d1-943f-88b0537bc24e · inbound

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes cites this paper.

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T22:37:39.015632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:37:39.015632Z digest=sha256:e9ee09be1632116bb34894a5cf8bd0c85abb75f9b4a3ecd19d2c22828c24840f

Observation a2fe6278-5c86-4b72-9064-55b7470acee5 · inbound

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 cites this paper.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:20:13.123695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:20:13.123695Z digest=sha256:36b35e7dffa3705f4e9596ad228d49043de3e4bc1d1f12be3338ce1a06d718ec

Observation 6f5abedc-1bfa-448b-bbb9-58b7c223be50 · inbound

MANTA: Cross-Modal Semantic Alignment and Information-Theoretic Optimization for Long-form Multimodal Understanding cites this paper.

MANTA: Cross-Modal Semantic Alignment and Information-Theoretic Optimization for Long-form Multimodal Understanding InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:00:07.336150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:00:07.336150Z digest=sha256:5c73e1fe3a672c3df7c825a97c6dffcc41d7ce1989d339fb9f17694e3d281ac2

Observation 8bfbe9db-ce8c-43be-b186-adac9a3252d1 · inbound

SciVid: Cross-Domain Evaluation of Video Models in Scientific Applications cites this paper.

SciVid: Cross-Domain Evaluation of Video Models in Scientific Applications InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:10.801532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:10.801532Z digest=sha256:e71b3a887110e69155741b2ce63e78c16c901bd5bdc065c641b0d13c92daafb2

Observation b9429fee-09fc-47cf-8c99-f2f74abd21c6 · inbound

Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning cites this paper.

Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:52.732569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:52.732569Z digest=sha256:fbe3c017a1c45ef274300690cf82cec15af627236880ae7bcfb11922c5e07c41

Observation 0aefedc9-f462-4c9a-b3b2-8c3d5401cc01 · inbound

Semantic Frame Interpolation cites this paper.

Semantic Frame Interpolation InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T19:39:51.612774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:39:51.612774Z digest=sha256:0f5c14415e1757c33150902056d934c64daec834d15ee596049bf85f84dcbea1

Observation dae05960-f941-4840-985e-c37e44273954 · inbound

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding cites this paper.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:04.192757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:04.192757Z digest=sha256:f60096398ff55f68f88595681f0ff1cf31a09211bbde2ab5c5bb4bb7c4b2d1f8

Observation 90af2435-5042-4201-ad4c-2e1abd6c5506 · inbound

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning cites this paper.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:22.840001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:22.840001Z digest=sha256:e7f5ebfee2a6614f4735b2de6e1582f1376373e8de2dac06e98f12b39ec25e73

Observation 970d892b-4ebf-4a2c-857f-c384ef6afc3b · inbound

LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering cites this paper.

LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T15:53:33.985536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:53:33.985536Z digest=sha256:6e6568a0efeebd305f3d2cfea7554dc207eabf2900cd7066719c274433bed716

Observation f72bac2e-7def-4f46-a900-25c955dfb1f9 · inbound

A Survey on Efficiency Optimization Techniques for DNN-based Video Analytics: Process Systems, Algorithms, and Applications cites this paper.

A Survey on Efficiency Optimization Techniques for DNN-based Video Analytics: Process Systems, Algorithms, and Applications InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 137

Resolution
unresolved
no resolver link, observed 2026-08-06T15:30:20.026745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:30:20.026745Z digest=sha256:6e61684baee8adf21621e28a0076c7afde5dc61086a8e922af6ff1536122618a

Observation 2228c1a3-fef8-4a05-9d27-bde4302af7f0 · inbound

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding cites this paper.

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:17.726835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:17.726835Z digest=sha256:e8c6dce1186c2981f66b0c393e23954fdf5a7b4894282a32a485a08142290262

Observation 6a9aaa68-a589-4f06-9676-baa83d4ef10e · inbound

Object-centric Video Question Answering with Visual Grounding and Referring cites this paper.

Object-centric Video Question Answering with Visual Grounding and Referring InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:37.661534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:37.661534Z digest=sha256:1a26e861ea6d0d622a12a03a47859503b1309c5ea93a827a2ddeff6f2487a892

Observation af83d3d0-c08e-4df0-801c-d839ffb69237 · inbound

Representation Shift: Unifying Token Compression with FlashAttention cites this paper.

Representation Shift: Unifying Token Compression with FlashAttention InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T10:17:37.471841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:17:37.471841Z digest=sha256:ca1757ab74ee0a16b3bff84ea1dc2c1a1454a8f006c5df97027f2f4efd064379

Observation d3969c71-fd3e-4fd9-a942-9c16d6e17c84 · inbound

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding cites this paper.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:36.666749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:36.666749Z digest=sha256:c718c6b17313d8956f3dc5f08431e92c925a7c616da5b65759dd482a61707485

Observation cdb02b9b-67d9-4830-be24-be0b0b3b54f8 · inbound

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering cites this paper.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:41.380659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:41.380659Z digest=sha256:94b498422c64c0ff56736b48491d8172a65e7082a6e088a07cb9395a6a95ab59

Observation b6867581-5b02-4758-9364-ed0c7a825638 · inbound

A Survey on Video Temporal Grounding with Multimodal Large Language Model cites this paper.

A Survey on Video Temporal Grounding with Multimodal Large Language Model InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.441242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.441242Z digest=sha256:895ae8cd2b160610184a0fb1680264439436a0afb48fc4bd73c746f949d1ae61

Observation 42adbfc4-cee6-4148-924c-fb69c2c67615 · inbound

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding cites this paper.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:01.537622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:01.537622Z digest=sha256:87c581d755076a5ec92ac77f64137626b7135add32928e6fc3aade2238c36a4b

Observation a0fc6081-e0ab-4d7b-aab9-69f4f9ca65f5 · inbound

Video Understanding by Design: How Datasets Shape Video Models cites this paper.

Video Understanding by Design: How Datasets Shape Video Models InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.645943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.645943Z digest=sha256:a36a819de9f96cc99a5420e278946434eee10009e66db378a8a8d19bdc048cd2

Observation 83e97f27-3ede-4322-9404-30dabd8d60ff · inbound

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs cites this paper.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:57.880798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:57.880798Z digest=sha256:2a75b6f64d803a368c95693f2988f7ead93e91d121d12d00817c1e14ffa70693

Observation 4cab872b-e3ab-488b-b404-c942bc03d25f · inbound

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency cites this paper.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:35.256568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:35.256568Z digest=sha256:96f8930c29c5af359e3ef60c3bc333b4c05ae1b7e27177db0997dc75d83ef8bf

Observation 24f8e7fc-8aca-4f6f-b9a0-9fe7d3c6f1cb · inbound

Privacy Beyond Pixels: Latent Anonymization for Privacy-Preserving Video Understanding cites this paper.

Privacy Beyond Pixels: Latent Anonymization for Privacy-Preserving Video Understanding InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:15:26.841249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T23:12:48.977906Z digest=sha256:68071a5e7a1ef28d0cca5a7bbf07efc4e3c0f6c715eb2ff9204ece5cde580069

Observation 54858588-607d-4546-a315-28a6f8808edb · inbound

Calibrated Multimodal Representation Learning with Missing Modalities cites this paper.

Calibrated Multimodal Representation Learning with Missing Modalities InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:10:22.648940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T22:08:07.217659Z digest=sha256:ec5528c64fcdaed7e76b057e5ae6d0577782bf6695a32702f7d82ef2d9d1fdb8

Observation 0bd9ef4e-cd8a-4dfd-8f4b-c25aa71efffe · inbound

MemVerse: Multimodal Memory for Lifelong Learning Agents cites this paper.

MemVerse: Multimodal Memory for Lifelong Learning Agents InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T18:49:01.911861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:49:01.911861Z digest=sha256:d1ef824be25b372b509561cca9669a8a4945e8ed0b46df41be0e052f8e93fb82

Observation a8c5d2f2-9c0e-49b3-b0e0-b42e0379d4fd · inbound

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning cites this paper.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.306944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:7f291d1e32a99367b364db6fc0dd74259b6ea916c6d63caf669ae0e326d32b88

Observation 37572fa0-5626-4f10-9f5a-258f3dde406e · inbound

Streaming Video Instruction Tuning cites this paper.

Streaming Video Instruction Tuning InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.574468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T19:44:11.032898Z digest=sha256:e147a9d92c97697b0394602947e7eebe3ed899ddd4bb3c90492e8e08968da1af

Observation bfaab06c-71e3-421d-9609-00e72cfd5b0b · inbound

A Paradigm Shift: Fully End-to-End Training for Temporal Sentence Grounding in Videos cites this paper.

A Paradigm Shift: Fully End-to-End Training for Temporal Sentence Grounding in Videos InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.574468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T19:43:57.298338Z digest=sha256:4107339d047be35d611e569370d462704ac7feb44d973ec4886d6759ea0e5538

Observation cb0ed3f3-8518-488b-8bc7-54393e26bd03 · inbound

UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding cites this paper.

UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T00:36:53.574468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T17:55:56.385801Z digest=sha256:9955415c030d5cdead15f49248668065b9448253ba0a6b953968468f61e72751

Observation 2f3da1ce-7e24-4264-83c9-dc24e3758d6c · inbound

InstrAct: Towards Action-Centric Understanding in Instructional Videos cites this paper.

InstrAct: Towards Action-Centric Understanding in Instructional Videos InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.574468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T17:07:49.334104Z digest=sha256:d0bfdba53e2574d92fc7fa5f929806d0bc9c7069e2e2cfc224294adcc3ce9595

Observation 5e9056e4-6d5b-4ad1-a422-e1c4b943801b · inbound

Efficient Spatial-Temporal Focal Adapter with SSM for Temporal Action Detection cites this paper.

Efficient Spatial-Temporal Focal Adapter with SSM for Temporal Action Detection InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.574468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T17:31:28.091617Z digest=sha256:abab6592b7aea4352a980f5f6a1e65b2c4e670fd7b80df262a43766fef01004b

Observation 80c999f0-f809-4892-80c9-05670df423ee · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:84a9ee8ab8339be8e66f6f567f6b280e56c77150ce26bfe60a8e2c77efbf749f

Observation 4c3bf265-d301-486d-a0dd-5db372730b9f · inbound

V-Nutri: Dish-Level Nutrition Estimation from Egocentric Cooking Videos cites this paper.

V-Nutri: Dish-Level Nutrition Estimation from Egocentric Cooking Videos InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.574468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T15:00:21.173667Z digest=sha256:d28cbc7202ddc961d6e6223c24255c0b0c7b43708e9e251a063e2740e47252d3

Observation 5abab684-03bd-4f0f-94bf-1dca71245f25 · inbound

Training-Free Semantic Multi-Object Tracking with Vision-Language Models cites this paper.

Training-Free Semantic Multi-Object Tracking with Vision-Language Models InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.574468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:12:08.080882Z digest=sha256:c1c3ba04386109dd405a0d4f95f48f324fb4f11ac0e62e355acb1fc3ba156149

Observation 9b6d375d-c00b-435d-a506-c1877d4b37fa · inbound

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding cites this paper.

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.574468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:28:58.920442Z digest=sha256:f7d98bf4f37a53df1f5d1a1e60acb44f759c23bb4c74204432fa555fbff0c8e1

Observation d83e73e6-b449-424e-a5af-6322efd980fc · inbound

FreqFormer: Hierarchical Frequency-Domain Attention with Adaptive Spectral Routing for Long-Sequence Video Diffusion Transformers cites this paper.

FreqFormer: Hierarchical Frequency-Domain Attention with Adaptive Spectral Routing for Long-Sequence Video Diffusion Transformers InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.574468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T15:54:34.742037Z digest=sha256:c7a13d2b2677e69f1bc3427241d0316f2a725879c88b19f58d42c5caf5802cd8

Observation 87782323-821d-4a64-9cac-a087cc63c553 · inbound

TransVLM: A Vision-Language Framework and Benchmark for Detecting Any Shot Transitions cites this paper.

TransVLM: A Vision-Language Framework and Benchmark for Detecting Any Shot Transitions InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.574468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-07T07:23:19.988849Z digest=sha256:48e33f3ab8c0787de0055b7f8519803af6659e7015b3f590d2975c51d9590e27

Observation 6bcde318-897e-4397-a625-9d49d956b91c · inbound

LoViF 2026 The First Challenge on Holistic Quality Assessment for 4D World Model (PhyScore) cites this paper.

LoViF 2026 The First Challenge on Holistic Quality Assessment for 4D World Model (PhyScore) InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.574468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T16:20:40.024442Z digest=sha256:58d980f35959a5ebfcc50667123289f6c1044ecee8a68b5bf26dd2ee625c9679

Observation 12b3b07f-e37d-4890-a123-a673735728e3 · inbound

CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives cites this paper.

CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.574468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T05:38:21.587653Z digest=sha256:5023bbe36b55b8c0aa12fc51200947d62696f7fde4b58f5202dc2b54ea83e61e

Observation 540ed8ae-eb16-47ec-84c2-4491799d41fd · inbound

Towards Unified Vision-Language Models with Incomplete Multi-Modal Inputs cites this paper.

Towards Unified Vision-Language Models with Incomplete Multi-Modal Inputs InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T13:43:28.717924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T13:39:55.838296Z digest=sha256:8d0b23330e462c711ff43cb5612dac1d8968c394ddd6bd7782d9faa512777945

Observation 548a9304-72f2-43f2-a259-ec377032e5c8 · inbound

Masked Diffusion Vision-Language Models for Temporal Action Localization cites this paper.

Masked Diffusion Vision-Language Models for Temporal Action Localization InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:03:13.675969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T07:56:52.524815Z digest=sha256:0167efab176da30dd4de05178e8e02f8871cbac64fb153bf25304b7ec9e2ffc6

Observation c32974c2-d111-475b-8892-da0514548ff6 · inbound

SlotMemory: Object-Centric KV Memory for Streaming Long-Video Generation cites this paper.

SlotMemory: Object-Centric KV Memory for Streaming Long-Video Generation InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:25:59.915450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T22:44:14.089224Z digest=sha256:b742072ed48e98ff8a842d5d8414650716982552c690667a18e8e417fe889ed0

Observation 15d7c3b1-9c1e-4b97-8b57-739a3057391c · inbound

Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval cites this paper.

Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:06:16.675454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T15:44:59.586877Z digest=sha256:a19cfaae237df41819529d8c8261bc5244832460c366cb11556726fad87b8de5

Observation 3bb252a5-28db-4320-8be4-cd59e35bf2ed · inbound

Hand Trajectory Fusion for Egocentric Natural Language Query Grounding cites this paper.

Hand Trajectory Fusion for Egocentric Natural Language Query Grounding InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:16:23.361936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T14:35:59.044093Z digest=sha256:e5888602fb904be617437150ee40a74b6d11736db2c10ea3ee328afe18864617

Observation 02c8d935-0e67-4f5c-bb43-a3de99cb64fd · inbound

Where Do We (Not) Need Temporal Context in Low-Resource Video Task Adaptation? cites this paper.

Where Do We (Not) Need Temporal Context in Low-Resource Video Task Adaptation? InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:56:28.973260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T10:31:10.409674Z digest=sha256:ee7f45164ebfd4e241fdeb0d19bea9db825e9b7611641cc08d93e3f44443a345

Observation 783c631f-84aa-4d63-91a5-94c250fdfd21 · inbound

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding cites this paper.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:07:12.866731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:96bd0f5a0abd15f0cb55c75d348826380d254259203fe63f40755b79b65e7b03

Observation 80ea8512-3fdd-4a73-a8a0-9b80926e8809 · inbound

OmniGen-AR: AutoRegressive Any-to-Image Generation cites this paper.

OmniGen-AR: AutoRegressive Any-to-Image Generation InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:47:29.719943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T17:05:16.883488Z digest=sha256:53a0e5b6f851a4f48919776507fbc383dd4063408065a696be9fd670ea4712c5

Observation 819f3b21-28cb-44ed-80f0-95472100bc9e · inbound

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning cites this paper.

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:17:29.029387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T17:21:38.543724Z digest=sha256:73d3ed2d6470f926551e7e784fdce3392d10362a153086761c4ea0eae0e7e82b

Observation 0370a05a-f20c-4ca0-81a3-e3dabcc01d6a · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 118

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:03.081920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:ff0789b624d39cd229955537b9ac1bcffa5f908b5827ac2521c2855fae886979

Observation c65d055d-8457-4dd4-ad6e-f5aea84e5f40 · inbound

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers cites this paper.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 259

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:28:32.148227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:e0ab8deed24df7382f63d290e52ecc63ab5f82bfdf2fb30ff9a9dceb9f3f6456

Observation 298abd9b-abef-4d23-bef4-ee332f59169d · inbound

CineOrchestra: Unified Entity-Centric Conditioning for Cinematic Video Generation cites this paper.

CineOrchestra: Unified Entity-Centric Conditioning for Cinematic Video Generation InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:58:32.699908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T06:51:31.625416Z digest=sha256:3c2600b24c2174a5790414f74f90eb0df252e70fb7f6c9cff2be7cde95aa699b

Observation 896845f4-3d43-4c4e-b595-bd8a8b081d21 · inbound

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams cites this paper.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:38:56.183828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:d7bbb3b938a63522043552da8818dbfe8ca0389d338f21eca39188959c9021e8

Observation cbf220b3-9e5e-4017-9a65-c6385dd757a5 · inbound

ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval cites this paper.

ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 55

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T05:49:36.553029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T15:34:55.062016Z digest=sha256:4e197bee9307d1f8dba810165c88a72e1e94be0f043cf533a9768e2560029ccc

Observation ba56b6ce-d30e-499e-a94c-ef4e93ca4736 · inbound

T-MOR: Learning Motion-Aware Skeleton Representations for Human Action Recognition cites this paper.

T-MOR: Learning Motion-Aware Skeleton Representations for Human Action Recognition InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:09:37.770438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:44:41.975446Z digest=sha256:d8443a53b49e07aceabcd5043a526a38ab3d15da085fc18f5e911ad8468e7b03

Observation 5144b89b-7a7f-47d7-a87e-287d799f017e · inbound

TuringViT: Making SOTA Vision Transformers Accessible to All cites this paper.

TuringViT: Making SOTA Vision Transformers Accessible to All InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:29:57.892760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T00:29:41.291832Z digest=sha256:4a115e5605f7c697f384e6a4140caccbbf19cc5b5bd76a6d19809e5ad227a773

Observation 51a988bd-d8ea-4162-a98f-522828bf77c1 · inbound

TuringViT: Making SOTA Vision Transformers Accessible to All cites this paper.

TuringViT: Making SOTA Vision Transformers Accessible to All InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-06-29T15:03:32.165860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T05:32:26.746776Z digest=sha256:05af20d35d3f4e1043b3fadf7906291f87d49bd017c1890d86b072bfc45355e9

Observation 10770e63-f3c8-4620-824e-8f9dae2936e3 · inbound

Toward Low-Latency Vision-Language Models with Doubly-Correct Predictions in Egocentric Visual Understanding cites this paper.

Toward Low-Latency Vision-Language Models with Doubly-Correct Predictions in Egocentric Visual Understanding InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-04T17:20:00.467194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-25T23:47:12.032082Z digest=sha256:0e945a111b6123d4d446774b79302b8c29cea03dfaecaf45c4b91d5fb2e34a1f

Observation 8e3dfb14-5527-4d05-8da4-b95867e89cb9 · inbound

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction cites this paper.

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:54:22.348256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T07:48:01.719339Z digest=sha256:0c1a008e9b61bfff8b40c4124445e0b95544a6126976ea8d53eba2f4aca17af0

Observation 6c80b62f-cce4-4116-aa83-1fa1c792e7a0 · inbound

Prompting-MammAlps: Fine-Grained Text-to-Video Retrieval for Camera-Trap Data cites this paper.

Prompting-MammAlps: Fine-Grained Text-to-Video Retrieval for Camera-Trap Data InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-07-14T14:53:56.693464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:53:56.693464Z digest=sha256:5b3619ca38e4b852187bc4e3a3efc55d0edd6aeee46baf83475938d4ee9d9394

Observation 9a0f75f7-93bf-4245-8d69-58171fbfda0b · inbound

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory cites this paper.

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 81

Resolution
unresolved
no resolver link, observed 2026-07-14T12:26:27.446079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:26:27.446079Z digest=sha256:b9bdca37423113608a1c55848261d3a6cfa40b8674179b49ddcbb5506f0e2eb3

Observation ab0e695a-a4a4-444a-a26d-3d23e63848d7 · inbound

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory cites this paper.

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-02T07:21:32.526822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:21:32.526822Z digest=sha256:019fbd8f224389d222681a2f8956737ada6d64a798c63dd620feccf50308b8b3

Observation 1713cb65-6caf-4f72-b2fc-2735c535ed09 · inbound

VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders cites this paper.

VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T02:51:46.752509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:51:46.752509Z digest=sha256:2d9cce8c88975566d9cf09deab71b64911d4983df1be406773eccdf832e70e91