Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T00:36:53.235740Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 100 of 111 outbound references and 100 inbound Pith citation observations for arXiv:2212.03191.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T00:36:53.235740Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T19:59:13.473248Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
100 of 111 outbound references displayed
External citation measurements
93
pith, observed 2026-08-05T02:28:24.338817Z
Observation 3904b6a7-d161-4169-b3a8-96a507d5c664 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Nsnet: Non-saliency suppression sampler for efficient video recognition
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cbffbc11-917c-4ad1-94dd-385cab12fc97 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Learn to cycle: Time-consistent feature discovery for action recognition
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dd35ef7d-c1ef-4b64-bb92-257903cfba94 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Self-supervising action recognition by statistical moment and subspace descriptors
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4551cede-8f60-46b9-b39f-c6229ae96a31 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Actionformer: Localizing moments of actions with transformers
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e626acfb-cab5-4f20-a700-3e5a8e8b8b37 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 57446b8b-8a03-4a00-80e3-ad5434efd26c · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Masked feature prediction for self-supervised visual pre-training
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 041299c2-4743-4bdf-a2d4-5450b4b6993a · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning CoCa: Contrastive Captioners are Image-Text Foundation Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation afaab448-fd03-4875-8d07-d6e643c521f0 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Multiview transformers for video recognition
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 66281555-6776-43b3-9d94-1a1e57adb8a1 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Merlot: Multimodal neural script knowledge models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ff7620d4-1be5-4c82-843f-88a18b392ccb · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning On the Opportunities and Risks of Foundation Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 72dfe7fc-0cf1-4782-9d65-7bbbd4f9d610 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1df9fb52-dc1e-4933-bc0d-c08c2db0eec2 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning INTERN: A New Learning Paradigm Towards General Vision
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e4da4d3c-8d47-45b6-b5ff-48a6252475e3 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Learning transferable visual models from natural language supervision
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9b2314a5-8cee-4968-9d4b-1cc13321daf3 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Scaling up visual and vision-language representation learning with noisy text supervision
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 109f2a8e-3b42-4a52-a230-15d8417d5650 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Florence: A New Foundation Model for Computer Vision
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4f1d64d9-2c27-4878-a18f-787c3dc6ce92 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Unified contrastive learning in image-text-label space
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ba08958b-c143-46ad-8cf5-da2b9005a221 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning SimVLM: Simple visual language model pretraining with weak supervision
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a06c1fe5-9bea-4dad-a1f9-59470a7f62ec · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f2dbac5c-c1d7-41a8-9f03-5ea139503240 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 398788ad-629b-4001-b807-112e6bf10745 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning BEit: BERT pre-training of image transformers
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f564919a-d241-4b27-8dc5-089f0e45c941 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Pathways: Asynchronous distributed dataflow for ml
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 10b34b71-aa82-49fb-893b-7c79dec90fe5 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning X-clip: End-to-end multi-grained contrastive learning for video-text retrieval
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e835e721-e8bb-49eb-917e-7d8738d2da5b · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d592edaf-3d7a-463e-a2e4-c647938fae95 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning All in One: Exploring Unified Video-Language Pre-training
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ed71fbf0-5b96-40a4-8947-b5943e26dab4 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Masked autoencoders are scalable vision learners
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f6427790-add4-4399-ba0e-9dd4f2a67756 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Learning Spatiotemporal Features via Video and Text Pair Discrimination
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a1d8d62b-3f27-4931-b883-56bd2c53aaa5 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Quo vadis, action recognition? a new model and the kinetics dataset
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1e26c95d-0036-43ad-9c94-acc008c81bca · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning An image is worth 16x16 words: Transformers for image recognition at scale
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 97522ffd-971f-482e-92f4-fa7cc729d036 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning MultiMAE: Multi-modal Multi-task Masked Autoencoders
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f1b3dfd1-d7f2-4246-a5ad-9a97e7ede817 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 43d26c3f-fb83-487a-9878-515c56343167 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning LAVENDER: Unifying Video-Language Understanding as Masked Language Modeling
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 37845baf-0dba-4d3b-915d-9899e0d124a2 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Merlot reserve: Neural script knowledge through vision and language and sound
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ae004afd-d140-46f7-b35a-d7c0c2c74132 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Unsupervised visual representation learning by context prediction
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4723c39e-4cc3-4f7d-8350-a19b0649b9f5 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Unsupervised learning of visual representations using videos
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fb44d790-7eca-4b01-8e83-85d44d9d1a76 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Unsupervised learning of visual representations by solving jigsaw puzzles
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f4619852-2280-447c-8202-519597bcd303 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Colorful image colorization
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 08324d3f-53f6-433a-b574-a1e9211bf325 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Masked Autoencoders As Spatiotemporal Learners
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 769a66e5-c946-46e2-abf9-94ed606d4775 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Unsupervised feature learning via non-parametric instance discrimination
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2f4b1f9b-58dc-4d18-9ce9-698117bd4380 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Momentum contrast for unsupervised visual representation learning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 54313e2e-2c6d-4fa7-bfb8-31add93d391e · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning A simple framework for contrastive learning of visual representations
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a17948b0-4a8a-49c2-a591-67fba7f18663 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Bootstrap your own latent-a new approach to self-supervised learning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e98d55ec-8417-4f5b-b7ca-863d1716455b · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Exploring simple siamese representation learning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 077e4be3-83a6-4924-8e52-f8386b83e1b4 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Generative pretraining from pixels
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cadc7453-4fd4-4dc2-ad4b-8fb9d1ee65e8 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Zero-shot text-to-image generation
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 125ac4f9-8d1b-4e65-b57b-9c292d5d28cf · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Bevt: Bert pretraining of video transformers
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 25cf0191-21ab-4879-bf31-b02cacfb6501 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning End-to-end learning of visual representations from uncurated instructional videos
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1eaf6c1f-7f3f-4dbd-8fe1-793a40eb840d · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0d455d99-3861-4248-9217-1b71bfe27ee0 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Scaling up vision- language pre-training for image captioning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 45bb48ee-1534-460b-a89a-e4629a7ed4aa · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning An empirical study of training end-to-end vision-and-language transformers
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cb05c3d7-b7f6-4b6a-aace-20d7ebb73939 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning How Much Can CLIP Benefit Vision-and-Language Tasks?
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 394409b5-09ee-46ff-b7a7-49f06cd927e2 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 94c2227e-51bf-41c5-9628-3d813875b9b9 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Murphy, and Cordelia Schmid
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a3db7f46-f5c6-47cf-ac64-c8eacf823504 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Actbert: Learning global-local video-text representations
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ab0c1b3e-9a54-42b4-9d74-cbac6ed77c38 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Less is more: Clipbert for video-and-language learning via sparse sampling
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ed286fd8-33ec-4656-841a-605e03c91755 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Frozen in time: A joint video and image encoder for end-to-end retrieval
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation df6acf2d-ce08-4040-9031-ec1f5337f94f · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation daf2cd83-44df-4d94-9636-da67314f460f · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning InternVideo-Ego4D: A Pack of Champion Solutions to Ego4D Challenges
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 125f031c-56af-4fb1-b9c2-54c26841c724 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Vivit: A video vision transformer
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a0f9a437-972d-46b2-8949-b337cfbdb66f · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Video swin transformer
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3c835898-0ae6-401d-baf2-dc100efc610b · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Align before fuse: Vision and language representation learning with momentum distillation
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cb220de4-fd55-4eb4-88c9-1ec6f7541301 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Bsn: Boundary sensitive network for temporal action proposal generation
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f153cb14-1f77-4a63-9bc8-deed1fe2a383 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Bmn: Boundary-matching network for temporal action proposal generation
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cfb36e00-385d-4d5f-b291-c5053f1307ec · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Augment your batch: Improving generalization through instance repetition
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a556e19c-29e9-48f8-9636-437f237f3bcf · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Uniformer: Unified transformer for efficient spatial-temporal representation learning
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 553d1d0d-c341-41f6-b7f2-3cc8e6a5c298 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Temporal segment networks: Towards good practices for deep action recognition
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4c47372a-b8a2-42a4-8a05-2c885d9bebc2 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 06eee4ec-f8b8-422c-8a69-cea7ef01bac1 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Ava: A video dataset of spatio-temporally localized atomic visual actions
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f365664f-37db-4ab0-9990-5cf0a2a5d10d · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning something something
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a79c3217-3bd8-4419-ac62-458aabe245d0 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning A Short Note about Kinetics-600
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 25073e3d-7956-4d69-b3ea-815b437f8a5d · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning A Short Note on the Kinetics-700 Human Action Dataset
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a7473c7a-7460-4718-ae07-d8ac09ecd542 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 49c0baeb-7464-4264-972c-23a6bdcbc89b · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Flamingo: a Visual Language Model for Few-Shot Learning
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 28f5ce3e-010a-4d07-bc33-0c600825708b · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Ean: event adaptive network for enhanced action recognition
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e6fb0d7a-bf64-464e-b8bd-c12b02062df7 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Slowfast networks for video recognition
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fe0f9f6c-549b-44d7-891d-afa458f2b9ed · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Temporal context aggregation network for temporal action proposal refinement
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b520d146-266c-40ca-933f-6f726ee4d751 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Tsp: Temporally-sensitive pretraining of video encoders for localization tasks
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2ab8484b-ddaf-4b2d-9574-27f03a4e1f38 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Actor-context-actor relation network for spatio-temporal action localization
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1cda90a1-64f6-429e-9734-f7175806b97a · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Relation Modeling in Spatio-Temporal Action Localization
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1bbe1b0d-3e33-46ad-94c2-9f3cbdf14c85 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Activitynet: A large-scale video benchmark for human activity understanding
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b058bb58-df2e-4099-800b-aac3f4266817 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Hacs: Human action clips and segments dataset for recognition and temporal localization
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 21385d48-09e2-42cc-a87e-aa3ed4f73611 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Hmdb: a large video database for human motion recognition
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a441b6ad-3c9a-4642-9523-12edc9633dfb · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning in the wild
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e58262bb-91a9-4abe-9fe1-5ecb4a666e31 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Fineaction: A fine-grained video dataset for temporal action localization
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 305de34c-d1fc-4b85-9f6a-546e3a9cb41d · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning The AVA-Kinetics Localized Human Actions Video Dataset
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 95ff95e0-14bf-4c78-96b8-15bb76c2dd87 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Microsoft coco: Common objects in context
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4bca6380-1687-4bec-a8b5-85130b0ec985 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Mask r-cnn
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 395d12ae-d898-47ed-ac60-b6fc9cfe6bb8 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Asynchronous interaction aggregation for action detection
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4dc41450-fea2-4490-b286-f4a94198cb14 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Ts2-net: Token shift and selection transformer for text-video retrieval
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2fab92fb-d8e8-4498-bf7c-4907c4130824 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Bridging the gap between learning in discrete and continuous environments for vision-and-language navigation
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6a182d2a-0e0a-4dae-af1e-d1df9e27c009 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning 1st Place Solutions for RxR-Habitat Vision-and-Language Navigation Competition (CVPR 2022)
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 779a5061-293a-4889-b7ae-50a2d3404ad7 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6b002104-5b7f-4200-8540-ec9801a7d1a6 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Msr-vtt: A large video description dataset for bridging video and language
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 35ac8685-0577-4375-8322-4b7df5fdb57b · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Deep learning for video classification and captioning
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dc73f63c-03e0-4c40-b633-dcca7f1df363 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Localizing moments in video with natural language
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4e29b554-3fed-44f5-b3c5-7e01be66b7ac · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning A dataset for movie description
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b8dc8f73-21ef-41d4-a06e-5d4b69662908 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Vatex: A large-scale, high-quality multilingual dataset for video-and-language research
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0ab77ef9-76a9-42c5-95ec-bc1c2c80ac1b · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Collecting highly parallel data for paraphrase evaluation
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dc7b93fb-238a-4ac4-9d47-c0f4619a54a2 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Tgif: A new dataset and benchmark on animated gif description
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation afc4b203-517b-48ba-9b13-feab2632f974 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 355c2f4b-df10-4abb-bffc-11ef4e25dae4 · outbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Beyond the nav-graph: Vision-and-language navigation in continuous environments
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 01890349-770e-4c2c-9b11-0c15604e7c9d · inbound
VideoChat: Chat-Centric Video Understanding InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dbc082f8-2ef1-41da-babd-92fd7526efe7 · inbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d0657f60-4e87-4833-8595-c8a29637bd6c · inbound
LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 195
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5508df41-5c95-4f58-b5a3-c3927c243577 · inbound
LRM: Large Reconstruction Model for Single Image to 3D InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 107448f9-83d7-444d-b0ef-af0b68ae0758 · inbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cf9a3d9f-a4be-44ef-8b03-2752a1e04304 · inbound
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 152
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 26568750-392f-4f45-b9e1-0dc79ca3b760 · inbound
Revisiting Feature Prediction for Learning Visual Representations from Video InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 287
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 69b91bbb-8d7a-4fc6-9e55-b0d424f983e9 · inbound
Towards Long Video Understanding via Fine-detailed Video Story Generation InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f529530-4988-48b5-a831-8c247a5de911 · inbound
Multimodal Contextualized Support for Enhancing Video Retrieval System InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 516ba0a7-066f-4d67-ba4b-452eaab05ab0 · inbound
Gramian Multimodal Representation Learning and Alignment InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31e0d036-f164-4b30-bb86-d74b0d0d4d93 · inbound
ShotVL: Human-Centric Highlight Frame Retrieval via Language Queries InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a417401b-6c6e-49ee-a758-1f05bf83348f · inbound
Do Language Models Understand Time? InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 192
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ce6ac81-ac32-4b5c-ac18-878f9ed0616d · inbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 527a5034-1726-4fac-984e-db07468774e6 · inbound
VidCtx: Context-aware Video Question Answering with Image Models InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47ee8998-e4aa-4b7f-befe-9a345cd6ab89 · inbound
Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9489c8d4-380e-4e33-aee0-489b66a22283 · inbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8645d375-12b2-447d-bb3b-d448f37c19fa · inbound
LongViTU: Instruction Tuning for Long-Form Video Understanding InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c63646bf-f532-4508-ba8d-8831ca267467 · inbound
An Empirical Study of Autoregressive Pre-training from Videos InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a3acdbf-0840-45e7-9e6a-ed3fd685817b · inbound
Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1bf4bc0-51da-4e06-92df-b148982abc98 · inbound
Admitting Ignorance Helps the Video Question Answering Models to Answer InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c5358da-12bd-4211-8196-7cbcf3e17854 · inbound
SMART-Vision: Survey of Modern Action Recognition Techniques in Vision InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 173
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d55281f1-d88e-46cc-a9fb-13cecc4bdc98 · inbound
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dc38dd8-50f9-41c9-8692-9e850f9b14ca · inbound
SpatioTemporal Learning for Human Pose Estimation in Sparsely-Labeled Videos InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ead0cae-eefb-4090-bfb7-abaf41a42baa · inbound
Understanding Long Videos via LLM-Powered Entity Relation Graphs InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e1eaf42-8b49-4753-b98c-53e90cf689e3 · inbound
XRF V2: A Dataset for Action Summarization with Wi-Fi Signals, and IMUs in Phones, Watches, Earbuds, and Glasses InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18a5d58a-d8dc-486a-bcbf-96dcf678f6e4 · inbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc9bac64-023a-4caa-b918-1bec63136a6a · inbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d980b77-377e-436e-b647-a44a6c4e628a · inbound
Generative Ghost: Investigating Ranking Bias Hidden in AI-Generated Videos InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa1a084a-51b0-4e57-94fe-f26759cc1cfd · inbound
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c92fe79b-13d1-4bc5-a59a-c97df7637a2d · inbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8401b309-0b16-4b9d-ac1b-30f5586f3016 · inbound
HuMoCon: Concept Discovery for Human Motion Understanding InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a19f1bcc-2c7e-4e11-9527-3b1c0426aa55 · inbound
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 348e432f-cd03-416b-8af0-23bf91de4165 · inbound
Improving Keystep Recognition in Ego-Video via Dexterous Focus InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00fe6806-3f7c-4d4b-9759-d075e8f50405 · inbound
Large-scale Self-supervised Video Foundation Model for Intelligent Surgery InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ba4e3ee-4df0-48c8-b7f6-ca213a40175a · inbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e067d0b9-cc70-41ab-840e-dcdd53d4e93d · inbound
EgoM2P: Egocentric Multimodal Multitask Pretraining InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 109
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 300d6d22-2ed4-4a14-89bc-5e3f9fe69607 · inbound
DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1afc1f04-1945-446e-8181-a52b3e8865f1 · inbound
Vision Generalist Model: A Survey InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 177
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90ff68e9-ab1c-4fff-90a3-784606db2304 · inbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5d6b164-6f58-4f03-a66e-b3565b0bde8d · inbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b084f60e-1e9e-47d1-943f-88b0537bc24e · inbound
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2fe6278-5c86-4b72-9064-55b7470acee5 · inbound
DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f5abedc-1bfa-448b-bbb9-58b7c223be50 · inbound
MANTA: Cross-Modal Semantic Alignment and Information-Theoretic Optimization for Long-form Multimodal Understanding InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bfbe9db-ce8c-43be-b186-adac9a3252d1 · inbound
SciVid: Cross-Domain Evaluation of Video Models in Scientific Applications InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9429fee-09fc-47cf-8c99-f2f74abd21c6 · inbound
Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0aefedc9-f462-4c9a-b3b2-8c3d5401cc01 · inbound
Semantic Frame Interpolation InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dae05960-f941-4840-985e-c37e44273954 · inbound
Sparse-Dense Side-Tuner for efficient Video Temporal Grounding InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90af2435-5042-4201-ad4c-2e1abd6c5506 · inbound
OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 970d892b-4ebf-4a2c-857f-c384ef6afc3b · inbound
LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f72bac2e-7def-4f46-a900-25c955dfb1f9 · inbound
A Survey on Efficiency Optimization Techniques for DNN-based Video Analytics: Process Systems, Algorithms, and Applications InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 137
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2228c1a3-fef8-4a05-9d27-bde4302af7f0 · inbound
VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a9aaa68-a589-4f06-9676-baa83d4ef10e · inbound
Object-centric Video Question Answering with Visual Grounding and Referring InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af83d3d0-c08e-4df0-801c-d839ffb69237 · inbound
Representation Shift: Unifying Token Compression with FlashAttention InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3969c71-fd3e-4fd9-a942-9c16d6e17c84 · inbound
TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdb02b9b-67d9-4830-be24-be0b0b3b54f8 · inbound
VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6867581-5b02-4758-9364-ed0c7a825638 · inbound
A Survey on Video Temporal Grounding with Multimodal Large Language Model InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42adbfc4-cee6-4148-924c-fb69c2c67615 · inbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0fc6081-e0ab-4d7b-aab9-69f4f9ca65f5 · inbound
Video Understanding by Design: How Datasets Shape Video Models InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83e97f27-3ede-4322-9404-30dabd8d60ff · inbound
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cab872b-e3ab-488b-b404-c942bc03d25f · inbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24f8e7fc-8aca-4f6f-b9a0-9fe7d3c6f1cb · inbound
Privacy Beyond Pixels: Latent Anonymization for Privacy-Preserving Video Understanding InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 54858588-607d-4546-a315-28a6f8808edb · inbound
Calibrated Multimodal Representation Learning with Missing Modalities InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0bd9ef4e-cd8a-4dfd-8f4b-c25aa71efffe · inbound
MemVerse: Multimodal Memory for Lifelong Learning Agents InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8c5d2f2-9c0e-49b3-b0e0-b42e0379d4fd · inbound
TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 37572fa0-5626-4f10-9f5a-258f3dde406e · inbound
Streaming Video Instruction Tuning InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bfaab06c-71e3-421d-9609-00e72cfd5b0b · inbound
A Paradigm Shift: Fully End-to-End Training for Temporal Sentence Grounding in Videos InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cb0ed3f3-8518-488b-8bc7-54393e26bd03 · inbound
UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2f3da1ce-7e24-4264-83c9-dc24e3758d6c · inbound
InstrAct: Towards Action-Centric Understanding in Instructional Videos InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5e9056e4-6d5b-4ad1-a422-e1c4b943801b · inbound
Efficient Spatial-Temporal Focal Adapter with SSM for Temporal Action Detection InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 80c999f0-f809-4892-80c9-05670df423ee · inbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c3bf265-d301-486d-a0dd-5db372730b9f · inbound
V-Nutri: Dish-Level Nutrition Estimation from Egocentric Cooking Videos InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5abab684-03bd-4f0f-94bf-1dca71245f25 · inbound
Training-Free Semantic Multi-Object Tracking with Vision-Language Models InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9b6d375d-c00b-435d-a506-c1877d4b37fa · inbound
One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d83e73e6-b449-424e-a5af-6322efd980fc · inbound
FreqFormer: Hierarchical Frequency-Domain Attention with Adaptive Spectral Routing for Long-Sequence Video Diffusion Transformers InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 87782323-821d-4a64-9cac-a087cc63c553 · inbound
TransVLM: A Vision-Language Framework and Benchmark for Detecting Any Shot Transitions InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6bcde318-897e-4397-a625-9d49d956b91c · inbound
LoViF 2026 The First Challenge on Holistic Quality Assessment for 4D World Model (PhyScore) InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 12b3b07f-e37d-4890-a123-a673735728e3 · inbound
CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 540ed8ae-eb16-47ec-84c2-4491799d41fd · inbound
Towards Unified Vision-Language Models with Incomplete Multi-Modal Inputs InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 548a9304-72f2-43f2-a259-ec377032e5c8 · inbound
Masked Diffusion Vision-Language Models for Temporal Action Localization InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c32974c2-d111-475b-8892-da0514548ff6 · inbound
SlotMemory: Object-Centric KV Memory for Streaming Long-Video Generation InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 15d7c3b1-9c1e-4b97-8b57-739a3057391c · inbound
Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3bb252a5-28db-4320-8be4-cd59e35bf2ed · inbound
Hand Trajectory Fusion for Egocentric Natural Language Query Grounding InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 02c8d935-0e67-4f5c-bb43-a3de99cb64fd · inbound
Where Do We (Not) Need Temporal Context in Low-Resource Video Task Adaptation? InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 783c631f-84aa-4d63-91a5-94c250fdfd21 · inbound
Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 80ea8512-3fdd-4a73-a8a0-9b80926e8809 · inbound
OmniGen-AR: AutoRegressive Any-to-Image Generation InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 819f3b21-28cb-44ed-80f0-95472100bc9e · inbound
CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0370a05a-f20c-4ca0-81a3-e3dabcc01d6a · inbound
InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 118
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c65d055d-8457-4dd4-ad6e-f5aea84e5f40 · inbound
HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 259
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 298abd9b-abef-4d23-bef4-ee332f59169d · inbound
CineOrchestra: Unified Entity-Centric Conditioning for Cinematic Video Generation InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 896845f4-3d43-4c4e-b595-bd8a8b081d21 · inbound
LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cbf220b3-9e5e-4017-9a65-c6385dd757a5 · inbound
ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ba56b6ce-d30e-499e-a94c-ef4e93ca4736 · inbound
T-MOR: Learning Motion-Aware Skeleton Representations for Human Action Recognition InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5144b89b-7a7f-47d7-a87e-287d799f017e · inbound
TuringViT: Making SOTA Vision Transformers Accessible to All InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 51a988bd-d8ea-4162-a98f-522828bf77c1 · inbound
TuringViT: Making SOTA Vision Transformers Accessible to All InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 10770e63-f3c8-4620-824e-8f9dae2936e3 · inbound
Toward Low-Latency Vision-Language Models with Doubly-Correct Predictions in Egocentric Visual Understanding InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8e3dfb14-5527-4d05-8da4-b95867e89cb9 · inbound
Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6c80b62f-cce4-4116-aa83-1fa1c792e7a0 · inbound
Prompting-MammAlps: Fine-Grained Text-to-Video Retrieval for Camera-Trap Data InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a0f75f7-93bf-4245-8d69-58171fbfda0b · inbound
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab0e695a-a4a4-444a-a26d-3d23e63848d7 · inbound
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1713cb65-6caf-4f72-b2fc-2735c535ed09 · inbound
VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.