Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T18:40:22.275159Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 0 inbound Pith citation observations for arXiv:2412.07704.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T18:40:22.275159Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
78 of 78 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 675d14c4-d3a2-45a6-a40d-716825cb5614 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Flamingo: a Visual Language Model for Few-Shot Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b47675f9-c508-4379-bb3e-d4d290741054 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Vivit: A video vision transformer
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38d5c2e2-3795-476a-b98c-d71670add427 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Hiervl: Learning hierarchical video- language embeddings
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b74dc9e9-fa4a-41c3-9209-ef3c5f19b86d · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Frozen in time: A joint video and image encoder for end-to-end retrieval
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5fc93dc-7da2-4886-927a-590bf0e0459a · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Is space-time attention all you need for video understanding? In ICML, volume 2, page 4, 2021
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 34c8f8ad-7228-4930-8442-ff00864656f5 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Blinkdl/rwkv-lm
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 172b787c-158d-4dea-9fd9-a028f716f0fc · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Revisiting the” video” in video-language understanding
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4bcf2b93-2c01-4852-ae96-741a45dfa03e · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Activitynet: A large-scale video benchmark for human activity understanding
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0f99526-3ada-45a8-923f-01379c9e66b3 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Locvtp: Video-text pre-training for temporal localization
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 51d10e7b-49f6-44d8-ac88-ea40bf174a17 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Metaxas, and Hongxia Yang
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 936c7b90-1738-482d-b743-2c656a516604 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Gonzalez, Ion Stoica, and Eric P
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 83564cf4-7ebe-42db-ba93-4a78643a4965 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Unsupervised and semi-supervised domain adaptation for action recognition from drones
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e0c386f9-b55b-4cae-92fc-1303633d524b · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Free dolly: Introducing the world’s first truly open instruction-tuned llm
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 18e3b570-4d85-432f-bdcf-5c61fa1298c8 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Prompt switch: Efficient clip adaptation for text-video re- trieval
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ea1e73ca-d9b1-40f0-abc4-c90b0ae38730 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Imagenet: A large-scale hierarchical image database
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44711116-d431-413a-b5ab-30eafca01d55 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Text-guided video masked autoencoder
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 87dcb7e5-5f7f-4c77-8885-fc4d1ed888e5 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Slowfast networks for video recognition
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a19e7fb1-7fe3-4e50-992a-7320ce51d2b9 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18b4b722-8935-494a-9cf8-8deb5bf16990 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Multi-modal transformer for video retrieval
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 080c0049-af7c-4442-88b9-29b1076e3965 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Openllama: An open repro- duction of llama
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 858f3263-772b-419f-bf27-ab7e57e7adc8 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Ego4d: Around the world in 3,000 hours of egocentric video
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dc6f2b7-473c-420b-83a7-58515e2278ce · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Ma-lmm: Memory-augmented large multimodal model for long-term video understanding
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0ec6fe47-169b-45ab-a838-17919f451dad · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Long movie clip classification with state-space video models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation df4f55af-fa3a-40f4-8e0a-65cc1978df08 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Video re- cap: Recursive captioning of hour-long videos
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 363a9f56-e2ba-483b-ae32-7a3e6db037f0 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Perceiver: General perception with iterative attention
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 812c4982-dced-47e4-85a8-6fd4b60f6468 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c8dccd7a-5203-4e85-a82d-112f67b55425 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Dense-captioning events in videos
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 986c2357-dba8-4c17-8413-b75dd961f300 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Video Token Merging for Long-form Video Understanding
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55c1a9b1-a921-4da6-a56e-5fc612464cf6 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Less is more: Clipbert for video-and-language learning via sparse sampling
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3fdf41e7-e82d-41c5-a7b3-2187ec78968e · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Gonzalez, Ion Stoica, Xuezhe Ma, and Hao Zhang
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1a756374-a7bb-4303-8d5e-f8fd96dfb66f · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Hero: Hierarchical encoder for video+ lan- guage omni-representation pre-training
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 22634687-060e-427a-bff1-b6c0f8143e31 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Lavender: Unifying video- language understanding as masked language modeling
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f49ba053-96f9-450c-9b42-0083c9bdfcd9 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Ego-exo: Transferring visual representations from third-person to first-person videos
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0de33a2d-d776-4a50-a48b-0f03f32bb807 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Rouge: A package for automatic evaluation of summaries
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e17bd6b1-8794-4626-8b68-7eb329baeab3 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Egocentric video-language pretraining
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8c889a7f-4fe5-4376-b0fc-a079288a3226 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Retinex-inspired unrolling with cooperative prior architecture search for low-light image enhancement
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ee9993b1-19a1-49d4-89a9-784672706038 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Sgdr: Stochastic gradient descent with warm restarts
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 78a68820-7c24-4602-8e0c-86c4c0a60891 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Decoupled weight de- cay regularization
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a23ca64a-d0e1-4f27-9215-6301de439714 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ee2f2e22-5a24-44d1-a06c-883deee273bf · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning X-clip: End-to-end multi-grained con- trastive learning for video-text retrieval
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e2abef21-70a8-4a9c-bfad-dc77572e9c03 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning End-to-end learning of visual representations from uncurated instruc- tional videos
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8e2154c4-c4e3-4eb8-8a58-997ada5779eb · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation dc123eb7-8a58-492a-a5d1-cb67a5a964a5 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cef34cec-a4e9-4292-85e6-6e4926a25ecf · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Slip: Self-supervision meets language-image pre- training
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 79a35dd4-7a64-46fe-a56b-0bc4207a3e97 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Keeping your eye on the ball: Tra- jectory attention in video transformers
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 201ee4fa-1134-4e8b-86f7-fa24e8dc2366 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Learning transferable visual models from natural language supervi- sion
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5caf0299-3444-4b8c-a6ab-42d316724101 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Movie description
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 97dabd7e-7888-4e23-873a-3567774d950d · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Actor and observer: Joint modeling of first and third-person videos
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1eccf011-45ad-4296-a59b-3a6c2129d3e3 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Charades-Ego: A Large-Scale Dataset of Paired Third and First Person Videos
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bcd8913-db3f-42d7-bdf0-bd88170d5c2e · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Videobert: A joint model for video and language representation learning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 035e5820-11a0-45af-b166-02439be83035 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Long-form video-language pre- training with multimodal temporal contrastive learning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4b7db35b-33bb-4347-990f-56ed75b03109 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Coin: A large-scale dataset for comprehensive instructional video analysis
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5354a18e-c39b-4408-b1c7-d3eba3df0e71 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Perceiver-vl: Efficient vision-and-language modeling with iterative latent attention
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 12a33b53-98b0-43d8-b5df-db60e4f26e4f · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Internlm: A multilingual language model with progressively enhanced capabilities
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1d96876-a14c-44df-889e-60819aed4134 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Yfcc100m: The new data in multimedia research
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2bcaf4c2-a33b-4e66-84dc-5b170c1722b7 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Attention is all you need
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab75073f-9ea2-4f39-bbb0-234ffc962ceb · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Selective structured state-spaces for long-form video understanding
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4a3faca8-1c04-4fe0-827a-19d99d598132 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73d944d8-079b-4d01-a357-36b6102fedcc · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning DHP Benchmark: Are LLMs Good NLG Evaluators?
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce0bbc7b-c8f3-4d86-9372-7f6bd640d9bc · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Unified coarse-to-fine alignment for video-text retrieval
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0217e50e-9914-4bfe-a08f-03f1b55941d3 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Towards long-form video understanding
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation af4a49fd-e24f-4298-a8e9-0d1986020330 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8634cbef-9b21-4947-afe3-a6ee2c5d0267 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Msr-vtt: A large video description dataset for bridging video and language
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 12f45cdf-17c8-487c-b5ef-6ee4bd4c4d86 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Ad- vancing high-resolution video-language representation with large-scale video transcriptions
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c95dff12-9dd3-4d14-a1da-b3f0b253e691 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Clip-vip: Adapting pre- trained image-text model to video-language alignment
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fe072311-6922-424c-8289-b740a55c263e · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Just ask: Learning to answer ques- tions from millions of narrated videos
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f430c376-b266-4521-b89d-a5dc611e5127 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Zero-shot video question answering via frozen bidirectional language models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb140d43-9d0b-4397-9990-cf86e245ea77 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Taco: Token-aware cascade contrastive learning for video-text alignment
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 37292d9b-0b3b-4440-8f52-346e291a77df · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Scaling White-Box Transformers for Vision
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b577bacb-e363-45ca-a5be-e2f1cfe3a8c4 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Filip: Fine-grained interactive language-image pre-training
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a2ed1c0f-67f9-43c4-815c-4c1447826737 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Hitea: Hierarchical temporal- aware video-language pre-training
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5dcc4eb2-2123-4225-b95d-e4cab5cdf4cb · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Self-chained image-language model for video localization and question answering
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45c29bbd-b6da-4944-a25c-ea7bac2ae2ef · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Learning from inside: Self- driven siamese sampling and reasoning for video question answering
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 78cc4f21-adba-4b55-b2e0-e883298bb3e7 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning White-box transformers via sparse rate reduction
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f07426f0-cd41-48c7-9ff9-d8d981f15060 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning BERTScore: Evaluating Text Generation with BERT
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1946b448-e8db-4ff6-9c31-a3e64bfe55ab · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Videoprism: A foundational visual encoder for video understanding
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c9a9235e-43b2-49bc-a573-5842e84d09f8 · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Cen- terclip: Token clustering for efficient text-video retrieval
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d237a457-aeb9-4f15-a5ed-e7d2a9bc79ae · outbound
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning P Xing, Hao Zhang, Joseph E
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
No inbound Pith citation observations are available.