Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-24T00:22:35.635679Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 2 inbound Pith citation observations for arXiv:2406.05615.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-24T00:22:35.635679Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-14T09:11:59.440912Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-05-23T07:25:28.511257Z
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 976b9646-8b96-49e8-925e-a3f348526e55 · outbound
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives A CLIP-Hitchhiker's Guide to Long Video Retrieval
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 123c8b5e-ffc0-47de-aec6-9e0e9d898a3a · outbound
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives A Short Note about Kinetics-600
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0302d9cb-47c8-42b9-b052-6c0b97b4e89c · outbound
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1c8d84ce-ab50-46cd-a837-aef0d9ea6995 · outbound
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives Multimodal Pretraining for Dense Video Captioning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9058bebf-c827-460b-af80-512bfe559eb3 · outbound
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives Temporal Tessellation: A Unified Approach for Video Analysis
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1fde12c8-86c0-49aa-87e6-857c917fffef · outbound
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives In Proceedings of the 2018 Con- ference on Empirical Methods in Natural Language Processing, pages 1369–1379
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bd69b2c4-4d4e-44ed-91b8-551ffb3f7747 · outbound
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives VideoChat: Chat-Centric Video Understanding
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f61f62ae-f6e4-414c-9216-2a8e3757797d · outbound
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives Video Swin Transformer
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation babc48cd-1c10-4681-aff7-5fdd4ed290fd · outbound
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives KDMCSE: Knowledge Distillation Multimodal Sentence Embeddings with Adaptive Angular margin Contrastive Learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8ed671ab-8510-49f7-84d3-584a2422b938 · outbound
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18983–18992
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a6a1c026-ba9f-49ab-a19a-b13543fa9066 · outbound
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives In Proceedings of the 29th ACM International Conference on Multimedia, pages 2871– 2879
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bc340d84-edd0-475c-a159-daf7183c25e8 · outbound
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives How2: A Large-scale Dataset for Multimodal Language Understanding
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8141bcfc-d6f8-429f-a758-d68c80b39b17 · outbound
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1207–1216
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e7f0a07b-598f-4cc0-a715-39259373732c · outbound
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives Weakly Supervised Dense Video Captioning via Jointly Usage of Knowledge Distillation and Cross-modal Matching
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3a21679c-6baf-4c27-8972-71413f860338 · outbound
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives In International Conference on Machine Learning, pages 3891–3900
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d0697f62-7cba-45cb-8c4a-b15b5003c3ca · outbound
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives VideoGLUE: Video General Understanding Evaluation of Foundation Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 09739cbe-1798-4c2e-95e3-13223114382e · outbound
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives In Proceedings of the IEEE/CVF international conference on com- puter vision, pages 6023–6032
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a23216ec-1ea4-4366-aa10-9b5c3ca7fbb5 · outbound
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives Advances in Neural Information Processing Systems, 34:23634–23651
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a348eece-2361-4d7f-a2e4-20785c3b50c4 · outbound
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives MV-Adapter: Multimodal Video Transfer Learning for Video Text Retrieval
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 16710594-e48d-4b97-b214-637010526ebf · outbound
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives set up the stand ✔ 2
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ddf9681e-4bdf-4e58-86b2-ad5f53b17c5b · outbound
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives Video moment retrieval 38s 48s 60s 64s Q: People in scuba gear are swimming around
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 76956ff7-933b-42b2-a40a-ee85b8d72a96 · inbound
Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph Generation Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0e4e6dec-ef78-4322-abe8-37b149c03e8d · inbound
Compositional Context Fine-Tuning Vision-Language Model for Complex Assembly Action Understanding from Videos Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.