Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T10:51:29.261016Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 2 inbound Pith citation observations for arXiv:2502.02885.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T10:51:29.261016Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T18:49:03.038872Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T10:41:06.814014Z
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 43ea0d3f-6a88-400a-8816-1faee030fb74 · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval Un- masked teacher: Towards training-efficient video foundation models,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e24eff71-66be-46ee-a89f-960ce4464e67 · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation fbc34ced-5b6b-44ba-9302-d687cb966a61 · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d885a4a9-5386-4a2a-8d18-2715cd315e63 · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval Clip-vip: Adapting pre-trained image-text model to video-language representation alignment,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 084cc9dd-47b9-45a1-8ea7-a539e7277e10 · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval Long-form video- language pre-training with multimodal temporal contrastive learning,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a78c93a6-f973-4b2f-9e37-08aa319a33ae · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval Litevl: Efficient video-language learning with enhanced spatial-temporal mod- eling,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4b77a9ed-dfc2-457b-b4ea-295080229ff5 · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval CLIP2TV: Align, Match and Distill for Video-Text Retrieval
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a78511e2-5632-4827-841d-41e5e5b8285a · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval Less is more: Clipbert for video-and-language learning via sparse sampling,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 20f71575-29c3-45a2-bc5c-663fd5312806 · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval X-clip: End-to-end multi-grained con- trastive learning for video-text retrieval,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 62e86228-f5fd-47bd-80ff-50fe959a3636 · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval Advancing high-resolution video-language representation with large-scale video transcriptions,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bb633cea-6a1b-40ce-a1a7-d7509da2db60 · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval Bidirectional cross-modal knowledge exploration for video recognition with pre-trained vision-language models,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5064f4a4-2741-41f2-89d4-e4df90d2c81a · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval She-net: Syntax-hierarchy-enhanced text-video retrieval,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ceec9b38-2584-4c6e-b585-bf67748092d1 · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval Teachtext: Crossmodal generalized distillation for text-video retrieval,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4c3a5699-381f-46f9-95e4-016a677f26a3 · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval Text is mass: Modeling as stochastic embedding for text-video retrieval,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3d7c1bb0-8baa-4bf0-9706-2bb8e2502ec9 · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval Cap4video++: Enhancing video understanding with auxiliary captions,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8db09856-898d-47d9-8897-847583b1d86b · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval The neglected tails in vision-language models,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9a71ca16-5362-4b0d-9687-fc31138a502a · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval Verbs in action: Improving verb understanding in video-language models,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 67a74f46-f828-4c84-8e47-d1ccd88e212b · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval Videocon: Robust video-language alignment via contrast captions,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2c8ee2f4-1d81-4f21-82a9-86326b8c9edf · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval DREAM: Improving Video-Text Retrieval Through Relevance-Based Augmentation Using Large Foundation Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7afff887-8912-42b5-8552-76931bf2b82b · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval Automatic prompt optimization with
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0492c361-dc26-425d-bea4-609ac165b528 · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval Dynamic prompt learning: Addressing cross-attention leakage for text-based image edit- ing,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 18a5d58a-d8dc-486a-bcbf-96dcf678f6e4 · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 850cc9ea-5a95-4a69-9af1-084fa2ceb6fb · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval Learning transferable visual models from natural language supervision,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 39fef332-2f04-4f03-87f3-699628cf4cdc · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval Ts2-net: Token shift and selection trans- former for text-video retrieval,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a6b10a50-a0a4-4918-a58f-63ffde2acf0d · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval Mv-adapter: Multimodal video transfer learning for video text retrieval,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d60ca170-ae0b-446a-8142-de76b40260b7 · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval Holistic features are almost sufficient for text-to-video retrieval,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bb58eef9-2878-47e0-af92-cd278f40d59f · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval Disentangled Representation Learning for Text-Video Retrieval
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1b0a681-8aec-459c-9e21-b7a8eb9cd8a4 · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval LLaMA: Open and Efficient Foundation Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c37c38c7-ec0d-4d56-b704-a6a85e782de2 · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval Reversed in time: A novel temporal- emphasized benchmark for cross-modal video-text retrieval,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4789df69-4273-4670-8185-82ec70e6dc28 · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval Visual instruction tuning,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9d3b8af-35c7-40a5-93c1-28f4693e274a · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval Improving clip training with language rewrites,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 99361dea-d931-4b7d-947d-87196362e4d6 · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval mplug- 2: A modularized multi-modal foundation model across text, image and video,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a98214e6-1284-46e5-9904-eed78540db68 · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval Dual-modal attention-enhanced text- video retrieval with triplet partial margin contrastive learning,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 22e7c9ee-d26f-49db-a4e7-d6b331c666b5 · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval Msr-vtt: A large video description dataset for bridging video and language,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c5b5af83-5b05-4bd8-b152-d263ee9880ce · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval Use What You Have: Video Retrieval Using Representations From Collaborative Experts
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 566b2468-af15-439a-8062-096dc8610593 · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval Collecting highly parallel data for paraphrase evaluation,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 70739db6-90e6-4f3f-8833-9c40e6eedc32 · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval Localizing moments in video with natural language,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b034008b-686c-4a8c-b0cc-f03527e36b19 · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 914bb2f5-da5f-4acd-8c7d-14a79bc37ffa · outbound
Expertized Caption Auto-Enhancement for Video-Text Retrieval Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c435583e-9907-42d8-9e71-7104318873c8 · inbound
MemVerse: Multimodal Memory for Lifelong Learning Agents Expertized Caption Auto-Enhancement for Video-Text Retrieval
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 721cd6cd-30bd-4b95-9f61-3a8e8471d1b0 · inbound
EmbodiedGovBench: A Benchmark for Governance, Recovery, and Upgrade Safety in Embodied Agent Systems Expertized Caption Auto-Enhancement for Video-Text Retrieval
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.