Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T12:10:17.560221Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2608.11013.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T12:10:17.560221Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4d7ac9b3-8ceb-4ae3-87d9-7bf9299bb4cc · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Learning to compose topic-aware mixture of experts for zero-shot video captioning,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f533bef9-e475-42cf-936d-a43eb374ae32 · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning DeCap: Decoding CLIP Latents for Zero-Shot Captioning via Text-Only Training
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19c991cc-536b-47da-9036-bc94ea65deb0 · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da62bf44-5334-46a7-a68d-666868ee2414 · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 64294703-8e7a-4721-ae98-2d64fe6fc190 · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Improving cross-modal alignment with synthetic pairs for text-only image captioning,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 507ee31f-a271-46ee-9576-c776fcacef4d · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Retta: Retrieval-enhanced test-time adaptation for zero-shot video cap- tioning,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 624c36ab-f79f-4326-8ab3-91b525f424ff · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Text-only training for image captioning using noise-injected CLIP,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ebb39956-b509-4919-8638-3a7cb53d208b · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Sequence to sequence-video to text,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 87b366b9-07bc-4024-99ce-0f2f9d010979 · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Bidirectional long- short term memory for video description,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 351d23f9-35a5-44a6-bddd-c03b3be6856e · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Graph convolutional network meta- learning with multi-granularity pos guidance for video captioning,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 825e202e-8fe2-4ee3-b98c-b07524bac78e · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Describing videos by exploiting temporal structure,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fde556e8-833d-4bdd-b1f7-9c59505c39da · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Icocap: Improving video captioning by compounding images,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2b072796-1967-427d-9af9-b0957ac999b3 · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Memory- based augmentation network for video captioning,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 51c5fec0-3cc5-4a59-8530-6b7ade97ba1b · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Attention is all you need,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 73b61cdb-cd4a-451d-9b67-0ce755e24ddf · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Hierarchical modular network for video captioning,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 841c5f3c-b9ce-4f86-9946-389a85bccb3b · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Swinbert: End-to-end transformers with sparse attention for video captioning,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bc11fe1d-6959-434e-bc19-90eb76ead7d6 · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cabbb01-3849-49c3-a761-3b511f14332c · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Ma-lmm: Memory-augmented large multimodal model for long-term video understanding,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a9dfdb99-f0a2-475d-b9ba-b17f6bc85d72 · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Text-Only Training for Image Captioning using Noise-Injected CLIP
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86c4d617-89d4-4634-bf51-4ffe29d951be · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Language Models Can See: Plugging Visual Controls in Text Generation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0f70ae5-8cda-4d16-8160-57cd5ed326ae · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Zerocap: Zero-shot image-to-text generation for visual-semantic arithmetic,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4241089a-f193-4dc9-8fb1-ad9b694a9576 · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Learning transferable visual models from natural language supervision,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 45b26696-9dfd-472a-bc6e-f1d27156b202 · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning From Association to Generation: Text-only Captioning by Unsupervised Cross-modal Mapping
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8131c68b-a5d3-46e6-b0ef-e539345c6f38 · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 986d18f8-79b1-4ef9-9ab4-04171e9586e4 · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c2ce2d1-fb0e-4d80-b6c6-f90609bcdebc · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Language models are unsupervised multitask learners,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 013a8cef-d9a6-40d7-a267-57d1516bd0d8 · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Delving Deeper into the Decoder for Video Captioning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3301e15f-319a-4d25-8899-1ec6326b9f3c · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Improving video captioning with temporal composition of a visual-syntactic embedding,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d3429b60-cc3d-4cf0-a7cf-308a47018340 · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Zero-Shot Video Captioning with Evolving Pseudo-Tokens
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce135537-b377-42d9-81ff-aa36e85e9a18 · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning MultiCapCLIP: Auto-encoding prompts for zero-shot multilingual visual captioning,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 785c6cf8-21b3-4bff-8e1a-04dec522b27b · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Visual instruction tuning,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ec9bfde8-4dd7-4aca-8880-33a90c8adda8 · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2fbcb5a-fe08-4b75-8113-53e73c86bfd5 · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Msr-vtt: A large video description dataset for bridging video and language,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5c7b881-8e63-4578-b5d8-2f4ea125bd26 · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Collecting highly parallel data for paraphrase evaluation,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b88c278-12d7-4e40-8221-9daaf3a9ee04 · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Vatex: A large-scale, high-quality multilingual dataset for video-and-language research,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1673f31e-d513-4bdb-a491-338327be8df2 · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Bleu: a method for automatic evaluation of machine translation,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 34951392-678c-4a7e-b98b-3cd0a53ea77f · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Meteor universal: Language specific translation evaluation for any target language,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3db6d68b-92f4-4f6b-99bb-bc988827387c · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Rouge: A package for automatic evaluation of summaries,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 242319e9-d6db-45e9-86f0-0eb8885f8a5f · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Cider: Consensus- based image description evaluation,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ee924bf3-d658-4431-97e2-b98db66504c4 · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Decoupled Weight Decay Regularization
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11e8d915-f8fd-4515-ba2d-437559911d98 · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Expanding language-image pretrained models for general video recognition,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 25108cc2-59b5-4d44-8fa3-526165b86e40 · outbound
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Wan: Open and Advanced Large-Scale Video Generative Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.