Pith. sign in

Paper Citation Record · LEDGER

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning

As of 12 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2608.11013.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11013 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:10:17.560221Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact2
  • verified fuzzy23
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4d7ac9b3-8ceb-4ae3-87d9-7bf9299bb4cc · outbound

This paper cites Learning to compose topic-aware mixture of experts for zero-shot video captioning,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Learning to compose topic-aware mixture of experts for zero-shot video captioning,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:18.171021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:10:17.374700Z digest=sha256:f9c3319f3dce839c4b647f78fed0f5a7914cc9cf19b537e8eb68f380a428fb27

Observation f533bef9-e475-42cf-936d-a43eb374ae32 · outbound

This paper cites DeCap: Decoding CLIP Latents for Zero-Shot Captioning via Text-Only Training.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning DeCap: Decoding CLIP Latents for Zero-Shot Captioning via Text-Only Training

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.379818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.379818Z digest=sha256:7514662144360627ce7e15cdf9fdff5e25123ecc017eae4f41e0f0d11d828349

Observation 19c991cc-536b-47da-9036-bc94ea65deb0 · outbound

This paper cites Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.384647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.384647Z digest=sha256:91d8fbfa537cdf960869859efd71fdd28fca8029e2c2240f25ffba52460754c4

Observation da62bf44-5334-46a7-a68d-666868ee2414 · outbound

This paper cites IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-12T12:10:17.754368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:10:17.389553Z digest=sha256:19375b05d4ce77833835e69f9550da6a07ff451f990dafc3e478d2aca6254c32

Observation 64294703-8e7a-4721-ae98-2d64fe6fc190 · outbound

This paper cites Improving cross-modal alignment with synthetic pairs for text-only image captioning,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Improving cross-modal alignment with synthetic pairs for text-only image captioning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:18.156961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:10:17.394603Z digest=sha256:6b82f5d76ca9df8b40cce41c78ae45bdec45edba0f36af6a9806e2ef243f304b

Observation 507ee31f-a271-46ee-9576-c776fcacef4d · outbound

This paper cites Retta: Retrieval-enhanced test-time adaptation for zero-shot video cap- tioning,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Retta: Retrieval-enhanced test-time adaptation for zero-shot video cap- tioning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:18.143378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:10:17.399148Z digest=sha256:32abd10346634d6ff06c81ecabec737a091c08cd3ae119cb4a59efdc93ab0511

Observation 624c36ab-f79f-4326-8ab3-91b525f424ff · outbound

This paper cites Text-only training for image captioning using noise-injected CLIP,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Text-only training for image captioning using noise-injected CLIP,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:18.129207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:10:17.404015Z digest=sha256:45c4038705802916fea58a34a25de4e955a60cd05d88ac068c4f7a4e8e395ece

Observation ebb39956-b509-4919-8638-3a7cb53d208b · outbound

This paper cites Sequence to sequence-video to text,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Sequence to sequence-video to text,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:18.115003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:10:17.408247Z digest=sha256:75fa0824204ed2097d57607a7a8790b132bcb4875b3fa166896fcba02bf11641

Observation 87b366b9-07bc-4024-99ce-0f2f9d010979 · outbound

This paper cites Bidirectional long- short term memory for video description,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Bidirectional long- short term memory for video description,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:18.100003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:10:17.412383Z digest=sha256:4866673422b2f07d6e869b15f35693d37996a6bd372f0e16e76dc2d4d4d25beb

Observation 351d23f9-35a5-44a6-bddd-c03b3be6856e · outbound

This paper cites Graph convolutional network meta- learning with multi-granularity pos guidance for video captioning,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Graph convolutional network meta- learning with multi-granularity pos guidance for video captioning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:18.084924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:10:17.416548Z digest=sha256:3a78bf1a8ca95efb6d263ce7c2b59617a6680b570578e682388e90b5dbc3835f

Observation 825e202e-8fe2-4ee3-b98c-b07524bac78e · outbound

This paper cites Describing videos by exploiting temporal structure,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Describing videos by exploiting temporal structure,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:18.070508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:10:17.420779Z digest=sha256:adc7e3041072f836e744eb89fb5a28adbecdeb6df176ac7ddc78985d163ccb22

Observation fde556e8-833d-4bdd-b1f7-9c59505c39da · outbound

This paper cites Icocap: Improving video captioning by compounding images,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Icocap: Improving video captioning by compounding images,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:18.056297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:10:17.425189Z digest=sha256:caa34ed33f713d519ee97f793cff78b15cb4ec076391493dee00e349d382c8fa

Observation 2b072796-1967-427d-9af9-b0957ac999b3 · outbound

This paper cites Memory- based augmentation network for video captioning,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Memory- based augmentation network for video captioning,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:18.041638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:10:17.429108Z digest=sha256:03e3d1d8dc09eb2687d342b263aa3556e390445e7bfd645bd5264bc0e0d55e21

Observation 51c5fec0-3cc5-4a59-8530-6b7ade97ba1b · outbound

This paper cites Attention is all you need,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Attention is all you need,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:18.026051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:10:17.433335Z digest=sha256:96ff6bac5f856f49408fc8df3d235c6b186180ef4a5e9aba0d96ce16ef2b731f

Observation 73b61cdb-cd4a-451d-9b67-0ce755e24ddf · outbound

This paper cites Hierarchical modular network for video captioning,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Hierarchical modular network for video captioning,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:18.010478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:10:17.437339Z digest=sha256:39c98837186be387ba716bbed6867b46861f9e5b4fca0eae3c50737ed330004a

Observation 841c5f3c-b9ce-4f86-9946-389a85bccb3b · outbound

This paper cites Swinbert: End-to-end transformers with sparse attention for video captioning,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Swinbert: End-to-end transformers with sparse attention for video captioning,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:17.995895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:10:17.441498Z digest=sha256:1cd54c877f83d87648a63e7bcf08dea0be9d307994ed6305b332468ab2446246

Observation bc11fe1d-6959-434e-bc19-90eb76ead7d6 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.446549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.446549Z digest=sha256:f9457b5ba08588e137f5c72e25cd3ca35a333bdff68a6afbe2e0e69d1222a798

Observation 8cabbb01-3849-49c3-a761-3b511f14332c · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video understanding,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Ma-lmm: Memory-augmented large multimodal model for long-term video understanding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:17.981085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:10:17.452167Z digest=sha256:aa653d6181dced673e6d3473232c7d0a248bd61b01a7a28a0d6b1455974c0280

Observation a9dfdb99-f0a2-475d-b9ba-b17f6bc85d72 · outbound

This paper cites Text-Only Training for Image Captioning using Noise-Injected CLIP.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Text-Only Training for Image Captioning using Noise-Injected CLIP

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.456353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.456353Z digest=sha256:beeab1c551475089df51b3d883430342f4ac561763b9d1505ae1c98bb32a6e7a

Observation 86c4d617-89d4-4634-bf51-4ffe29d951be · outbound

This paper cites Language Models Can See: Plugging Visual Controls in Text Generation.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Language Models Can See: Plugging Visual Controls in Text Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.461499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.461499Z digest=sha256:a8657cad19417124b91aa94f4508c84a84b13da07e4c9f87a3c0b2168338c04b

Observation e0f70ae5-8cda-4d16-8160-57cd5ed326ae · outbound

This paper cites Zerocap: Zero-shot image-to-text generation for visual-semantic arithmetic,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Zerocap: Zero-shot image-to-text generation for visual-semantic arithmetic,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:17.967231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:10:17.466408Z digest=sha256:37e314821a150c82c5fc39a3bd127aeb98f5ab51558be7d43a23bd88c20d55e5

Observation 4241089a-f193-4dc9-8fb1-ad9b694a9576 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Learning transferable visual models from natural language supervision,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:17.953545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:10:17.471210Z digest=sha256:20e4d890c4058f2f6f96012084827b160dc41fa03c1515e4de6f5c9de7e9de72

Observation 45b26696-9dfd-472a-bc6e-f1d27156b202 · outbound

This paper cites From Association to Generation: Text-only Captioning by Unsupervised Cross-modal Mapping.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning From Association to Generation: Text-only Captioning by Unsupervised Cross-modal Mapping

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.475874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.475874Z digest=sha256:c75c103df878119a513dfec46b2fe3dcb99fc889e90f83f98aece6b10be34f6f

Observation 8131c68b-a5d3-46e6-b0ef-e539345c6f38 · outbound

This paper cites Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.480381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.480381Z digest=sha256:85a0f542223665e660d67e99055710ac8cdf663ee6946b70cb150a9076c76827

Observation 986d18f8-79b1-4ef9-9ab4-04171e9586e4 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.485268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.485268Z digest=sha256:207c344ba6daae0de50c8879a5bab1dd579363f655dbd9ba06046ce6039580f9

Observation 2c2ce2d1-fb0e-4d80-b6c6-f90609bcdebc · outbound

This paper cites Language models are unsupervised multitask learners,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Language models are unsupervised multitask learners,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.490054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.490054Z digest=sha256:6d5bf49d5c50f5cffa66fdcd4803631bce28ebe144263df51a756d754ae41253

Observation 013a8cef-d9a6-40d7-a267-57d1516bd0d8 · outbound

This paper cites Delving Deeper into the Decoder for Video Captioning.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Delving Deeper into the Decoder for Video Captioning

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-12T12:10:17.658495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:10:17.494306Z digest=sha256:1a4d4e9ac1af7f0d736b7df3ce051975abcfd9af0cc04e99c3abb4a078333943

Observation 3301e15f-319a-4d25-8899-1ec6326b9f3c · outbound

This paper cites Improving video captioning with temporal composition of a visual-syntactic embedding,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Improving video captioning with temporal composition of a visual-syntactic embedding,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:17.921709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:10:17.499050Z digest=sha256:3c8e84fa56297c40564931a3ef4713132510c0ab8d39b157ec9bc3d82076e531

Observation d3429b60-cc3d-4cf0-a7cf-308a47018340 · outbound

This paper cites Zero-Shot Video Captioning with Evolving Pseudo-Tokens.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Zero-Shot Video Captioning with Evolving Pseudo-Tokens

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.503202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.503202Z digest=sha256:5c557d8000f27a5c4bdb027cc757535aff7cc0c51577c5aa9c188890a1c16410

Observation ce135537-b377-42d9-81ff-aa36e85e9a18 · outbound

This paper cites MultiCapCLIP: Auto-encoding prompts for zero-shot multilingual visual captioning,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning MultiCapCLIP: Auto-encoding prompts for zero-shot multilingual visual captioning,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:17.907312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:10:17.507689Z digest=sha256:9730cda3bffdf2d288d8e62f15888b7d11183811f6c79262920e6fea709befe9

Observation 785c6cf8-21b3-4bff-8e1a-04dec522b27b · outbound

This paper cites Visual instruction tuning,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Visual instruction tuning,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:17.892387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:10:17.511806Z digest=sha256:46016f4ed3a85ad9e5e75fb4f28b4adf5e7a3dd132f2581039bc7f69e61905c5

Observation ec9bfde8-4dd7-4aca-8880-33a90c8adda8 · outbound

This paper cites AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.515974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.515974Z digest=sha256:d2d7a618d877d08643c91a3bad899a8f830b2e12d1415d48c5b3d8dcaf3e2fcd

Observation e2fbcb5a-fe08-4b75-8113-53e73c86bfd5 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Msr-vtt: A large video description dataset for bridging video and language,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.520437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.520437Z digest=sha256:d08c88b17057b59b362bacedc59eec1dbdef51dcfdf09fc79e245bdfba74a445

Observation a5c7b881-8e63-4578-b5d8-2f4ea125bd26 · outbound

This paper cites Collecting highly parallel data for paraphrase evaluation,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Collecting highly parallel data for paraphrase evaluation,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.525100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.525100Z digest=sha256:6c8855cfdb45c0673f90fea4b4ded00a64feda08f510f6a46265adc3b1a6c66c

Observation 6b88c278-12d7-4e40-8221-9daaf3a9ee04 · outbound

This paper cites Vatex: A large-scale, high-quality multilingual dataset for video-and-language research,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Vatex: A large-scale, high-quality multilingual dataset for video-and-language research,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.529439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.529439Z digest=sha256:c6b45441c673fae2cef1b76e528fafda6b1959f2df62840eab65608ad652d3a6

Observation 1673f31e-d513-4bdb-a491-338327be8df2 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Bleu: a method for automatic evaluation of machine translation,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:17.851757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:10:17.534038Z digest=sha256:5a3e34c07c90dc021e44d6d66766ee1b11d21ebf98257b714abd36f9ece86d65

Observation 34951392-678c-4a7e-b98b-3cd0a53ea77f · outbound

This paper cites Meteor universal: Language specific translation evaluation for any target language,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Meteor universal: Language specific translation evaluation for any target language,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:17.836508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:10:17.538541Z digest=sha256:11ceb3c0b7fa4df7b5bf8e9e08af0003ee4c7bf3af3f7b7fc6b0fe0ce230e5eb

Observation 3db6d68b-92f4-4f6b-99bb-bc988827387c · outbound

This paper cites Rouge: A package for automatic evaluation of summaries,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Rouge: A package for automatic evaluation of summaries,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.542918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.542918Z digest=sha256:cf52bee1d0101311b4fbf69cb4f830c931e05c566a9fa496bd9575d6b6b286d0

Observation 242319e9-d6db-45e9-86f0-0eb8885f8a5f · outbound

This paper cites Cider: Consensus- based image description evaluation,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Cider: Consensus- based image description evaluation,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:17.813288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:10:17.547305Z digest=sha256:8e077166318eb86fb03df764bf765490e3cb194306bcbdb7e137858029db398d

Observation ee924bf3-d658-4431-97e2-b98db66504c4 · outbound

This paper cites Decoupled Weight Decay Regularization.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Decoupled Weight Decay Regularization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.551480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.551480Z digest=sha256:c6dad6bfb0b1059d42cb1e251cd0fc4dc492a754ebd9709961e7a35634f1fab0

Observation 11e8d915-f8fd-4515-ba2d-437559911d98 · outbound

This paper cites Expanding language-image pretrained models for general video recognition,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Expanding language-image pretrained models for general video recognition,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:17.799454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:10:17.555961Z digest=sha256:a84a1416e3c4eb464ffca420a3c99595491082f3222f31f976a70e9c8a2b6d9e

Observation 25108cc2-59b5-4d44-8fa3-526165b86e40 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Wan: Open and Advanced Large-Scale Video Generative Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.560221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.560221Z digest=sha256:c1fc0eef30408c1f37dfdd5cc3bc3708755d92663bb3d6a63e6b05a4867636b2

Pith citing papers

No inbound Pith citation observations are available.