Pith. sign in

Paper Citation Record · LEDGER

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning

As of 13 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2608.11013.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11013 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:10:17.560221Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact2
  • verified fuzzy23
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4d7ac9b3-8ceb-4ae3-87d9-7bf9299bb4cc · outbound

This paper cites Learning to compose topic-aware mixture of experts for zero-shot video captioning,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Learning to compose topic-aware mixture of experts for zero-shot video captioning,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:18.171021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:10:17.374700Z digest=sha256:56d5eac3e5168a70f83331b3dfe07e8d9b81181f20b0bfe4420b67262d17b47c

Observation f533bef9-e475-42cf-936d-a43eb374ae32 · outbound

This paper cites DeCap: Decoding CLIP Latents for Zero-Shot Captioning via Text-Only Training.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning DeCap: Decoding CLIP Latents for Zero-Shot Captioning via Text-Only Training

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.379818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.379818Z digest=sha256:4817b646cc78a512c231698b46d1fc5cf1abbb8af4acac79f391059399e9c408

Observation 19c991cc-536b-47da-9036-bc94ea65deb0 · outbound

This paper cites Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.384647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.384647Z digest=sha256:b6a61c9215e7d4658092bc4208aa81627cbc141691eee25e8349ffceaf2dd61d

Observation da62bf44-5334-46a7-a68d-666868ee2414 · outbound

This paper cites IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-12T12:10:17.754368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:10:17.389553Z digest=sha256:c90084599082a47214b095c370c5d246959fc21872eed19a552a72c4cf1451ad

Observation 64294703-8e7a-4721-ae98-2d64fe6fc190 · outbound

This paper cites Improving cross-modal alignment with synthetic pairs for text-only image captioning,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Improving cross-modal alignment with synthetic pairs for text-only image captioning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:18.156961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:10:17.394603Z digest=sha256:c749faea2312a708be164ca2f2cf9679728ef9dfbe9d7b26f63bcebc615cab1e

Observation 507ee31f-a271-46ee-9576-c776fcacef4d · outbound

This paper cites Retta: Retrieval-enhanced test-time adaptation for zero-shot video cap- tioning,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Retta: Retrieval-enhanced test-time adaptation for zero-shot video cap- tioning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:18.143378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:10:17.399148Z digest=sha256:de775372b0d7655ed96840414d7101ab98451a9f2a70227073e8546a4b41a768

Observation 624c36ab-f79f-4326-8ab3-91b525f424ff · outbound

This paper cites Text-only training for image captioning using noise-injected CLIP,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Text-only training for image captioning using noise-injected CLIP,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:18.129207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:10:17.404015Z digest=sha256:366cfd30dd47330d6224baea8aa6d47e3fe63701a7b71c866eb3f1d5af99a7dd

Observation ebb39956-b509-4919-8638-3a7cb53d208b · outbound

This paper cites Sequence to sequence-video to text,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Sequence to sequence-video to text,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:18.115003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:10:17.408247Z digest=sha256:474583be050dcb6446682272e63e1f6ca986838dcd84c03dd6cf84068dbbdbdb

Observation 87b366b9-07bc-4024-99ce-0f2f9d010979 · outbound

This paper cites Bidirectional long- short term memory for video description,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Bidirectional long- short term memory for video description,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:18.100003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:10:17.412383Z digest=sha256:f717a19436e277921c53385d7c168bbce75d2b3b1af68889778d1ce5ae0c29f7

Observation 351d23f9-35a5-44a6-bddd-c03b3be6856e · outbound

This paper cites Graph convolutional network meta- learning with multi-granularity pos guidance for video captioning,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Graph convolutional network meta- learning with multi-granularity pos guidance for video captioning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:18.084924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:10:17.416548Z digest=sha256:51018450f64cdf645b31ace04f74e72960b0455655351a07543da35867f36d52

Observation 825e202e-8fe2-4ee3-b98c-b07524bac78e · outbound

This paper cites Describing videos by exploiting temporal structure,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Describing videos by exploiting temporal structure,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:18.070508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:10:17.420779Z digest=sha256:ddff3d65013a5b6954b6e2dde691fd4223cc52557f450230077a3d9c4efea947

Observation fde556e8-833d-4bdd-b1f7-9c59505c39da · outbound

This paper cites Icocap: Improving video captioning by compounding images,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Icocap: Improving video captioning by compounding images,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:18.056297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:10:17.425189Z digest=sha256:d1853dd016635b51536db474ede69b6546b8a6df19a313547a943b0e0511538e

Observation 2b072796-1967-427d-9af9-b0957ac999b3 · outbound

This paper cites Memory- based augmentation network for video captioning,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Memory- based augmentation network for video captioning,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:18.041638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:10:17.429108Z digest=sha256:d46bbf2bc1c638b93d01e563cf962ae90a9acd901f5741d12d89af24debad671

Observation 51c5fec0-3cc5-4a59-8530-6b7ade97ba1b · outbound

This paper cites Attention is all you need,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Attention is all you need,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:18.026051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:10:17.433335Z digest=sha256:30c31490c877e0d1871a7e9acf8631dddc7e94eb2a25ad72ec0accf21187c22f

Observation 73b61cdb-cd4a-451d-9b67-0ce755e24ddf · outbound

This paper cites Hierarchical modular network for video captioning,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Hierarchical modular network for video captioning,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:18.010478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:10:17.437339Z digest=sha256:ba217d66813f712a651802627102a1d1707fa13578a969e42ffa4ef55b5c2989

Observation 841c5f3c-b9ce-4f86-9946-389a85bccb3b · outbound

This paper cites Swinbert: End-to-end transformers with sparse attention for video captioning,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Swinbert: End-to-end transformers with sparse attention for video captioning,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:17.995895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:10:17.441498Z digest=sha256:60e7202ce3dc07683d7e3350b32abfb3c779d5ef190ede8bc34006a1e696e8eb

Observation bc11fe1d-6959-434e-bc19-90eb76ead7d6 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.446549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.446549Z digest=sha256:006a1884e33ff272afad7aa44765215c70fe843f8e086909173f443d41b6a5d9

Observation 8cabbb01-3849-49c3-a761-3b511f14332c · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video understanding,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Ma-lmm: Memory-augmented large multimodal model for long-term video understanding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:17.981085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:10:17.452167Z digest=sha256:a60ecf993ca53098b44dd2c68e6a0823e05a2432956cef11f8cc91df5eb9b1b4

Observation a9dfdb99-f0a2-475d-b9ba-b17f6bc85d72 · outbound

This paper cites Text-Only Training for Image Captioning using Noise-Injected CLIP.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Text-Only Training for Image Captioning using Noise-Injected CLIP

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.456353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.456353Z digest=sha256:da1a8fe0657b62a6c842257366299d8e89ccf06fb07f2fe5dd3b22c7935d5941

Observation 86c4d617-89d4-4634-bf51-4ffe29d951be · outbound

This paper cites Language Models Can See: Plugging Visual Controls in Text Generation.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Language Models Can See: Plugging Visual Controls in Text Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.461499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.461499Z digest=sha256:120555c03961a4404588f22582179392f8682f3abd0812d9eddbb8cc06644307

Observation e0f70ae5-8cda-4d16-8160-57cd5ed326ae · outbound

This paper cites Zerocap: Zero-shot image-to-text generation for visual-semantic arithmetic,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Zerocap: Zero-shot image-to-text generation for visual-semantic arithmetic,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:17.967231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:10:17.466408Z digest=sha256:e138b38a7b26e96178964809207ee13db618ced246e5f0dd6fe6c2cb4eab8b2e

Observation 4241089a-f193-4dc9-8fb1-ad9b694a9576 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Learning transferable visual models from natural language supervision,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:17.953545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:10:17.471210Z digest=sha256:32af3a632a54711515a83967808f6466224816b9258163a101c3aa4f31b48b64

Observation 45b26696-9dfd-472a-bc6e-f1d27156b202 · outbound

This paper cites From Association to Generation: Text-only Captioning by Unsupervised Cross-modal Mapping.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning From Association to Generation: Text-only Captioning by Unsupervised Cross-modal Mapping

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.475874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.475874Z digest=sha256:f1bf009daa09d86df2e96e2355a6c2e7b601752a9c8716dd5409ca78de50decb

Observation 8131c68b-a5d3-46e6-b0ef-e539345c6f38 · outbound

This paper cites Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.480381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.480381Z digest=sha256:954fe32d77efec4203a8b5b6742b9e72ddcf42ad045c1a3ca915dff3d87f789c

Observation 986d18f8-79b1-4ef9-9ab4-04171e9586e4 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.485268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.485268Z digest=sha256:29180a96cb6848a3b07fe7112b401e25969455c2d79d3c43bc1879b22de1b35c

Observation 2c2ce2d1-fb0e-4d80-b6c6-f90609bcdebc · outbound

This paper cites Language models are unsupervised multitask learners,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Language models are unsupervised multitask learners,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.490054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.490054Z digest=sha256:f92ada8160e612681314d6f0e92ca7269a3c0b8dfde348009b3ddbf4ce04c2d8

Observation 013a8cef-d9a6-40d7-a267-57d1516bd0d8 · outbound

This paper cites Delving Deeper into the Decoder for Video Captioning.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Delving Deeper into the Decoder for Video Captioning

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-12T12:10:17.658495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:10:17.494306Z digest=sha256:6d2abb2f59f67e9f4b538ea5d77e5c85fb3e684a02bd64e3d2d2e59cae12e1f8

Observation 3301e15f-319a-4d25-8899-1ec6326b9f3c · outbound

This paper cites Improving video captioning with temporal composition of a visual-syntactic embedding,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Improving video captioning with temporal composition of a visual-syntactic embedding,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:17.921709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:10:17.499050Z digest=sha256:97d8ec80bc59b77cbe9512484b0b240d8c19b84dd55fb5aaee36855e27a0da30

Observation d3429b60-cc3d-4cf0-a7cf-308a47018340 · outbound

This paper cites Zero-Shot Video Captioning with Evolving Pseudo-Tokens.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Zero-Shot Video Captioning with Evolving Pseudo-Tokens

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.503202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.503202Z digest=sha256:9f7e643627e620ec3d53797644148a4aeec638e12e680eba0826b388f60c0427

Observation ce135537-b377-42d9-81ff-aa36e85e9a18 · outbound

This paper cites MultiCapCLIP: Auto-encoding prompts for zero-shot multilingual visual captioning,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning MultiCapCLIP: Auto-encoding prompts for zero-shot multilingual visual captioning,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:17.907312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:10:17.507689Z digest=sha256:a684069cad57b61a33cfcfa4f77daedd1ff32cf0a0d75921299f2374a868b57b

Observation 785c6cf8-21b3-4bff-8e1a-04dec522b27b · outbound

This paper cites Visual instruction tuning,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Visual instruction tuning,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:17.892387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:10:17.511806Z digest=sha256:fa74727cd2b06f598b073aae731888b90bbf9242b6b4641b16cd0075d300c453

Observation ec9bfde8-4dd7-4aca-8880-33a90c8adda8 · outbound

This paper cites AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.515974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.515974Z digest=sha256:6b1885704eb17fa5c239d0e27c2c0a944688e7a8ae01f103b35f44b04fae1ca3

Observation e2fbcb5a-fe08-4b75-8113-53e73c86bfd5 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Msr-vtt: A large video description dataset for bridging video and language,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.520437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.520437Z digest=sha256:8cc6df5e4e3e36f00d3eb6b235bcde897053783048280019c58dfea8feb6595a

Observation a5c7b881-8e63-4578-b5d8-2f4ea125bd26 · outbound

This paper cites Collecting highly parallel data for paraphrase evaluation,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Collecting highly parallel data for paraphrase evaluation,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.525100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.525100Z digest=sha256:e24ba78161515d790fd091d3cd33ffd64b9edb91de603ff5931c93a468d79f74

Observation 6b88c278-12d7-4e40-8221-9daaf3a9ee04 · outbound

This paper cites Vatex: A large-scale, high-quality multilingual dataset for video-and-language research,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Vatex: A large-scale, high-quality multilingual dataset for video-and-language research,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.529439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.529439Z digest=sha256:92f0bdd02875caed939620b74daf9dbfc515571c7bdbc1c1aaf26bfe0d4f6b3f

Observation 1673f31e-d513-4bdb-a491-338327be8df2 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Bleu: a method for automatic evaluation of machine translation,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:17.851757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:10:17.534038Z digest=sha256:fb315d251f962e27ad314b5f333db3a0aaca0931febfe230e8727255615b13da

Observation 34951392-678c-4a7e-b98b-3cd0a53ea77f · outbound

This paper cites Meteor universal: Language specific translation evaluation for any target language,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Meteor universal: Language specific translation evaluation for any target language,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:17.836508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:10:17.538541Z digest=sha256:24f10685d53826047979156faf8f66fda079e08bf7bf7454b851f2a4210a522d

Observation 3db6d68b-92f4-4f6b-99bb-bc988827387c · outbound

This paper cites Rouge: A package for automatic evaluation of summaries,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Rouge: A package for automatic evaluation of summaries,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.542918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.542918Z digest=sha256:db8a55d267bfc00bdad2f7313e1e4e11038845a7897130c242072f57d76cbc11

Observation 242319e9-d6db-45e9-86f0-0eb8885f8a5f · outbound

This paper cites Cider: Consensus- based image description evaluation,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Cider: Consensus- based image description evaluation,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:17.813288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:10:17.547305Z digest=sha256:d4695ea49b6fe3c00e4012f5cf59ad6c908a94a427d8f58c5d3b7a57f70801ae

Observation ee924bf3-d658-4431-97e2-b98db66504c4 · outbound

This paper cites Decoupled Weight Decay Regularization.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Decoupled Weight Decay Regularization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.551480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.551480Z digest=sha256:d813667118bc8a646519b9b60e6d92f51712c185efd4c423807e1b1fcb3adc42

Observation 11e8d915-f8fd-4515-ba2d-437559911d98 · outbound

This paper cites Expanding language-image pretrained models for general video recognition,.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Expanding language-image pretrained models for general video recognition,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:17.799454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:10:17.555961Z digest=sha256:86407067016f4df1b8bcfdec522ccd46fdd600865b7bf8c03a5b4f1d182c6314

Observation 25108cc2-59b5-4d44-8fa3-526165b86e40 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning Wan: Open and Advanced Large-Scale Video Generative Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:17.560221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:17.560221Z digest=sha256:7cf1ec85aa60fef10a348aa915b04d16d57a730e9651f8a621570ba22ce0647f

Pith citing papers

No inbound Pith citation observations are available.