Pith. sign in

Paper Citation Record · LEDGER

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation

As of 17 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 2 inbound Pith citation observations for arXiv:2412.10768.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10768 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:43:01.110656Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:38:08.339998Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T20:27:03.913459Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact1
  • verified fuzzy35
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dd153951-d89a-4a29-895b-f5070adee5ff · outbound

This paper cites The Foley grail: The art of perform- ing sound for film, games, and animation.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation The Foley grail: The art of perform- ing sound for film, games, and animation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:02.018261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:00.831775Z digest=sha256:67e9113759cd3ca5cd58143e116160d253be76957ce619f6897855388d245e80

Observation cc69f13e-4ca3-4858-a248-4a172f5a6604 · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation In- structpix2pix: Learning to follow image editing instructions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:02.003279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:00.837425Z digest=sha256:c43f0164d29c6ace9d80dfcd0cadd52d9bae043057b7036afefb78549a6cfa84

Observation 394a27a7-aabb-408b-ba07-37715ce35c5e · outbound

This paper cites Video generation models as world simulators.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Video generation models as world simulators

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.842354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.842354Z digest=sha256:cbf72a01ad06cc6f242a2ec5f6023c68a42b453491db2dbaca448b38bac5dcac

Observation 0963105e-5194-4970-b34c-0f821db9a0f4 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Vggsound: A large-scale audio-visual dataset

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.978661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:00.847396Z digest=sha256:e2e9082f21e53069b54a0bb8d89606dbb25f4b2623b8b726ddfe9f417f0f7c71

Observation 1f456e97-c8d9-4bed-a688-96db9796cdd3 · outbound

This paper cites Gentron: Diffusion trans- formers for image and video generation.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Gentron: Diffusion trans- formers for image and video generation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.963716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:00.852162Z digest=sha256:bfee2d505281276da6cde6e103e243812d07e5c5985fc1d8ce6bd577124e96d8

Observation 7504d809-a4a5-4610-a200-d4f02bb44391 · outbound

This paper cites Audio-vision: sound on screen.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Audio-vision: sound on screen

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.948849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:00.856979Z digest=sha256:20c8e5e7dc846f54f283aa6e7184b43ecc68b843ac4429a0c531975e414fed69

Observation 42357cef-9a0e-4976-b4d7-37d60620d868 · outbound

This paper cites Scaling instruction- finetuned language models.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Scaling instruction- finetuned language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.933714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:00.861754Z digest=sha256:dd654d8142755ce349d91442899fb408def7d8fbc66941cab2dabc7281a84d99

Observation 70a28615-c939-412f-addc-4ce07851924a · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.867260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.867260Z digest=sha256:a695cd7dbdc81e47867e638282e39cd09cce05ad1684214b795790d49e2e6196

Observation b5e1b2c8-5007-4e4c-a5a6-da577476013d · outbound

This paper cites Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.877098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.877098Z digest=sha256:9c8a82dc8975d19ef283ef8528dbf2cc2ebfdabd9167db9e97c9a0673906e013

Observation f9bd243d-5c75-4739-afa0-c72c3a3200c4 · outbound

This paper cites Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.881954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.881954Z digest=sha256:5f35b390f334718f43d11b2de9ed146dbac0c6b85aad440b68f110241b94fb5c

Observation b0a8ec89-425b-4bb6-824d-e9428d88aac6 · outbound

This paper cites Imagebind: One embedding space to bind them all.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Imagebind: One embedding space to bind them all

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.917393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:00.887089Z digest=sha256:b480e7c9d724c5ed16bb378cf599c2f1ef6d522909466668f2f59af2356700a1

Observation fe58a50c-5079-4ba4-9d18-971491a11c80 · outbound

This paper cites Determining op- tical flow.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Determining op- tical flow

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.902604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:00.891837Z digest=sha256:7a50321cc6a31dd063ba6237b5532ef78a07a71e2fa7806caed51f9e8f5becd2

Observation 81536570-f0b8-4950-ab39-2777285f32f3 · outbound

This paper cites Make-an-audio: Text-to-audio gen- eration with prompt-enhanced diffusion models.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Make-an-audio: Text-to-audio gen- eration with prompt-enhanced diffusion models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.887430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:00.896690Z digest=sha256:541404af9903627beb62e50edd033814df93f2c1432984ef2202b2f08564a543

Observation 3a79237d-0f2d-4eb2-b0a3-fc343e6218b1 · outbound

This paper cites Captivating sound.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Captivating sound

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.871039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:00.901474Z digest=sha256:afb388741e8f1725ba37142c9dd8f3470b93aa042172df3bdaef87134ad2dad9

Observation 43da5ea3-8cb9-4f97-b8a6-e16b0aee59ba · outbound

This paper cites Taming visually guided sound generation.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Taming visually guided sound generation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.855536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:00.906290Z digest=sha256:4532119ec18574d2d562ff26372b323ddfc6446caa8d23922026a4949275c213

Observation 7bbc4dc9-61b1-4ad3-812f-df8ad7362d2c · outbound

This paper cites Mixing audio: concepts, practices, and tools.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Mixing audio: concepts, practices, and tools

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.839857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:00.911140Z digest=sha256:f09470082f6daa52b54c2357cfb56d35520c051b6c0a41e9b01a952e8d8609b8

Observation 2432f03f-54a6-4d14-abc8-427a56b07f0c · outbound

This paper cites Read, watch and scream! sound generation from text and video.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Read, watch and scream! sound generation from text and video

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.825042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:00.915875Z digest=sha256:7b3508be4002a461aa18d5a84e242cfdf281dbfba2addd160f0179e2ea38c953

Observation f01f30cf-82ef-45fb-bb91-d7747fc80aba · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.920633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.920633Z digest=sha256:4b0e17f488b4cd1add180655bc728c2a6397edfdd7e857e8d6acf6542893e64c

Observation 3133d9e4-9647-48bd-bcd2-f71ab0d354f0 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fi- delity speech synthesis.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Hifi-gan: Generative adversarial networks for efficient and high fi- delity speech synthesis

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.810087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:00.925539Z digest=sha256:ccd6d43f2ff0993e407df6cb8577d46a8494388bdb0004b971fc808fcb6bc521

Observation 75b0328f-621c-491d-a3ed-870cadf9102b · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation AudioGen: Textually Guided Audio Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.930103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.930103Z digest=sha256:28c96fc348e62f28e114cbc4a36d25be3f5ba3fb1245e0d8265f6be75310e09b

Observation 38d1f91e-193d-48cb-a0d1-73980d2d4881 · outbound

This paper cites Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-11T15:43:01.305284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:00.935561Z digest=sha256:799f32165ad91b2a708e9da778668940feae0408ace3eca7b20d9e7dc170647c

Observation b8b58e80-68ad-478f-a8f0-d83a2c93cfa3 · outbound

This paper cites V oice- box: Text-guided multilingual universal speech generation at scale, 2023.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation V oice- box: Text-guided multilingual universal speech generation at scale, 2023

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.793935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:00.940625Z digest=sha256:a9d82f72843b223a4b294b3ac757474219cb7af4a94688082adb31b4eb51dd0c

Observation 67ef12d0-af5a-4b55-9d23-416a97f5b652 · outbound

This paper cites Flow Matching for Generative Modeling.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Flow Matching for Generative Modeling

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.945415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.945415Z digest=sha256:f2885b92c83123eb77b1c745a424159921d9cc5b7db3adfaf87c583ff065dfb0

Observation bddc6b1e-a04d-4b05-a33e-ea5dabcbdd7d · outbound

This paper cites Audi- oLDM: Text-to-audio generation with latent diffusion mod- els.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Audi- oLDM: Text-to-audio generation with latent diffusion mod- els

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.776837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:00.950472Z digest=sha256:44a624212fd7e5e8cc749af13211e97ec19b5938aaed6884fcbda50af13815fe

Observation 78e9addd-a28d-476c-9917-4230865d43b7 · outbound

This paper cites FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio Generation.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.955420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.955420Z digest=sha256:af5fdb60b8b96df721e48067efc9355fd45415b91261a41e280b2e326ae19615

Observation a8892f56-768c-46bc-bfca-c8d4807d0c24 · outbound

This paper cites Plumbley.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Plumbley

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.759012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:00.960680Z digest=sha256:8c6a994e6a0b9c6f139995b6217ea47608d5004255ca65978b82c4efb1305ec2

Observation 8fcadfb4-164f-4c07-b2e9-6d9c8ec496b8 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.965632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.965632Z digest=sha256:94e057d147017fcf11386e5649e07ecbb1b0a736dd33198cbd9cd1837de7bf5d

Observation 78d0d2b7-7e5e-410d-935a-c9cac5453cd8 · outbound

This paper cites Separate Anything You Describe.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Separate Anything You Describe

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.970707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.970707Z digest=sha256:2f356254f84cd1c958117f824df50ea53139db8228ce23b0fba7c9c41cf08e60

Observation 8d314d68-d31e-46f1-b28a-60ad9c7dd7e9 · outbound

This paper cites Diff-foley: Synchronized video-to-audio synthesis with la- tent diffusion models, 2023.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Diff-foley: Synchronized video-to-audio synthesis with la- tent diffusion models, 2023

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.743876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:00.975676Z digest=sha256:0b6cd20f03b853c6c0e1290fca248221ba88984cf59c7b950e63640713af2bb1

Observation 19bb9429-a598-4d77-8b51-38262cfc3520 · outbound

This paper cites Albergo, Nicholas M.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Albergo, Nicholas M

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.727746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:00.980418Z digest=sha256:7b13a49808d346ecb290069718d8429e7d59928a52366f1b5aea4a964c4d8756

Observation 3c02fe20-b9e2-4a7f-9b7a-4668f091d418 · outbound

This paper cites Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization, 2024.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.710578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:00.985150Z digest=sha256:eec130b472cca148ede87ee0f71e2dd382304629be084006fd095c7d5e30f68d

Observation d496671b-d7de-463a-87d5-87e22465082e · outbound

This paper cites SampleRNN: An Unconditional End-to-End Neural Audio Generation Model.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation SampleRNN: An Unconditional End-to-End Neural Audio Generation Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.989949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.989949Z digest=sha256:60cdef7bf4d3e8b44e077afd65c44f9c3f8925b41b2a33265b45aec90fe0156c

Observation b6b1c667-3311-4ac6-be88-5c8557f19490 · outbound

This paper cites Text-to-Audio Generation Synchronized with Videos.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Text-to-Audio Generation Synchronized with Videos

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.994918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.994918Z digest=sha256:2f294768e77386abd360ff25c4f18672654e8c62f4460d07fc258dede09a466a

Observation c193d648-be82-4cc0-9ed6-b7975fab2e0d · outbound

This paper cites Bal- ancing act: Distribution-guided debiasing in diffusion mod- els.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Bal- ancing act: Distribution-guided debiasing in diffusion mod- els

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.695014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:01.000235Z digest=sha256:416c82d5ed13d5083ae6b14b8ad63da75a1e318150ea321301a9c70d587e0756

Observation 060323ee-af08-4265-8a17-867538b01346 · outbound

This paper cites Scalable diffusion models with transformers, 2023.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Scalable diffusion models with transformers, 2023

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.679620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:01.004938Z digest=sha256:0230147051821b1c2e5b9195c7b2f1715c56bca3100d5dab185ae4b815e18125

Observation 341338be-00a1-4f5a-9cce-1df26449df12 · outbound

This paper cites Film: Visual reasoning with a general conditioning layer.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Film: Visual reasoning with a general conditioning layer

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.663487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:01.009584Z digest=sha256:3a7759f12c629e37c7b057cf8ae45a0c2d52b3c9cf5cb54540aec6c7375bd9dc

Observation 91a4fadc-1435-4907-8ed2-f3cc48242973 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Learning transferable visual models from natural language supervi- sion

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.647834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:01.014105Z digest=sha256:299a4dc58451d28cfee9be891215ac011f7f90f29dc0a77ad03a1fef4ea7b678

Observation 3b585338-d1ed-4056-9f83-e666efe31037 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:01.018952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:01.018952Z digest=sha256:f4447f46d478677d97c5a673b771ef722a5b4dfd1915498c3e2502e2f56f2827

Observation 3fe98423-9360-4269-92b1-00befebb6bbe · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation High-resolution image synthesis with latent diffusion models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:01.024145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:01.024145Z digest=sha256:c7558d44ed2610761fc7ff48da595281073a5204f960f38c9ed58b6eba2e5de3

Observation 757ba374-dd75-4de7-a979-d360d2eb4d65 · outbound

This paper cites I hear your true colors: Image guided audio generation, 2022.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation I hear your true colors: Image guided audio generation, 2022

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.608310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:01.029283Z digest=sha256:6af15bc7969bf10b7f198a98c18699a3f206530a310a33319dcbde9ca6ad6b06

Observation 42c483e9-14f8-407e-bf7b-ef23dc79e883 · outbound

This paper cites Auto- acd: A large-scale dataset for audio-language representation learning.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Auto- acd: A large-scale dataset for audio-language representation learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.591925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:01.034089Z digest=sha256:4137fa625b33981611f6b3b94ee416ef108d225ae749ae29aa5e7e8ccd78d9d3

Observation b5e2a8e2-082f-4ce4-b38f-11c8bb590245 · outbound

This paper cites Learning from Between-class Examples for Deep Sound Recognition.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Learning from Between-class Examples for Deep Sound Recognition

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:01.039003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:01.039003Z digest=sha256:03f561bc934ec0777586ed674e66416cfb59c7ad7557f2738e6da6636d415b0f

Observation 9159ca6c-bf3d-4acf-b2aa-ae3f5fb14cb0 · outbound

This paper cites Attention is all you need.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Attention is all you need

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:01.044268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:01.044268Z digest=sha256:7ea0472db05b1b43f7d59e1d7f21bbfd8d3244b7875d7c735c3f2a7dd4d4c572

Observation dff84b15-706b-4178-a31c-c2e78cff4f93 · outbound

This paper cites Audiobox: Unified audio generation with natural language prompts, 2023.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Audiobox: Unified audio generation with natural language prompts, 2023

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.562666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:01.049825Z digest=sha256:6dc083fb7c6008397a7454f46939678c0120555590daf321f7bc35f556e5f54f

Observation 44a95349-babf-4205-b7ce-e6d28e03dc31 · outbound

This paper cites V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foun- dation models.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foun- dation models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.545952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:01.054754Z digest=sha256:529e81118df58c8a117e121db44f9f542504056fd17914c1102934e636c6a355

Observation 5a5265e9-f206-4211-991e-1848bed6af07 · outbound

This paper cites ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:01.059533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:01.059533Z digest=sha256:d9549f89c886fec87c7d94d8df261120dd38e3f01fb904ae24a9e827b5bf8508

Observation 559861e4-be63-48f3-a037-448bc3afc367 · outbound

This paper cites Wav2clip: Learning robust audio repre- sentations from clip.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Wav2clip: Learning robust audio repre- sentations from clip

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.529275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:01.065138Z digest=sha256:26bbbe0843b4b4969f86f3c4d04791aa2b74c5caa3fb78f5dbde2c70a4e9257e

Observation e8c49149-27bc-4322-8696-04682191ead8 · outbound

This paper cites Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.512058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:01.069941Z digest=sha256:b9f054a83d4be48c5e056fc5d7dc122f5bc579e3cc9d4138e49173ee9dace27e

Observation d7192d4f-08a4-4c10-b83c-894b0be99496 · outbound

This paper cites Son- icvisionlm: Playing sound with vision language models.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Son- icvisionlm: Playing sound with vision language models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.494889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:01.074951Z digest=sha256:d6a18ed189fc18a99219d6c1a387497128c2e5b9e00a34a0d3ddc549b504aa52

Observation 436724e7-41b5-4930-a16c-78d0a5a465e6 · outbound

This paper cites Seeing and hearing: Open-domain visual- audio generation with diffusion latent aligners.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Seeing and hearing: Open-domain visual- audio generation with diffusion latent aligners

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:01.079975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:01.079975Z digest=sha256:19752c76533e812c2ceec8e822bb60970da5501059c8f5477483c8f827342034

Observation 6b8802d4-dd19-427d-b44a-c1a00998902c · outbound

This paper cites Diffsound: Discrete diffusion model for text-to-sound generation.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Diffsound: Discrete diffusion model for text-to-sound generation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.466318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:01.085034Z digest=sha256:2ddc604a941736995f850245c4eeaabad1a5e8d5a64d7fc9ac1702c594b50fc7

Observation 1403e556-2132-4faa-b35f-1318ecaba71c · outbound

This paper cites Diverse and aligned audio-to-video genera- tion via text-to-video model adaptation.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Diverse and aligned audio-to-video genera- tion via text-to-video model adaptation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.450177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:01.090156Z digest=sha256:a82b334df53aeec352836f37eaa842bd057bdad42905d14babe95afc5b488c27

Observation 079eb143-1244-4a61-9e47-b3b193f10348 · outbound

This paper cites Foley- crafter: Bring silent videos to life with lifelike and synchro- nized sounds, 2024.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Foley- crafter: Bring silent videos to life with lifelike and synchro- nized sounds, 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.434140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:01.095147Z digest=sha256:33b89d8dcddedf6e40f4dfc30b2c62885b00ce644a4e5f3f2043d96805c8c923

Observation 8bafe478-14c6-4828-83f8-99f37de51dc7 · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Llava- next: A strong zero-shot video understanding model, 2024

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.417921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:01.099877Z digest=sha256:6d3d4041a0abe279d2d10ba0ec4e43a4874f7df4059d40605bc2cc941d5cc384

Observation c84ec568-408e-435f-8846-805b45ad41f6 · outbound

This paper cites an unresolved cited work.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:43:01.402261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:43:01.104419Z digest=sha256:df41d02ed3661db3921ccecdbf8fb964e671df05ecce7a4856b5539729f3f9e6

Observation 10fc6dca-0337-4a2c-b5dd-4542c2b61058 · outbound

This paper cites Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:01.110656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:01.110656Z digest=sha256:c99636d136b36c245c1ab03114ef3528dcea512f1793e1845a5f57192ecb4392

Pith citing papers

Observation 52d52cb8-48a5-4be9-989e-07211047f4bf · inbound

AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation cites this paper.

AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:08.339998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:08.339998Z digest=sha256:c37e14b1be1d2ac96704dbb8f89b733b714df5be62e8802afff8ee97563d3563

Observation eaa0d4c0-e44a-4942-bc74-9a9d04cb9093 · inbound

Sound Scene Synthesis at the DCASE 2024 Challenge cites this paper.

Sound Scene Synthesis at the DCASE 2024 Challenge VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:27:03.921470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T20:27:03.773063Z digest=sha256:dcb9cfa0eb745ec9ef1d67a5a19e830c38630b15c0a0cde5924d8c9db680b26f