Pith. sign in

Paper Citation Record · LEDGER

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation

As of 12 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 2 inbound Pith citation observations for arXiv:2412.10768.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10768 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:43:01.110656Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:38:08.339998Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T20:27:03.913459Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact1
  • verified fuzzy35
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dd153951-d89a-4a29-895b-f5070adee5ff · outbound

This paper cites The Foley grail: The art of perform- ing sound for film, games, and animation.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation The Foley grail: The art of perform- ing sound for film, games, and animation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:02.018261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.831775Z digest=sha256:9ad669713faef13a50e1d2bd89cc50474df712c9d0d3bdc570b85f92e76796a2

Observation cc69f13e-4ca3-4858-a248-4a172f5a6604 · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation In- structpix2pix: Learning to follow image editing instructions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:02.003279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.837425Z digest=sha256:c73600b8d21c83912ea0c4b14b56d77e57a0b44ff5b6c347cfef35b37313321f

Observation 394a27a7-aabb-408b-ba07-37715ce35c5e · outbound

This paper cites Video generation models as world simulators.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Video generation models as world simulators

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.842354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.842354Z digest=sha256:1c2d0c80f9e55edd641ad3b10c227e38cce9ab445a95e666ce1685ffc7b36ca7

Observation 0963105e-5194-4970-b34c-0f821db9a0f4 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Vggsound: A large-scale audio-visual dataset

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.978661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.847396Z digest=sha256:6a160acf9d4e5d8183b58213cb3e8092b2316fe15ea2d160bfc0c249a59afa58

Observation 1f456e97-c8d9-4bed-a688-96db9796cdd3 · outbound

This paper cites Gentron: Diffusion trans- formers for image and video generation.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Gentron: Diffusion trans- formers for image and video generation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.963716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.852162Z digest=sha256:2a7e75e075e6a9bbd209a78bc574f4c8ce47cfeeb2aea3f4d48a7d6fcc74698c

Observation 7504d809-a4a5-4610-a200-d4f02bb44391 · outbound

This paper cites Audio-vision: sound on screen.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Audio-vision: sound on screen

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.948849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.856979Z digest=sha256:7d5f3c024324e30f8ce3ca87efda8c1f695d82395e0532b1fa3919da5e28bbc6

Observation 42357cef-9a0e-4976-b4d7-37d60620d868 · outbound

This paper cites Scaling instruction- finetuned language models.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Scaling instruction- finetuned language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.933714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.861754Z digest=sha256:c3a56410d38b7087a208b96e0147b1f07624fdd098d3841481a5c97766517a26

Observation 70a28615-c939-412f-addc-4ce07851924a · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.867260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.867260Z digest=sha256:c5db15366b8a2c73189219e6170e3e3533163a73cbf63bfa136568a43daf1bf2

Observation b5e1b2c8-5007-4e4c-a5a6-da577476013d · outbound

This paper cites Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.877098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.877098Z digest=sha256:45580a421acfbde3b9ac731a5e2139a97f8611dc1acb0237cc910d79a3f870b6

Observation f9bd243d-5c75-4739-afa0-c72c3a3200c4 · outbound

This paper cites Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.881954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.881954Z digest=sha256:b2dd7f242b6db9d387d429b90e722815732f6fdb083f7adc3d20946ab7b31cab

Observation b0a8ec89-425b-4bb6-824d-e9428d88aac6 · outbound

This paper cites Imagebind: One embedding space to bind them all.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Imagebind: One embedding space to bind them all

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.917393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.887089Z digest=sha256:3505c4bde23225fed68843a261e88fba14c5f51c9c91b7de1c5e7ee9ea480fb9

Observation fe58a50c-5079-4ba4-9d18-971491a11c80 · outbound

This paper cites Determining op- tical flow.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Determining op- tical flow

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.902604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.891837Z digest=sha256:c6dd50a49d736147f80123e10ba06d6262b26c27753adbd3389acbfd605b6bdb

Observation 81536570-f0b8-4950-ab39-2777285f32f3 · outbound

This paper cites Make-an-audio: Text-to-audio gen- eration with prompt-enhanced diffusion models.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Make-an-audio: Text-to-audio gen- eration with prompt-enhanced diffusion models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.887430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.896690Z digest=sha256:6da6f80a9c142a298a69467ef6644b8a1e8efdba48273f2558485c980d5a1501

Observation 3a79237d-0f2d-4eb2-b0a3-fc343e6218b1 · outbound

This paper cites Captivating sound.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Captivating sound

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.871039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.901474Z digest=sha256:5929932040b4f48c3c2188feffcdfbbc91786ab96467bbd1a56743577f86fb4c

Observation 43da5ea3-8cb9-4f97-b8a6-e16b0aee59ba · outbound

This paper cites Taming visually guided sound generation.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Taming visually guided sound generation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.855536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.906290Z digest=sha256:6fe6505ef092c8749af44ebb591431a3724f1a1591b1b7149de00193647625a8

Observation 7bbc4dc9-61b1-4ad3-812f-df8ad7362d2c · outbound

This paper cites Mixing audio: concepts, practices, and tools.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Mixing audio: concepts, practices, and tools

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.839857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.911140Z digest=sha256:9334eb37b53e0dbfe32331a8de46dfb9ecc99783d55017854bb0e4d5f44d4a22

Observation 2432f03f-54a6-4d14-abc8-427a56b07f0c · outbound

This paper cites Read, watch and scream! sound generation from text and video.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Read, watch and scream! sound generation from text and video

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.825042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.915875Z digest=sha256:b374cff299424bac5d6bbcbe0f2f46fe5b7bab1ff8547f5d617eb7c774fd2fb3

Observation f01f30cf-82ef-45fb-bb91-d7747fc80aba · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.920633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.920633Z digest=sha256:be1911c05d1f2f4367245f75614ef9c02dc46c750c1490bee6a0b6aeb4ba3951

Observation 3133d9e4-9647-48bd-bcd2-f71ab0d354f0 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fi- delity speech synthesis.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Hifi-gan: Generative adversarial networks for efficient and high fi- delity speech synthesis

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.810087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.925539Z digest=sha256:aff3787c6cd7537b574599f31010c144a0386c26cc28b754ba281570f88b9613

Observation 75b0328f-621c-491d-a3ed-870cadf9102b · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation AudioGen: Textually Guided Audio Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.930103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.930103Z digest=sha256:72bba6b19e77bd4e45dc76125b14b6e14c43ffd070dcfa18b927efce6eab7b58

Observation 38d1f91e-193d-48cb-a0d1-73980d2d4881 · outbound

This paper cites Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-11T15:43:01.305284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.935561Z digest=sha256:6ffd8026d66d5e7bafa2083c5c3e814c7a8c7da9b2d3abd91ccfc1527f8d40a2

Observation b8b58e80-68ad-478f-a8f0-d83a2c93cfa3 · outbound

This paper cites V oice- box: Text-guided multilingual universal speech generation at scale, 2023.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation V oice- box: Text-guided multilingual universal speech generation at scale, 2023

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.793935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.940625Z digest=sha256:9754ae7ed920ab4a839cedaf8d74acc86770f512032b3577198454acc3941c95

Observation 67ef12d0-af5a-4b55-9d23-416a97f5b652 · outbound

This paper cites Flow Matching for Generative Modeling.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Flow Matching for Generative Modeling

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.945415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.945415Z digest=sha256:63296cf061804e202e95cf7d3cddd5d9e5883eafddb26842898497f30b51ae61

Observation bddc6b1e-a04d-4b05-a33e-ea5dabcbdd7d · outbound

This paper cites Audi- oLDM: Text-to-audio generation with latent diffusion mod- els.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Audi- oLDM: Text-to-audio generation with latent diffusion mod- els

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.776837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.950472Z digest=sha256:eb8beb3f15a93b05ec79499d321feba8a8b6efbca2b26e52db41433f90fc1860

Observation 78e9addd-a28d-476c-9917-4230865d43b7 · outbound

This paper cites FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio Generation.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.955420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.955420Z digest=sha256:f3a54cce93c8237da92ae63b38532cd7969dc4714ba4253f00a02da46f2b618e

Observation a8892f56-768c-46bc-bfca-c8d4807d0c24 · outbound

This paper cites Plumbley.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Plumbley

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.759012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.960680Z digest=sha256:91c981ba80846040b82f10c54adff15c6da3e92b1697ecd540115e6d38c2bc6f

Observation 8fcadfb4-164f-4c07-b2e9-6d9c8ec496b8 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.965632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.965632Z digest=sha256:dbf15da10cb466301b13aeea5f57050c9e13dbfc90cbaf4f58ab0065ca145bd8

Observation 78d0d2b7-7e5e-410d-935a-c9cac5453cd8 · outbound

This paper cites Separate Anything You Describe.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Separate Anything You Describe

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.970707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.970707Z digest=sha256:2cb4986623ad205a97599bc30c2cd0ba6c1266ec0c49df66c305f3de53835418

Observation 8d314d68-d31e-46f1-b28a-60ad9c7dd7e9 · outbound

This paper cites Diff-foley: Synchronized video-to-audio synthesis with la- tent diffusion models, 2023.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Diff-foley: Synchronized video-to-audio synthesis with la- tent diffusion models, 2023

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.743876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.975676Z digest=sha256:b81b08bacc2e0e04022b9965c5ce8d623b12d05cd6c3edb17fd6dea8d951ea38

Observation 19bb9429-a598-4d77-8b51-38262cfc3520 · outbound

This paper cites Albergo, Nicholas M.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Albergo, Nicholas M

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.727746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.980418Z digest=sha256:5f62743fd3e667545fa163e714f318f4ecd42d87e65cd15701e5f8fb1c67dfe5

Observation 3c02fe20-b9e2-4a7f-9b7a-4668f091d418 · outbound

This paper cites Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization, 2024.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.710578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.985150Z digest=sha256:388f11146b0c297e9943642ff6eaf00df19679fa94b1dc9130adabbb13a4c7b2

Observation d496671b-d7de-463a-87d5-87e22465082e · outbound

This paper cites SampleRNN: An Unconditional End-to-End Neural Audio Generation Model.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation SampleRNN: An Unconditional End-to-End Neural Audio Generation Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.989949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.989949Z digest=sha256:531857b5072cc3800414fd13dd1fa6e4338eb17b9248e9e3c240bd117b7cc9c0

Observation b6b1c667-3311-4ac6-be88-5c8557f19490 · outbound

This paper cites Text-to-Audio Generation Synchronized with Videos.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Text-to-Audio Generation Synchronized with Videos

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.994918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.994918Z digest=sha256:fedfebc3fc474ee4e6a6abd3655d685c81ce52dcfe9f2f4f15d58e0aa85af6cd

Observation c193d648-be82-4cc0-9ed6-b7975fab2e0d · outbound

This paper cites Bal- ancing act: Distribution-guided debiasing in diffusion mod- els.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Bal- ancing act: Distribution-guided debiasing in diffusion mod- els

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.695014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.000235Z digest=sha256:cc1db60fcfbec669338caa7a3965712b7fef1f4181d4e83f4adfa450d60dc9d6

Observation 060323ee-af08-4265-8a17-867538b01346 · outbound

This paper cites Scalable diffusion models with transformers, 2023.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Scalable diffusion models with transformers, 2023

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.679620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.004938Z digest=sha256:892a882161916e46f044ec3737588ad78f4830d8d2d9cafc3aad232346976606

Observation 341338be-00a1-4f5a-9cce-1df26449df12 · outbound

This paper cites Film: Visual reasoning with a general conditioning layer.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Film: Visual reasoning with a general conditioning layer

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.663487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.009584Z digest=sha256:65e2e5ad75371bf3b1e3aa4dfcb5dbc98aa1af62600bb628370868be1af8085a

Observation 91a4fadc-1435-4907-8ed2-f3cc48242973 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Learning transferable visual models from natural language supervi- sion

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.647834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.014105Z digest=sha256:5d09c082d4bb75c639b1013ed4cf962dfab604c56450bad447a723d7f7f483ec

Observation 3b585338-d1ed-4056-9f83-e666efe31037 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:01.018952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:01.018952Z digest=sha256:f3565268d0e683ced35c735d83f745bff8210238e25335bcce6628d4beb9cd8e

Observation 3fe98423-9360-4269-92b1-00befebb6bbe · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation High-resolution image synthesis with latent diffusion models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:01.024145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:01.024145Z digest=sha256:77e83b0753113a653543eae4bcd3069aa660ec88017bfead4d3c3d6c0ee0c71f

Observation 757ba374-dd75-4de7-a979-d360d2eb4d65 · outbound

This paper cites I hear your true colors: Image guided audio generation, 2022.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation I hear your true colors: Image guided audio generation, 2022

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.608310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.029283Z digest=sha256:ba52ccaf5dad6c13081cd98aec141ac37d597ed2f822f4f2ccf2d58e7d94d50d

Observation 42c483e9-14f8-407e-bf7b-ef23dc79e883 · outbound

This paper cites Auto- acd: A large-scale dataset for audio-language representation learning.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Auto- acd: A large-scale dataset for audio-language representation learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.591925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.034089Z digest=sha256:622e4cb81eb41622ae8d51dbf6730a6347da485675ad56ed32f3d5146d8aa170

Observation b5e2a8e2-082f-4ce4-b38f-11c8bb590245 · outbound

This paper cites Learning from Between-class Examples for Deep Sound Recognition.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Learning from Between-class Examples for Deep Sound Recognition

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:01.039003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:01.039003Z digest=sha256:9d858a14ad39961803f43c75ffcb2e0b2675c3d8c2e7b68456d62f9f455ea026

Observation 9159ca6c-bf3d-4acf-b2aa-ae3f5fb14cb0 · outbound

This paper cites Attention is all you need.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Attention is all you need

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:01.044268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:01.044268Z digest=sha256:1ff884f256dd8e1bd9f6164f7b421380fac523ea137d55ed35fa4d9c5dffd8d4

Observation dff84b15-706b-4178-a31c-c2e78cff4f93 · outbound

This paper cites Audiobox: Unified audio generation with natural language prompts, 2023.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Audiobox: Unified audio generation with natural language prompts, 2023

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.562666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.049825Z digest=sha256:acadc68e3038c763e597ba729679d8a110fb61ace3ec11cc4b5254978bfcd1e1

Observation 44a95349-babf-4205-b7ce-e6d28e03dc31 · outbound

This paper cites V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foun- dation models.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foun- dation models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.545952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.054754Z digest=sha256:4085cac9fb9958680bc6625f72d297abe39b75817a47cca1be37fcd6f8b94637

Observation 5a5265e9-f206-4211-991e-1848bed6af07 · outbound

This paper cites ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:01.059533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:01.059533Z digest=sha256:65462c67b3e694abca230c0d9ad4f7afb783535716fa8b3206e4c59bbf741d8d

Observation 559861e4-be63-48f3-a037-448bc3afc367 · outbound

This paper cites Wav2clip: Learning robust audio repre- sentations from clip.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Wav2clip: Learning robust audio repre- sentations from clip

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.529275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.065138Z digest=sha256:6ce691b0193693abd29ff1d96e36362360dcd10feec0129807197947df5be828

Observation e8c49149-27bc-4322-8696-04682191ead8 · outbound

This paper cites Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.512058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.069941Z digest=sha256:3d82d69091a8ef8efcc113c2e3ad026d1dffb619f79e6dcbee7c07311d7fc0bc

Observation d7192d4f-08a4-4c10-b83c-894b0be99496 · outbound

This paper cites Son- icvisionlm: Playing sound with vision language models.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Son- icvisionlm: Playing sound with vision language models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.494889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.074951Z digest=sha256:ef5a071a36b0a06ea52f1dfe6edaa9a38b868f396bc200ce944f30e0daf9d0f7

Observation 436724e7-41b5-4930-a16c-78d0a5a465e6 · outbound

This paper cites Seeing and hearing: Open-domain visual- audio generation with diffusion latent aligners.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Seeing and hearing: Open-domain visual- audio generation with diffusion latent aligners

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:01.079975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:01.079975Z digest=sha256:23ad63aafc37d39f7700fa609b7563232ac8766f46ff33a0dfb6de2f2ab4c1aa

Observation 6b8802d4-dd19-427d-b44a-c1a00998902c · outbound

This paper cites Diffsound: Discrete diffusion model for text-to-sound generation.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Diffsound: Discrete diffusion model for text-to-sound generation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.466318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.085034Z digest=sha256:f7bc02cc6d0110a70db7528987d02b5d20f1a24ff1803fa217646e904cc751f4

Observation 1403e556-2132-4faa-b35f-1318ecaba71c · outbound

This paper cites Diverse and aligned audio-to-video genera- tion via text-to-video model adaptation.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Diverse and aligned audio-to-video genera- tion via text-to-video model adaptation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.450177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.090156Z digest=sha256:f79ca42519af6ef1288c4f48c89173f5dc9e5502f33de4d344acfc8a0fca8fe9

Observation 079eb143-1244-4a61-9e47-b3b193f10348 · outbound

This paper cites Foley- crafter: Bring silent videos to life with lifelike and synchro- nized sounds, 2024.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Foley- crafter: Bring silent videos to life with lifelike and synchro- nized sounds, 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.434140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.095147Z digest=sha256:6519d40c5e88eeae8fa9cea1b5d988045d1b01a46eb8b331a99246720b49be96

Observation 8bafe478-14c6-4828-83f8-99f37de51dc7 · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Llava- next: A strong zero-shot video understanding model, 2024

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.417921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.099877Z digest=sha256:82f50888a31039b9af8546f73277510ac449b2806fd8d35a1c2663bb5efaad6e

Observation c84ec568-408e-435f-8846-805b45ad41f6 · outbound

This paper cites an unresolved cited work.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:43:01.402261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.104419Z digest=sha256:af895deb61b33ef8dcffebf3ae7570b82b0d198c919404ee9687fa488c95a6c6

Observation 10fc6dca-0337-4a2c-b5dd-4542c2b61058 · outbound

This paper cites Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:01.110656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:01.110656Z digest=sha256:c330296a72ceb79ddb81d55e787b834ba682849ba209a9cfcde4e912c1bc3598

Pith citing papers

Observation 52d52cb8-48a5-4be9-989e-07211047f4bf · inbound

AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation cites this paper.

AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:08.339998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:08.339998Z digest=sha256:c752700675794768344810e85fe609e1134313473c213b5d7fb6c2d7114232ab

Observation eaa0d4c0-e44a-4942-bc74-9a9d04cb9093 · inbound

Sound Scene Synthesis at the DCASE 2024 Challenge cites this paper.

Sound Scene Synthesis at the DCASE 2024 Challenge VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:27:03.921470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:27:03.773063Z digest=sha256:c9602d92a1f894a738abc65cd2799f775294eaf6b501cb5de30c0ce08f424adb