Pith. sign in

Paper Citation Record · LEDGER

Video-Guided Foley Sound Generation with Multimodal Controls

As of 17 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 7 inbound Pith citation observations for arXiv:2411.17698.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17698 v4

Coverage vector

measured 100 of 105 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:00:01.822180Z

measured 107 of 107 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:43:11.816985Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T13:28:18.784433Z

Reference resolution

100 of 105 outbound references displayed

  • verified exact1
  • verified fuzzy23
  • unresolved76
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2457b27e-5a00-4393-a477-a81d7049c7c0 · outbound

This paper cites The Foley grail: The art of perform- ing sound for film, games, and animation.

Video-Guided Foley Sound Generation with Multimodal Controls The Foley grail: The art of perform- ing sound for film, games, and animation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:00.730927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:00.730927Z digest=sha256:94928937f61cb41df9ba03f3881607ebab85da221566e00c1940fd1346991db4

Observation 6399721d-6540-4938-9beb-33c895b73929 · outbound

This paper cites SegDiff: Image Segmentation with Diffusion Probabilistic Models.

Video-Guided Foley Sound Generation with Multimodal Controls SegDiff: Image Segmentation with Diffusion Probabilistic Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:00.752898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:00.752898Z digest=sha256:6fba458ac2fab7a6745876cb0da1287c74fcce277ba0a6c4c8c78f2833b80f32

Observation 6eff200d-d3cd-48b8-bd78-a8037408c108 · outbound

This paper cites Look, listen and learn.

Video-Guided Foley Sound Generation with Multimodal Controls Look, listen and learn

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:00.765177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:00.765177Z digest=sha256:41b276023e1582345841c3e2403925a404b02ad4f0086644fd40cac9255b1d06

Observation 21142e2c-6c01-40a1-8f3d-c9104b5a3261 · outbound

This paper cites Objects that sound.

Video-Guided Foley Sound Generation with Multimodal Controls Objects that sound

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:00.778529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:00.778529Z digest=sha256:55afc437802e309ec0fae9d5df5c4232e5b2d47159eb6a3ab14e92940351d332

Observation 41c3fa47-e7da-45f2-862f-eebbbdf28c13 · outbound

This paper cites Labelling unlabelled videos from scratch with multi-modal self-supervision.

Video-Guided Foley Sound Generation with Multimodal Controls Labelling unlabelled videos from scratch with multi-modal self-supervision

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:00.789490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:00.789490Z digest=sha256:a1ce8279c2b323d13544a63c628ab15071e000f9a5973de13c4aeba5756da58c

Observation 9aa756f4-7f6d-4be1-bd0e-a2c9ab5c6a5e · outbound

This paper cites MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation.

Video-Guided Foley Sound Generation with Multimodal Controls MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:00.800938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:00.800938Z digest=sha256:dccb79068749535706da3fca4f3834680e125df33fe399eeb238da4dcebaef5a

Observation 865b4c70-96ed-49a1-bc8d-72a70d482094 · outbound

This paper cites Lumiere: A Space-Time Diffusion Model for Video Generation.

Video-Guided Foley Sound Generation with Multimodal Controls Lumiere: A Space-Time Diffusion Model for Video Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:00.812243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:00.812243Z digest=sha256:4d8e87b7235b5fb938424d50b586fe3044adcfac68d32f10bec431f93be536fd

Observation 7837a425-a4b3-4db8-aa80-9b020e5a8138 · outbound

This paper cites Meta 3D Gen.

Video-Guided Foley Sound Generation with Multimodal Controls Meta 3D Gen

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:00.823447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:00.823447Z digest=sha256:fce5108c1a41eec50d0711931cd46f00528a4b8862fe67c103e22ed26e098676

Observation 68f0fb4e-0f3e-47e5-978b-170837a6ee2a · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

Video-Guided Foley Sound Generation with Multimodal Controls In- structpix2pix: Learning to follow image editing instructions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:00.829523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:00.829523Z digest=sha256:db8c49e1f2a08b239cf9cb713cbe1bc0b017fd322a6f0a71726d5f168335a741

Observation 41e91dda-99d6-416b-8597-89c07e10bca5 · outbound

This paper cites Maskgit: Masked generative image transformer.

Video-Guided Foley Sound Generation with Multimodal Controls Maskgit: Masked generative image transformer

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:00.840847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:00.840847Z digest=sha256:1a893a35277255686be94b386a0cedb6c315dcf909b94333c4d2f9cfe10af692

Observation 6f4e0a40-29c0-4081-8e35-f156a83a34c6 · outbound

This paper cites Ac- tion2sound: Ambient-aware generation of action sounds from egocentric videos.

Video-Guided Foley Sound Generation with Multimodal Controls Ac- tion2sound: Ambient-aware generation of action sounds from egocentric videos

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:00.849163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:00.849163Z digest=sha256:e9c9b77de103a1c4d06607b3f8e8b2ac3ef4eddc478c13f2cef192e950f61b9d

Observation 66f64726-4722-4c97-ae03-39ccc738f4c7 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

Video-Guided Foley Sound Generation with Multimodal Controls Vggsound: A large-scale audio-visual dataset

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:00.859848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:00.859848Z digest=sha256:a88a6510cfe7f09ae151a1e7c3b983fe887d79a1019637afcd421d5003f214bc

Observation 5a413fc4-e2bd-47d6-ab20-17c36b237601 · outbound

This paper cites Audio-Visual Synchronisation in the wild.

Video-Guided Foley Sound Generation with Multimodal Controls Audio-Visual Synchronisation in the wild

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:00.867723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:00.867723Z digest=sha256:bbe01018511e1bd85b5dcbc13c73c79bc2b931e733dcd01aa93744e3cdfa23d2

Observation 0aad9924-188b-417a-b6b4-d564c4b3c5ac · outbound

This paper cites Be everywhere- hear everything (bee): Audio scene reconstruction by sparse audio-visual samples.

Video-Guided Foley Sound Generation with Multimodal Controls Be everywhere- hear everything (bee): Audio scene reconstruction by sparse audio-visual samples

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:00.879818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:00.879818Z digest=sha256:9d225603bd2c11d9d16cc90d471d392495dd0bffde83b55f9d44941d8093a9f3

Observation c8777f78-477d-465c-abee-0cb956ba4e40 · outbound

This paper cites Structure from silence: Learning scene structure from ambient sound.

Video-Guided Foley Sound Generation with Multimodal Controls Structure from silence: Learning scene structure from ambient sound

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:00.896577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:00.896577Z digest=sha256:ff00634b31894c0415ac377f2d2be1745d59d33b08c7f0a4af693e1adb3ee386

Observation 0c305d74-b8bc-432b-9141-5c5a7da58ad8 · outbound

This paper cites Sound localization by self-supervised time delay estimation.

Video-Guided Foley Sound Generation with Multimodal Controls Sound localization by self-supervised time delay estimation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:00.906863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:00.906863Z digest=sha256:fbe2bb4afa97a1e1c52c1af64b0451878d63a386a785e75582e72458ad64b7a8

Observation f9d45e0f-e30c-4b2a-97ed-b07f13f2b331 · outbound

This paper cites Sound lo- calization from motion: Jointly learning sound direction and camera rotation.

Video-Guided Foley Sound Generation with Multimodal Controls Sound lo- calization from motion: Jointly learning sound direction and camera rotation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:00.914490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:00.914490Z digest=sha256:717ea695bd71777626ba012f7c1467efe8ed22aa3c02cbf6151a128028cd7f6f

Observation ea3186e6-239b-4983-babf-b85671558d9d · outbound

This paper cites Gebru, Christian Richardt, Anurag Kumar, William Laney, Andrew Owens, and Alexander Richard.

Video-Guided Foley Sound Generation with Multimodal Controls Gebru, Christian Richardt, Anurag Kumar, William Laney, Andrew Owens, and Alexander Richard

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:00.923570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:00.923570Z digest=sha256:26464925c953e9c464093b4bb398b6f0bca88057557f45573a831c0683c3c4f8

Observation f7835a5f-aa77-4744-8226-377641f52728 · outbound

This paper cites Images that sound: Composing images and sounds on a single canvas.

Video-Guided Foley Sound Generation with Multimodal Controls Images that sound: Composing images and sounds on a single canvas

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:00.931271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:00.931271Z digest=sha256:2860af7425be34a38a38f8f16b27c7c64a7f2c11757af75796ca6cbd46d65258

Observation e9f52ae8-f0b4-4f32-a162-6368742719e8 · outbound

This paper cites MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis.

Video-Guided Foley Sound Generation with Multimodal Controls MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:00.939431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:00.939431Z digest=sha256:b69481318f54f370f366cbd7a12bccee260aeac29e9da9c48b708bf1f50355f5

Observation 74b1bc94-6c01-4f4f-9039-68f6820aeeab · outbound

This paper cites T-foley: A controllable waveform-domain diffusion model for temporal- event-guided foley sound synthesis.

Video-Guided Foley Sound Generation with Multimodal Controls T-foley: A controllable waveform-domain diffusion model for temporal- event-guided foley sound synthesis

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:00.951776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:00.951776Z digest=sha256:e016463c8a68eb13d3577fc013d3b94caf179114728d050f35bf9effbe141c22

Observation 0e21fad2-1f2a-43b3-aadb-01f4d8686d53 · outbound

This paper cites Syncfusion: Multimodal onset-synchronized video- to-audio foley synthesis.

Video-Guided Foley Sound Generation with Multimodal Controls Syncfusion: Multimodal onset-synchronized video- to-audio foley synthesis

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:00.961945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:00.961945Z digest=sha256:e871b25d388373e0a59ed8bd7dbee85173925091701d9e46a70cea6ae40125bf

Observation 0c819244-6ae7-49d7-ae35-78231caee072 · outbound

This paper cites Diffusion models beat gans on image synthesis.

Video-Guided Foley Sound Generation with Multimodal Controls Diffusion models beat gans on image synthesis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:00.972648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:00.972648Z digest=sha256:87fd1d5e586e9af2ddd90a75b9017be2853567287e1dd570afa0f75cc013b1cd

Observation a8343c95-aadf-4aaf-9193-587b5ca97103 · outbound

This paper cites Conditional generation of audio from video via foley analogies.

Video-Guided Foley Sound Generation with Multimodal Controls Conditional generation of audio from video via foley analogies

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:00.987190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:00.987190Z digest=sha256:55b79216212054e51294770342c30f99747e2202b06029afed4634570969b2a4

Observation 1daaf8f0-678f-46ba-ae64-c7c36ee83198 · outbound

This paper cites Fast Timing-Conditioned Latent Audio Diffusion.

Video-Guided Foley Sound Generation with Multimodal Controls Fast Timing-Conditioned Latent Audio Diffusion

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:00.994912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:00.994912Z digest=sha256:be239e370ea4262cfaca2b19d5f9cfc3f9516a47244fb7556404af7da03911ce

Observation 2e90d6e4-672c-40f7-9697-8dbad70dd4fd · outbound

This paper cites Stable Audio Open.

Video-Guided Foley Sound Generation with Multimodal Controls Stable Audio Open

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.001957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.001957Z digest=sha256:a62e3052cf2659ccedfbaa2ee152fec54bdec363272ccf1fefe001e15cad969a

Observation de4f8942-43be-4710-9cd5-508af27863c9 · outbound

This paper cites Self- supervised video forensics by audio-visual anomaly detec- tion.

Video-Guided Foley Sound Generation with Multimodal Controls Self- supervised video forensics by audio-visual anomaly detec- tion

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.010744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.010744Z digest=sha256:e6c0d5e1ee4bec507200871d91233e3593c366867c0bd78ea8f050b87ba87a32

Observation fbffc427-d301-42de-bb6b-af287ff07fc4 · outbound

This paper cites Visualechoes: Spatial visual rep- resentation learning through echolocation.

Video-Guided Foley Sound Generation with Multimodal Controls Visualechoes: Spatial visual rep- resentation learning through echolocation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.017871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.017871Z digest=sha256:33793717da15c14aa69a38b76a13b43fd32da6d180a8fd8c39b9ca819ee91bc7

Observation c64f2dcd-6e0b-43d4-8aae-b5b26ed4d687 · outbound

This paper cites VampNet: Music Generation via Masked Acoustic Token Modeling.

Video-Guided Foley Sound Generation with Multimodal Controls VampNet: Music Generation via Masked Acoustic Token Modeling

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.025959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.025959Z digest=sha256:de9efeb9b06db66839508d8f53a9a0f9793a542187e17a1bc8f5e3c184cc9227

Observation 7781c1bb-edbf-48e1-8301-b8e794bf0369 · outbound

This paper cites Motion Guidance: Diffusion-Based Image Editing with Differentiable Motion Estimators.

Video-Guided Foley Sound Generation with Multimodal Controls Motion Guidance: Diffusion-Based Image Editing with Differentiable Motion Estimators

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.036572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.036572Z digest=sha256:fe1485e3512054ad6659cf5dcf6d2da43bd68daa0918ed9495faf73afb76dc80

Observation a447525d-0437-4ac3-9834-07ded9146db6 · outbound

This paper cites Visual anagrams: Generating multi-view optical illusions with dif- fusion models.

Video-Guided Foley Sound Generation with Multimodal Controls Visual anagrams: Generating multi-view optical illusions with dif- fusion models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.053177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.053177Z digest=sha256:866ce9b266256279c41ca2f5c7acb1ddcfe089c1e8bc107ff4583de7d6b4de37

Observation 53e5c949-1c0b-48da-ac29-25c02b87c15b · outbound

This paper cites Factorized Diffusion: Perceptual Illusions by Noise Decomposition.

Video-Guided Foley Sound Generation with Multimodal Controls Factorized Diffusion: Perceptual Illusions by Noise Decomposition

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.066856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.066856Z digest=sha256:3c17d447b4215bc538ba1d7783985d76ab129608518cd98ececd6a2e3d7a8bca

Observation 3e68ac05-b2a1-40e8-a640-5c209a62b089 · outbound

This paper cites Imagebind: One embedding space to bind them all.

Video-Guided Foley Sound Generation with Multimodal Controls Imagebind: One embedding space to bind them all

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.074945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.074945Z digest=sha256:f9423f8db6b686b31493d789aaea0518da45036511153f0475fe5f7a30623116

Observation e8d0ba44-3b67-4b21-a455-a9981a39633c · outbound

This paper cites Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning.

Video-Guided Foley Sound Generation with Multimodal Controls Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.084536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.084536Z digest=sha256:4f3f65c54030811d0e4b225b7ef97f8996d3c3c631088d5f69a4db73ed33435d

Observation e6ea8779-1e42-4d65-9ab3-cec559bdadf3 · outbound

This paper cites Cnn ar- chitectures for large-scale audio classification.

Video-Guided Foley Sound Generation with Multimodal Controls Cnn ar- chitectures for large-scale audio classification

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.092454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.092454Z digest=sha256:7022f21c6b69e75b1ba5dd78045d0a3dccbe312db0ee3c0bbdc7cfc0e3e7fd4b

Observation 034b3616-9135-4f56-972f-a41cebaa5375 · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

Video-Guided Foley Sound Generation with Multimodal Controls Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.109333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.109333Z digest=sha256:9fe2a51f6d98ed7a024e3d02f667782b682f52260d50d1f3b451dc2271134e28

Observation dedb1bc9-2234-4311-b19b-d16cc0c04f2d · outbound

This paper cites Classifier-Free Diffusion Guidance.

Video-Guided Foley Sound Generation with Multimodal Controls Classifier-Free Diffusion Guidance

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.116744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.116744Z digest=sha256:3c2728d119a6ee4c5a772adcad8e97ecd4b5278ec4092a2a171583db511d032b

Observation c6a38678-5100-431e-9abb-30823c122d1a · outbound

This paper cites Denoising diffu- sion probabilistic models.

Video-Guided Foley Sound Generation with Multimodal Controls Denoising diffu- sion probabilistic models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.127537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.127537Z digest=sha256:9fc5e6f1f7133928e5831ffe90ff60c2f87dbe0576cd9a6159b4c4e153f86fdf

Observation 627fea7e-fbaa-49b0-907c-c852a72483dc · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

Video-Guided Foley Sound Generation with Multimodal Controls Imagen Video: High Definition Video Generation with Diffusion Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.135291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.135291Z digest=sha256:68d9573d4f9988f2c3374846ac4e7de1c97a6b19261646d2e61367e5039c1287

Observation 07159305-f677-4e1a-8447-b58c91d1eb38 · outbound

This paper cites Text2room: Extracting textured 3d meshes from 2d text-to-image models.

Video-Guided Foley Sound Generation with Multimodal Controls Text2room: Extracting textured 3d meshes from 2d text-to-image models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.147208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.147208Z digest=sha256:7479dbe82e66f6ed470e0ffb9e2cf8d163839ad937c8b50f78aae34f97114563

Observation 9cacad05-223d-434a-9ba0-9d512b25943d · outbound

This paper cites Mix and local- ize: Localizing sound sources in mixtures.

Video-Guided Foley Sound Generation with Multimodal Controls Mix and local- ize: Localizing sound sources in mixtures

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:00:05.511326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:00:01.155052Z digest=sha256:c7329220f091809c635e6530cb9ff5eda8faf55b5f316ab98cced1d5fddc6a1e

Observation ecdaeefd-2578-4082-98f9-8fa6c3106483 · outbound

This paper cites Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis.

Video-Guided Foley Sound Generation with Multimodal Controls Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.171144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.171144Z digest=sha256:72ede557fa4c1105ed780bd70cfdd7a20489e2a2143e721892b2d6e6d186d7af

Observation 3261e9c0-5f93-4d01-bbeb-342fe2014754 · outbound

This paper cites Taming visually guided sound generation.

Video-Guided Foley Sound Generation with Multimodal Controls Taming visually guided sound generation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:00:05.472309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:00:01.182798Z digest=sha256:22e5f94e42a5cb72b8e98ec5b5ae43aa8ac59480425778e5f1978fe4f5a8d85f

Observation 5baf00bf-e453-4d87-8ff4-49d0abc9b4a1 · outbound

This paper cites Synchformer: Efficient Synchronization from Sparse Cues.

Video-Guided Foley Sound Generation with Multimodal Controls Synchformer: Efficient Synchronization from Sparse Cues

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.193205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.193205Z digest=sha256:1b2693b0a83eb455fe06a7dedb070ae8920fc8d1c0daa15abadf0b608e815418

Observation 3b381f57-861f-4979-aab6-ab0f04ece6e4 · outbound

This paper cites Read, Watch and Scream! Sound Generation from Text and Video.

Video-Guided Foley Sound Generation with Multimodal Controls Read, Watch and Scream! Sound Generation from Text and Video

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.207428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.207428Z digest=sha256:35f4b86d7b1d24e9575403a7d6e804997420cdd7ec2e1fc845b84b91ec95d7b2

Observation cc4ec938-f565-4686-92e4-e54e4c869aba · outbound

This paper cites Repurpos- ing diffusion-based image generators for monocular depth estimation.

Video-Guided Foley Sound Generation with Multimodal Controls Repurpos- ing diffusion-based image generators for monocular depth estimation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.221295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.221295Z digest=sha256:4aa9bd712873e7b8417cf51139eb594a9fa24a7b50dacb533a8dfaa7e4ffc518

Observation 3bc047e1-feb8-4058-b5f2-bd19d317fb30 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

Video-Guided Foley Sound Generation with Multimodal Controls Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.230914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.230914Z digest=sha256:2bfb35f45dd5d69a3764c617349d98a7f61df41241924f8a51c95fdbb5cf2aa5

Observation afa46d4b-b38e-48a7-99ca-ad9babc012d0 · outbound

This paper cites Adam: A method for stochastic optimization.

Video-Guided Foley Sound Generation with Multimodal Controls Adam: A method for stochastic optimization

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:00:05.391835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:00:01.260218Z digest=sha256:f61ff8c46de26cdc9f76ba4da492e066dff059d1eb47a86ff7aed5256bf0f19b

Observation a63dae68-5789-42d6-bc30-ef0df350b60f · outbound

This paper cites Auto-Encoding Variational Bayes.

Video-Guided Foley Sound Generation with Multimodal Controls Auto-Encoding Variational Bayes

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.270750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.270750Z digest=sha256:bc6f46afb47a43870b7131232588d42f95c25186d93009bc655e7bf7eed7d96f

Observation 0beb49cb-7cd4-4df9-ac30-973738e6b204 · outbound

This paper cites Efficient Training of Audio Transformers with Patchout.

Video-Guided Foley Sound Generation with Multimodal Controls Efficient Training of Audio Transformers with Patchout

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.282601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.282601Z digest=sha256:e601a71e415fb0e454480a8e08edc218eb9e7f488c9af8d2a8d24b5ba578b7c8

Observation fb4d3e8e-7735-420e-a8c5-3e214c70d075 · outbound

This paper cites On information and sufficiency.

Video-Guided Foley Sound Generation with Multimodal Controls On information and sufficiency

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:00:05.359803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:00:01.297076Z digest=sha256:280f921bebeca5e17f70fb2603d1d389e1b64d487b675914dc8a4094a39e6d9f

Observation 14416ecd-43cc-4ded-ab6c-c1f004f716ea · outbound

This paper cites High-fidelity audio compres- sion with improved rvqgan.

Video-Guided Foley Sound Generation with Multimodal Controls High-fidelity audio compres- sion with improved rvqgan

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:00:05.311640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:00:01.311941Z digest=sha256:2111884255f56a2b082c9887259ab83f1575e0bb533929dc780355a57d99e31c

Observation 0c2b5cc3-476d-4e17-a9df-663e814145d6 · outbound

This paper cites Video-foley: Two-stage video-to-sound generation via tem- poral event condition for foley sound.

Video-Guided Foley Sound Generation with Multimodal Controls Video-foley: Two-stage video-to-sound generation via tem- poral event condition for foley sound

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.319404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.319404Z digest=sha256:8b62246315baf783c6e7d51da29a89abbb02a9c77ee552b2bd0abfa9c3f66fc2

Observation 335f7ab1-be73-4536-8409-ee79b62ccac3 · outbound

This paper cites Self-supervised audio-visual soundscape stylization.

Video-Guided Foley Sound Generation with Multimodal Controls Self-supervised audio-visual soundscape stylization

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:00:05.258203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:00:01.328283Z digest=sha256:0f838704d6b1663d48a6cd17a7b1473ce90883dcf3c80696a0b8e8d6a25bbec3

Observation eb4a188d-6522-45eb-872d-06d2461b1682 · outbound

This paper cites Siamese Vision Transformers are Scalable Audio-visual Learners.

Video-Guided Foley Sound Generation with Multimodal Controls Siamese Vision Transformers are Scalable Audio-visual Learners

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.333545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.333545Z digest=sha256:de485db6fe299528f6fca3a7fbd819393f12901d9e38a0c5eed00be09d9b6409

Observation 9b6b7a8f-8a7a-4a41-8a6f-878ddf0d54a4 · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

Video-Guided Foley Sound Generation with Multimodal Controls AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.363135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.363135Z digest=sha256:3f1ea5a370dd3b182728a2d32aa5f1d06b3f8f6fd4029066d9ac38c98f181734

Observation 578cee9a-84fe-44de-b88b-b7e6035cb438 · outbound

This paper cites AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining.

Video-Guided Foley Sound Generation with Multimodal Controls AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.380274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.380274Z digest=sha256:24aa027f95e34eedba0ec0cd60ac5ed4037187bc36892e8d53905db8db3b029b

Observation 5cdd5ec3-9a3a-4cf4-9ab6-06a1a0877a9d · outbound

This paper cites Audio-visual segmentation via unlabeled frame exploitation.

Video-Guided Foley Sound Generation with Multimodal Controls Audio-visual segmentation via unlabeled frame exploitation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:00:05.210566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:00:01.392005Z digest=sha256:2e72e246a9a8707f940e5e16d5034d59d0cf0fea285c18496a6ef2a3bc3af5e7

Observation 60550093-2dce-440f-aca4-f8b9325323f0 · outbound

This paper cites Compositional visual generation with composable diffusion models.

Video-Guided Foley Sound Generation with Multimodal Controls Compositional visual generation with composable diffusion models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.407284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.407284Z digest=sha256:6e142ee763a78fe29f20bfe437b37a838a21188d1af2445da921216ae075683c

Observation 1853b062-636d-4e9b-b9fe-416459dcf4e7 · outbound

This paper cites Zero-1-to- 3: Zero-shot one image to 3d object.

Video-Guided Foley Sound Generation with Multimodal Controls Zero-1-to- 3: Zero-shot one image to 3d object

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.421120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.421120Z digest=sha256:17e846c90228f84aaf96f0a988340520bdb61b3ecb72cfaea652f97b6db67cdd

Observation d76e0fd0-569f-4dc4-9655-53f2635f8e89 · outbound

This paper cites Tell what you hear from what you see – video to audio generation through text.

Video-Guided Foley Sound Generation with Multimodal Controls Tell what you hear from what you see – video to audio generation through text

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:00:05.112463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:00:01.430629Z digest=sha256:aa1fd937def80fc50df46554ccab709a82cabd6dde71101c0b8446163d396c7c

Observation acbd7d65-5a16-4ad4-945e-4b0340b6f7ca · outbound

This paper cites Decoupled Weight Decay Regularization.

Video-Guided Foley Sound Generation with Multimodal Controls Decoupled Weight Decay Regularization

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.437352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.437352Z digest=sha256:99ed58cfc9041f37d664062e50dbaa96df06ec66f54196687c3ce17b3f7db36d

Observation ee902bcb-93e0-4ed7-bcd7-1245476532a7 · outbound

This paper cites Diff-foley: Synchronized video-to-audio synthesis with la- tent diffusion models.

Video-Guided Foley Sound Generation with Multimodal Controls Diff-foley: Synchronized video-to-audio synthesis with la- tent diffusion models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:00:05.070502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:00:01.443761Z digest=sha256:f6e828289ba42a51d3a8cc637ffb5fa305ed43c97b9b4f1c2612818d209c03a0

Observation 96b606ba-603e-47e8-81f1-a3df3daac4a7 · outbound

This paper cites Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization.

Video-Guided Foley Sound Generation with Multimodal Controls Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.449443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.449443Z digest=sha256:28857a7eafd05de0dda7dfb59f279531bc8a806d47d87e962a68a01b55ba8d04

Observation 51655e22-5f83-4cf4-9502-951f94ea95f7 · outbound

This paper cites Learning Spatial Features from Audio-Visual Correspondence in Egocentric Videos.

Video-Guided Foley Sound Generation with Multimodal Controls Learning Spatial Features from Audio-Visual Correspondence in Egocentric Videos

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-08-12T12:00:02.722549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:00:01.462978Z digest=sha256:557048ef3464e159cac684e6ba35e614dc893a9b2e2b229cf3ce59623d48ef8e

Observation 7c3068cf-5269-4cb4-977b-ecac4aa4a979 · outbound

This paper cites FoleyGen: Visually-Guided Audio Generation.

Video-Guided Foley Sound Generation with Multimodal Controls FoleyGen: Visually-Guided Audio Generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.471769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.471769Z digest=sha256:22bb02813dcbfb6deb08f0acbdab9ece48339a229e47e9a9a2fe46384abd72e6

Observation f894264a-9b14-40e4-b8ef-e04132987274 · outbound

This paper cites SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations.

Video-Guided Foley Sound Generation with Multimodal Controls SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.477920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.477920Z digest=sha256:d2149e794f4f2205530653a1c58a612b70d9dff824609a688e831012d4a072d9

Observation 8fa936ad-ab39-45fc-bffa-b99ff18775e1 · outbound

This paper cites Exponential moving average of weights in deep learning: Dynamics and benefits.

Video-Guided Foley Sound Generation with Multimodal Controls Exponential moving average of weights in deep learning: Dynamics and benefits

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:00:05.016194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:00:01.489629Z digest=sha256:f2a3364165668e72f444199bcc21ad0dc6e4b20ebadf3dfc68452fd4f8157f00

Observation 762e7bc2-1d65-421f-b83e-e1682b77948d · outbound

This paper cites Audio- visual instance discrimination with cross-modal agreement.

Video-Guided Foley Sound Generation with Multimodal Controls Audio- visual instance discrimination with cross-modal agreement

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:00:04.978471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:00:01.499680Z digest=sha256:b3df60bd2b60e57515f9f64bd4828d85eb291551ea54df75c9d39d2d8fdfa821

Observation 74f97cd6-ef80-4bfa-8742-568379e4347f · outbound

This paper cites Glide: Towards photorealistic image generation and editing with text-guided diffusion models, 2021.

Video-Guided Foley Sound Generation with Multimodal Controls Glide: Towards photorealistic image generation and editing with text-guided diffusion models, 2021

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:00:04.935300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:00:01.508338Z digest=sha256:0abd7005086d6fbfa9d2e6e9a01aa3f5f99730df4af7e4c68c02051f67ed6ff9

Observation 25b46985-c3c8-4696-b4f1-127740650adf · outbound

This paper cites Audio-visual glance network for efficient video recognition.

Video-Guided Foley Sound Generation with Multimodal Controls Audio-visual glance network for efficient video recognition

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:00:04.899916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:00:01.516612Z digest=sha256:62bdc1c9a23b215fa3d91d400706ff2161647fe4baae465143b6b7464a162f06

Observation a8925498-d880-4840-a9a3-d239e8af9124 · outbound

This paper cites Audio-visual scene analysis with self-supervised multisensory features.

Video-Guided Foley Sound Generation with Multimodal Controls Audio-visual scene analysis with self-supervised multisensory features

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:00:04.861646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:00:01.527275Z digest=sha256:222590a40e25e4d62cc405823652c7fe65558a5e0a8141716556a967efd789a6

Observation 083287b6-00c5-4469-bcfc-938af0354143 · outbound

This paper cites Visually indicated sounds.

Video-Guided Foley Sound Generation with Multimodal Controls Visually indicated sounds

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.539494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.539494Z digest=sha256:9dfb263c420786a15b37eed92895e90919f525b627363e0331b8f63867ca5154

Observation 048b05b2-4652-404d-a99f-235a8e1684e0 · outbound

This paper cites Can clip help sound source localization? In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 5711–5720, 2024.

Video-Guided Foley Sound Generation with Multimodal Controls Can clip help sound source localization? In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 5711–5720, 2024

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:00:04.807400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:00:01.551749Z digest=sha256:c756e9be65f9e2dc56cd245cad7abe460744d560a34ea4cd563605b2971b0c22

Observation 2bddf225-85ce-45be-837b-6fb7ba14a963 · outbound

This paper cites Masked Generative Video-to-Audio Transformers with Enhanced Synchronicity.

Video-Guided Foley Sound Generation with Multimodal Controls Masked Generative Video-to-Audio Transformers with Enhanced Synchronicity

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.559306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.559306Z digest=sha256:8af96ca42da7a2486631ea29750e9342673d1f354fbcffee531451c9c4c346ac

Observation d91db69b-2f1a-4036-87ea-e2c18e86efe8 · outbound

This paper cites Scalable diffusion models with transformers.

Video-Guided Foley Sound Generation with Multimodal Controls Scalable diffusion models with transformers

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:00:04.758825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:00:01.566933Z digest=sha256:757cad453363e591343f032c7fa16369229b91d930bc705733e4971c9905a581

Observation 909eddfa-e886-4208-ba21-9b32558460cb · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Video-Guided Foley Sound Generation with Multimodal Controls Movie Gen: A Cast of Media Foundation Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.579080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.579080Z digest=sha256:0bffb0225ceb097f45438e38a7eff51ba76f0bcf887cef1bf38b920dc484d170

Observation d03a2af5-319d-456c-b2fc-2bb163a20aee · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Video-Guided Foley Sound Generation with Multimodal Controls Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.590757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.590757Z digest=sha256:6423ea1d0819853ea8233e79b3f10c31c66817d481a0105f45bbea14f7c3d62e

Observation 27c7bed2-3bab-4696-8b88-13bdaa5e3b64 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Video-Guided Foley Sound Generation with Multimodal Controls High-resolution image synthesis with latent diffusion models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.599919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.599919Z digest=sha256:5fb91e2cab84f6e2101a3b930c2c5b3f52ce70adf0657e7edb7f3ad9e53e581b

Observation 2bf05cea-294a-467f-a113-1b90eea6ab09 · outbound

This paper cites Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J Fleet, and Mohammad Norouzi.

Video-Guided Foley Sound Generation with Multimodal Controls Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J Fleet, and Mohammad Norouzi

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.612856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.612856Z digest=sha256:20cd1706c6b8b080e405f4894c4877d3fe60907e82b42a181adedda54933878b

Observation 11149f29-8619-4d45-96a6-f7562c4cc6f8 · outbound

This paper cites Monocular Depth Estimation using Diffusion Models.

Video-Guided Foley Sound Generation with Multimodal Controls Monocular Depth Estimation using Diffusion Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.621085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.621085Z digest=sha256:a90905a1d63a248289e72c058f43af59826dc9279bcc963b7c1773aacd6693b5

Observation 06099ec4-7b50-40e6-af08-8de45ff2e717 · outbound

This paper cites I hear your true colors: Im- age guided audio generation.

Video-Guided Foley Sound Generation with Multimodal Controls I hear your true colors: Im- age guided audio generation

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:00:04.618357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:00:01.633886Z digest=sha256:34805c3b9f829fa318aa67d7b64cb340097fbdfdfd70032bb64de1523b8c3ad1

Observation f71f6a1c-3823-4af5-bbbb-a0387ab6afa7 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

Video-Guided Foley Sound Generation with Multimodal Controls Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.644779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.644779Z digest=sha256:ef532e569d47736ea9b48b9127ed743a1cc23536190639f143ba22f19378515d

Observation 58c2f75c-532e-4609-b2ff-f28109cc2b8f · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

Video-Guided Foley Sound Generation with Multimodal Controls Deep unsupervised learning using nonequilibrium thermodynamics

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.652691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.652691Z digest=sha256:545d0c466b1f4a407392fef0ab86f75d9f0490ee01bd668a64617d2f7dcc165d

Observation 305745b6-f73d-41d3-b916-693346e4d417 · outbound

This paper cites Denoising Diffusion Implicit Models.

Video-Guided Foley Sound Generation with Multimodal Controls Denoising Diffusion Implicit Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.667769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.667769Z digest=sha256:e3bb228467c123eb9be2d50b92c85b4ce8adaddb49d9875f287f20615b976f19

Observation abf455fe-aece-4d7b-9f44-67400bae8176 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

Video-Guided Foley Sound Generation with Multimodal Controls Score-Based Generative Modeling through Stochastic Differential Equations

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.675667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.675667Z digest=sha256:07f2f6beef0f2b9681894a9134446fd31963a9df49215731a2f959e2cfb4f436

Observation 2005c276-6ce6-41bc-8e85-e498bde233e3 · outbound

This paper cites From vision to au- dio and beyond: A unified model for audio-visual representa- tion and generation.

Video-Guided Foley Sound Generation with Multimodal Controls From vision to au- dio and beyond: A unified model for audio-visual representa- tion and generation

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:00:04.536768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:00:01.682711Z digest=sha256:cbeabfde6e8c9a3f538367793966a77e2c25a5f123552a5c494e54b888144422

Observation e6f0771f-aaf4-46bf-9110-9800c9861ec0 · outbound

This paper cites Eventfulness for interactive video alignment.

Video-Guided Foley Sound Generation with Multimodal Controls Eventfulness for interactive video alignment

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:00:04.496559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:00:01.690605Z digest=sha256:a08a98a612b5dbc127d46ed64ed8eff2d54ccebc4e6665b8367d3e1fa0df0346

Observation e6e8909a-634d-486e-bc02-f373299d398c · outbound

This paper cites Temporally Aligned Audio for Video with Autoregression.

Video-Guided Foley Sound Generation with Multimodal Controls Temporally Aligned Audio for Video with Autoregression

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.699988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.699988Z digest=sha256:27d63e5c954a5be801f8ac330295e1d604afc865fc8f838ca4fb2717f64c8e75

Observation 6f929645-ec96-4214-b62b-3bde401a7794 · outbound

This paper cites V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foun- dation models.

Video-Guided Foley Sound Generation with Multimodal Controls V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foun- dation models

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:00:04.451699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:00:01.716511Z digest=sha256:d803e4f779d4b20a85a6f367e1c0e7e578c04e901c7edd2fb5f09d0f58832015

Observation c5169d54-d480-4e57-ad90-a8a88d14bcf6 · outbound

This paper cites Posediffusion: Solving pose estimation via diffusion-aided bundle adjustment.

Video-Guided Foley Sound Generation with Multimodal Controls Posediffusion: Solving pose estimation via diffusion-aided bundle adjustment

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.724961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.724961Z digest=sha256:76d838b28b3fe27fc85dedaf584ea9c6eadf75ef12546dedfe249cca5e116af5

Observation 6b3a4b8f-4aeb-4e6c-8d18-c85d08fd2129 · outbound

This paper cites Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching.

Video-Guided Foley Sound Generation with Multimodal Controls Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.732227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.732227Z digest=sha256:dbd20422f24820d38a7d5de2c1e45e96bb71295cf6b5d8564e7897dcc35ca72c

Observation 3423a090-0db6-4bd2-aaf6-41a9c251e7b3 · outbound

This paper cites Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

Video-Guided Foley Sound Generation with Multimodal Controls Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:00:04.385160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:00:01.745874Z digest=sha256:31af80240ff13efabf132c1f12aa6ebd301d3ca4717f4108f720b477f4cc94bc

Observation 93173244-823e-478b-871c-1d56198e661a · outbound

This paper cites Son- icvisionlm: Playing sound with vision language models.

Video-Guided Foley Sound Generation with Multimodal Controls Son- icvisionlm: Playing sound with vision language models

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.758011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.758011Z digest=sha256:d847b376f25388fdd7abc7d61073b216fa9fd3dc0cf60f3ef882159daa3404ca

Observation 45d6a05e-d6af-433c-8f04-e444c9c13f89 · outbound

This paper cites Open-vocabulary panop- tic segmentation with text-to-image diffusion models.

Video-Guided Foley Sound Generation with Multimodal Controls Open-vocabulary panop- tic segmentation with text-to-image diffusion models

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.764701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.764701Z digest=sha256:a628a5c1feac1370be8dbaf60a40a80f101072eaeed84fad993c3365a2a933c6

Observation 0991ced3-6aa9-450d-a670-69c14db4695a · outbound

This paper cites Video-to-Audio Generation with Hidden Alignment.

Video-Guided Foley Sound Generation with Multimodal Controls Video-to-Audio Generation with Hidden Alignment

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.772811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.772811Z digest=sha256:d347697177f536dce1e9e5fdd16238d957388dd36b72442a5b7a5c60fcf04a3f

Observation 7be71e43-d6b5-4cfa-bc7e-135b8d5c202b · outbound

This paper cites Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation.

Video-Guided Foley Sound Generation with Multimodal Controls Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.780503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.780503Z digest=sha256:581e18d64405a0f372f583ca707e3023554a108e029298e0b2a2d7b1432462ca

Observation ff11c86d-9e46-4ddd-a7ca-ecbadf08c434 · outbound

This paper cites Telling left from right: Learning spatial correspondence of sight and sound.

Video-Guided Foley Sound Generation with Multimodal Controls Telling left from right: Learning spatial correspondence of sight and sound

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:00:04.281736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:00:01.792953Z digest=sha256:249d6c9728dcef7fb428128e12b45f52c8cf608a5e498400f3e63af70479b16b

Observation 7534a639-d017-43e6-8267-7b9f7e7c7c04 · outbound

This paper cites Cameras as rays: Pose estimation via ray diffusion.

Video-Guided Foley Sound Generation with Multimodal Controls Cameras as rays: Pose estimation via ray diffusion

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:00:04.230876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:00:01.810365Z digest=sha256:d3d880cfc522e9c3d22eb4ab0d7109af40913b89b41476fe1675c6f81322c907

Observation d0b00aba-8be9-4d19-b1ff-ccebc15b3b6c · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

Video-Guided Foley Sound Generation with Multimodal Controls FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.822180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.822180Z digest=sha256:c406f8be4b826da925a4ea00ec004739212754a199baf47ce2def375feaafbd7

Pith citing papers

Observation 0bf5ea14-9876-48ba-a115-d84c7c909ef7 · inbound

AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation cites this paper.

AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Video-Guided Foley Sound Generation with Multimodal Controls

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:08.186627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:08.186627Z digest=sha256:a3970b5ce0c3bec602ad4de77dc385f336df6dc38c8154eeaa628664fc9c2643

Observation 2dd8e76a-9b05-4864-af2c-d36c1d7586b2 · inbound

Sound Scene Synthesis at the DCASE 2024 Challenge cites this paper.

Sound Scene Synthesis at the DCASE 2024 Challenge Video-Guided Foley Sound Generation with Multimodal Controls

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T20:27:03.768169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:27:03.768169Z digest=sha256:036c399e510dcc459f9c88e1a7657b202c52a17bf4a0d420bfdb2ebf4bbec879

Observation ebef3684-b017-4923-be0b-8838c65c39b1 · inbound

Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT cites this paper.

Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT Video-Guided Foley Sound Generation with Multimodal Controls

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T14:26:46.637418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:26:46.637418Z digest=sha256:ac404a027f5d5dbd2f38b5abdcd1f74fb0300bf043182db60bffbda10604fbd7

Observation 5001ff69-54ec-4cfe-8d99-1f5e3c17a8b5 · inbound

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks cites this paper.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Video-Guided Foley Sound Generation with Multimodal Controls

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:09.819840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:09.819840Z digest=sha256:6b0e8a4a05d1a71839ec1be026fc73850639c848d3b3d6325855b2a5eca607cf

Observation 83671cc3-1f8c-46ab-adbb-f180111af7f3 · inbound

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance cites this paper.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Video-Guided Foley Sound Generation with Multimodal Controls

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:05.422297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:05.422297Z digest=sha256:1f56ceed706bfb59e89baaa71a706b27693bd7c512eae3e201bb3c7da6b0d63e

Observation d9381e5b-49e4-451c-b97a-762ff65b9d3b · inbound

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters cites this paper.

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters Video-Guided Foley Sound Generation with Multimodal Controls

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:43:11.816985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:43:11.816985Z digest=sha256:0518cd7893a39b8c8884fcbd289cfde58bb6e34cdabb32864b9ad74ecf7a3ee2

Observation cecbaf3d-0d77-4f51-9c13-55545ca37267 · inbound

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation cites this paper.

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation Video-Guided Foley Sound Generation with Multimodal Controls

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:28:18.786232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T08:04:48.283908Z digest=sha256:cc503dcf03f89074e42dc3cc47f684440e5f1aaf4b1c72776a60b9baaa7645ac