Pith. sign in

Paper Citation Record · LEDGER

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance

As of 9 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2506.20995.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20995 v4

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:45:08.841789Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy48
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b2efef86-b747-425b-86c4-8cc021f95666 · outbound

This paper cites The Foley Grail: The Art of Performing Sound for Film, Games, and Animation.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance The Foley Grail: The Art of Performing Sound for Film, Games, and Animation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.351577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:04.880574Z digest=sha256:b6c70ca049ab123295300c8d8becb5e75525eb14f828b32443877debf684d566

Observation bf501168-58c8-491e-88bf-c0087732d0c7 · outbound

This paper cites Qwen2.5-VL Technical Report.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:04.943204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:04.943204Z digest=sha256:10ca119cb032ab708cb123df6a3fd8b4cca44af44d5375565f75bfe9a0d609e2

Observation 138ac0b1-663d-4891-ae36-377b6e6f1b38 · outbound

This paper cites Erasedraw: Learning to insert objects by erasing them from images.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Erasedraw: Learning to insert objects by erasing them from images

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.340230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:05.024306Z digest=sha256:fdadb4908e455e8f809c863cab1633acfe93839edb547d326e06caff0651b720

Observation 9cc5593b-09cc-4f8e-9f1e-3ef031053a27 · outbound

This paper cites Action2sound: Ambient-aware generation of action sounds from egocentric videos.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Action2sound: Ambient-aware generation of action sounds from egocentric videos

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.327905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:05.113307Z digest=sha256:02f0ac758de8d8d2ff0e75a79abfb2c5f9663a0709e155f4e36e2dff8475363a

Observation f19cb447-815d-43e3-b6e9-057ba58a388e · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Vggsound: A large-scale audio-visual dataset

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.316890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:05.234622Z digest=sha256:0990dc0375d6e83f9764bae0eca1bac2b4c9e339ef92d3a006e84c06b2d3c20c

Observation 50d1729f-30ce-4e2c-9a55-874ca54bf8db · outbound

This paper cites Generating visually aligned sound from videos.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Generating visually aligned sound from videos

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.305399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:05.353789Z digest=sha256:04f988f60664b94544562be88c2c3ded3c169e83e9d6aa052d275918b86a9f04

Observation 83671cc3-1f8c-46ab-adbb-f180111af7f3 · outbound

This paper cites Video-Guided Foley Sound Generation with Multimodal Controls.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Video-Guided Foley Sound Generation with Multimodal Controls

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:05.422297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:05.422297Z digest=sha256:8f3f0ee42b29d1a657db6d1da6d33d933b2a14705b6cacd80b86dae32c63e363

Observation 465ce467-2a71-45bb-94bb-5bd783ab6dd0 · outbound

This paper cites Mmaudio: Taming multimodal joint training for high-quality video-to-audio synthesis.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Mmaudio: Taming multimodal joint training for high-quality video-to-audio synthesis

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.293192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:05.510328Z digest=sha256:2d86136fae4e605e7f63a16757c74bc461a52d25a10efdb5f39a262ac42f3b50

Observation f6f79a6e-9c2b-497d-9ccb-aa341c3f8f5e · outbound

This paper cites Clotho: An audio captioning dataset.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Clotho: An audio captioning dataset

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.281452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:05.606265Z digest=sha256:050586e2d9ebe57a9b3f737ef1de1848d33d10d0af4918297368f3d8b8f00c35

Observation c48835b7-ca7e-4973-9b99-a865742c2709 · outbound

This paper cites Compositional visual generation with energy based models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Compositional visual generation with energy based models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.269734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:05.722732Z digest=sha256:294ffdcda6224414c6b40b2860219ccc9a8252e11a5ef0c58d2425df19dc308c

Observation 8fcf95f3-9723-4d9e-8778-5fdb8415bf20 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Scaling rectified flow transformers for high-resolution image synthesis

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.258064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:05.808806Z digest=sha256:9157da647e3316cbc3014793749da94a592081a659cc8050e14ecc6cd2b33a34

Observation 5a745c0a-b954-408e-bd54-993eacde0289 · outbound

This paper cites Sketch2sound: Controllable audio generation via time-varying signals and sonic imitations.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Sketch2sound: Controllable audio generation via time-varying signals and sonic imitations

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.245835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:05.904797Z digest=sha256:f708d705410803aa9629e9cfc0dc91d5d7314dc2bea3c01bbb64c4a620b2b150

Observation e2701f7c-992a-43d0-bba1-362698e35d76 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Audio set: An ontology and human-labeled dataset for audio events

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.233666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:05.968972Z digest=sha256:76daf5f4acb243a53628253841566d42b9bd22d5273970faf10403c4e3714465

Observation 8ab04e99-e18d-484c-9a57-60fc2b2b7602 · outbound

This paper cites Imagebind: One embedding space to bind them all.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Imagebind: One embedding space to bind them all

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.222069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.021697Z digest=sha256:a33e163437d10fa36102980aa09528c8a0cfa2b7b6852575c5b42b82e3e48143

Observation d2359448-9249-44de-8680-c523e1dad7d8 · outbound

This paper cites Instructme: An instruction guided music edit and remix framework with latent diffusion models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Instructme: An instruction guided music edit and remix framework with latent diffusion models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.209188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.075730Z digest=sha256:35e44f9314c3e079f2a6b91680906c25da73e5c38ccb61190d72f7a89adb7291

Observation 784a3cab-fd12-4243-b191-259de7ea9baf · outbound

This paper cites Classifier-free diffusion guidance.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Classifier-free diffusion guidance

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.198403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.151487Z digest=sha256:f1b7de103f72157e63b43d4fddc34f57d1f40ef0a75ced719df2fdd11cea9359

Observation 791d89c6-44a6-45a4-9721-f24578894e16 · outbound

This paper cites Taming visually guided sound generation.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Taming visually guided sound generation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.186932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.203790Z digest=sha256:7560559e17c11f7cdb91935e5d420313a3de790c8f987105557fe51e014dc5d5

Observation 3f00ebf4-1299-48dc-809f-7d36449e9505 · outbound

This paper cites Synchformer: Efficient synchronization from sparse cues.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Synchformer: Efficient synchronization from sparse cues

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.174622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.275002Z digest=sha256:3d7dad5fb785600ff9e866a7454607a65d041131ecaa579e3b1d1602afc06981

Observation 76d83d63-5816-4a20-bcb7-456bbf0e3d3b · outbound

This paper cites Audioeditor: A training-free diffusion-based audio editing framework.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Audioeditor: A training-free diffusion-based audio editing framework

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.163649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.323102Z digest=sha256:31733f026ae1b5aed71ea6af6e1cfc3b048b1f7a17e8c75f8a0865f8e2e4d533

Observation a6d67126-3422-4ee6-b220-50bc93246462 · outbound

This paper cites Simultaneous music separation and generation using multi-track latent diffusion models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Simultaneous music separation and generation using multi-track latent diffusion models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.151760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.376824Z digest=sha256:d71f172c75b217c8c26c1aa63f6008d648f8b4977f96ebd5fb90ac6d2f3dbcd5

Observation 543f135c-bde9-4a08-907e-32953996f07c · outbound

This paper cites Analyzing and improving the training dynamics of diffusion models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Analyzing and improving the training dynamics of diffusion models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.140505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.432024Z digest=sha256:8aabef550b9f944cc2bfc0f272cbff9cf1a97a873d0b7f870af79590341948f9

Observation f0c3b862-20cb-440b-8185-4c5cc2213f44 · outbound

This paper cites Audiocaps: Generating captions for audios in the wild.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Audiocaps: Generating captions for audios in the wild

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.128711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.488469Z digest=sha256:e413a9823514a645252acc42a232b10b998e8d4835ea6076eb9db1bb30bd9cdc

Observation ee2e0e48-327f-416e-93d1-9878a97a85f3 · outbound

This paper cites Panns: Large-scale pretrained audio neural networks for audio pattern recognition.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Panns: Large-scale pretrained audio neural networks for audio pattern recognition

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.116398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.539100Z digest=sha256:07bb46a98b34e2ba539f19197c63a7169e349d7cc4cd76c1abe1512ea6944459

Observation 2285d8a5-0e8d-4396-94bf-2d20c51f43b8 · outbound

This paper cites Vintage: Joint video and text conditioning for holistic audio generation.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Vintage: Joint video and text conditioning for holistic audio generation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.104521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.608604Z digest=sha256:a6062d22e6dd443e7499d5be750c975220e0ec194d74c261da9b6a1a3b4568a2

Observation 3b64cfc8-9b49-4187-b718-f8eddc76d437 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Gonzalez, Hao Zhang, and Ion Stoica

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.091741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.675323Z digest=sha256:8e636a9133c3df7007631967c74e09da0de1e1d0e77426572a0636a4bf625f10

Observation 8d5aed84-561b-4ff3-b3da-80fc48acb6f0 · outbound

This paper cites Flow matching for generative modeling.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Flow matching for generative modeling

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.080943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.722269Z digest=sha256:4ee2bb123218b18a3e8fd3896628dbe87033fa11baf2a5a41f4645ab518fb1f9

Observation 72f18168-e4f9-43bd-bd23-35e8754133ee · outbound

This paper cites Compositional visual generation with composable diffusion models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Compositional visual generation with composable diffusion models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.069743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.789909Z digest=sha256:839e0c928aa2cc82dc60cd83e425d8bdaa304e091afd89a12488069df62de90f

Observation fbac445a-9263-4075-8796-713d4175d007 · outbound

This paper cites Tell what you hear from what you see-video to audio generation through text.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Tell what you hear from what you see-video to audio generation through text

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.058915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.849457Z digest=sha256:424181e9ce52094f0fc32b141ad3315dc4ab8559739f1baf0e525892cff937ab

Observation 5c27eb4f-97b7-4233-b673-0f8f00d8fd96 · outbound

This paper cites Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.047856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.916636Z digest=sha256:125dfeb7555e3adf34b9bef5406eabd74ec340698b297544ad38cf7dce884ee7

Observation 79b5cffe-d872-4a6b-a594-afe71466b990 · outbound

This paper cites Multi-source diffusion models for simultaneous music generation and separation.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Multi-source diffusion models for simultaneous music generation and separation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.037038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.972516Z digest=sha256:0a5104d4449fb61c74a67b5b025826113696e6619c74303a6b17d67a432fdba0

Observation 1c51d7c4-9601-4119-a8e5-7b6e238e7efa · outbound

This paper cites Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.024863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:07.046130Z digest=sha256:44d092afe30b22eaed492b32588160a84649f0893dd8b451075736b0815b6e44

Observation 20fca138-03ed-415c-958c-0173fb57af3a · outbound

This paper cites Audio-visual scene analysis with self-supervised multisensory features.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Audio-visual scene analysis with self-supervised multisensory features

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.012958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:07.096950Z digest=sha256:6ae9f1d0ff5154ab572959e461e549b4e8135054acf09c78acfbc95aeb7c59b8

Observation 0f77170d-dce8-464e-84cd-b1b0bb67c57f · outbound

This paper cites Stemgen: A music generation model that listens.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Stemgen: A music generation model that listens

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.001536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:07.138164Z digest=sha256:6ffc532babac9cbaa48178c47bfadba7fc5dcb75846bd588316b319ccb94f01b

Observation 186907ed-1502-45d5-82a0-79d9c0153b15 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Movie Gen: A Cast of Media Foundation Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:07.206934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:07.206934Z digest=sha256:709722e7432bec90caf29a64532f2725ea7e4a091f64ee7aeb7b1ff1dcfc2b1e

Observation 463dd510-0502-4450-9692-436cc26a425b · outbound

This paper cites Generalized multi-source inference for text conditioned music diffusion models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Generalized multi-source inference for text conditioned music diffusion models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.990280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:07.274186Z digest=sha256:77f5723ba696e1bf15791374cc11ccad01425769c1a42f55ce3fefbc09510732

Observation b96150f7-b8f9-4748-a4de-37565fc32655 · outbound

This paper cites Improved techniques for training gans.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Improved techniques for training gans

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.978310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:07.347645Z digest=sha256:702b8e3ab1187d4e4bd67ffbc307084eb22484398435b09fd00ddb2eb6176f36

Observation e65fd7f1-2874-4163-a217-7e3bb9e57f15 · outbound

This paper cites Smartmask: context aware high-fidelity mask generation for fine-grained object insertion and layout control.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Smartmask: context aware high-fidelity mask generation for fine-grained object insertion and layout control

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.960730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:07.473056Z digest=sha256:274de457b5a749e58394566a4a3092c5918befb1ca5409a8bc35d2d292f1202c

Observation a4f55f75-e5ed-4ed4-bb7c-d95b5b065c55 · outbound

This paper cites Visually guided sound source separation with audio-visual predictive coding.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Visually guided sound source separation with audio-visual predictive coding

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.832404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:07.585707Z digest=sha256:b290b53d12d23d027d79036527b375db00e32fda2dbc7a51b4f7d33088402994

Observation 4e31c034-6b9a-4f01-93ab-ae7ff2779e8c · outbound

This paper cites sd3.5, 2024.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance sd3.5, 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.502321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:07.682202Z digest=sha256:45e8eaaccea9487f8d0cb607ca566dc215f4be9eb8e887d99cf3c08b0ea83994

Observation 37145041-e787-421d-b6f3-a1e17e010e53 · outbound

This paper cites Steinmetz and Joshua D.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Steinmetz and Joshua D

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.318638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:07.779691Z digest=sha256:7c7ad563596f0abf649b9ac8cef4902994bed4e7432adf46f4d24798dbe19422

Observation 03c38e60-10be-4a19-ae61-8e74395b6276 · outbound

This paper cites Add-it: Training-free object insertion in images with pretrained diffusion models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Add-it: Training-free object insertion in images with pretrained diffusion models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.045041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:07.883778Z digest=sha256:0aaca6ab354b1d4ba4937e60fd5d64c3b9bc56c18ca53bdf354efe51730d062a

Observation 0bc16ca5-f94f-4c75-ba2a-1b4004b179b2 · outbound

This paper cites Liu, Kevin J.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Liu, Kevin J

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:10.900034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:08.011468Z digest=sha256:5a196210838cc5e6ac8771a4800b041f525e021ea52be41e57e0d44500318da5

Observation 9777209a-2c83-403a-82c9-35f0b3cdcafe · outbound

This paper cites Temporally aligned audio for video with autoregression.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Temporally aligned audio for video with autoregression

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:10.758908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:08.098102Z digest=sha256:2bb49e95aca0c31c1eb7626a513e213b451c75d26c472086855f2d89f748629d

Observation f6e8b40c-660b-4d33-9552-06e095b8af5c · outbound

This paper cites V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:10.554435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:08.178693Z digest=sha256:00842fe2c268710c8c049de2a0048603a938add1a3b26b237b567e976df5a9ea

Observation 5ba41fe0-c0b5-41dc-88ec-625b962f1694 · outbound

This paper cites Frieren: Efficient video-to-audio generation network with rectified flow matching.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Frieren: Efficient video-to-audio generation network with rectified flow matching

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:10.395064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:08.285868Z digest=sha256:acc7431aaa2985234723736718b811c2b8eef951ce382d172d3e8ab224fd8d7c

Observation e0dde940-a687-4083-9f75-0920aa92541e · outbound

This paper cites Audit: Audio editing by following instructions with latent diffusion models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Audit: Audio editing by following instructions with latent diffusion models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:10.226759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:08.365573Z digest=sha256:c99ff8face7beded99ae3e44702f40e13e242d19b8e14b0ad2ec7e9d107fa58e

Observation d4d331b0-7589-46d7-a134-6f921ba9eae9 · outbound

This paper cites Stable diffusion 2.0 and the importance of negative prompts for good results, 2022.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Stable diffusion 2.0 and the importance of negative prompts for good results, 2022

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.978231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:08.474046Z digest=sha256:d2eb157a0189d2e7aa800c5b0390eeb976fdd383c49c5488473762eca2d2cd82

Observation fbf2a7b1-9d47-4e00-a5ab-eed923cef885 · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.732295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:08.535917Z digest=sha256:c35fc31e1e387a92df444f80803e40c19c0c920154a904c48635f4562acc11be

Observation 0102a54a-4938-4296-8fd9-88cd98f822af · outbound

This paper cites Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.528185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:08.546056Z digest=sha256:84e1fe7401a0fc9f2efd6e876833a5ecd8aa3ff1c73082be6d732e93702a8205

Observation 6d026ab3-f7b1-4d29-9af6-1d3023bc86fd · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Adding conditional control to text-to-image diffusion models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.304523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:08.621106Z digest=sha256:dc0ffe2840be81850d3ce27359c628f1064698565c13a93d9aad05191f4bdd3a

Observation f86c4cf0-2d82-4547-8a55-783902287549 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:08.730625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:08.730625Z digest=sha256:49c054135900be3329a6d29e372ea66624d96fc398179461ae43b56208bd5acb

Observation b754503b-876d-4787-b826-07e6136fc464 · outbound

This paper cites Visually guided sound source separation using cascaded opponent filter network.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Visually guided sound source separation using cascaded opponent filter network

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.086947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T22:45:08.841789Z digest=sha256:5aa4d423731648fecc651215fe93f38a6642f97c93980705f9fcd6cc4cd8a536

Pith citing papers

No inbound Pith citation observations are available.