Pith. sign in

Paper Citation Record · LEDGER

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance

As of 20 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2412.18157.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18157 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:01:38.747896Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5dfb7121-e594-4dc0-93c3-7e824b842549 · outbound

This paper cites Generating visually aligned sound from videos,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Generating visually aligned sound from videos,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.315954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:01:38.603327Z digest=sha256:42a77cc2ebf6b50308ff02ffa5f15092aa3c84c2f991dc676e60bf3854a05b22

Observation 7f37af99-e3dc-4270-8fd2-c3016a9657a5 · outbound

This paper cites Taming visually guided sound generation,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Taming visually guided sound generation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.301234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:01:38.608027Z digest=sha256:d81a5ca76d80a5079a41d40cf1057e96fa15efacf551ab323d79433794076ebd

Observation c4581f99-f988-4b81-9220-c5044e0f6b6a · outbound

This paper cites Attention is all you need,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Attention is all you need,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.612589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.612589Z digest=sha256:c77fd444a41b8dd1ac85d251a36172f1f9417647c97a82be13797e2382e16306

Observation 46a1da61-e450-446f-8cfc-a394399df2be · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Learning transferable visual models from natural language supervision,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.618248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.618248Z digest=sha256:a695178949ff52212fe4fda69b5bf85aba6f73bf432e29b3451e4b79e8d0630c

Observation a744e841-2c08-404c-8b66-45e555bc176e · outbound

This paper cites Imagebind: One embedding space to bind them all,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Imagebind: One embedding space to bind them all,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.623065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.623065Z digest=sha256:f793983b19f74b679cc4f06d5819d88db91c3aae053f691c1006975328bda39a

Observation c8e4a828-487c-4ba8-a2c2-779ae22072e6 · outbound

This paper cites Denoising diffusion probabilistic models,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Denoising diffusion probabilistic models,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.628058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.628058Z digest=sha256:8434e03c75a24a3128a482f8f3dfeee2994a7ab0d3182b186f3ce310e4703526

Observation f99302b1-5d3d-4c06-ab10-26ca27acd214 · outbound

This paper cites Conditional generation of audio from video via foley analogies,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Conditional generation of audio from video via foley analogies,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.249712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:01:38.633488Z digest=sha256:aa0d68e2924f6bd2d16267926d52885ebeae5bb2f925a176743c3d5bc154fdf9

Observation 0de05aa5-1667-4472-bad8-014d5c258bf7 · outbound

This paper cites Varietysound: Timbre-controllable video to sound generation via unsupervised information disentanglement,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Varietysound: Timbre-controllable video to sound generation via unsupervised information disentanglement,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.235676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:01:38.637922Z digest=sha256:3225f57528ea1fab815fec416297d855c895d9eef2df993e484b3a74513b23f6

Observation bb71c8fd-1a60-497d-bd44-2f50c84b26e6 · outbound

This paper cites Diff-foley: Synchronized video- to-audio synthesis with latent diffusion models,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Diff-foley: Synchronized video- to-audio synthesis with latent diffusion models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.221071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:01:38.642369Z digest=sha256:36a0af1b01dd1708d5dbdaae191c80d0f1b5d420f5d65f393d6f75f876f410e4

Observation b0baf4ea-d77a-4717-8936-db8d6dda74ef · outbound

This paper cites I hear your true colors: Image guided audio generation,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance I hear your true colors: Image guided audio generation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.206553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:01:38.647528Z digest=sha256:2999c38a914d4a764672ef6fdf6225b1a3ceb5383eeca25678dbeda377bd69a5

Observation 82f2665c-0c2a-40d2-aeef-9889bc041184 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.651937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.651937Z digest=sha256:bb9136cca914c2ceac54b138d78ae3477081f08b03fe53d78b8022b3cfc71d97

Observation 5b036e71-af05-4ca9-8827-5dfeebc4e2ea · outbound

This paper cites Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.656647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.656647Z digest=sha256:28f3ee910077254c96e8011a1668197bf4d546e57f833b68628830b1e3535c63

Observation ebc93496-4ccc-4660-8eb4-33adb2f3926e · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.661167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.661167Z digest=sha256:1de548bc6d451c8e4f7f2314acbc1c45ab3d84bdfec0201121541fd5eb0de03f

Observation 66576f66-c5e7-4e9b-ae71-754e911cb725 · outbound

This paper cites Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.667045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.667045Z digest=sha256:fa2d493deb0e7542457e95c3241c8fd3759c258afc87fa15a6b9a8ed259d636b

Observation dc1c3c51-5603-4238-89e0-28dd1de65d9c · outbound

This paper cites Predicting deep zero-shot con- volutional neural networks using textual descriptions,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Predicting deep zero-shot con- volutional neural networks using textual descriptions,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.183592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:01:38.671276Z digest=sha256:a9ea0c24a68797959aaaf48eb14993a3e8ed1503cf2a892e52a503fbbff5a721

Observation 7f9402c1-e700-4568-98b8-c32d1ff27321 · outbound

This paper cites Integrating language guidance into vision-based deep metric learning,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Integrating language guidance into vision-based deep metric learning,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.168740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:01:38.675614Z digest=sha256:270919d0649a76a8d7e1ecd148229e6964ea877652eadc699bde8f6b040f6838

Observation 9f53574b-4bc0-45cc-9953-91bb9bf5360b · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.679692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.679692Z digest=sha256:c8e77c9cd71360bd5f1e0ebdfd5ecfbbfe0005af276e09a4b3e2fb28de080258

Observation ec336e35-c0ad-496c-ac41-f1363c5c8ddb · outbound

This paper cites Adding conditional control to text-to-image diffusion models,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Adding conditional control to text-to-image diffusion models,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.684397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.684397Z digest=sha256:48cfc49362315c6bd9f55f798477145fbbea815dad8d75f9d38697ba57f90593

Observation 5c52080e-0a92-40c9-9e0b-7bfaf9c2c1e2 · outbound

This paper cites The benefit of temporally-strong labels in audio event classification,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance The benefit of temporally-strong labels in audio event classification,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.145445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:01:38.688822Z digest=sha256:e4e0500f5527a009950bf183b1c940e4656413f60b8e67de55985d0c7c03d917

Observation 828f7af1-5619-4181-8f10-f5d78f21b61c · outbound

This paper cites Vggsound: A large- scale audio-visual dataset,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Vggsound: A large- scale audio-visual dataset,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.130794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:01:38.693180Z digest=sha256:4edc188bed78fba7bd55e2e26ecb910c7a1bea670694eea1ee6688cce88354ca

Observation daa08241-334d-499d-8abf-6c34faf50bf4 · outbound

This paper cites Text-to-audio grounding: Build- ing correspondence between captions and sound events,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Text-to-audio grounding: Build- ing correspondence between captions and sound events,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.116072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:01:38.697764Z digest=sha256:1bbf26918c0aee3338901fecf8d238000b80497205d11be0298cd3c02f76e045

Observation 7d0daffe-6166-48ba-a9fb-836f29a9d52f · outbound

This paper cites Towards Weakly Supervised Text-to-Audio Grounding.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Towards Weakly Supervised Text-to-Audio Grounding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.702238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.702238Z digest=sha256:50a51c966dbb40a28cfeb1bf8d4d6ad7e3abdacf93010f20ec4d004c1ba21790

Observation a8e132c8-e7ce-42df-8f5e-1e48c1b3d1a3 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.707509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.707509Z digest=sha256:6576e3e4cc5720db1f0c16f835391f6ece5211b1093e4629bdc811300fd67053

Observation 52cf62e5-1d30-42d6-8671-3b9ea584e164 · outbound

This paper cites Correlation of Fr\'echet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Correlation of Fr\'echet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.712224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.712224Z digest=sha256:a069c0ed351e7115935eb95f44995e9dc9a3066de44899b7c03562a674f4a0e4

Observation 51b0d9ab-a0c1-456f-aa1c-75de199c2702 · outbound

This paper cites Panns: Large-scale pretrained audio neural networks for audio pattern recognition,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Panns: Large-scale pretrained audio neural networks for audio pattern recognition,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.716900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.716900Z digest=sha256:d2e6d00ca5625ecf24e8734c7652cd3fcb53ffea31d3622e502adc17571049ef

Observation dd755cc4-22e3-4c89-81ff-0e3ab5c6b53d · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.091317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:01:38.721295Z digest=sha256:b54082ad7d526a6dfc3205c5f59dc8c1ff0e7451549cf2fcfad25f2c505f871d

Observation 1e89bd65-4bed-4f43-b4ec-6e44e35a5939 · outbound

This paper cites Cnn architectures for large-scale audio classification,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Cnn architectures for large-scale audio classification,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.076628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:01:38.725435Z digest=sha256:0124917a35494f7fc16992105e00211d818694de23abf3c27eeb655e6593e4aa

Observation 13e56c13-8629-482d-85ad-4487522ced56 · outbound

This paper cites V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.061764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:01:38.729633Z digest=sha256:3fb05a051b8170165fb731dc848c409627c35ebf6876814bb6153018e9ea4c05

Observation abc35f88-6c8d-4329-ac81-183c089e950c · outbound

This paper cites Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.734637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.734637Z digest=sha256:b798595fa51acca9e0a5b8782f352933b3616881da0e5c4eeb93244432853fa7

Observation 4df78d47-57ae-4bef-acd9-d64e21b250ed · outbound

This paper cites Video-foley: Two-stage video-to- sound generation via temporal event condition for foley sound,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Video-foley: Two-stage video-to- sound generation via temporal event condition for foley sound,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.739388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.739388Z digest=sha256:08f6f1508af776ce9a380234c814fafa6cb1fd4eef295189a75170ff4111e977

Observation 78d8003f-9395-4bc1-90d8-f339f896c700 · outbound

This paper cites Wav2clip: Learning robust audio representations from clip,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Wav2clip: Learning robust audio representations from clip,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.047018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:01:38.743732Z digest=sha256:a7739658ff0b82c6f36c57fc7b63f4daa07880cc1d2cb91f4fc0c7baa4312455

Observation fb97627d-ccf9-465d-8054-401f6636873f · outbound

This paper cites Fr ´echet audio distance: A reference-free metric for evaluating music enhancement algorithms,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Fr ´echet audio distance: A reference-free metric for evaluating music enhancement algorithms,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.031776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:01:38.747896Z digest=sha256:7150315c034c1218df541599663071b9af42b22e70fad024fb7be3a7654d69aa

Pith citing papers

No inbound Pith citation observations are available.