Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T05:01:38.747896Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2412.18157.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T05:01:38.747896Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
32 of 32 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5dfb7121-e594-4dc0-93c3-7e824b842549 · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Generating visually aligned sound from videos,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7f37af99-e3dc-4270-8fd2-c3016a9657a5 · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Taming visually guided sound generation,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c4581f99-f988-4b81-9220-c5044e0f6b6a · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Attention is all you need,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46a1da61-e450-446f-8cfc-a394399df2be · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Learning transferable visual models from natural language supervision,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a744e841-2c08-404c-8b66-45e555bc176e · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Imagebind: One embedding space to bind them all,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8e4a828-487c-4ba8-a2c2-779ae22072e6 · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Denoising diffusion probabilistic models,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f99302b1-5d3d-4c06-ab10-26ca27acd214 · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Conditional generation of audio from video via foley analogies,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0de05aa5-1667-4472-bad8-014d5c258bf7 · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Varietysound: Timbre-controllable video to sound generation via unsupervised information disentanglement,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation bb71c8fd-1a60-497d-bd44-2f50c84b26e6 · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Diff-foley: Synchronized video- to-audio synthesis with latent diffusion models,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b0baf4ea-d77a-4717-8936-db8d6dda74ef · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance I hear your true colors: Image guided audio generation,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 82f2665c-0c2a-40d2-aeef-9889bc041184 · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b036e71-af05-4ca9-8827-5dfeebc4e2ea · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebc93496-4ccc-4660-8eb4-33adb2f3926e · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66576f66-c5e7-4e9b-ae71-754e911cb725 · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc1c3c51-5603-4238-89e0-28dd1de65d9c · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Predicting deep zero-shot con- volutional neural networks using textual descriptions,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7f9402c1-e700-4568-98b8-c32d1ff27321 · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Integrating language guidance into vision-based deep metric learning,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9f53574b-4bc0-45cc-9953-91bb9bf5360b · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance AudioLDM: Text-to-Audio Generation with Latent Diffusion Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec336e35-c0ad-496c-ac41-f1363c5c8ddb · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Adding conditional control to text-to-image diffusion models,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c52080e-0a92-40c9-9e0b-7bfaf9c2c1e2 · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance The benefit of temporally-strong labels in audio event classification,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 828f7af1-5619-4181-8f10-f5d78f21b61c · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Vggsound: A large- scale audio-visual dataset,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation daa08241-334d-499d-8abf-6c34faf50bf4 · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Text-to-audio grounding: Build- ing correspondence between captions and sound events,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7d0daffe-6166-48ba-a9fb-836f29a9d52f · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Towards Weakly Supervised Text-to-Audio Grounding
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8e132c8-e7ce-42df-8f5e-1e48c1b3d1a3 · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52cf62e5-1d30-42d6-8671-3b9ea584e164 · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Correlation of Fr\'echet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51b0d9ab-a0c1-456f-aa1c-75de199c2702 · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Panns: Large-scale pretrained audio neural networks for audio pattern recognition,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd755cc4-22e3-4c89-81ff-0e3ab5c6b53d · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1e89bd65-4bed-4f43-b4ec-6e44e35a5939 · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Cnn architectures for large-scale audio classification,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 13e56c13-8629-482d-85ad-4487522ced56 · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation abc35f88-6c8d-4329-ac81-183c089e950c · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4df78d47-57ae-4bef-acd9-d64e21b250ed · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Video-foley: Two-stage video-to- sound generation via temporal event condition for foley sound,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78d8003f-9395-4bc1-90d8-f339f896c700 · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Wav2clip: Learning robust audio representations from clip,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation fb97627d-ccf9-465d-8054-401f6636873f · outbound
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Fr ´echet audio distance: A reference-free metric for evaluating music enhancement algorithms,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
No inbound Pith citation observations are available.