Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:27:03.800014Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 3 inbound Pith citation observations for arXiv:2501.08587.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:27:03.800014Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:24:02.180525Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-10T20:27:04.247482Z
43 of 43 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 24fce64e-375e-45ed-bee2-ad952755a333 · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 77ceb030-86f3-478e-9d32-62ed4ca956c6 · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge This is a more flexible setup than the category- based generation used in the last year [2]
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2f7e9928-baca-466e-bc81-f6ec4e9c8ef6 · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Sound Scene Synthesis at the DCASE 2024 Challenge
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c19976b5-8a2d-4213-a290-7d4f0a482791 · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Objective Evaluation We employed the Fr ´echet Audio Distance (FAD) [6] with PANN- Wavegram-Logmel [7] embeddings as our primary objective metric
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d022a7f9-bc37-4c29-a1d0-fd48f1a9fa16 · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge System Performance Table 1 summarizes the evaluation results
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 01984bcb-c90b-4f23-995f-a26ac8cb3f52 · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge First, the generative aspect of organizing this challenge has been costly and labor intensive
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d6d2267f-0b13-465c-9cca-5ebaddf896d5 · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge While the submit- ted systems demonstrated promising capabilities, the significant gap between synthetic and reference audio quality indicates substantial room for improvement
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 896b9a5b-301a-4fa0-bb13-5126a48df30a · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a2bfb407-3084-446d-ae3c-33759162b60a · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge A Proposal for Foley Sound Synthesis Challenge
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c0476e70-cdb1-4b4c-bc49-f867c67df3e0 · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Foley sound synthesis at the dcase 2023 challenge,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3ea3db78-67ed-4cba-86e0-0e8a7cee0414 · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge AudioLDM: Text-to-Audio Generation with Latent Diffusion Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdb65355-b319-4204-a4ba-312a42b89132 · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Audiocaps: Gen- erating captions for audios in the wild,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6a829246-5d5c-43f3-bdd3-6d6bfaff6bd4 · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Audio set: An ontology and human-labeled dataset for audio events,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b25d28e6-1320-4048-9135-74800fd6d976 · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Fr ´echet audio distance: A reference-free metric for evaluating music enhancement algorithms
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ae928ba1-6638-4f65-a538-1356c3f45ac3 · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Panns: Large-scale pretrained audio neural net- works for audio pattern recognition,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ae6d1d63-b213-49e9-9052-efc718594ebb · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Correlation of fr´echet audio dis- tance with human perception of environmental audio is em- bedding dependent,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 65f02729-3396-4aa0-8e56-43c264ec4f13 · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Sound scene synthesis with audioldm and tango2 for dcase 2024 task7,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5dc700f2-4240-4786-b517-66c3ffcf2c77 · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Sound scene synthesis based on gan using contrastive learning and effective time-frequency swap cross attention mechanism,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2de08bad-89f7-49d7-bd1e-af93a9bebf24 · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Dif- fusion based sound scene synthesis for dcase challenge 2024 task 7,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 51529f4f-4c18-46b9-a575-eae1a0b49120 · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Sound scene synthesis based on fine-tuned latent diffusion model for dcase challenge 2024 task 7,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6aff4a8f-529d-4cb2-b20a-dc6f31b2ab39 · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Challenge on sound scene synthesis: Evaluating text-to-audio generation,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6bad9403-92a0-4be7-898a-fc683ea926f7 · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge T-foley: A controllable waveform-domain diffusion model for temporal-event-guided foley sound synthesis,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9b085ddf-f36b-4e14-8b47-932f52b08be0 · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge MambaFoley: Foley Sound Generation using Selective State-Space Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4792880d-b8f9-4c3e-b55c-e6a4eb4f081a · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Audio generation with multiple conditional diffu- sion model,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c442015e-633a-4f80-b227-07b2e24fd41f · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge PicoAudio: Enabling Precise Timestamp and Frequency Controllability of Audio Events in Text-to-audio Generation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8ad09ea-ca13-482f-b2c5-23ecf02bc56f · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Audioldm 2: Learn- ing holistic audio generation with self-supervised pretraining,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 310b16f1-67d9-462d-8ce4-46edf18bd4ea · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 629502ad-7164-4f64-b0bd-b10ca7ec0d24 · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fce6b11-ee5f-457c-ab71-a405aeb3c26d · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Fugatto 1: Foundational generative audio transformer opus 1,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 603edb4f-0b3c-4112-bda5-bf8f02ab24c6 · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Stable Audio Open
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6257149f-256c-4648-b54d-2de605277656 · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Improving Text-To-Audio Models with Synthetic Captions
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0cf9b71-9bd3-40d9-b780-85d5570f34b5 · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Syncfusion: Multi- modal onset-synchronized video-to-audio foley synthesis,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 04198b01-4384-4d8c-a34d-6f06b9467c6c · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Sonicvisionlm: Play- ing sound with vision language models,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f75205d1-a832-4e17-a211-7090fc7e4ca9 · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Video-foley: Two-stage video-to-sound generation via temporal event condition for fo- ley sound,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fe1a964-75ef-4ecf-808b-b314da4bc19c · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Read, Watch and Scream! Sound Generation from Text and Video
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae0ef7f6-0387-4fa6-b606-d8f6fae36ee0 · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Movie Gen: A Cast of Media Foundation Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dd8e76a-9b05-4864-af2c-d36c1d7586b2 · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Video-Guided Foley Sound Generation with Multimodal Controls
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaa0d4c0-e44a-4942-bc74-9a9d04cb9093 · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d656700e-3452-4743-a68e-9de0cdf232db · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f8e29d9-2b0f-4cb6-987b-c9d8ec0fee72 · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Masked gener- ative video-to-audio transformers with enhanced synchronic- ity,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7ca429fd-5320-47ba-add0-d721791b7105 · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Temporally Aligned Audio for Video with Autoregression
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e762b05d-a28f-45fe-a7d5-0e7fd3fc0cdf · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge Gotta Hear Them All: Towards Sound Source Aware Audio Generation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cecb5b1-87b6-46a2-932a-2cc195f06d3d · outbound
Sound Scene Synthesis at the DCASE 2024 Challenge MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f7e9928-baca-466e-bc81-f6ec4e9c8ef6 · inbound
Sound Scene Synthesis at the DCASE 2024 Challenge Sound Scene Synthesis at the DCASE 2024 Challenge
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c708edc0-fc08-4182-8948-978c0913df5e · inbound
SoundscapeAgent: Agentic Soundscape Construction for Controllable Synthesis and Scalable Audio-Language Supervision Sound Scene Synthesis at the DCASE 2024 Challenge
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58a07c9c-cbe5-4c96-96df-705333699a58 · inbound
Scene2Sound: Auditory-Grounded Soundscape Generation for 3D Gaussian Worlds Sound Scene Synthesis at the DCASE 2024 Challenge
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.