Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

Sound Scene Synthesis at the DCASE 2024 Challenge

T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Text-to-sound-scene systems still trail expert recordings by 36 percent, and FAD tracks human judgment closely enough to serve as a proxy.

desk verdict Honest, useful challenge report with clean evaluation protocol; the FAD-human correlation is real but rests on 5 points and a single engineered reference, and the paper itself flags the small sample. read the letter →

arxiv 2501.08587 v1 pith:AIMB6Q4L submitted 2025-01-15 cs.AI cs.SDeess.AS

classification cs.AIcs.SDeess.AS
keywords soundscenesynthesistext-to-audiogenerationDCASE2024challengeFréchetAudioDistanceperceptualevaluationgenerativebenchmarkqualityassessment
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Text-to-audio systems can now synthesize environmental sound scenes, but there has been no agreed way to compare them. This paper reports the DCASE 2024 challenge Task 7, which evaluated four submitted systems against a fixed set of sound-engineer reference recordings using the Fréchet Audio Distance (FAD) and human perceptual ratings. The organizers found that the best system scores 36 percent below the reference on a weighted perceptual score, that FAD correlates with human foreground and background fit at $r=0.94$ and with audio quality at $r=0.77$, and that the manual evaluation cost made the challenge unsustainable to repeat. The paper's contribution is a reusable evaluation recipe and a quantified measure of how far synthetic sound scenes are from professional quality.

What carries the argument

The central object is the Fréchet Audio Distance computed over PANN-Wavegram-Logmel embeddings, paired with a perceptual score built from Foreground Fit, Background Fit, and Audio Quality ratings: $\text{Perceptual Score} = (2FF + BF + AQ)/4$. FAD compares the mean ($\mu$) and covariance ($\Sigma$) of the embedding distributions of the reference and generated audio sets, and the specific embedding was chosen because prior work showed that FAD's agreement with human perception depends on the embedding. The reference side of both measures is a 250-caption evaluation set recorded by a single sound engineer, and the weighted perceptual formula gives foreground accuracy the largest say in the final ranking.

What would settle it

Have a second sound engineer record a fresh reference set for the same 250 captions, recompute FAD for the four submitted systems against both reference sets, and compare with the existing human ratings; if the FAD-based ranking or the 36 percent gap shifts while human ratings stay stable, the single-reference assumption is the source of the instability.

Watch

Extended reading notes

Core claim

The paper's central claim is that a standardized framework combining FAD computed on PANN-Wavegram-Logmel embeddings with ratings from 14 expert listeners gives a stable, interpretable comparison of sound scene synthesis systems. Using that framework, the best submitted system reaches an average perceptual score of 5.832 against the sound engineer reference's 8.793, a gap of about 36 percent, while FAD values range from 35.985 for the best system to 53.728 for the lowest-ranked one. The objective and subjective measures agree strongly on foreground fit ($r=0.94$) and background fit ($r=0.94$) and less strongly on overall audio quality ($r=0.77$), which the paper treats as useful but weak evidence because only five systems were compared.

Load-bearing premise

The evaluation assumes that a single set of reference recordings made by one sound engineer is the correct target for each caption, so a system is judged by how close it comes to that one distribution.

Editorial extensions

If this is right

  • FAD can serve as an inexpensive screen for future sound scene synthesis comparisons, since it tracks human foreground and background fit at $r=0.94$ across the systems tested.
  • The 36 percent gap between the best system and the reference quantifies the headroom remaining for generative sound models, giving later work a concrete improvement target.
  • The 'Foreground with Background in the background' caption structure and the separate FF, BF, and AQ ratings allow future evaluations to diagnose failure by scene layer rather than by overall quality alone.
  • The organizers' accounting of roughly 120 hours of expert effort plus platform and compute costs explains why the task was discontinued, implying that sustainable generative-audio benchmarks will need cheaper reference and rating protocols.
  • The drop from 32 submissions in the 2023 edition to 4 in 2024, paired with the removal of training-data constraints, suggests that evaluator overhead and reliance on large pre-existing models shape participation as much as synthesis skill.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The strong FAD-human correlation suggests a testable two-stage benchmark design: use FAD to pre-screen many systems, then spend limited human rating effort only on the top FAD candidates.
  • The single-reference design is the fragile part of the framework, because for open-ended captions many acoustic realizations can legitimately fit the same text and would be penalized for departing from one engineer's choices.
  • The 36 percent gap might shrink or grow if the reference set were expanded to multiple engineers' recordings, and checking that sensitivity would clarify whether the gap reflects model weakness or reference idiosyncrasy.
  • The difficulty of sustaining annual human evaluation points toward automated or semi-automated proxies, but the paper's own $r=0.77$ correlation on audio quality warns that FAD alone would misrank systems when overall quality is the criterion.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper reports on DCASE 2024 Challenge Task 7, a text-to-sound generation task for sound scenes. The authors describe a standardized evaluation framework combining an objective metric, Fréchet Audio Distance (FAD) with PANN-Wavegram-Logmel embeddings, and human perceptual ratings on three scales (Foreground Fit, Background Fit, Audio Quality), aggregated into a weighted Perceptual Score. The dataset consists of 310 audio-caption pairs (60 development, 250 evaluation) created by a sound engineer, and the task constrains outputs to 4-second mono clips without music or intelligible speech. Four submitted systems and an AudioLDM baseline were evaluated. The main reported findings are a substantial gap between the sound-engineer reference and the best submitted system, and strong correlations between FAD and the subjective metrics (0.94, 0.94, 0.77). The paper also discusses the decision to discontinue the task in 2025 due to cost and shifting research scope.

Significance. If the evaluation framework is valid, it provides a reusable protocol for comparing text-to-audio models in a constrained sound-scene setting. The manuscript has concrete strengths: it specifies the prompt structure, dataset sizes, rater blinding, self-rating removal, inter-rater agreement (Cronbach's alpha = 0.959), and it releases official evaluation software. The authors also explicitly acknowledge that the small number of systems limits the strength of the FAD-human correlation evidence. However, the paper's central claim that FAD is validated as a perceptual quality measure is threatened by the use of a single-engineer reference set and by the embedding-selection history in the authors' prior work. The paper also contains a numerical inconsistency in the headline performance-gap figure. These issues are load-bearing and require revision before the framework can be considered established.

major comments (3)
  1. [§3.1, §4.1, Eq. (1)] The FAD reference set in Eq. (1) is built from audio created by a single sound engineer, apparently one recording per prompt (§3.1). For open-ended text-to-audio generation, many acoustically different recordings can legitimately satisfy the same caption, so FAD(r,g) penalizes any valid output that is far in embedding space from that engineer's specific rendition. This threatens the paper's interpretation of FAD as a perceptual quality measure and could change system rankings if a different engineer's recordings were used as the reference. The authors should test the stability of the reported correlations and rankings by constructing reference sets from multiple independent engineers (or by using a larger set of references per prompt) and should state this limitation explicitly in the paper.
  2. [§5.2, Table 1] The headline claim of a "substantial 36% performance gap" between the reference (8.793) and the best submitted system (5.832) is not supported by the numbers in Table 1: the relative gap is (8.793 − 5.832)/8.793 ≈ 33.7%, not 36%. Please either correct the percentage or explain the calculation; this number appears in the Abstract and Section 5.2 and is reported as a key result.
  3. [§5.2, §4.1] The reported FAD-human correlations (0.94, 0.94, 0.77) are computed on only five systems (four submissions plus the baseline). Moreover, the PANN-Wavegram-Logmel embedding used for FAD was chosen in the authors' prior work [8] specifically to maximize correlation with human perception. Because that prior selection is not independent of the present validation, the correlations should be interpreted with caution. The paper acknowledges the small sample but does not discuss this selection issue; please add a discussion of this limitation and, ideally, provide confidence intervals or a leave-one-system-out analysis to assess robustness.
minor comments (4)
  1. [§1] In the Introduction, "motivated by the recent advances generative models" is missing "in" after "advances"; it should read "recent advances in generative models."
  2. [§3.2] In the sentence about Room Tone 1, there is a stray space before the period and "sounds" may be intended to be part of the category name; please rephrase for clarity (e.g., "Room Tone 1 (labeled as 'Nothing') sounds").
  3. [§5.1] The paper states that 24 evaluation captions were used for subjective rating but does not specify whether the FAD scores in Table 1 and Figure 1 were computed on those 24 captions or on the full 250-caption evaluation set; please clarify this to ensure the correlation analysis is interpretable.
  4. [§5.2] The phrase "weak evidence" for the FAD-human correlation is slightly misleading: with n=5, a correlation of 0.94 is large in magnitude but statistically fragile. Consider reporting exact p-values, confidence intervals, or a permutation test to make the strength of the evidence precise.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the challenge evaluation is an independent comparison of FAD against human ratings, not a fitted prediction.

full rationale

This paper reports the results of a DCASE challenge; it contains no derivation of a predicted quantity from fitted inputs. The objective metric, FAD, is defined in Eq. (1) using an external pretrained embedding, and the reference audio sets are fixed by the challenge design (Sec. 3.1, Sec. 4.1). The human perceptual scores are collected independently by a panel of raters (Sec. 4.2), and the reported correlations (Sec. 5.2) compare two separately measured quantities. The only near-circular element is that the FAD embedding was selected in the authors' prior work [8] to maximize FAD-human correlation; however, the challenge data and ratings used in Sec. 5.2 are not shown to be the same data used for that selection, so the correlation is an out-of-sample check rather than a fitted result. The paper itself cautions that the correlation is weak evidence due to the small number of systems (Sec. 5.2). The reference-set assumption flagged by the skeptic—one sound engineer's recordings as target—is a validity limitation for open-ended text-to-audio, not a circularity, because FAD is not defined in terms of the human ratings and no result is equivalent to its own input. No load-bearing self-citation chain or uniqueness theorem is invoked. Therefore the paper is self-contained as an evaluation report and receives a score of 0.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The ledger captures the free parameters (perceptual weights and embedding choice) and domain assumptions (validity of human ratings, representativeness of the 24 captions, and single-reference ground truth). No new entities are introduced.

free parameters (2)
  • Perceptual score weights = 2, 1, 1 for foreground fit, background fit, audio quality
    The final perceptual score is defined as (2FF + BF + AQ)/4 in Eq. (2); the weights are selected by the organizers, not derived from data or a stated objective, and they affect the official ranking.
  • FAD embedding model = PANN-Wavegram-Logmel
    The FAD metric in Eq. (1) is used with PANN-Wavegram-Logmel features; the choice was made in the authors' prior work [8] to maximize perceptual correlation, so it is a hand-selected evaluation parameter that the central correlation claim depends on.
assumptions (4)
  • domain assumption Human ratings on 0-10 scales for foreground fit, background fit, and audio quality are valid ground truth for sound scene synthesis quality.
    Invoked in Section 4.2; the perceptual score is treated as the reference standard for ranking, and the FAD correlation is measured against it.
  • domain assumption FAD with PANN-Wavegram-Logmel embeddings is a valid objective proxy for perceptual quality.
    Invoked in Section 4.1; the embedding is selected based on the authors' prior study [8], and the paper's claim that FAD correlates with human perception depends on this choice.
  • domain assumption The 24 evaluation captions selected for subjective rating are representative of the 250-caption evaluation set.
    Section 5.1 states 24 captions were selected with equal representation from six foreground categories; the paper does not measure whether these captions are representative of the full evaluation set.
  • domain assumption The reference audio clips produced by one sound engineer are the correct target for each text prompt.
    Sections 3.1 and 4.1 use the engineer-crafted audio as the reference distribution for FAD and as the human benchmark; open-ended text-to-audio generation may have multiple valid renditions not captured by this single reference.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Sound Scene Synthesis at the DCASE 2024 Challenge." pith.science (2026). https://pith.science/paper/AIMB6Q4L

@misc{pith2026250108587,
  author       = {Pith},
  title        = {Pith review of: Sound Scene Synthesis at the DCASE 2024 Challenge},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AIMB6Q4L}},
  note         = {Machine review of arXiv:2501.08587}
}
read the original abstract

This paper presents Task 7 at the DCASE 2024 Challenge: sound scene synthesis. Recent advances in sound synthesis and generative models have enabled the creation of realistic and diverse audio content. We introduce a standardized evaluation framework for comparing different sound scene synthesis systems, incorporating both objective and subjective metrics. The challenge attracted four submissions, which are evaluated using the Fr\'echet Audio Distance (FAD) and human perceptual ratings. Our analysis reveals significant insights into the current capabilities and limitations of sound scene synthesis systems, while also highlighting areas for future improvement in this rapidly evolving field.

Figures

Figures reproduced from arXiv: 2501.08587 by the authors.

Figure 1
Figure 1. Correlation between FAD scores on evaluation set and other indicators, computed on the 4 submitted systems and the baseline [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SoundscapeAgent: Agentic Soundscape Construction for Controllable Synthesis and Scalable Audio-Language Supervision

    cs.SD 2026-07 conditional novelty 6.0 of 10

    An agentic pipeline that plans, retrieves/generates, and deterministically renders multi-event soundscapes, and shows those structured outputs improve audio-language model reasoning over real-only data.

Reference graph

Works this paper leans on

43 extracted references · 29 canonical work pages · cited by 1 Pith paper

  1. [8]

    ACKNOWLEDGEMENTS We thank all the raters who did the subjective evaluation: Xie Zhi- Dong, Li XinYu, Liu HaiCheng, Zou XiaoYan, Sun Yu, Hae Chun Chung, Jae Hoon Jung, Yi Yuan, Haohe Liu, Xubo Liu, Mark D. Plumbley, Wenwu Wang, Sagnik Ghosh, Gaurav Verma, Sid- dharath Narayan Shakya, Shubham Sharma, Shivesh Singh, Urszula Oszczapinska, Paige Brady, Angjeli...

  2. [1]

    INTRODUCTION This paper presents Task 7 at the DCASE 2024 Challenge: sound scene synthesis. The challenge is motivated by the recent advances generative models for the creation of realistic and diverse audio con- tent, as proposed in [1] and following the last year’s version [2]

  3. [2]

    This is a more flexible setup than the category- based generation used in the last year [2]

    PROBLEM AND TASK DEFINITION We defined the challenge as a text-to-sound generation task, where systems must generate realistic environmental audio based on tex- tual descriptions. This is a more flexible setup than the category- based generation used in the last year [2]. Each prompt follows the following structure: ”Foregroundwith Background in the backg...

  4. [3]

    Sound Scene Synthesis at the DCASE 2024 Challenge

    DATASET AND BASELINE 3.1. Dataset Creation The challenge dataset contains 310 audio-captions in total, with 60 samples designated for development and 250 for evaluation. All audio content was carefully designed by a sound engineer to match specific prompts, ensuring high-quality and consistent sound scenes. The audio samples were sourced from Freesound.or...

  5. [4]

    Objective Evaluation We employed the Fr ´echet Audio Distance (FAD) [6] with PANN- Wavegram-Logmel [7] embeddings as our primary objective metric

    EV ALUATION METHODOLOGY 4.1. Objective Evaluation We employed the Fr ´echet Audio Distance (FAD) [6] with PANN- Wavegram-Logmel [7] embeddings as our primary objective metric. The embedding was chosen to maximize the correlation between the FAD score and the human perception [8]. The FAD computation is defined as: FAD(r, g) = ∥µr − µg∥2 + Tr(Σr + Σg − 2 p...

  6. [5]

    System Performance Table 1 summarizes the evaluation results

    RESULTS 5.1. System Performance Table 1 summarizes the evaluation results. The evaluation process encompassed four submitted systems [9, 10, 11, 12] assessed by a 1Room tone is a recorded sound with no specific sound event and used to capture natural noise of a recording environment. 2https://freesound.org/ 3https://sound-effects.bbcrewind.co.uk/search 4h...

  7. [6]

    First, the generative aspect of organizing this challenge has been costly and labor intensive

    DISCONTINUATION OF THE TASK It is worth mentioning why the organizers decided not to continue the DCASE challenge in 2025 despite the successful challenges in 2023 and 2024. First, the generative aspect of organizing this challenge has been costly and labor intensive. In this year’s challenge, it took (a) about 40 hours to create and refine the evaluation...

  8. [7]

    While the submit- ted systems demonstrated promising capabilities, the significant gap between synthetic and reference audio quality indicates substantial room for improvement

    CONCLUSION The DCASE 2024 Challenge Task 7 has provided valuable insights into the current state of sound scene synthesis while highlighting several crucial areas for future development. While the submit- ted systems demonstrated promising capabilities, the significant gap between synthetic and reference audio quality indicates substantial room for improv...

Show all 43 references
  1. [9]

    A proposal for foley sound synthesis challenge,

    K. Choi, S. Oh, M. Kang, and B. McFee, “A proposal for foley sound synthesis challenge,”arXiv preprint arXiv:2207.10760, 2022

  2. [10]

    Foley sound synthesis at the dcase 2023 challenge,

    K. Choi, J. Im, L. Heller, B. Mcfee, K. Imoto, Y . Okamoto, M. Lagrange, and S. Takamichi, “Foley sound synthesis at the dcase 2023 challenge,” in 2023 Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE 2023) , 2023

  3. [11]

    AudioLDM: Text-to-audio generation with latent diffusion models,

    H. Liu, Z. Chen, Y . Yuan, X. Mei, X. Liu, D. Mandic, W. Wang, and M. D. Plumbley, “AudioLDM: Text-to-audio generation with latent diffusion models,” arXiv preprint arXiv:2301.12503, 2023

  4. [12]

    Audiocaps: Gen- erating captions for audios in the wild,

    C. D. Kim, B. Kim, H. Lee, and G. Kim, “Audiocaps: Gen- erating captions for audios in the wild,” in Proceedings of the 2019 Conference of the North American Chapter of the As- sociation for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Paper...

  5. [13]

    Audio set: An ontology and human-labeled dataset for audio events,

    J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 2017, p...

  6. [14]

    Fr ´echet audio distance: A reference-free metric for evaluating music enhancement algorithms

    K. Kilgour, M. Zuluaga, D. Roblek, and M. Sharifi, “Fr ´echet audio distance: A reference-free metric for evaluating music enhancement algorithms.” in INTERSPEECH, 2019

  7. [15]

    Panns: Large-scale pretrained audio neural net- works for audio pattern recognition,

    Q. Kong, Y . Cao, T. Iqbal, Y . Wang, W. Wang, and M. D. Plumbley, “Panns: Large-scale pretrained audio neural net- works for audio pattern recognition,” IEEE/ACM Transac- tions on Audio, Speech, and Language Processing , vol. 28, pp. 2880–2894, 2020

  8. [16]

    Correlation of fr´echet audio dis- tance with human perception of environmental audio is em- bedding dependent,

    M. Tailleur, J. Lee, M. Lagrange, K. Choi, L. M. Heller, K. Imoto, and Y . Okamoto, “Correlation of fr´echet audio dis- tance with human perception of environmental audio is em- bedding dependent,” in 2024 32nd European Signal Process- ing Conference (EUSIPCO). IEEE, 2024

  9. [17]

    Sound scene synthesis with audioldm and tango2 for dcase 2024 task7,

    X. ZhiDong, L. XinYu, L. HaiCheng, Z. XiaoYan, and S. Yu, “Sound scene synthesis with audioldm and tango2 for dcase 2024 task7,” Samsung Research China-Nanjing, Nanjing, China, Tech. Rep., July 2024

  10. [18]

    Sound scene synthesis based on gan using contrastive learning and effective time-frequency swap cross attention mechanism,

    H. C. Chung and J. H. Jung, “Sound scene synthesis based on gan using contrastive learning and effective time-frequency swap cross attention mechanism,” KT Corporation, Seoul, Re- public of Korea, Tech. Rep., July 2024

  11. [19]

    Dif- fusion based sound scene synthesis for dcase challenge 2024 task 7,

    Y . Yuan, H. Liu, X. Liu, M. D. Plumbley, and W. Wang, “Dif- fusion based sound scene synthesis for dcase challenge 2024 task 7,” University of Surrey, Guildford, United Kingdom, Tech. Rep., July 2024

  12. [20]

    Sound scene synthesis based on fine-tuned latent diffusion model for dcase challenge 2024 task 7,

    S. Ghosh, G. Verma, S. N. Shakya, S. Sharma, and S. Singh, “Sound scene synthesis based on fine-tuned latent diffusion model for dcase challenge 2024 task 7,” Indian Institute of Technology Mandi, Kamand, Mandi, India, Tech. Rep., July 2024

  13. [21]

    Challenge on sound scene synthesis: Evaluating text-to-audio generation,

    J. Lee, M. Tailleur, L. M. Heller, K. Choi, M. Lagrange, B. McFee, K. Imoto, and Y . Okamoto, “Challenge on sound scene synthesis: Evaluating text-to-audio generation,” in Au- dio Imagination: NeurIPS 2024 Workshop AI-Driven Speech, Music, and Sound Generation

  14. [22]

    T-foley: A controllable waveform-domain diffusion model for temporal-event-guided foley sound synthesis,

    Y . Chung, J. Lee, and J. Nam, “T-foley: A controllable waveform-domain diffusion model for temporal-event-guided foley sound synthesis,” in ICASSP 2024-2024 IEEE Interna- tional Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024, pp. 6820–6824

  15. [23]

    Mambafoley: Foley sound generation using selec- tive state-space models,

    M. F. Colombo, F. Ronchini, L. Comanducci, and F. An- tonacci, “Mambafoley: Foley sound generation using selec- tive state-space models,” arXiv preprint arXiv:2409.09162 , 2024

  16. [24]

    Audio generation with multiple conditional diffu- sion model,

    Z. Guo, J. Mao, R. Tao, L. Yan, K. Ouchi, H. Liu, and X. Wang, “Audio generation with multiple conditional diffu- sion model,” in Proceedings of the AAAI Conference on Arti- ficial Intelligence, vol. 38, no. 16, 2024, pp. 18 153–18 161

  17. [25]

    Picoaudio: En- abling precise timestamp and frequency controllability of audio events in text-to-audio generation,

    Z. Xie, X. Xu, Z. Wu, and M. Wu, “Picoaudio: En- abling precise timestamp and frequency controllability of audio events in text-to-audio generation,” arXiv preprint arXiv:2407.02869, 2024

  18. [26]

    Audioldm 2: Learn- ing holistic audio generation with self-supervised pretraining,

    H. Liu, Y . Yuan, X. Liu, X. Mei, Q. Kong, Q. Tian, Y . Wang, W. Wang, Y . Wang, and M. D. Plumbley, “Audioldm 2: Learn- ing holistic audio generation with self-supervised pretraining,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2024

  19. [27]

    Auffusion: Leveraging the power of diffusion and large language models for text-to- audio generation,

    J. Xue, Y . Deng, Y . Gao, and Y . Li, “Auffusion: Leveraging the power of diffusion and large language models for text-to- audio generation,” arXiv preprint arXiv:2401.01044, 2024

  20. [28]

    Ezaudio: Enhancing text-to-audio gen- eration with efficient diffusion transformer,

    J. Hai, Y . Xu, H. Zhang, C. Li, H. Wang, M. Elhi- lali, and D. Yu, “Ezaudio: Enhancing text-to-audio gen- eration with efficient diffusion transformer,” arXiv preprint arXiv:2409.10819, 2024

  21. [29]

    Fugatto 1: Foundational generative audio transformer opus 1,

    Anonymous, “Fugatto 1: Foundational generative audio transformer opus 1,” in Submitted to The Thirteenth International Conference on Learning Representations, 2024, under review. [Online]. Available: https://openreview.net/ forum?id=B2Fqu7Y2cd

  22. [30]

    Stable audio open,

    Z. Evans, J. D. Parker, C. Carr, Z. Zukowski, J. Tay- lor, and J. Pons, “Stable audio open,” arXiv preprint arXiv:2407.14358, 2024

  23. [31]

    Improving text- to-audio models with synthetic captions,

    Z. Kong, S.-g. Lee, D. Ghosal, N. Majumder, A. Mehrish, R. Valle, S. Poria, and B. Catanzaro, “Improving text- to-audio models with synthetic captions,” arXiv preprint arXiv:2406.15487, 2024

  24. [32]

    Syncfusion: Multi- modal onset-synchronized video-to-audio foley synthesis,

    M. Comunit `a, R. F. Gramaccioni, E. Postolache, E. Rodol `a, D. Comminiello, and J. D. Reiss, “Syncfusion: Multi- modal onset-synchronized video-to-audio foley synthesis,” in ICASSP 2024-2024 IEEE International Conference on Acous- tics, Speech and Signal Processing (ICASSP) ...

  25. [33]

    Sonicvisionlm: Play- ing sound with vision language models,

    Z. Xie, S. Yu, Q. He, and M. Li, “Sonicvisionlm: Play- ing sound with vision language models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 26 866–26 875

  26. [34]

    Video-foley: Two-stage video-to-sound generation via temporal event condition for fo- ley sound,

    J. Lee, J. Im, D. Kim, and J. Nam, “Video-foley: Two-stage video-to-sound generation via temporal event condition for fo- ley sound,” arXiv preprint arXiv:2408.11915, 2024

  27. [35]

    Read, watch and scream! sound generation from text and video,

    Y . Jeong, Y . Kim, S. Chun, and J. Lee, “Read, watch and scream! sound generation from text and video,”arXiv preprint arXiv:2407.05551, 2024

  28. [36]

    Movie gen: A cast of media foundation models,

    A. Polyak, A. Zohar, A. Brown, A. Tjandra, A. Sinha, A. Lee, A. Vyas, B. Shi, C.-Y . Ma, C.-Y . Chuang, et al. , “Movie gen: A cast of media foundation models,” arXiv preprint arXiv:2410.13720, 2024

  29. [37]

    Video-guided foley sound generation with multimodal controls,

    Z. Chen, P. Seetharaman, B. Russell, O. Nieto, D. Bour- gin, A. Owens, and J. Salamon, “Video-guided foley sound generation with multimodal controls,” arXiv preprint arXiv:2411.17698, 2024

  30. [38]

    Vintage: Joint video and text conditioning for holistic audio generation,

    S. S. Kushwaha and Y . Tian, “Vintage: Joint video and text conditioning for holistic audio generation,” arXiv preprint arXiv:2412.10768, 2024

  31. [39]

    Frieren: Efficient video-to-audio generation with rectified flow matching,

    Y . Wang, W. Guo, R. Huang, J. Huang, Z. Wang, F. You, R. Li, and Z. Zhao, “Frieren: Efficient video-to-audio generation with rectified flow matching,” arXiv preprint arXiv:2406.00320, 2024

  32. [40]

    Masked gener- ative video-to-audio transformers with enhanced synchronic- ity,

    S. Pascual, C. Yeh, I. Tsiamas, and J. Serr `a, “Masked gener- ative video-to-audio transformers with enhanced synchronic- ity,” inEuropean Conference on Computer Vision. Springer, 2025, pp. 247–264

  33. [41]

    Temporally aligned audio for video with autoregression,

    I. Viertola, V . Iashin, and E. Rahtu, “Temporally aligned audio for video with autoregression,” arXiv preprint arXiv:2409.13689, 2024

  34. [42]

    Gotta hear them all: Sound source aware vision to audio generation,

    W. Guo, H. Wang, W. Cai, and J. Ma, “Gotta hear them all: Sound source aware vision to audio generation,” arXiv preprint arXiv:2411.15447, 2024

  35. [43]

    Taming multimodal joint train- ing for high-quality video-to-audio synthesis,

    H. K. Cheng, M. Ishii, A. Hayakawa, T. Shibuya, A. Schwing, and Y . Mitsufuji, “Taming multimodal joint train- ing for high-quality video-to-audio synthesis,” arXiv preprint arXiv:2412.15322, 2024

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.