Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models

T0 review · 4 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read TTA-Bench evaluates text-to-audio models on seven dimensions, from accuracy to toxicity, using 2,999 prompts and 118,314 human annotations; it finds models degrade on rare, complex, or perturbed prompts and show gender bias and toxicity.

desk verdict A substantial open-sourced benchmark for text-to-audio evaluation that fills a real gap, but the OOD generalization claim is overreached and the reporting needs cleanup. read the letter →

arxiv 2509.02398 v1 pith:VVQ5WEXL submitted 2025-09-02 cs.SD eess.AS

classification cs.SDeess.AS
keywords text-to-audiogenerationevaluationbenchmarkhumanannotationgeneralizationrobustnessfairnessbiastoxicity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to replace the narrow, quality-only evaluation of text-to-audio (TTA) generation with a single benchmark that also measures reliability and social responsibility. It constructs 2,999 prompts plus an evaluation protocol combining objective metrics (CLAP, AES, real-time factor) with 118,314 human ratings from expert and general listeners, and applies them to ten open TTA models. The central finding is that current models handle familiar, simple prompts well but degrade noticeably on rare/unseen descriptions, multi-event compositional prompts, and surface-level input perturbations, while also showing measurable gender bias and a nonzero rate of toxic audio generation. If adopted as a standard, the benchmark would make robustness, fairness, and safety as trackable as fidelity, and it would push model development toward generalization and responsibility rather than in-distribution quality alone. The paper also documents that experts and general users diverge systematically on four of five subjective metrics, so how a model is rated depends on who is listening.

What carries the argument

The load-bearing object is the benchmark's construction pipeline rather than any single formula: a 50-scene taxonomy with annotated event counts and temporal relations (none, parallel, sequential, complex) for accuracy; common/rare label pools sampled from the AudioSet ontology, plus LLM-written 'never heard in the real world' scenes, for generalization; six surface perturbation types (uppercase, synonym, misspelling, whitespace, rewrite, punctuation) for robustness; gender-neutralized prompts and demographic substitution pairs for bias and fairness; and I2P-adapted plus manually written sound-level toxic prompts for toxicity. The protocol's named measures include the robustness score RSp (m

What would settle it

Check the 300 generalization prompts (and their event-label combinations) for exact or near-verbatim n-gram overlap with the training caption sets of the ten models (AudioSet, AudioCaps, Freesound, Audio-Alpaca). If overlap is found and correlates with model generalization scores—or if re-running the comparison with verified out-of-distribution prompts changes which models lead the ranking—the out-of-distribution premise and the measured generalization gaps are called into question.

Watch

Extended reading notes

Core claim

The paper claims that TTA-Bench is the first evaluation framework to treat text-to-audio models as systems that must be accurate, reliable, and socially responsible at once. It backs the claim with 2,999 prompts spanning seven dimensions—accuracy, efficiency, generalization, robustness, fairness, bias, toxicity—and a protocol pairing CLAP, AES, and latency metrics with 118,314 expert and general-user ratings. On ten models, the paper reports strong in-distribution accuracy but a drop of up to about one point of quality and alignment on rare/unseen prompts, declines as event count and relation complexity rise, model-specific gender skew in generated voices, and toxic outputs in all five categ

Load-bearing premise

The generalization measurement assumes the LLM-written 'rare or unseen' prompts and their component audio labels are truly absent from the training data of all ten benchmarked models, but the paper never checks the prompts against the corpora those models were trained on.

Editorial extensions

If this is right

  • TTA development targets can shift from in-distribution fidelity to verifiable generalization: the benchmark quantifies, per model, how much quality and alignment are lost on rare, multi-event, or perturbed prompts.
  • Robustness, fairness, and toxicity become normal, comparable metrics rather than one-off study topics, so a model release can be accompanied by a consistent reliability and safety scorecard.
  • Because expert and general raters disagree significantly on alignment, usefulness, enjoyment, and complexity (but not on raw quality), future evaluations should report both perspectives separately; a single average hides a systematic listener bias.
  • The toxicity protocol supplies an operational, sound-level definition of audio toxicity (screams, violence, sexual, shocking, illegal), usable for red-teaming and content-safety audits of generative audio systems.
  • The finding that I2P-adapted and manually written toxic prompts trigger different per-category toxicity rates implies safety claims should be built on multiple prompt sources, not one.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The generalization numbers rest on an unverified premise: that LLM-written 'rare/unseen' prompts (e.g., 'crystalline ice flute resonance') are out-of-distribution for all ten models. Searching each model's training corpora (AudioSet, AudioCaps, Freesound, Audio-Alpaca) for these phrases or their event-label components would test this; partial overlap would distort the reported generalization gap.
  • The bias dimension depends on a commercial gender-recognition API applied to generated speech, which adds an unquantified measurement error. An extension would cross-validate API labels against human perception of voice gender, and would probe other attributes such as accent or vocal age.
  • The robustness test perturbs only surface text (typos, case, synonyms). An audio-side extension—noise, compression, resampling, or adversarial waveforms—would show whether text-robust models remain robust once the perturbation moves into the acoustic domain.
  • The expert/lay divergence pattern hints at two separate subjective factors—a shared perception of raw quality and a divergent appreciation for expressiveness—testable by factor analysis or inter-rater reliability modeling on the 118,314 annotation matrix.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper introduces TTA-Bench, a text-to-audio evaluation benchmark covering seven dimensions organized under functional quality, reliability, and social responsibility: accuracy, efficiency, generalization, robustness, fairness, bias, and toxicity. The benchmark comprises 2,999 prompts constructed through dataset extraction, LLM-assisted generation, manual writing, and adaptation of image-to-prompts for audio, and it proposes a unified protocol combining objective metrics (AES, CLAP, RTF) with 118,314 human annotations from expert and general raters. Ten TTA models are benchmarked, and the paper reports dimension-wise results, concluding that current models perform well on in-distribution accuracy but struggle with generalization, robustness, bias, and toxicity. The dataset and evaluation tools are stated to be open-sourced.

Significance. If the methodological concerns are addressed, this is a valuable and potentially standard-setting resource. The scale of annotation (118,314 human judgments), the breadth across seven dimensions, the inclusion of expert and non-expert raters, and the adaptation of the I2P toxicity taxonomy to the audio domain are concrete strengths. The dataset and tools are open-sourced, which supports reproducibility and follow-up work. The paper is also honest about the limitations of existing TTA evaluation and makes a credible case that robustness, fairness, bias, and toxicity are understudied. However, the benchmark's reliability as a holistic instrument depends on two load-bearing issues: the unverified out-of-distribution status of the generalization prompts, and the absence of uncertainty quantification in most model comparisons. Both are fixable but currently weaken the central claims.

major comments (4)
  1. [Data Construction, Generalization Prompt Collection; Table 2; Figure 3/Table 6] The generalization dimension is premised on 'rare/unseen' events. The rare label pool is derived from the AudioSet ontology, and Table 2 shows that nearly every evaluated model is trained on AudioSet or AudioCaps. The paper provides no overlap check between the 300 final prompt texts and the models' training captions, no component-label membership check, and no analysis linking the measured quality/alignment drop to lexical or embedding similarity with training data. Consequently, the headline conclusion that 'models struggle to generalize beyond seen domains' is not established; the drop in Figure 3/Table 6 could be driven by LLM phrasing style or by label combinations that are novel as sentences but in-distribution at the event-label level. Please add a quantitative OOD verification (e.g., n-gram or CLAP-embedding similarity against AudioCaps and other training corpora, or a held-out e
  2. [Experimental Results, Tables 5-7; Appendix 3] The paper makes fine-grained model comparisons (e.g., Tango 2 vs AudioLDM 2 in Table 5; 'Tango 2 maintains strong performance' in Generalization Results) but reports no confidence intervals, bootstrap estimates, or significance tests for any model-level difference. Subjective scores are based on 3 expert and 10 crowd raters per clip, so system means have nontrivial uncertainty. The only significance tests in the paper (Table 12) compare expert vs non-expert preferences, not model performance. The absence of uncertainty quantification is load-bearing because many claims in the text are ranking statements, and adjacent systems often differ by fractions of a point. Please add per-system CIs and pairwise significance tests or bootstrap intervals for Tables 5-7, and temper claims where differences are within noise.
  3. [Evaluation Method, Bias; Bias Results; Figure 21] The bias analysis relies on an unnamed 'commercial system API' for gender detection, with no validation of its accuracy on synthesized audio. When AudioLDM is excluded from 75.3% of outputs and Stable Audio Open from about 40%, the MAD values in Table 7 are computed on small, non-representative subsets, yet are compared across systems as if commensurate. Figure 21 reports gender proportions without any uncertainty or sample-size information. Please identify the API or use an open-source detector, report its accuracy on TTA-generated audio, give the per-system sample sizes behind each MAD, and report bias estimates with confidence intervals.
  4. [Appendix 3, Toxicity Annotation Protocol; Table 7] Toxicity labels are produced by five crowd participants using a majority-vote procedure, but no inter-annotator agreement (e.g., Krippendorff's alpha) is reported, and the treatment of 'Uncertain' labels in the denominator of the toxicity rate is not specified. Table 7's toxicity rates are used to rank systems and to claim that 'sexual content generally has lower toxicity rates' and that 'shocking content and hate speech categories tend to have higher toxicity rates.' With roughly 30 prompt instances per category per system, differences of a few percentage points are likely within annotation noise. Please report agreement statistics, CIs for the rates, and a clear handling rule for 'Uncertain' labels.
minor comments (5)
  1. [Appendix 4, Expert vs Non-Expert Preferences] The text refers to 'Figure X' when discussing the visualization of preference gaps; this placeholder must be replaced with the actual figure number.
  2. [Figure 4] The caption says the x-axis represents 10 systems in alphabetical order, but the plot shows numeric ticks (1-10) and no model names. Readers cannot tell which bar corresponds to which model.
  3. [Throughout] Model-name capitalization is inconsistent: 'TANGO 2' appears in Table 7 and the toxicity results, while 'Tango 2' is used elsewhere; 'Stable-Audio' vs 'Stable Audio Open'; and 'Aurrfusion' appears in the Figure 21 caption.
  4. [Evaluation Method, Bias] The median absolute deviation formula is garbled in the text ('MAD = 1/Nb sum | bNb - 1/Nb |'), and the citation to Pearson (1894) does not seem to define this quantity. Please rewrite the formula and use a standard reference or define it explicitly.
  5. [Data Construction, Robustness] The uppercase perturbation includes five fixed capitalization ratios (5%, 25%, 50%, 75%, 100%). The paper does not analyze sensitivity to this choice; a sentence on whether the robustness ranking is stable across ratios would improve the discussion.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: TTA-Bench is an empirical benchmark whose reported metrics and findings are measured directly, not derived from fitted parameters or self-citations.

full rationale

TTA-Bench is not a derivation paper; it contributes new prompts, human annotations, and measurements. The central quantities (MOS scores, RTF, robustness ratio RSp, fairness pairwise difference, MAD, toxicity rate) are defined directly from human/objective data and computed without fitting any parameter from the benchmark outputs, so no reported 'prediction' is forced by construction. The generalization dimension does define its test set using the AudioSet ontology and LLM-written 'unseen' sentences (Section Data Construction: 'we construct a Common/Rare Sound Event Pool Subset based on the AudioSet category ontology ... an LLM transforms each label set into a coherent yet implausible sound scene, generating 300 prompts'), so the out-of-distribution status of those prompts is a data-validity question, not a circularity: the measured drop is not an identity, as shown by Stable Audio Open's small gap in Figure 3. The paper's self-citations (Wang et al. 2024/2023, RAMP/RAMP+, MusicEval) appear in the introduction and related work as background for evaluation practices and are not used as the premise from which any benchmark result is derived. Thus there is no circular step.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The benchmark's measurement validity rests on untested domain assumptions about rater reliability, prompt out-of-distribution status, and the gender-detection API. These are not fabricated entities but unvalidated instrumental assumptions.

free parameters (3)
  • Uppercase perturbation ratios = 5%, 25%, 50%, 75%, 100%
    Hand-chosen levels to simulate varying degrees of capitalization noise; no sensitivity analysis is provided.
  • Probe annotation validity threshold = score difference <= 2
    Used to accept or reject rater responses on 30 probe samples; chosen ad hoc without sensitivity analysis.
  • Rater deviation adjustment threshold = deviation > 4 points from mean
    Scores deviating by more than 4 points from the mean are adjusted; threshold appears arbitrary.
assumptions (4)
  • domain assumption Subjective 10-point scores validly measure the five constructs defined in the rubric (quality, complexity, enjoyment, usefulness, alignment).
    No validation against known reference audio or construct-validity study is provided beyond rater guidelines (Appendix 3).
  • domain assumption LLM-generated rare/unseen prompts are out-of-distribution for all evaluated models.
    No check of whether the prompt phrases or component labels appear in each model's training data (Generalization Prompt Collection).
  • domain assumption The commercial gender-detection API used in the bias analysis is accurate on generated audio.
    The API is unnamed and unvalidated; 5.7% to 75.3% of clips are excluded because no gender is detected (Table 7).
  • domain assumption Non-expert raters can consistently label audio toxicity from written criteria.
    No inter-rater agreement statistics are reported despite the majority-vote protocol (Toxicity Annotation Protocol).

how reviews work

0 comments
Cite this review

Pith. "Pith review of TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models." pith.science (2026). https://pith.science/paper/VVQ5WEXL

@misc{pith2026250902398,
  author       = {Pith},
  title        = {Pith review of: TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VVQ5WEXL}},
  note         = {Machine review of arXiv:2509.02398}
}
read the original abstract

Text-to-Audio (TTA) generation has made rapid progress, but current evaluation methods remain narrow, focusing mainly on perceptual quality while overlooking robustness, generalization, and ethical concerns. We present TTA-Bench, a comprehensive benchmark for evaluating TTA models across functional performance, reliability, and social responsibility. It covers seven dimensions including accuracy, robustness, fairness, and toxicity, and includes 2,999 diverse prompts generated through automated and manual methods. We introduce a unified evaluation protocol that combines objective metrics with over 118,000 human annotations from both experts and general users. Ten state-of-the-art models are benchmarked under this framework, offering detailed insights into their strengths and limitations. TTA-Bench establishes a new standard for holistic and responsible evaluation of TTA systems. The dataset and evaluation tools are open-sourced at https://nku-hlt.github.io/tta-bench/.

Figures

Figures reproduced from arXiv: 2509.02398 by the authors.

Figure 1
Figure 1. The data construction overview of the six key dimensions. Prompts corresponding to efficiency tasks are excluded, as [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. We analyze model performance from three per [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 5
Figure 5. Model’s robustness score and its performance under various perturbations. [PITH_FULL_IMAGE:figures/full_fig_p007_5.png] view at source ↗
Figures from the paper (16 more)
Figure 6
Figure 6. Figure 6: Bar chart illustrating the composition of sound [PITH_FULL_IMAGE:figures/full_fig_p010_6.png]
Figure 7
Figure 7. Figure 7: The prompt used in the dataset extraction method. [PITH_FULL_IMAGE:figures/full_fig_p012_7.png]
Figure 8
Figure 8. Figure 8: The prompt used in the LLM generation method. [PITH_FULL_IMAGE:figures/full_fig_p012_8.png]
Figure 9
Figure 9. Figure 9: Instruction to get common and rare audio labels. [PITH_FULL_IMAGE:figures/full_fig_p013_9.png]
Figure 10
Figure 10. Figure 10: Instruction to introduce unseen audio labels. [PITH_FULL_IMAGE:figures/full_fig_p013_10.png]
Figure 11
Figure 11. Figure 11: Instruction to obtain final generalization prompt from combinations of audio labels. [PITH_FULL_IMAGE:figures/full_fig_p014_11.png]
Figure 12
Figure 12. Figure 12: The prompt used in the synonym method. Rewrite the following input text to ensure that the generated text remains consistent in its core content, but has a different linguistic expression [PITH_FULL_IMAGE:figures/full_fig_p014_12.png]
Figure 13
Figure 13. Figure 13: The prompt used in the rewrite method. Based on the semantics, insert punctuation marks such as ",/./!/?" between the words of the provided sentence. For short sentences, inserting one punctuation mark is sufficient, while for long sentences, you may insert 2-3 punctu…
Figure 14
Figure 14. Figure 14: The prompt used in the punctual method [PITH_FULL_IMAGE:figures/full_fig_p014_14.png]
Figure 15
Figure 15. Figure 15: The prompt used by LLM in language fairness. [PITH_FULL_IMAGE:figures/full_fig_p016_15.png]
Figure 16
Figure 16. Figure 16: Sample LLM Interaction for I2P Adaptation [PITH_FULL_IMAGE:figures/full_fig_p016_16.png]
Figure 17
Figure 17. Figure 17: Complete LLM Interaction for Toxic Prompt Creation [PITH_FULL_IMAGE:figures/full_fig_p016_17.png]
Figure 18
Figure 18. Figure 18: Annotation UI for raters to annotate quality scores. [PITH_FULL_IMAGE:figures/full_fig_p017_18.png]
Figure 19
Figure 19. Figure 19: Annotation UI for raters to annotate toxicity. [PITH_FULL_IMAGE:figures/full_fig_p019_19.png]
Figure 20
Figure 20. Figure 20: Differences in Evaluation Preferences Between Experts and Non-Experts. [PITH_FULL_IMAGE:figures/full_fig_p019_20.png]
Figure 21
Figure 21. Figure 21: The proportion of gender across all systems. [PITH_FULL_IMAGE:figures/full_fig_p020_21.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation

    cs.CV 2025-12 conditional novelty 5.0 of 10

    T2AV-Compass is a new 500-prompt benchmark that evaluates text-to-audio-video models on quality, alignment, instruction following, and realism, exposing a persistent audio-realism bottleneck across 11 systems.

Reference graph

Works this paper leans on

48 extracted references · 36 canonical work pages · cited by 1 Pith paper

  1. [1]

    , " * write output.state after.block = add.period write newline

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al

    Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774

  4. [4]

    M.; Sun, P.; Shen, X.; Khan, F

    Bakr, E. M.; Sun, P.; Shen, X.; Khan, F. F.; Li, L. E.; and Elhoseiny, M. 2023. HRS-Bench: Holistic, Reliable and Scalable Benchmark for Text-to-Image Models. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), 19984--19996. IEEE

  5. [5]

    Barratt, S.; and Sharma, R. 2018. A note on the inception score. arXiv preprint arXiv:1801.01973

  6. [6]

    Chen, S.; Wang, C.; Wu, Y.; Zhang, Z.; Zhou, L.; Liu, S.; Chen, Z.; Liu, Y.; Wang, H.; Li, J.; He, L.; Zhao, S.; and Wei, F. 2025. Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers. IEEE Transactions on Audio, Speech and Language Processing, 33: 705--718

  7. [7]

    W.; Hou, L.; Longpre, S.; Zoph, B.; Tay, Y.; Fedus, W.; Li, Y.; Wang, X.; Dehghani, M.; Brahma, S.; et al

    Chung, H. W.; Hou, L.; Longpre, S.; Zoph, B.; Tay, Y.; Fedus, W.; Li, Y.; Wang, X.; Dehghani, M.; Brahma, S.; et al. 2024. Scaling instruction-finetuned language models. Journal of Machine Learning Research, 25(70): 1--53

  8. [8]

    Cooper, E.; Huang, W.-C.; Tsao, Y.; Wang, H.-M.; Toda, T.; and Yamagishi, J. 2024. A review on subjective and objective evaluation of synthetic speech. Acoustical Science and Technology, 45(4): 161--183

Show all 48 references
  1. [9]

    Costa-juss \`a , M.; Meglioli, M.; Andrews, P.; Dale, D.; Hansanti, P.; Kalbassi, E.; Mourachko, A.; Ropers, C.; and Wood, C. 2024. M u T ox: Universal MU ltilingual Audio-based TOX icity Dataset and Zero-shot Detector. In Ku, L.-W.; Martins, A.; and Srikumar, V., eds., Findin...

  2. [10]

    D.; Gangal, V.; Gehrmann, S.; Gupta, A.; Li, Z.; Mahamood, S.; Mahendiran, A.; Mille, S.; Srivastava, A.; Tan, S.; Wu, T.; Sohl-Dickstein, J.; Choi, J

    Dhole, K. D.; Gangal, V.; Gehrmann, S.; Gupta, A.; Li, Z.; Mahamood, S.; Mahendiran, A.; Mille, S.; Srivastava, A.; Tan, S.; Wu, T.; Sohl-Dickstein, J.; Choi, J. D.; Hovy, E.; Dusek, O.; Ruder, S.; Anand, S.; Aneja, N.; Banjade, R.; Barthe, L.; Behnke, H.; Berlot-Attwell, I.; ...

  3. [11]

    Drossos, K.; Lipping, S.; and Virtanen, T. 2020. Clotho: An audio captioning dataset. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 736--740. IEEE

  4. [12]

    Elizalde, B.; Deshmukh, S.; and Wang, H. 2023. Natural Language Supervision for General-Purpose Audio Representations. arXiv:2309.05767

  5. [13]

    Elizalde, B.; Deshmukh, S.; and Wang, H. 2024. Natural language supervision for general-purpose audio representations. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 336--340. IEEE

  6. [14]

    H.; and Pons, J

    Evans, Z.; Carr, C.; Taylor, J.; Hawley, S. H.; and Pons, J. 2024. Fast timing-conditioned latent audio diffusion. In Forty-first International Conference on Machine Learning

  7. [15]

    Ghosal, D.; Majumder, N.; Mehrish, A.; and Poria, S. 2023. Text-to-Audio Generation using Instruction Tuned LLM and Latent Diffusion Model. arXiv preprint arXiv:2304.13731

  8. [16]

    Guan, W.; Wang, K.; Zhou, W.; Wang, Y.; Deng, F.; Wang, H.; Li, L.; Hong, Q.; and Qin, Y. 2024. LAFMA: A Latent Flow Matching Model for Text-to-Audio Generation. In Interspeech 2024, 4813--4817

  9. [17]

    He, Y.; Jain, Y.; Liu, X.; Markham, A.; and Vineet, V. 2024. RiTTA: Modeling Event Relations in Text-to-Audio Generation. arXiv preprint arXiv:2412.15922

  10. [18]

    Huang, J.; Ren, Y.; Huang, R.; Yang, D.; Ye, Z.; Zhang, C.; Liu, J.; Yin, X.; Ma, Z.; and Zhao, Z. 2023 a . Make-an-audio 2: Temporal-enhanced text-to-audio generation. arXiv preprint arXiv:2305.18474

  11. [19]

    Huang, K.; Duan, C.; Sun, K.; Xie, E.; Li, Z.; and Liu, X. 2025. T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-Image Generation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 47(5): 3563--3579

  12. [20]

    Huang, R.; Huang, J.; Yang, D.; Ren, Y.; Liu, L.; Li, M.; Ye, Z.; Liu, J.; Yin, X.; and Zhao, Z. 2023 b . Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models. arXiv:2301.12661

  13. [21]

    Kilgour, K.; Zuluaga, M.; Roblek, D.; and Sharifi, M. 2019. Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms. arXiv:1812.08466

  14. [22]

    D.; Kim, B.; Lee, H.; and Kim, G

    Kim, C. D.; Kim, B.; Lee, H.; and Kim, G. 2019 a . A udio C aps: Generating Captions for Audios in The Wild. In Burstein, J.; Doran, C.; and Solorio, T., eds., Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: H...

  15. [23]

    D.; Kim, B.; Lee, H.; and Kim, G

    Kim, C. D.; Kim, B.; Lee, H.; and Kim, G. 2019 b . Audiocaps: Generating captions for audios in the wild. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short...

  16. [24]

    Kreuk, F.; Synnaeve, G.; Polyak, A.; Singer, U.; Défossez, A.; Copet, J.; Parikh, D.; Taigman, Y.; and Adi, Y. 2023. AudioGen: Textually Guided Audio Generation. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. Ope...

  17. [25]

    Kumar Nandwana, M.; He, Y.; Liu, J.; Yu, X.; Shang, C.; Du Bois, E.; McGuire, M.; and Bhat, K. 2024. Voice Toxicity Detection Using Multi-Task Learning. In ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 331--335

  18. [26]

    S.; Gupta, A.; Zhang, Y.; Narayanan, D.; Teufel, H

    Lee, T.; Yasunaga, M.; Meng, C.; Mai, Y.; Park, J. S.; Gupta, A.; Zhang, Y.; Narayanan, D.; Teufel, H. B.; Bellagente, M.; Kang, M.; Park, T.; Leskovec, J.; Zhu, J.-Y.; Fei-Fei, L.; Wu, J.; Ermon, S.; and Liang, P. 2023. Holistic Evaluation of Text-To-Image Models. arXiv:2311.04287

  19. [27]

    Li, B.; Qi, X.; Lukasiewicz, T.; and Torr, P. 2019. Controllable text-to-image generation. Advances in neural information processing systems, 32

  20. [28]

    Liu, C.; Wang, H.; Zhao, J.; Zhao, S.; Bu, H.; Xu, X.; Zhou, J.; Sun, H.; and Qin, Y. 2025. MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation. In ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Pro...

  21. [29]

    Liu, H.; Chen, Z.; Yuan, Y.; Mei, X.; Liu, X.; Mandic, D.; Wang, W.; and Plumbley, M. D. 2023. AudioLDM : Text-to-Audio Generation with Latent Diffusion Models. Proceedings of the International Conference on Machine Learning, 21450--21474

  22. [30]

    Liu, H.; Yuan, Y.; Liu, X.; Mei, X.; Kong, Q.; Tian, Q.; Wang, Y.; Wang, W.; Wang, Y.; and Plumbley, M. D. 2024. AudioLDM 2: Learning Holistic Audio Generation With Self-Supervised Pretraining. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 32: 2871--2883

  23. [31]

    Majumder, N.; Hung, C.-Y.; Ghosal, D.; Hsu, W.-N.; Mihalcea, R.; and Poria, S. 2024. Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization. In Proceedings of the 32nd ACM International Conference on Multimedia, 564--572

  24. [32]

    Meng, F.; Shao, W.; Luo, L.; Wang, Y.; Chen, Y.; Lu, Q.; Yang, Y.; Yang, T.; Zhang, K.; Qiao, Y.; and Luo, P. 2024. PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models. CoRR, abs/2406.11802

  25. [33]

    Pearson, K. 1894. Contributions to the mathematical theory of evolution. Philosophical Transactions of the Royal Society of London. A, 185: 71--110

  26. [34]

    Ramesh, A.; Pavlov, M.; Goh, G.; Gray, S.; Voss, C.; Radford, A.; Chen, M.; and Sutskever, I. 2021. Zero-shot text-to-image generation. In International conference on machine learning, 8821--8831. Pmlr

  27. [35]

    Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 a . High-Resolution Image Synthesis With Latent Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 10684--10695

  28. [36]

    Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 b . High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 10684--10695

  29. [37]

    Schramowski, P.; Brack, M.; Deiseroth, B.; and Kersting, K. 2023. Safe latent diffusion: Mitigating inappropriate degeneration in diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 22522--22531

  30. [38]

    Tjandra, A.; Wu, Y.-C.; Guo, B.; Hoffman, J.; Ellis, B.; Vyas, A.; Shi, B.; Chen, S.; Le, M.; Zacharov, N.; Wood, C.; Lee, A.; and Hsu, W.-N. 2025. Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

  31. [39]

    Wang, H.; Liu, S.; Meng, L.; Li, J.; Yang, Y.; Zhao, S.; Sun, H.; Liu, Y.; Sun, H.; Zhou, J.; et al. 2025 a . FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching. arXiv preprint arXiv:2502.11128

  32. [40]

    Wang, H.; Zhao, S.; Zheng, X.; and Qin, Y. 2023. RAMP: Retrieval-Augmented MOS Prediction via Confidence-based Dynamic Weighting. In INTERSPEECH 2023, 1095--1099

  33. [41]

    Wang, H.; Zhao, S.; Zheng, X.; Zhou, J.; Wang, X.; and Qin, Y. 2025 b . RAMP+: Retrieval-Augmented MOS Prediction With Prior Knowledge Integration. IEEE Transactions on Audio, Speech and Language Processing, 33: 1520--1534

  34. [42]

    Wang, H.; Zhao, S.; Zhou, J.; Zheng, X.; Sun, H.; Wang, X.; and Qin, Y. 2024. Uncertainty-Aware Mean Opinion Score Prediction. In Interspeech 2024, 1215--1219

  35. [43]

    Wang, H.; Zheng, X.; and Qin, Y. 2023. Intermediate-Task Learning with Pretrained Model for Synthesized Speech MOS Prediction. In 2023 IEEE International Conference on Multimedia and Expo (ICME), 378--383

  36. [44]

    Xie, Z.; Xu, X.; Wu, Z.; and Wu, M. 2025. AudioTime: A Temporally-aligned Audio-text Benchmark Dataset. In ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1--5

  37. [45]

    Xue, J.; Deng, Y.; Gao, Y.; and Li, Y. 2024. Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 32: 4700--4712

  38. [46]

    Yang, D.; Yu, J.; Wang, H.; Wang, W.; Weng, C.; Zou, Y.; and Yu, D. 2023. Diffsound: Discrete diffusion model for text-to-sound generation. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31: 1720--1733

  39. [47]

    Zeghidour, N.; Luebs, A.; Omran, A.; Skoglund, J.; and Tagliasacchi, M. 2021. Soundstream: An end-to-end neural audio codec. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30: 495--507

  40. [48]

    L.; Remez, T.; Kreuk, F.; Défossez, A.; Copet, J.; Synnaeve, G.; and Adi, Y

    Ziv, A.; Gat, I.; Lan, G. L.; Remez, T.; Kreuk, F.; Défossez, A.; Copet, J.; Synnaeve, G.; and Adi, Y. 2024. Masked Audio Generation using a Single Non-Autoregressive Transformer. arXiv:2401.04577

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.