Pith. sign in

REVIEW 2 major objections 6 minor 40 references

A production framework shows AudioX best balances reference identity and diversity for SFX variation, while other models suit specific edits.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 01:20 UTC pith:AG22CHGT

load-bearing objection Solid, usable evaluation protocol for reference-guided SFX variation; AudioX ranking holds inside the stated ESC-50 ATA setup, with the main soft spot already flagged by the authors. the 2 major comments →

arxiv 2607.09973 v1 pith:AG22CHGT submitted 2026-07-10 cs.SD cs.AIcs.SYeess.ASeess.SPeess.SY

A Production-Oriented Framework for Evaluation of SFX Generation

classification cs.SD cs.AIcs.SYeess.ASeess.SPeess.SY
keywords sound effects generationreference-guided variationaudio-to-audio evaluationproduction requirementsidentity preservationESC-50SFX morphingAudioX
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Industrial sound design needs generators that keep a reference clip's identity, allow controlled variation, and stay efficient, yet most audio models are scored only on text-to-audio or narrow editing tasks. This paper defines nine production requirements and a two-stage protocol: every method is lightly adapted and tested on the same ESC-50 reference-guided audio-to-audio variation task, then examined on its native strengths such as morphing, temporal control, inpainting, or targeted editing. Objective metrics (FAD, ImageBind alignment and diversity) plus a human identity study reveal complementary trade-offs rather than a single winner. Among full-generation baselines, AudioX gives the strongest joint balance of fidelity, alignment, and perceptual identity while still supporting morphing; the others remain preferable for specialized operations. The framework therefore supplies a practical decision protocol for choosing and designing industrial SFX pipelines.

Core claim

Under a shared reference-conditioned audio-to-audio variation task on ESC-50, AudioX supplies the strongest overall trade-off among full-generation baselines between reference alignment, diversity, quality (FAD), and human identity scores, while still supporting SFX morphing; the remaining baselines are most suitable for their native editing operations rather than full variation.

What carries the argument

The two-stage production-oriented evaluation framework: a common ESC-50 ATA variation task that equalizes conditioning, plus capability-specific analyses of each model's native operations, all scored against nine explicit production requirements.

Load-bearing premise

The small ESC-50 environmental soundbank, used with only light fine-tuning and class-name plus reference conditioning, is a realistic enough proxy for industrial SFX libraries and workflows.

What would settle it

Re-run the identical shared ATA protocol and human study on a larger, production-style SFX library; if another full-generation model then dominates the joint alignment-diversity-S-MOS trade-off, the AudioX claim is falsified.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Practitioners can select a model by matching production need (full variation vs. local repair vs. temporal control) instead of a single quality ranking.
  • Future unified industrial pipelines can be designed around the observed complementary strengths rather than forcing one architecture to do everything.
  • Shared ATA protocols with identity-preserving metrics become a practical benchmark for reference-guided SFX tools.
  • Capability profiles make the cost of domain shift and limited pretraining explicit, guiding when fine-tuning or hybrid systems are required.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same requirement set could be reused to stress-test real-time or video-conditioned SFX systems that the current static-clip protocol does not cover.
  • ImageBind-style embedding alignment may still miss production-critical cues such as transient sharpness that only the paper's separate onset diagnostics capture.
  • If industrial libraries contain far more out-of-manifold classes than ESC-50, the ranking among full-generation models could reverse once domain-shift penalties are larger.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. The paper proposes a production-oriented evaluation framework for reference-guided SFX variation. It defines nine production requirements (R1–R9), selects five heterogeneous baselines (AudioLDM, T-Foley, ThinkSound, AudioX, A2SB), and evaluates them with a two-stage protocol: (i) a shared reference-conditioned ATA variation task on ESC-50 after light fine-tuning (N=10 variants per reference, 4000 outputs per model), and (ii) capability-specific analyses of native operations (SFX morphing, temporal/energy control, inpainting, object-centric editing). Metrics include FAD, ImageBind alignment and diversity (with explicit formulas), fixed transient diagnostics (FWHM ratio, pre-onset Δ, onset error), and a 15-rater S-MOS study with 95% CIs. Under the shared ATA setting, AudioX is reported as the strongest full-generation trade-off (FAD 9.34, alignment 0.59, diversity 0.27, S-MOS 3.37); A2SB is correctly separated as local inpainting; other methods are positioned for specific native strengths. The framework is presented as a decision protocol rather than a single ranking.

Significance. If the result holds under the stated scope, the paper supplies a reusable, requirement-driven protocol that makes heterogeneous audio generation and editing methods comparable under a shared production objective—something currently missing when methods are only scored in their original TTA/VTA/inpainting settings. Strengths include: a large shared ATA bake-off (4000 outputs/model), off-the-shelf FAD and ImageBind metrics with explicit alignment/diversity formulas, fixed transient diagnostics, human S-MOS with CIs, explicit separation of A2SB’s local regime, Pareto-style trade-off plots, and an accompanying demo page with capability profiles. This is a useful empirical and methodological contribution for industrial SFX evaluation and for designing future unified pipelines, even though external validity is limited by ESC-50.

major comments (2)
  1. Sec. 3.3 (Dataset / Fine-Tuning Protocol) and Sec. 5: The central ranking claim is scoped to ESC-50 after deliberately light, method-specific fine-tuning. That is a reasonable few-shot proxy, but industrial SFX soundbanks differ in class structure, duration, and production constraints. The paper already flags this as a limitation; it should still state more explicitly in the abstract/conclusion that AudioX’s lead is under this ESC-50 ATA setup, and ideally add a short sensitivity note (e.g., whether the multi-metric ordering is stable under alternative fine-tuning budgets or a second small soundbank) so the production-decision protocol is not over-read as domain-general.
  2. Table 2 and R9 (Efficiency): Inference costs are reported under heterogeneous hardware, batch sizes, and sampling steps, and the text correctly calls them diagnostic rather than comparable. Because R9 is one of the nine production requirements and efficiency enters the Pareto-inspired discussion (Fig. 6b, appendix), the manuscript should either normalize more carefully (e.g., variants per GPU-hour on a common device, or FLOPs/steps) or demote R9 from cross-method ranking language so that efficiency does not appear load-bearing for the production recommendation.
minor comments (6)
  1. Table 3 footnote: A2SB metrics are computed only on inpainted regions; keep this separation equally visible in the abstract and in Fig. 2 so readers do not treat A2SB as a full-generation peer.
  2. Sec. 3.4 / Appendix: ImageBind is justified over CLAP for ATA; a short note on known ImageBind audio limitations (or a small CLAP sanity check) would strengthen the metric choice.
  3. Transient diagnostics: peak prominence ρ=0.10, 40 ms spacing, and δ_max=150 ms are fixed; state briefly that they were not tuned per method (already implied) and whether results are sensitive to small changes.
  4. Fig. 3 / Figs. 7–8: white/black boxes help, but captions could name the failure modes (smearing, high-frequency loss) more explicitly for non-specialist readers.
  5. Notation: A2SB appears as A 2SB / A2SB inconsistently; unify. Minor typos (e.g., “EV ALUATION”, “ener gy”) should be cleaned.
  6. Listening study: 15 raters × 100 trials/method is solid; report whether raters were audio professionals or general listeners, as that affects S-MOS interpretation for production use.

Circularity Check

0 steps flagged

No significant circularity: empirical bake-off of external models under shared public metrics and protocol.

full rationale

This paper proposes an evaluation framework and reports an empirical ranking of heterogeneous pretrained SFX/audio models (AudioLDM, AudioX, T-Foley, ThinkSound, A2SB) under a shared ESC-50 reference-guided ATA protocol plus capability-specific native analyses. The strongest claim—that AudioX offers the best full-generation trade-off among alignment, diversity, FAD, and S-MOS—is a measured outcome on fixed off-the-shelf metrics (FAD via AudioLDM-eval, ImageBind cosine alignment/diversity, onset-envelope transient diagnostics, and a 15-rater S-MOS study), not a quantity derived from free parameters fitted to the same target. Baselines are external released checkpoints; fine-tuning is lightweight and not used to force the ranking. There is no self-definitional loop (requirements R1–R9 are operational definitions with independent signals, not tautologies of the conclusion), no fitted-input-called-prediction, no load-bearing uniqueness theorem imported from the authors, and no renaming of a known result as a first-principles derivation. Self-citation is absent or non-load-bearing; the work is self-contained against public data and human ratings. Score 0 is therefore appropriate.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 2 invented entities

The central ranking rests on standard audio-ML metrics and a domain modeling choice that ESC-50 + light fine-tuning stands in for production SFX. Free parameters are analysis thresholds and generation counts, not fitted to force the AudioX conclusion. No new physical entities are postulated; the ‘invented’ items are the evaluation constructs themselves.

free parameters (5)
  • N variants per reference = 10
    Fixed at N=10 for all models in the shared ATA protocol; affects diversity and alignment statistics.
  • Onset peak prominence ρ and min spacing = ρ=0.10, 40 ms, 150 ms
    Hand-chosen transient detection settings (ρ=0.10, 40 ms spacing, δmax=150 ms) used for FWHM, pre-onset Δ, and onset error; fixed across methods but not derived from first principles.
  • Fine-tuning step budgets per model = model-specific (see Sec. 6.3)
    Lightweight adaptation lengths (e.g., AudioLDM 200 steps, ThinkSound 10×150, AudioX 20×200, T-Foley 25 epochs) chosen to avoid overfitting; different budgets could change relative rankings.
  • Morphing noise level σ = low/high (qualitative sweep)
    Controls transfer strength in AudioLDM/AudioX SFX morphing ablations; low vs high σ define the reported alignment–diversity trade-off.
  • A2SB mask duration regimes = 0.3–1.0 s and 0.3–2.0 s
    0.3–1.0 s vs 0.3–2.0 s masks define the inpainting diagnostics; longer masks degrade FAD/diversity by construction of the task.
axioms (5)
  • domain assumption ImageBind audio embeddings provide a valid proxy for reference identity preservation and inter-variant diversity in reference-conditioned SFX tasks.
    Sec. 3.4 and Appendix 6.2.1 justify ImageBind over CLAP because the protocol is ATA not text-driven; this is standard practice but still an unproven perceptual equivalence.
  • domain assumption ESC-50 few-shot adaptation after light fine-tuning approximates practical industrial SFX soundbank deployment.
    Sec. 3.3 Dataset; authors themselves note limited scope vs real production libraries in Sec. 5.
  • domain assumption Lightweight fine-tuning with frozen encoders preserves each model’s native editing scope enough for fair shared-task comparison.
    Sec. 3.3 Fine-Tuning Protocol; heterogeneous pretraining remains a stated limitation.
  • domain assumption FAD on cropped 4 s clips measures fidelity/realism for short SFX events.
    Sec. 3.4; FAD is distributional and the paper correctly warns it is not direct perceptual fidelity (Table 7).
  • standard math Standard cosine similarity and pairwise distance formulas on ℓ2-normalized embeddings define alignment and diversity.
    Eqs. (1)–(2) in Appendix; ordinary embedding geometry.
invented entities (2)
  • Nine production requirements R1–R9 for reference-guided SFX variation no independent evidence
    purpose: Structure baseline selection, metrics, and decision profiles for industrial use.
    Table 1 and Sec. 3.1 introduce the requirement set as the paper’s organizing construct; not independently validated outside this study.
  • Two-stage evaluation protocol (shared ATA + capability-specific native analyses) no independent evidence
    purpose: Enable comparison of heterogeneous generation/editing methods under one production objective without erasing native strengths.
    Core methodological contribution (Sec. 3.3); evidence is internal to the paper’s experiments.

pith-pipeline@v1.1.0-grok45 · 25218 in / 3525 out tokens · 36716 ms · 2026-07-14T01:20:36.512131+00:00 · methodology

0 comments
read the original abstract

Industrial sound design requires audio generation systems that not only produce realistic audio, but also preserve the perceptual identity of a reference, support controllable variation, and remain efficient for practical workflows. Existing evaluations are usually tied to text-to-audio (TTA), unconditional, or task-specific settings, limiting assessment for reference-guided sound effects (SFX) variation. To address this gap, we present a production-oriented evaluation framework for structured comparison of heterogeneous audio generation and editing methods. Our framework identifies nine production requirements and explicitly accounts for differences in model capabilities, enabling comparison under a common production objective. A two-stage protocol is introduced: (1) a reference-guided audio-to-audio (ATA) variation task, in which all methods are evaluated under the same ESC-50 SFX adaptation setup, and (2) capability-specific analyses of native operations such as SFX morphing, temporal and energy alignment, inpainting, and targeted editing. This framework combines objective metrics (including FAD, ImageBind-based reference alignment, and diversity across generated variants), together with a human study of perceptual identity preservation and transient diagnosis. Our study reveals complementary strengths and trade-offs across baselines for different production needs. Among the full-generation baselines evaluated under a shared ATA setting, AudioX provides the strongest overall trade-off between reference alignment and diversity while still supporting SFX morphing. Other baselines remain most suitable for specific editing operations. Our framework establishes a structured evaluation and decision protocol for reference-guided SFX variation and provides a practical basis for designing future unified industrial audio generation pipelines. Audio demos are on the accompanying web page.

Figures

Figures reproduced from arXiv: 2607.09973 by Eric Granger, M\'elodie Desbos, Mohammadhadi Shateri, Yara Bahram.

Figure 1
Figure 1. Figure 1: Overview of the proposed production-oriented evaluation framework for reference-guided SFX variation. A. In production settings, a reference sound often requires multiple diverse variations (Sec. 3.1); B. Existing baselines address complementary capabilities (10; 1; 5; 13; 11), including full variation generation (Sec. 3.2); C. We propose an evaluation framework to assess models’ abilities at generating va… view at source ↗
Figure 2
Figure 2. Figure 2: Diversity–identity alignment (R3-R2) trade-off across reference-guided SFX variation methods. Each point represents a model positioned according to its ability to preserve reference identity and generate diverse variations. Higher diversity and alignment are better [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Qualitative comparison for the ATA variation task on the human-sound example laughing. For A 2 SB inpainting, the masked and regenerated sections between 0.3 s and 1 s are outlined. Similar texture to reference is outlined in white. Top: mel￾spectrogram. Bottom: energy curve. localized energy changes. Overall, the qualitative analysis com￾plements the quantitative comparison by revealing method-specific tr… view at source ↗
Figure 5
Figure 5. Figure 5: visually supports these results through light but direc￾tionally consistent fluctuations. Attenuation slightly reduces the target transient regions, enhancement increases transient strength, and reverberation introduces more diffuse repeated patterns. How￾ever, the editing scope remains limited in this setting. Since the prompts are deliberately simple and the evaluated references of￾ten contain a single d… view at source ↗
Figure 4
Figure 4. Figure 4: Ablations on SFX Morphing task for AudioLDM (10) and AudioX (11) on ESC-50 (16). From left to right: the reference audio (e.g., sheep, toilet_flush, cough) and four generated samples conditioned on the target text prompt with different initialization noise levels σ (transfer strength). source timbre and spectral structure more closely, consistent with its higher alignment (0.77 vs. 0.40) and lower diversit… view at source ↗
Figure 6
Figure 6. Figure 6: Pareto-Inspired Analysis across reference-guided SFX variation methods. We claim for pareto-inspired as for b) methods may not be direct equals in inference efficiency due to different backbones, batch size and audio representations. For all evaluated values; higher is better. 6.4. Additional Results and Evaluation [PITH_FULL_IMAGE:figures/full_fig_p012_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Additional qualitative comparison of ATA variations using mel-spectrograms on three representative ESC-50 classes: crow, chirping birds, and hand_saw. Each row corresponds to one baseline and each column to one reference class. The figure high￾lights differences in local texture preservation, temporal organization, and variation behavior across methods. For A 2 SB inpainting, the masked and regenerated reg… view at source ↗
Figure 8
Figure 8. Figure 8: Additional qualitative comparison of ATA variations using energy curves for the same three ESC-50 classes as in [PITH_FULL_IMAGE:figures/full_fig_p015_8.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

40 extracted references · 1 linked inside Pith

  1. [1]

    T-FOLEY: A controllable waveform- domain diffusion model for temporal-event-guided foley sound syn- thesis,

    Y . Chung, J. Lee, and J. Nam, “T-FOLEY: A controllable waveform- domain diffusion model for temporal-event-guided foley sound syn- thesis,” inICASSP, 2024

  2. [2]

    SmartDJ: Declarative audio editing with audio language model,

    Z. Lan, Y . Hao, and M. Zhao, “SmartDJ: Declarative audio editing with audio language model,” inICLR, 2026

  3. [3]

    Ac- foley: Reference-audio-guided video-to-audio synthesis with acous- tic transfer,

    P. Fang, Y . He, Y . Xing, Q. Chen, S.-N. Lim, and H. Yang, “Ac- foley: Reference-audio-guided video-to-audio synthesis with acous- tic transfer,” inICLR, 2026

  4. [4]

    Audiomorphix: Training-free audio editing with diffusion probabilistic models,

    J. Liang, Y . Chen, Y . Yuan, D. Jia, X. Zhuang, Z. Chen, Y . Wang, and Y . Wang, “Audiomorphix: Training-free audio editing with diffusion probabilistic models,”arXiv, 2025

  5. [5]

    Thinksound: Chain-of-thought reasoning in multimodal large lan- guage models for audio generation and editing,

    H. Liu, K. Luo, J. Wang, W. Wang, Q. Chen, Z. Zhao, and W. Xue, “Thinksound: Chain-of-thought reasoning in multimodal large lan- guage models for audio generation and editing,” inNeurIPS, 2025

  6. [6]

    Stable audio open,

    Z. Evans, J. D. Parker, C. Carr, Z. Zukowski, J. Taylor, and J. Pons, “Stable audio open,”ICASSP, 2025

  7. [7]

    AudioLDM 2: Learn- ing holistic audio generation with self-supervised pretraining,

    H. Liu, Y . Yuan, X. Liu, X. Mei, Q. Kong, Q. Tian, Y . Wang, W. Wang, Y . Wang, and M. D. Plumbley, “AudioLDM 2: Learn- ing holistic audio generation with self-supervised pretraining,”ACM TASLP, 2024

  8. [8]

    AudioGen: Textually guided audio generation,

    F. Kreuk, G. Synnaeve, A. Polyak, U. Singer, A. Défossez, J. Copet, D. Parikh, Y . Taigman, and Y . Adi, “AudioGen: Textually guided audio generation,” inICLR, 2023

  9. [9]

    Edmsound: Spec- trogram based diffusion models for efficient and high-quality audio synthesis,

    G. Zhu, Y . Wen, M.-A. Carbonneau, and Z. Duan, “Edmsound: Spec- trogram based diffusion models for efficient and high-quality audio synthesis,”NeurIPS workshop Machine Learning for audio, 2023

  10. [10]

    AudioLDM: Text-to-audio generation with latent diffusion models,

    H. Liu, Z. Chen, Y . Yuan, X. Mei, X. Liu, D. Mandic, W. Wang, and M. D. Plumbley, “AudioLDM: Text-to-audio generation with latent diffusion models,” inICML, 2023

  11. [11]

    AudioX: Diffusion transformer for anything-to-audio generation,

    Z. Tian, Z. Liu, Y . Jin, R. Yuan, L. Xue, X. Tan, Q. Chen, W. Xue, and Y . Guo, “AudioX: Diffusion transformer for anything-to-audio generation,” inICLR, 2026

  12. [12]

    SoundMorpher: Perceptually- uniform sound morphing with diffusion model,

    X. Niu, J. Zhang, and C. P. Martin, “SoundMorpher: Perceptually- uniform sound morphing with diffusion model,” 2024

  13. [13]

    A2sb: Audio-to-audio schrodinger bridges,

    Z. Kong, K. J. Shih, W. Nie, A. Vahdat, S. gil Lee, J. F. Santos, A. Ju- kic, R. Valle, and B. Catanzaro, “A2sb: Audio-to-audio schrodinger bridges,”arXiv, 2025

  14. [14]

    Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models,

    S. Luo, C. Yan, C. Hu, and H. Zhao, “Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models,” inNeurIPS, 2023

  15. [15]

    Video-guided foley sound generation with multimodal controls,

    Z. Chen, P. Seetharaman, B. Russell, O. Nieto, D. Bourgin, A. Owens, and J. Salamon, “Video-guided foley sound generation with multimodal controls,” inCVPR, 2025

  16. [16]

    ESC: Dataset for environmental sound classification,

    K. J. Piczak, “ESC: Dataset for environmental sound classification,” inACM MM, 2015

  17. [17]

    Tango 2: Aligning diffusion-based text-to-audio genera- tions through direct preference optimization,

    N. Majumder, C.-Y . Hung, D. Ghosal, W.-N. Hsu, R. Mihalcea, and S. Poria, “Tango 2: Aligning diffusion-based text-to-audio genera- tions through direct preference optimization,” inACM ICM, 2024

  18. [18]

    Sketch2sound: Controllable audio generation via time-varying sig- nals and sonic imitations,

    H. F. García, O. Nieto, J. Salamon, B. Pardo, and P. Seetharaman, “Sketch2sound: Controllable audio generation via time-varying sig- nals and sonic imitations,” inICASSP, 2025

  19. [19]

    Uniform: A unified multi-task diffusion transformer for audio-video generation,

    L. Zhaoet al., “Uniform: A unified multi-task diffusion transformer for audio-video generation,”arXiv, 2025

  20. [20]

    Sam audio: Segment anything in audio,

    B. Shi, A. Tjandra, J. Hoffman, H. Wang, Y .-C. Wu, L. Gao, J. Richter, M. Le, A. Vyas, S. Chen, C. Feichtenhofer, P. Dollár, W.- N. Hsu, and A. Lee, “Sam audio: Segment anything in audio,”arXiv, 2025

  21. [21]

    MMAudio: Taming multimodal joint training for high- quality video-to-audio synthesis,

    H. K. Cheng, M. Ishii, A. Hayakawa, T. Shibuya, A. Schwing, and Y . Mitsufuji, “MMAudio: Taming multimodal joint training for high- quality video-to-audio synthesis,” inCVPR, 2025

  22. [22]

    Onoma-to-wave: Environmental sound synthesis from onomatopoeic words,

    Y . Okamoto, K. Imoto, S. Takamichi, R. Yamanishi, T. Fukumori, and Y . Yamashita, “Onoma-to-wave: Environmental sound synthesis from onomatopoeic words,”APSIPA Transactions, 2021

  23. [23]

    Ddsp-sfx: Acoustically-guided sound effects generation with differentiable digital signal process- ing,

    Y . Liu, C. Jin, and D. Gunawan, “Ddsp-sfx: Acoustically-guided sound effects generation with differentiable digital signal process- ing,”DAFx, 2023

  24. [24]

    Loop copilot: Conducting ai ensembles for music generation and iterative editing,

    Y . Zhang, A. Maezawa, G. Xia, K. Yamamoto, and S. Dixon, “Loop copilot: Conducting ai ensembles for music generation and iterative editing,” 2023

  25. [25]

    Uniaudio: An audio founda- tion model toward universal audio generation,

    D. Yang, J. Tian, X. Tan, R. Huang, S. Liu, X. Chang, J. Shi, S. Zhao, J. Bian, Z. Zhao, X. Wu, and H. Meng, “Uniaudio: An audio founda- tion model toward universal audio generation,” 2024

  26. [26]

    Prompt-guided precise audio editing with diffusion models,

    L. Zhao, L. Feng, D. Ge, R. Chen, F. Yi, C. Zhang, X.-L. Zhang, and X. Li, “Prompt-guided precise audio editing with diffusion models,” inPMLR. PMLR, 2024

  27. [27]

    AUDIT: Audio editing by following instructions with latent diffusion mod- els,

    Y . Wang, Z. Ju, X. Tan, L. He, Z. Wu, J. Bian, and S. Zhao, “AUDIT: Audio editing by following instructions with latent diffusion mod- els,” inNeurIPS, 2023

  28. [28]

    Au- dioeditor: A training-free diffusion-based audio editing framework,

    Y . Jia, Y . Chen, J. Zhao, S. Zhao, W. Zeng, Y . Chen, and Y . Qin, “Au- dioeditor: A training-free diffusion-based audio editing framework,” arXiv preprint arXiv:2409.12466, 2024

  29. [29]

    Audiobox tta- rag: Retrieval-augmented generation for zero-shot and few-shot text- to-audio,

    M. Yang, B. Shi, M. Le, W.-N. Hsu, and A. Tjandra, “Audiobox tta- rag: Retrieval-augmented generation for zero-shot and few-shot text- to-audio,”Arxiv, 2025

  30. [30]

    Dreamaudio: Few-shot customized text-to-audio generation via natural language supervision,

    Y . Yuan, X. Liu, H. Liu, X. Kang, Z. Chen, Y . Wang, M. D. Plumb- ley, and W. Wang, “Dreamaudio: Few-shot customized text-to-audio generation via natural language supervision,”arXiv, 2025

  31. [31]

    Audiocaps: Generating captions for audios in the wild,

    C. D. Kim, B. Kim, H. Lee, and G. Kim, “Audiocaps: Generating captions for audios in the wild,” inNAACL-HLT, 2019, pp. 119–132

  32. [32]

    VGGSound: A large-scale audio-visual dataset,

    H. Chen, W. Xie, A. Vedaldi, and A. Zisserman, “VGGSound: A large-scale audio-visual dataset,” inICASSP, 2020

  33. [33]

    Movie gen: A cast of media foundation mod- els,

    T. M. G. team @Meta, “Movie gen: A cast of media foundation mod- els,”CVPR, 2024

  34. [34]

    Audiotime: A temporally-aligned audio-text benchmark dataset,

    Z. Xie, X. Xu, Z. Wu, and M. Wu, “Audiotime: A temporally-aligned audio-text benchmark dataset,”ICASSP, 2025

  35. [35]

    Tta-bench: A comprehensive benchmark for evaluating text-to-audio models,

    H. Wang, C. Liu, J. Chen, H. Liu, Y . Jia, S. Zhao, J. Zhou, H. Sun, H. Bu, and Y . Qin, “Tta-bench: A comprehensive benchmark for evaluating text-to-audio models,”AAAI, 2026

  36. [36]

    H. Liu, Y . Zhang, R. Mira, Y . Zang, and I. E. Ashimine. (2023) audi- oldm_eval: Audio generation evaluation. GitHub repository

  37. [37]

    ImageBind: One embedding space to bind them all,

    R. Girdhar, A. El-Nouby, Z. Liu, M. Singh, K. V . Alwala, A. Joulin, and I. Misra, “ImageBind: One embedding space to bind them all,” inCVPR, 2023

  38. [38]

    Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners,

    Y . Xing, Y . He, Z. Tian, X. Wang, and Q. Chen, “Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners,” inCVPR, 2024

  39. [39]

    Clap learning audio concepts from natural language supervision,

    M. A. I. Benjamin Elizalde, Soham Deshmukh and H. Wang, “Clap learning audio concepts from natural language supervision,” in ICASSP, 2023. DAFx.8 Proceedings of the 29th International Conference on Digital Audio Effects (DAFx26), Cambridge, MA, USA, 1–4 September 2026

  40. [40]

    For each trial: listen to the Reference, then the Candidate. Rate the identity fidelity (1—5), defined as the similarity to the reference event/source (excluding loudness)

    APPENDIX All material described in this appendix is available on the accom- panying web page1. 6.1. Capability profile of evaluated baselines Table 5:Capability profile on production requirements for reference-guided SFX variation across representative audio gen- eration and editing methods. Color encodes each method’s suit- ability per requirement: nativ...