REVIEW 4 major objections 5 minor 49 references
Simi-SFX: A similarity-based conditioning method for controllable sound effect synthesis
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A normalized Mahalanobis-distance similarity vector gives fine-grained, fader-like timbre control in a lightweight DDSP sound-effect synthesizer.
desk verdict A useful incremental contribution with a genuine new conditioning vector and released datasets, but the 'fine-grained' claim is over-sold and the controllability evidence needs a perceptual check. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the normalized Mahalanobis distance (MDN): for each sound, the Mahalanobis distance of its CLAP embedding to the mean and covariance of each class's embedding cluster, min-max normalized per channel to [0,1]. This vector is smoothed by an MLP and injected into the DDSP decoder through FiLM layers, and the conditioning layers are later fine-tuned with uniformly sampled pseudo-scores using an L1 and L2 regression loss against scores measured from the generated audio. A transient synthesizer based on the inverse discrete cosine transform of sinusoids, regularized by an L2 loss on a harmonic-percussive separated peak signal, carries the impulsive content that the noise filter alone cannot represent.
What would settle it
Re-run the interpolation test and measure the generated audio with a different pre-trained embedding model, or with human perceptual similarity ratings, instead of the metric used to compute fine-tuning losses. If the OLS $R^{2}$ falls well below the reported 0.4774 and 0.6041, the apparent timbre control is specific to the training metric rather than a general acoustic property.
Extended reading notes
Core claim
The central claim is that a normalized Mahalanobis distance (MDN) vector—the per-class Mahalanobis distance of a CLAP audio embedding, min-max scaled to [0,1]—serves as a continuous, bounded descriptor of timbre that a DDSP decoder can learn to follow. The paper demonstrates this through synthesis of footsteps and impact sounds with FAD and LSD values comparable to a larger filterbank model, and through an OLS regression showing an exponential relationship between the interpolated conditioning score and the measured normalized Mahalanobis distance of the generated audio. A fine-tuning stage trains the conditioning layers to match uniformly sampled pseudo-scores, and the authors show timbre interpolation between classes, such as board and gravel footsteps. They also report a limitation: the effective control range is concentrated near similarity scores of 0.8 to 1.0, leaving the lower range less influential.
Load-bearing premise
The load-bearing premise is that the embedding-based Mahalanobis-distance metric used to measure timbre similarity in the controllability test is independent of the metric the model was trained to match during fine-tuning, and the paper does not specify which embedding model is used for the fine-tuning measurement.
Editorial extensions
If this is right
- Timbre can be steered continuously between classes by interpolating the similarity vector, enabling sounds that mix characteristics of two categories, such as footsteps on board versus gravel.
- The guiding vector is derived from pre-trained embeddings rather than text prompts or one-hot labels, so the method extends to any dataset where class-level embedding clusters can be estimated.
- The effective control range is concentrated near similarity scores of 0.8 to 1.0, so flattening the score distribution or mapping scores to a more uniform scale could extend granularity across the full [0,1] range.
- Synthesis quality is comparable to NoisebandNet while using the smaller noise synthesizer of the original DDSP, suggesting a practical lightweight configuration for impulsive sound effects.
Reading between the lines
- If the same embedding family is used for both the fine-tuning loss and the controllability evaluation, the reported R^2 values may partly measure the model's fit to that specific metric rather than generalizable timbre control; a cross-embedding or human perceptual test would separate these.
- The similarity vector behaves like a soft multi-class membership, so sampling uniformly from the score space could synthesize novel timbres that do not correspond to any single recording in the dataset.
- The observed concentration of control in the 0.8–1.0 range suggests an asymmetric mapping; a monotonic rescaling of the score during fine-tuning, such as sampling from a beta distribution, could spread effective control more evenly.
- Because the method only needs class-level embedding statistics, it could be applied to fine-grained subcategories within a single broad class, enabling control over subtle material or surface differences beyond the two datasets introduced here.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes Simi-SFX, a DDSP-based sound effect synthesizer conditioned on a normalized Mahalanobis-distance similarity vector (MDN) computed from CLAP embeddings. An MLP plus FiLM layer projects this vector into the decoder, which drives noise and transient synthesizers. After main reconstruction training, the conditioning layers are fine-tuned by sampling uniform pseudo-scores in [0,1] and minimizing an L1+L2 loss against a similarity score measured from the generated audio. The authors introduce two Freesound-derived datasets (Footstep-set and Impact-set), report FAD/LSD reconstruction results against DDSP, NoisebandNet, and ICGAN, and evaluate controllability with OLS regression of the measured normalized Mahalanobis distance against the interpolated conditioning score. They report mean R^2 values of 0.4774 (Footstep-set) and 0.6041 (Impact-set), which they interpret as evidence of strong correlation and fine-grained timbre control, and they demonstrate a qualitative timbre interpolation example.
Significance. If the controllability claim survives scrutiny, the contribution has genuine practical merit: a compact [0,1] conditioning vector is an intuitive interface, the model is lightweight and trains in hours on a single GPU, the datasets are released on Zenodo, and the project page provides audio examples. The reconstruction comparison is a credible baseline study with multiple seeds and shows the proposed model is competitive with NoisebandNet while clearly improving over DDSP and ICGAN on the proposed datasets. However, the paper's distinctive claim of fine-grained, full-range timbre control currently rests on a regression metric whose independence from the fine-tuning objective is not demonstrated, on moderate R^2 values, and on the authors' own observation that effective control is concentrated near 0.8-1.0. The evaluation would be substantially strengthened by a fully specified third-party embedding metric or by perceptual judgments, and by an ablation that isolates the conditioning path.
major comments (4)
- [§3.3 and §4.3] The independence of the controllability evaluation is not established, and this is load-bearing for the central claim. Equation (5) is minimized using "the score measured from the generated audio," but the paper does not state which embedding model, class statistics, and normalization are used for that measurement. Section 4.3 says the evaluation deliberately switches to a PANNs-based embedding extractor "to minimize potential biases," which only makes sense if the fine-tuning measurement used a different pipeline; the reader cannot verify this. Moreover, the PANNs model used in §4.3 is itself fine-tuned to classify the model's synthesized sounds, so it is not an independent pre-trained evaluator. Please specify the exact fine-tuning measurement (embedding model, MD computation, normalization), and report R^2 with an embedding model that is neither the conditioning model nor fine-tuned on the generated test sounds.
- [§5.2 and Conclusion] The reported effect sizes do not support the abstract's "fine-grained timbre control." Mean R^2 values of 0.4774 and 0.6041, even if reliable, leave a large fraction of the variance unexplained, and the regression is performed on the log-transformed model in Eq. (6), so the reported R^2 is a fit in log space; no back-transformed or distribution-free measure is given. The authors also state that separation is concentrated in the range 0.8-1.0 and that values below 0.8 have limited influence on timbre, which is an explicit limitation on controllability. To support the headline claim, the manuscript needs either a demonstration of usable control across the whole [0,1] interval (for example, per-category monotonicity, perceptual ratings at several score values, or a remapped control curve) or a more modest claim restricted to boundary and extreme timbre separation.
- [§4.2, §4.3, §5.2] No ablation isolates the contribution of the similarity conditioning. The regression experiment tests the full model after fine-tuning, and the fine-tuning objective is precisely to make the measured MDN match the conditioning score, so the reported R^2 is at least partly a self-consistency result. The paper does not compare against the same architecture with the similarity conditioning removed, with random or shuffled conditioning vectors, or with class-label conditioning under the same decoder and loss. Such an ablation is needed to show that the MDN vector, rather than the decoder capacity or the fixed spectral centroid and loudness features (which are held constant during interpolation), is responsible for the observed timbre shifts.
- [§5.2 and §6] The qualitative timbre interpolation example does not compensate for the weak quantitative controllability evidence. Figure 7 shows spectrograms for one pair of channels but provides no listening ratings, no independent acoustic measurements, and no test of whether listeners can perceive the interpolated steps. Since the paper claims a practical fader/knob control on a [0,1] score, the authors should provide listening results or ABX discrimination at several interpolation positions, including positions below 0.8, where their own analysis indicates limited control.
minor comments (5)
- [Eq. (2)] Equation (2) should be parenthesized as (MD - MD_min)/(MD_max - MD_min); as printed, the normalization is ambiguous.
- [Eq. (3)] In Eq. (3), n is used both as the output sample index and as the summation variable for frames, which makes the equation unreadable; please use m for the frame index.
- [§3.2 and author byline] There are several typos: "of of" in the training paragraph, "Eletrical" in the affiliation, and "Y unyi Liu" in the author byline.
- [§5.2] The mean R^2 is reported only as a dataset-level number; a per-category table with standard deviations or confidence intervals would help identify which classes drive the effect and whether the result is uniform or dominated by a few categories.
- [Figures 8 and 9] Figures 8 and 9 are placed in the appendix but are the central evidence in Section 5.2; they should be integrated into the main text or referenced consistently with their appendix location.
Circularity Check
No significant circularity: the controllability evaluation deliberately switches to PANNs embeddings, making the reported R^2 a cross-model transfer check rather than a fitted-input tautology.
full rationale
The paper's central claim is that the normalized Mahalanobis-distance similarity vector gives timbre control, supported by OLS regression between the input score c and a measured normalized Mahalanobis distance y of the synthesized audio (Sec. 5.2, Eqs. 6-8). The potential circularity would arise if the same embedding-based measurement were used both to fine-tune the conditioning layers (Eq. 5) and to evaluate them. The paper explicitly breaks that loop: 'To measure the similarity of each output sound relative to each category, we opted to abandon the pre-trained CLAP model for embedding extraction to minimize potential biases. Instead, we used PANNs [38]' (Sec. 4.3). The conditioning score is defined with CLAP embeddings (Sec. 2.1), so the evaluation with PANNs is an independent, cross-model check; the moderate R^2 values (0.4774 and 0.6041) reflect partial transfer, not a construction. The fine-tuning loss in Eq. 5 does train the conditioning layers to match a measured score of the generated audio, but that is a legitimate training objective; it would become circular only if the identical measurement were reused as the evaluation metric, which the paper says it avoids. The manuscript does not state which embedder computes the measured score in Eq. 5, and this is a reproducibility gap, not demonstrated circularity. Self-citations to DDSP-SFX and ICGAN supply architecture and baselines, but the similarity-conditioning claim does not rest on those citations. Synthesis quality is benchmarked externally against DDSP, NoisebandNet, and ICGAN with FAD/LSD. No step in the claimed derivation reduces by construction to its own input.
Assumptions & free parameters
free parameters (5)
- Diagonal loading epsilon =
unspecified 'small value'
- HPSS kernel size and margin =
31, (1,3)
- Number of transient sinusoids K =
f/2 (128 at f=256)
- Frequency clipping range =
(0,128)
- Fine-tuning score sampling distribution =
Uniform [0,1]
assumptions (5)
- domain assumption CLAP embeddings encode perceptually relevant timbre information for sound effects.
- domain assumption Mahalanobis distance to class-conditional Gaussian clusters is a meaningful measure of timbre similarity.
- ad hoc to paper The PANNs-based embedding similarity after fine-tuning is an independent and valid measure of timbre for evaluating controllability.
- domain assumption HPSS with the given parameters and local peak estimation correctly identify perceptually relevant transient onsets.
- ad hoc to paper Uniform sampling of similarity scores during fine-tuning improves coverage of the control space.
Cite this review
Pith. "Pith review of Simi-SFX: A similarity-based conditioning method for controllable sound effect synthesis." pith.science (2026). https://pith.science/paper/PLO4EDSS
@misc{pith2026241218710,
author = {Pith},
title = {Pith review of: Simi-SFX: A similarity-based conditioning method for controllable sound effect synthesis},
year = {2026},
howpublished = {\url{https://pith.science/paper/PLO4EDSS}},
note = {Machine review of arXiv:2412.18710}
}
read the original abstract
Generating sound effects with controllable variations is a challenging task, traditionally addressed using sophisticated physical models that require in-depth knowledge of signal processing parameters and algorithms. In the era of generative and large language models, text has emerged as a common, human-interpretable interface for controlling sound synthesis. However, the discrete and qualitative nature of language tokens makes it difficult to capture subtle timbral variations across different sounds. In this research, we propose a novel similarity-based conditioning method for sound synthesis, leveraging differentiable digital signal processing (DDSP). This approach combines the use of latent space for learning and controlling audio timbre with an intuitive guiding vector, normalized within the range [0,1], to encode categorical acoustic information. By utilizing pre-trained audio representation models, our method achieves expressive and fine-grained timbre control. To benchmark our approach, we introduce two sound effect datasets--Footstep-set and Impact-set--designed to evaluate both controllability and sound quality. Regression analysis demonstrates that the proposed similarity score effectively controls timbre variations and enables creative applications such as timbre interpolation between discrete classes. Our work provides a robust and versatile framework for sound effect synthesis, bridging the gap between traditional signal processing and modern machine learning techniques.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
J. Schouten, “The perception of timbre,” presented at the The 6th international congress on acoustics, Tokyo, Japan, August 21-28, 1968, pp. 35–44 (1968)
work page 1968
-
[2]
D. Moffat, R. Selfridge, and J. D. Reiss, “Sound ef- fect synthesis,” in Foundations in Sound Design for Inter- active Media, pp. 274–299 (Routledge, 2019)
work page 2019
-
[3]
L. Turchet, “Footstep sounds synthesis: Design, implementation, and evaluation of foot–floor interac- tions, surface materials, shoe types, and walkers’ fea- tures,” Applied Acoustics , vol. 107, pp. 46–68 (2016), doi:https://doi.org/10.1016/j.apacoust.2015.05.013, URL https://www.sciencedirect.com/science/ article/pii/S0003682X15001747
-
[4]
Physically Based Sound Synthesis and Control of Footsteps Sounds,
L. Turchet, S. Serafin, S. Dimitrov, and R. Nor- dahl, “Physically Based Sound Synthesis and Control of Footsteps Sounds,” presented at the Proceedings of the International Conference on Digital Audio Effects (DAFx) (2010), URL https://api.semanticscholar. org/CorpusID:35113931
work page 2010
-
[5]
Caesynth: Real-Time Tim- bre Interpolation and Pitch Control with Conditional Au- toencoders,
A. V . Puche and S. Lee, “Caesynth: Real-Time Tim- bre Interpolation and Pitch Control with Conditional Au- toencoders,” presented at the2021 IEEE 31st International Workshop on Machine Learning for Signal Processing (MLSP), pp. 1–6 (2021), doi:10.1109/MLSP52302.2021. 9596414
arXiv 2021
-
[6]
Farnell, Designing Sound (The MIT Press, 2010)
A. Farnell, Designing Sound (The MIT Press, 2010)
work page 2010
-
[7]
To- wards Lightweight Controllable Audio Synthesis with Conditional Implicit Neural Representations,
J. Zuiderveld, M. Federici, and E. J. Bekkers, “To- wards Lightweight Controllable Audio Synthesis with Conditional Implicit Neural Representations,” presented at the NeurIPS 2021 Workshop on Deep Generative Mod- els and Downstream Applications (2021), URL https: //openreview.net/forum?id=-e7kA-HhM3
work page 2021
-
[8]
Flow Synthesizer: Uni- versal Audio Synthesizer Control with Normalizing Flows,
P. Esling, N. Masuda, A. Bardet, R. Despres, and A. Chemla-Romeu-Santos, “Flow Synthesizer: Uni- versal Audio Synthesizer Control with Normalizing Flows,” Applied Sciences (2019), URL https://api. semanticscholar.org/CorpusID:210905230
work page 2019
Show all 49 references
-
[9]
The control-synthesis approach for making expressive and controllable neural music synthe- sizers,
N. Jonason, “The control-synthesis approach for making expressive and controllable neural music synthe- sizers,” (2020)
2020
-
[10]
Physics-Driven Diffusion Models for Impact Sound Synthesis from Videos,
K. Su, K. Qian, E. Shlizerman, A. Torralba, and C. Gan, “Physics-Driven Diffusion Models for Impact Sound Synthesis from Videos,” presented at the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 9749–9759 (2023), doi:10.1109/ CVPR52729.2023.00940
2023
-
[11]
Conditional Sound Effects Generation with Regularized WGAN,
Y . Liu and C. Jin, “Conditional Sound Effects Generation with Regularized WGAN,” presented at the 20th Sound and Music Computing Conference, Stockholm (2023)
2023
-
[12]
Synthesis- ing Knocking Sound Effects Using Conditional Wave- GAN,
A. Barahona-Rios and S. Pauletto, “Synthesis- ing Knocking Sound Effects Using Conditional Wave- GAN,” presented at the Sound and Music Computing (2020), URL https://api.semanticscholar. org/CorpusID:229198283
2020
-
[13]
AudioLDM: Text-to-Audio Generation with Latent Diffusion Models,
H. Liu, Z. Chen, Y . Yuan, X. Mei, X. Liu, D. P. Mandic, et al. , “AudioLDM: Text-to-Audio Generation with Latent Diffusion Models,” presented at the In- ternational Conference on Machine Learning (2023), URL https://api.semanticscholar.org/ CorpusID:256390486
2023
-
[14]
Diffsound: Discrete Diffusion Model for Text-to-Sound Generation,
D. Yang, J. Yu, H. Wang, W. Wang, C. Weng, Y . Zou, et al. , “Diffsound: Discrete Diffusion Model for Text-to-Sound Generation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 31, pp. 1720–1733 (2023), doi:10.1109/TASLP.2023.3268730
2023
-
[15]
Diff- Foley: Synchronized Video-to-Audio Synthesis with La- tent Diffusion Models,
S. Luo, C. Yan, C. Hu, and H. Zhao, “Diff- Foley: Synchronized Video-to-Audio Synthesis with La- tent Diffusion Models,” presented at the Neural Informa- tion Processing Systems (2023), URL https://api. semanticscholar.org/CorpusID:259309037
2023
-
[16]
FoleyGAN: Visu- ally Guided Generative Adversarial Network-Based Syn- chronous Sound Generation in Silent Videos,
S. Ghose and J. J. Prevost, “FoleyGAN: Visu- ally Guided Generative Adversarial Network-Based Syn- chronous Sound Generation in Silent Videos,”IEEE Trans- actions on Multimedia, vol. 25, pp. 4508–4519 (2023), doi: 10.1109/TMM.2022.3177894
2023
-
[17]
DDSP: Differentiable Digital Signal Processing,
J. Engel, L. Hantrakul, C. Gu, and A. Roberts, “DDSP: Differentiable Digital Signal Processing,” Inter- national Conference on Learning Representations (2020)
2020
-
[18]
Audio Style Transfer,
E. Grinstein, N. Q. K. Duong, A. Ozerov, and P. P´erez, “Audio Style Transfer,” 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 586–590 (2017), URL https://api. semanticscholar.org/CorpusID:23919379
2017
-
[19]
Style Trans- fer of Audio Effects with Differentiable Signal Process- ing,
C. Steinmetz, N. Bryan, and J. Reiss, “Style Trans- fer of Audio Effects with Differentiable Signal Process- ing,” Journal of the Audio Engineering Society, vol. 70, pp. 708–721 (2022 11), doi:10.17743/jaes.2022.0025
2022
-
[20]
Latent Timbre Synthesis,
K. Tatar, D. Bisig, and P. Pasquier, “Latent Timbre Synthesis,” Neural Computing and Applications , vol. 33, pp. 1–18 (2021 01), doi:10.1007/s00521-020-05424-2
2021 doi
-
[21]
A Style-Based Generator Architecture for Generative Adversarial Net- works,
T. Karras, S. Laine, and T. Aila, “A Style-Based Generator Architecture for Generative Adversarial Net- works,” 2019 IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition (CVPR) , pp. 4396–4405 (2018), URL https://api.semanticscholar. org/CorpusID:54482423
2018
-
[22]
StyleWave- GAN: Style-based synthesis of drum sounds using gener- ative adversarial networks for higher audio quality,
A. Lavault, A. Roebel, and M. V oiry, “StyleWave- GAN: Style-based synthesis of drum sounds using gener- ative adversarial networks for higher audio quality,” pre- sented at the 2022 30th European Signal Processing Con- ference (EUSIPCO) , pp. 234–238 (2022), doi:10.23919/ EUS...
2022
-
[23]
Generative timbre spaces with variational audio syn- thesis,
P. Esling, A. Chemla-Romeu-Santos, and A. Bit- ton, “Generative timbre spaces with variational audio syn- thesis,” presented at the Proceedings of the International Conference on Digital Audio Effects (DAFx), pp. 175–181 (2018)
2018
-
[24]
Real-time Timbre Remapping with Differentiable DSP,
J. Shier, C. Saitis, A. Robertson, and A. McPher- son, “Real-time Timbre Remapping with Differentiable DSP,” (2024), URL https://arxiv.org/abs/ 2407.04547
2024 arXiv
-
[25]
Continuous descriptor-based control for deep audio synthesis,
N. Devis, N. Demerl ´e, S. Nabi, D. Genova, and P. Esling, “Continuous descriptor-based control for deep audio synthesis,” presented at theICASSP 2023-2023 IEEE 10 Submitted to J. Audio Eng. Soc., 2024 December DRAFT International Conference on Acoustics, Speech and Signal Pro...
2023
-
[26]
Neural Audio Synthesis of Mu- sical Notes with WaveNet Autoencoders,
J. Engel, C. Resnick, A. Roberts, S. Dieleman, M. Norouzi, D. Eck,et al., “Neural Audio Synthesis of Mu- sical Notes with WaveNet Autoencoders,” presented at the International Conference on Machine Learning (2017)
2017
-
[27]
DDSP-SFX: Acoustically-guided sound effects generation with differ- entiable digital signal processing,
Y . Liu, C. Jin, and D. Gunawan, “DDSP-SFX: Acoustically-guided sound effects generation with differ- entiable digital signal processing,” presented at the Pro- ceedings of the International Conference on Digital Audio Effects (DAFx), pp. 216–221 (2023)
2023
-
[28]
NoiseBand- Net: Controllable Time-Varying Neural Synthesis of Sound Effects Using Filterbanks,
A. Barahona-R’ios and T. Collins, “NoiseBand- Net: Controllable Time-Varying Neural Synthesis of Sound Effects Using Filterbanks,” IEEE/ACM Trans- actions on Audio, Speech, and Language Processing , vol. 32, pp. 1573–1585 (2023), URL https://api. semanticscholar.org/CorpusID:259937076
2023
-
[29]
Neu- ral Synthesis of Footsteps Sound Effects with Generative Adversarial Networks,
M. Comunit `a, H. Phan, and J. D. Reiss, “Neu- ral Synthesis of Footsteps Sound Effects with Generative Adversarial Networks,” presented at the Audio Engineer- ing Society Convention 152 (2022), URL http://www. aes.org/e-lib/browse.cfm?elib=21796, con- vention Paper 10583
2022
-
[30]
DarkGAN: Exploiting Knowledge Distillation for Comprehensible Audio Synthesis with GANs,
J. Nistal, S. Lattner, and G. Richard, “DarkGAN: Exploiting Knowledge Distillation for Comprehensible Audio Synthesis with GANs,” presented at the Interna- tional Society for Music Information Retrieval Conference (2021), URL https://api.semanticscholar. org/CorpusID:236881319
2021
-
[31]
ICGAN: An implicit condition- ing method for interpretable feature control of neural au- dio synthesis,
Y . Liu and C. Jin, “ICGAN: An implicit condition- ing method for interpretable feature control of neural au- dio synthesis,” presented at theProceedings of the Interna- tional Conference on Digital Audio Effects (DAFx)(2024)
2024
-
[32]
A. Lundberg, Data-Driven Procedural Audio: Pro- cedural Engine Sounds Using Neural Audio Synthesis , Master’s thesis, KTH School of Electrical Engineering and Computer Science (EECS) (2020)
2020
-
[33]
An analysis/synthesis tool for transient signals that allows a flexible sines+transients+noise model for audio,
T. Verma and T. Meng, “An analysis/synthesis tool for transient signals that allows a flexible sines+transients+noise model for audio,” presented at the Proceedings of the 1998 IEEE International Con- ference on Acoustics, Speech and Signal Processing, ICASSP ’98 (Cat. No.98CH...
1998
-
[34]
Serrano, A Neural Analysis–Synthesis Approach to Learning Procedural Audio Models , Master’s thesis, New Jersey Institute of Technology, Department of Com- puter Science (2022)
D. Serrano, A Neural Analysis–Synthesis Approach to Learning Procedural Audio Models , Master’s thesis, New Jersey Institute of Technology, Department of Com- puter Science (2022)
2022
-
[35]
Rigid-Body Sound Synthesis with Differen- tiable Modal Resonators,
R. Diaz, B. Hayes, C. Saitis, G. Fazekas, and M. Sandler, “Rigid-Body Sound Synthesis with Differen- tiable Modal Resonators,” presented at the ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 1–5 (2023), doi: 10.1109/I...
2023
-
[36]
Fr ´echet Audio Distance: A Reference-Free Metric for Evaluating Music Enhancement Algorithms,
K. Kilgour, M. Zuluaga, D. Roblek, and M. Shar- ifi, “Fr ´echet Audio Distance: A Reference-Free Metric for Evaluating Music Enhancement Algorithms,” pre- sented at the Interspeech (2019), URL https://api. semanticscholar.org/CorpusID:202725406
2019
-
[37]
CNN ar- chitectures for large-scale audio classification,
S. Hershey, S. Chaudhuri, D. P. W. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, et al. , “CNN ar- chitectures for large-scale audio classification,” 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 131–135 (2016), URL https://api.se...
2016
-
[38]
PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recogni- tion,
Q. Kong, Y . Cao, T. Iqbal, Y . Wang, W. Wang, and M. D. Plumbley, “PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recogni- tion,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 2880–2894 (2019), URL https://api.semanticsc...
2019
-
[39]
Large-Scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation,
Y . Wu, K. Chen, T. Zhang, Y . Hui, T. Berg- Kirkpatrick, and S. Dubnov, “Large-Scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation,” ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (I...
2022
-
[40]
On the Generalised Distance in Statistics,
P. C. Mahalanobis, “On the Generalised Distance in Statistics,” Sankhya A, vol. 80, no. 1, pp. 1–7 (2018), doi: 10.1007/s13171-019-00164-5, reprint of: Mahalanobis, P.C. (1936)
2018 doi
-
[41]
Fader Networks: Manipulat- ing Images by Sliding Attributes,
G. Lample, N. Zeghidour, N. Usunier, A. Bordes, L. Denoyer, and M. Ranzato, “Fader Networks: Manipulat- ing Images by Sliding Attributes,” presented at the Neural Information Processing Systems , vol. abs/1706.00409 (2017), URL https://api.semanticscholar. org/CorpusID:27009824
2017 arXiv
-
[42]
Sine, Transient, Noise Neural Modeling of Piano Notes,
R. Simionato and S. Fasciani, “Sine, Transient, Noise Neural Modeling of Piano Notes,” (2024), URL https://arxiv.org/abs/2409.06513
2024 arXiv
-
[43]
Extend- ing Harmonic-Percussive Separation of Audio Signals,
J. Driedger, M. M ¨uller, and S. Disch, “Extend- ing Harmonic-Percussive Separation of Audio Signals,” presented at the International Society for Music Informa- tion Retrieval Conference (2014), URL https://api. semanticscholar.org/CorpusID:9367555
2014
-
[44]
FiLM: visual reasoning with a gen- eral conditioning layer,
E. Perez, F. Strub, H. de Vries, V . Dumoulin, and A. Courville, “FiLM: visual reasoning with a gen- eral conditioning layer,” presented at the Proceedings of the Thirty-Second AAAI Conference on Artificial Intel- ligence and Thirtieth Innovative Applications of Artifi- cial I...
2018
-
[45]
Freesound tech- nical demo,
F. Font, G. Roma, and X. Serra, “Freesound tech- nical demo,” Proceedings of the 21st ACM international conference on Multimedia (2013), URL https://api. semanticscholar.org/CorpusID:28550242
2013
-
[46]
Parallel Wavegan: A Fast Waveform Generation Model Based on Generative Adversarial Networks with Multi-Resolution Spectrogram,
R. Yamamoto, E. Song, and J.-M. Kim, “Parallel Wavegan: A Fast Waveform Generation Model Based on Generative Adversarial Networks with Multi-Resolution Spectrogram,” ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing Submitted to J. Aud...
2019
-
[47]
Adam: A Method for Stochastic Optimization,
D. P. Kingma and J. Ba, “Adam: A Method for Stochastic Optimization,” presented at the Proceedings of the 3rd International Conference on Learning Representa- tions (ICLR) (2015)
2015
-
[48]
D. C. Montgomery, E. A. Peck, and G. G. Vining, Introduction to Linear Regression Analysis, 5th ed. (Wiley, Hoboken, NJ, 2012)
2012
-
[49]
Audio Set: An ontology and human-labeled dataset for audio events,
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, et al., “Audio Set: An ontology and human-labeled dataset for audio events,” presented at the Proc. IEEE ICASSP 2017 (2017). APPENDIX 12 Submitted to J. Audio Eng. Soc., 2024 December DRAFT Fig. 8...
2017
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.