REVIEW 4 major objections 5 minor 1 cited by
Frequency Dynamic Convolutions for Sound Event Detection
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper claims that convolutions whose kernels adapt to the input's frequency content raise sound event detection on DESED by up to 10.98% over a baseline CRNN, and that a lighter variant matches the top score with 30% fewer parameters.
desk verdict Useful ablations around a known frequency-adaptive convolution family, but the headline PSDS1 gains are confounded with a 3-4x parameter increase and no matched-capacity static baseline. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the frequency-adaptive attention weight vector $\pi(f, x) \in \mathbb{R}^K$: a small learned subnetwork reads the input feature map, aggregates over time, and outputs, for every mel-frequency bin, $K$ weights that recombine a shared bank of $K$ basis kernels into a frequency-specific effective filter. Because the basis kernels are shared across frequencies, parameter growth stays modest; because the weights depend on the input, the same layer behaves differently for a broadband noise burst and a narrow tone. Each later variant changes how the basis kernels are organized (dilation per kernel, multiple dynamic branches, a static/dynamic channel split) or how time is aggregated (Temporal Attention Pooling combining time attention, velocity attention, and average pooling), but the input-dependent, per-frequency recombination of kernels is the mechanism that carries the whole argument.
What would settle it
Train the final MDFD or TFD model with per-frequency attention weights replaced by one shared weight across all frequency bins, keeping parameter count and the training recipe identical, and compare PSDS1 on DESED's strongly labeled real test recordings; if the gain over the baseline CRNN does not shrink, the reported improvement is not due to frequency adaptivity. Replication of the 0.455 PSDS1 under the paper's exact augmentation and post-processing settings, with standard errors across seeds, would also settle whether the 10.98% figure lies outside evaluation noise.
Extended reading notes
Core claim
The central claim is that the shift invariance built into standard 2D convolution is the wrong inductive bias for the frequency axis of audio, and that replacing it with frequency-adaptive convolution measurably improves sound event detection. In FDY conv, the convolution output at frequency bin $f$ is the attention-weighted sum of $K$ basis-kernel responses, $$y(f) = \sum_{i=1}^{K} \pi_i(f, x)\, (W_i * x + b_i),$$ where the weights $\pi_i$ depend on both the frequency bin $f$ and the input $x$, so each frequency bin receives a different effective filter matched to the content there. Extending this idea, DFD conv diversifies basis kernels by dilation, PFD conv mixes a static branch with a small dynamic branch to cut parameters, MDFD conv combines several dilated dynamic branches, and TFD conv replaces temporal average pooling with attention-based pooling over time. The dissertation reports that MDFD achieves the highest PSDS1, 0.455 versus 0.410 for the baseline CRNN (a 10.98% gain), and that TFD matches 0.455 with 12.703M parameters versus 18.157M for MDFD.
Load-bearing premise
The paper assumes the performance gain is caused by frequency adaptivity, while its comparisons also add parameters, a learned attention mechanism, and hyperparameters selected on the same validation set used for the headline numbers.
Editorial extensions
If this is right
- CRNN-style sound event detection systems can be upgraded by swapping standard 2D convolutions for FDY-type layers, with reported PSDS1 gains of 7.56% for FDY conv and up to 10.98% for MDFD conv over the baseline on DESED.
- The parameter-lean variants matter for deployment: PFD conv keeps roughly baseline-level accuracy while cutting parameters by 54.4% relative to FDY conv, and TFD conv matches the best PSDS1 at 12.703M parameters, about 30% fewer than MDFD conv.
- Different event classes favor different designs: FDY helps non-stationary events, DFD helps broad-spectrum events, PFD helps quasi-stationary events, and TFD helps transient events, so architecture choice can be guided by the target sound inventory.
- Combining the MDFD-CRNN with pretrained transformer encoders (ATST-frame plus BEATs) and change-detection-based event bounding raises the true PSDS1 to 0.577 without external pretraining data or ensembling.
- The paper's controlled comparison of 2D and 1D front-ends (PSDS1 0.410 versus 0.192) indicates that preserving the frequency axis as a spatial dimension matters more than raw model capacity in this setting.
Reading between the lines
- A natural experiment not reported in the paper is to hold capacity constant while collapsing the per-frequency attention weights to a single shared weight; that ablation would isolate how much of the 10.98% gain is frequency adaptivity itself rather than extra parameters, the attention mechanism, or architecture selection.
- The smooth, class-dependent clustering of attention weights along the frequency axis (the paper's PCA analysis) suggests FDY conv is learning an input-dependent filterbank, which connects this work to learnable front-end filter designs in speaker verification where the same mechanism could transfer.
- The class-wise specialization across variants points toward a learned router or mixture-of-experts design in which the model picks the dilated, partial, or temporal-attention adapter per event type; the paper does not explore that combination.
- Because the headline numbers come from a single validation set that also drove the architecture sweeps, multi-seed replication on a fresh test split is the prudent way to confirm the ordering of the variants before building on it.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The dissertation proposes a family of frequency-adaptive convolutions for sound event detection: Frequency Dynamic Convolution (FDY conv) and its extensions DFD, PFD, MDFD, and TFD. The core idea is to replace standard 2D convolutions with kernels that are dynamically combined using frequency-dependent attention weights, motivated by the argument that frequency-axis shift invariance is inappropriate for audio. The models are evaluated on DESED using PSDS1, with reported improvements over a CRNN baseline, parameter-efficiency claims for PFD and TFD, a class-wise analysis, and three engineering case studies. The paper claims MDFD is the best performer (PSDS1 0.455 vs 0.410 baseline) and that TFD matches it with 30% fewer parameters.
Significance. If the reported gains are real, the work is significant: it offers a systematic exploration of frequency-adaptive convolutions in a realistic SED benchmark and produces a parameter-efficient variant that matches the best accuracy. The paper also deserves credit for evaluating on the external DESED benchmark, using a standard mean-teacher pipeline, reporting parameter counts, and attempting class-wise and qualitative analyses. However, the central claim that frequency adaptivity causes the observed improvements is currently not established: the comparisons are not capacity-controlled, the headline model is selected from many configurations on the same validation set without variance estimates, and there are internal numeric inconsistencies in the baseline and parameter-efficiency numbers. These issues are load-bearing because the paper's contribution is precisely the attribution of the gains to the proposed mechanism.
major comments (4)
- [§4.2, Table 4.2 vs §3.2.4, Table 3.1] The baseline score is inconsistent across the manuscript: Table 3.1 reports a baseline PSDS1 of 0.396, while Table 4.2 (and Table 3.2) reports 0.410 for the same CRNN with 4.428M parameters. Correspondingly, Section 3.3.3 states that FDY conv improves over the baseline by 8.95% (0.434 vs 0.396), whereas the abstract and Table 4.2 imply a 7.56% improvement (0.441 vs 0.410). The headline 10.98% gain for MDFD conv is computed from 0.455 vs 0.410, so the choice of baseline directly changes the central quantitative claim. The authors should identify which configuration and post-processing setting each table refers to and make the numbers consistent.
- [Table 4.2] The main comparison is not capacity-controlled. MDFD conv uses 18.157M parameters and TFD conv uses 12.703M, versus 4.428M for the baseline CRNN, a 4.1x and 2.9x gap. No standard 2D CRNN with a matched parameter count is trained under the same mean-teacher, augmentation, and post-processing pipeline. The 7.6%-10.98% PSDS1 gains could therefore be partly or wholly due to extra model capacity rather than to frequency adaptivity. PFD conv at 5.041M/0.442 provides partial evidence, but even a static CRNN at 5M parameters is untested. The load-bearing condition for the paper's claim should be tested by training static CRNNs with matched parameter counts (for example, approximately 5M, 12M, and 18M parameters) under otherwise identical settings.
- [§3.3-§3.7 and §4.2-§4.3] All proposed modules are selected on the same DESED real validation set that is later used for the headline comparison. Tables 3.3-3.9 and 4.3-4.5 sweep dilation sets, branch proportions, channel widths, and pooling components, and the best configurations are then reported in Table 4.2. No repeated-seed experiments or variance estimates are provided, so differences of 0.004-0.007 PSDS1 (e.g., 0.448 vs 0.441, or 0.455 vs 0.451) are not shown to be above evaluation noise. The authors should report multiple runs (or confidence intervals) for at least the key comparisons, or otherwise demonstrate that the selected configurations are not artifacts of validation-set overfitting.
- [§3.5.3, §3.7.4, Table 4.2, Abstract] The parameter-efficiency claim is internally inconsistent. Section 3.5.3 states that PFD-CRNN (1/8) has 5.401M parameters and reduces parameters by 51.9% relative to FDY conv (11.061M), while Section 3.7.4 and Table 4.2 list PFD conv as 5.041M, and the abstract reports a 54.4% reduction. These numbers imply different model configurations or an arithmetic error. Since parameter efficiency is a stated contribution, this needs to be resolved with a single consistent set of model sizes and a clear definition of which proportion and channel configuration is being reported.
minor comments (5)
- [Abstract and Keywords] There are typos in the abstract and keywords: 'leadin to inconsistencies' should be 'leading to', 'translational equivariacne' should be 'translational equivariance', and 'temproal attention pooling' should be 'temporal attention pooling'.
- [§3.1, Eq. (3.1)] The text calls the property in Eq. (3.1) 'shift invariance' but the equation T(F(x)) = F(T(x)) actually defines translation equivariance. The authors should either fix the mathematical statement or adjust the terminology, since the distinction matters for the paper's motivation.
- [§3.5.3] The text refers to 'Table 3.7' when presenting PFD conv results, but the results are in Table 3.5. The cross-reference should be corrected.
- [§3.7.2, Eqs. (3.11)-(3.13)] In Eq. (3.11), the first two terms use xs,t but the third term uses xt; the subscript convention should be unified, and the dimensions of the attention weights should be stated explicitly.
- [§4.5, Table 4.8] The comparison with 1D CRNN is interesting but the 1D model is described as 'adapted from Wav2Vec2.0 and HuBERT' without a precise layer configuration; please provide the exact architecture (kernel sizes, strides, channels, pooling) so the result is reproducible.
Circularity Check
No circular derivation: all central claims are empirical DESED benchmark results; self-citations provide provenance, not load-bearing assumptions.
full rationale
The paper's central claims are measured, not derived from the definitions of the proposed modules. Table 4.2 reports PSDS1 scores for the baseline CRNN and each frequency-adaptive variant on the DESED real validation set using the official DESED evaluation toolkit (Section 4.1.8). Each module is defined by explicit equations (FDY conv in Eqs. 3.2-3.4; DFD conv in Eqs. 3.5-3.7; PFD conv in Eq. 3.9; MDFD conv in Eq. 3.10; TAP in Eqs. 3.11-3.13), and the reported gains are benchmark outcomes, not terms defined in terms of the target metric. No parameter is fitted to PSDS1 and then renamed as a prediction: the only learned quantities are network weights trained with the mean-teacher objective (Eqs. 4.1-4.10), and the evaluation is on held-out real recordings. The paper does cite the author's own prior publications for the origins of FDY/DFD/PFD/MDFD/TFD conv and for the efficient implementation trick in Section 3.3.2, but those citations are descriptive provenance: the equations and experimental protocol are reproduced inside the thesis and the results are independently measured against the external DESED benchmark. There is no uniqueness theorem imported from the authors, no ansatz smuggled in only through a citation, and no renaming of a known empirical pattern as a derivation. The main methodological risks, such as hyperparameter selection on the same validation set (Tables 3.3-3.9, 4.3-4.5) and the parameter-count gap between the baseline and MDFD conv (4.428M vs 18.157M), are experimental-control concerns rather than circularity; they do not make any equation reduce to its own input. I therefore find no circular step.
Assumptions & free parameters
free parameters (5)
- Number of basis kernels K =
4 (5 in early dilation experiments)
- Dilation size set for DFD/MDFD =
DFD best (1,1),(1,2),(1,3),(1,3); MDFD best (1)x5+(2,3)+(2,2,3)+(2,3,3)
- Dynamic branch proportion for PFD/MDFD =
1/8 for PFD; 5/8 for best TAP+PFD; 11/8 channel expansion for best MDFD
- TAP pooling component combination =
Average + Time Attention + Velocity Attention
- Post-processing parameters =
weak-prediction mask threshold (not stated); median filter window of 7 frames
assumptions (5)
- domain assumption Shifting a sound along frequency changes its perceptual meaning, so 2D convolution's frequency shift invariance is inappropriate for SED.
- domain assumption The log-mel spectrogram with 128 mel bins, window 2048, hop 256, and per-frequency normalization is a sufficient input representation.
- domain assumption Mean teacher with EMA updates and consistency losses is an effective semi-supervised training framework for DESED.
- domain assumption PSDS1 computed by the official DESED toolkit is a valid threshold-independent measure of SED quality.
- standard math Class-wise F1 comparisons are statistically meaningful under ANOVA and Tukey HSD assumptions.
invented entities (1)
-
FDY conv and variants (DFD, PFD, MDFD, TFD)
independent evidence
Cite this review
Pith. "Pith review of Frequency Dynamic Convolutions for Sound Event Detection." pith.science (2026). https://pith.science/paper/YE7XLY3Y
@misc{pith2026250612785,
author = {Pith},
title = {Pith review of: Frequency Dynamic Convolutions for Sound Event Detection},
year = {2026},
howpublished = {\url{https://pith.science/paper/YE7XLY3Y}},
note = {Machine review of arXiv:2506.12785}
}
abstract
Recent research in deep learning-based Sound Event Detection (SED) has primarily focused on Convolutional Recurrent Neural Networks (CRNNs) and Transformer models. However, conventional 2D convolution-based models assume shift invariance along both the temporal and frequency axes, leadin to inconsistencies when dealing with frequency-dependent characteristics of acoustic signals. To address this issue, this study proposes Frequency Dynamic Convolution (FDY conv), which dynamically adjusts convolutional kernels based on the frequency composition of the input signal to enhance SED performance. FDY conv constructs an optimal frequency response by adaptively weighting multiple basis kernels based on frequency-specific attention weights. Experimental results show that applying FDY conv to CRNNs improves performance on the DESED dataset by 7.56% compared to the baseline CRNN. However, FDY conv has limitations in that it combines basis kernels of the same shape across all frequencies, restricting its ability to capture diverse frequency-specific characteristics. Additionally, the $3\times3$ basis kernel size is insufficient to capture a broader frequency range. To overcome these limitations, this study introduces an extended family of FDY conv models. Dilated FDY conv (DFD conv) applies convolutional kernels with various dilation rates to expand the receptive field along the frequency axis and enhance frequency-specific feature representation. Experimental results show that DFD conv improves performance by 9.27% over the baseline. Partial FDY conv (PFD conv) addresses the high computational cost of FDY conv, which results from performing all convolution operations with dynamic kernels. Since FDY conv may introduce unnecessary adaptivity for quasi-stationary sound events, PFD conv integrates standard 2D convolutions with frequency-adaptive kernels to reduce computational complexity while maintaining performance. Experimental results demonstrate that PFD conv improves performance by 7.80% over the baseline while reducing the number of parameters by 54.4% compared to FDY conv. Multi-Dilated FDY conv (MDFD conv) extends DFD conv by addressing its structural limitation of applying the same dilation across all frequencies. By utilizing multiple convolutional kernels with different dilation rates, MDFD conv effectively captures diverse frequency-dependent patterns. Experimental results indicate that MDFD conv achieves the highest performance, improving the baseline CRNN performance by 10.98%. Furthermore, standard FDY conv employs Temporal Average Pooling, which assigns equal weight to all frames along the time axis, limiting its ability to effectively capture transient events. To overcome this, this study proposes TAP-FDY conv (TFD conv), which integrates Temporal Attention Pooling (TA) that focuses on salient features, Velocity Attention Pooling (VA) that emphasizes transient characteristics, and Average Pooling (AP) that captures stationary properties. TAP-FDY conv achieves the same performance as MDFD conv but reduces the number of parameters by approximately 30.01% (12.703M vs. 18.157M), achieving equivalent accuracy with lower computational complexity. Class-wise performance analysis reveals that FDY conv improves detection of non-stationary events, DFD conv is particularly effective for events with broad spectral features, and PFD conv enhances the detection of quasi-stationary events. Additionally, TFD conv (TFD-CRNN) demonstrates strong performance in detecting transient events. In the case studies, PFD conv effectively captures stable signal patterns in tank powertrain fault recognition, DFD conv recognizes wide harmonic spectral patterns on speed-varying motor fault recognition, while TFD conv outperforms other models in detecting transient signals in offshore arc detection. These results suggest that frequency-adaptive convolutions and their extended variants provide a robust alternative to conventional 2D convolutions in deep learning-based audio processing.
Figures
Figures from the paper (16 more)
Forward citations
Cited by 1 Pith paper
-
Auditory Intelligence: Understanding the World Through Sound
A position paper proposing four cognitively inspired audio tasks (ASPIRE, SODA, AUX, AUGMENT) to push machine hearing beyond recognition toward explanation, reasoning, and interaction.
Reference graph
Works this paper leans on
-
[1]
Virtanen, M
T. Virtanen, M. D. Plumbley, and D. Ellis, Computational Analysis of Sound Scenes and Events , Springer Publishing Company, Incorporated, 1st ed., 2017, pp. 3-11, 71-77, ISBN: 3319634496
2017
-
[2]
Sound event detection in domestic environ- ments with weakly labeled data and soundscape synthesis,
N. Turpault, R. Serizel, A. P. Shah, and J. Salamon, “Sound event detection in domestic environ- ments with weakly labeled data and soundscape synthesis,” in DCASE Workshop , 2019
2019
-
[3]
Convolutional Recurrent Neural Networks for Polyphonic Sound Event Detection,
E. C ¸ akır, G. Parascandolo, T. Heittola, H. Huttunen, and T. Virtanen, “Convolutional Recurrent Neural Networks for Polyphonic Sound Event Detection,” IEEE/ACM Transac- tions on Audio, Speech, and Language Processing , vol. 25, no. 6, pp. 1291-1303, 2017, doi:10.1109/TASLP.2017.2690575
arXiv 2017
-
[4]
Metrics for Polyphonic Sound Event Detection,
A. Mesaros, T. Heittola, and T. Virtanen, “Metrics for Polyphonic Sound Event Detection,” Applied Sciences, vol. 6, no. 6, article 162, 2016, doi:10.3390/app6060162
-
[6]
Scaper: A library for soundscape synthesis and augmentation
Salamon, J., MacConnell, D., Cartwright, M., Li, P. and Bello, J. “Scaper: A library for soundscape synthesis and augmentation”, 2017 IEEE Workshop On Applications Of Signal Processing To Audio And Acoustics (WASPAA), 2017
2017
-
[7]
Study on Frequency Dependent Convolution Methods for Sound Event Detection,
H. Nam, S.-H. Kim, B.-Y. Ko, D. Min, and Y.-H. Park, “Study on Frequency Dependent Convolution Methods for Sound Event Detection,” in Proc. INTER-NOISE, 2024
2024
-
[8]
SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition,
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition,” in Proc. Interspeech, 2019
2019
-
[9]
Conformer: Convolution-augmented Transformer for Speech Recognition,
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented Transformer for Speech Recognition,” in Proc. Interspeech, 2020
2020
Show all 145 references
-
[10]
Coherence-Based Phonemic Analysis on the Effect of Reverberation to Practical Automatic Speech Recognition,
H. Nam and Y.-H. Park, “Coherence-Based Phonemic Analysis on the Effect of Reverberation to Practical Automatic Speech Recognition,” Applied Acoustics, vol. 227, p. 110233, 2025
2025
-
[11]
wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,” in Advances in Neural Information Processing Systems, 2020
2020
-
[12]
Hu- BERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hu- BERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2021
2021
-
[13]
Attentive Statistics Pooling for Deep Speaker Embed- ding,
K. Okabe, T. Koshinaka, and K. Shinoda, “Attentive Statistics Pooling for Deep Speaker Embed- ding,” in Proc. Interspeech, 2018. 87
2018
-
[14]
Exploring the Encoding Layer and Loss Function in End-to-End Speaker and Language Recognition System,
W. Cai, J. Chen, and M. Li, “Exploring the Encoding Layer and Loss Function in End-to-End Speaker and Language Recognition System,” in Proc. Interspeech, 2018
2018
-
[15]
Analysis-Based Optimization of Temporal Dynamic Convo- lutional Neural Network for Text-Independent Speaker Verification,
S.-H. Kim, H. Nam, and Y.-H. Park, “Analysis-Based Optimization of Temporal Dynamic Convo- lutional Neural Network for Text-Independent Speaker Verification,” IEEE Access, vol. 11, 2023
2023
-
[16]
Integrating Frequency Translational Invariance in TDNNs and Frequency Positional Information in 2D ResNets to Enhance Speaker Verification,
J. Thienpondt, B. Desplanques, and K. Demuynck, “Integrating Frequency Translational Invariance in TDNNs and Frequency Positional Information in 2D ResNets to Enhance Speaker Verification,” in Proc. Interspeech, 2021
2021
-
[17]
Convolution-Based Channel-Frequency Attention for Text-Independent Speaker Verification,
J. Li, Y. Tian, and T. Lee, “Convolution-Based Channel-Frequency Attention for Text-Independent Speaker Verification,” in ICASSP, 2023, doi:10.1109/ICASSP49357.2023.10095415
2023
-
[18]
PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition,
Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumbley, “PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2020
2020
-
[19]
Deep Learning Based Cough Detection Camera Using Enhanced Features,
G.-T. Lee, H. Nam, S.-H. Kim, S.-M. Choi, Y. Kim, and Y.-H. Park, “Deep Learning Based Cough Detection Camera Using Enhanced Features,” Expert Systems with Applications , vol. 206, 2022, doi:10.1016/j.eswa.2022.117811
2022
-
[20]
Real-Time Sound Recognition System for Human Care Robot Considering Custom Sound Events,
S.-H. Kim, H. Nam, S.-M. Choi, and Y.-H. Park, “Real-Time Sound Recognition System for Human Care Robot Considering Custom Sound Events,” IEEE Access, vol. 12, 2024
2024
-
[21]
AST: Audio Spectrogram Transformer,
Y. Gong, Y.-A. Chung, and J. Glass, “AST: Audio Spectrogram Transformer,” in Proc. Interspeech, 2021
2021
-
[22]
BEATs: Audio Pre-Training with Acoustic Tokenizers,
S. Chen, Y. Wu, C. Wang, S. Liu, D. Tompkins, Z. Chen, W. Che, X. Yu, and F. Wei, “BEATs: Audio Pre-Training with Acoustic Tokenizers,” in ICML, 2023
2023
-
[23]
Overview and Evaluation of Sound Event Localization and Detection in DCASE 2019,
A. Politis, A. Mesaros, S. Adavanne, T. Heittola, and T. Virtanen, “Overview and Evaluation of Sound Event Localization and Detection in DCASE 2019,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, 2020
2019
-
[24]
STARSS22: A Dataset of Spatial Recordings of Real Scenes with Spatiotemporal Annotations of Sound Events,
A. Politis, K. Shimada, P. Sudarsanam, S. Adavanne, D. Krause, Y. Koyama, N. Takahashi, S. Takahashi, Y. Mitsufuji, and T. Virtanen, “STARSS22: A Dataset of Spatial Recordings of Real Scenes with Spatiotemporal Annotations of Sound Events,” in DCASE Workshop , 2022
2022
-
[25]
Data Augmentation and Squeeze-and-Excitation Network on Multiple Dimension for Sound Event Localization and Detection in Real Scenes,
B.-Y. Ko, H. Nam, S.-H. Kim, D. Min, S.-D. Choi, and Y.-H. Park, “Data Augmentation and Squeeze-and-Excitation Network on Multiple Dimension for Sound Event Localization and Detection in Real Scenes,” DCASE Challenge , 2022
2022
-
[26]
Automated Audio Captioning with Recurrent Neural Networks,
K. Drossos, S. Adavanne, and T. Virtanen, “Automated Audio Captioning with Recurrent Neural Networks,” in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics , 2017
2017
-
[27]
Clotho: An Audio Captioning Dataset,
K. Drossos, S. Lipping, and T. Virtanen, “Clotho: An Audio Captioning Dataset,” in ICASSP, 2020
2020
-
[28]
ChatGPT Caption Paraphrasing and FENSE-Based Caption Filtering for Automated Audio Captioning,
I. Choi, H. Nam, D. Min, S.-D. Choi, and Y.-H. Park, “ChatGPT Caption Paraphrasing and FENSE-Based Caption Filtering for Automated Audio Captioning,” DCASE Challenge , 2024. 88
2024
-
[29]
Mind the Domain Gap: A Systematic Analysis on Bioacoustic Sound Event Detection,
J. Liang, I. Nolasco, B. Ghani, H. Phan, E. Benetos, and D. Stowell, “Mind the Domain Gap: A Systematic Analysis on Bioacoustic Sound Event Detection,” arXiv preprint arXiv:2403.18638 , 2024
2024 arXiv
-
[30]
Few-Shot Bioacoustic Event Detection Utilizing Spectro- Temporal Receptive Field,
D. Min, H. Nam, and Y.-H. Park, “Few-Shot Bioacoustic Event Detection Utilizing Spectro- Temporal Receptive Field,” in Proc. INTER-NOISE, 2024
2024
-
[31]
PRTFNet: HRTF Individualization for Accurate Spectral Cues Using a Compact PRTF,
B.-Y. Ko, G.-T. Lee, H. Nam, and Y.-H. Park, “PRTFNet: HRTF Individualization for Accurate Spectral Cues Using a Compact PRTF,” IEEE Access, vol. 11, 2023
2023
-
[32]
FilterAugment: An Acoustic Environmental Data Augmentation Method,
B.-Y. Ko, Y.-H. Park, G.-T. Lee, and H. Nam, “FilterAugment: An Acoustic Environmental Data Augmentation Method,” in International Congress on Acoustics (ICA) , 2022
2022
-
[33]
AudioLDM: Text-to-Audio Generation with Latent Diffusion Models,
H. Liu, Z. Chen, Y. Yuan, X. Mei, X. Liu, D. Mandic, W. Wang, and M. D. Plumbley, “AudioLDM: Text-to-Audio Generation with Latent Diffusion Models,” in ICML, 2023
2023
-
[34]
AudioGen: Textually Guided Audio Generation,
F. Kreuk, G. Synnaeve, A. Polyak, U. Singer, A. D´ efossez, J. Copet, D. Parikh, Y. Taigman, and Y. Adi, “AudioGen: Textually Guided Audio Generation,” in International Conference on Learning Representations (ICLR), 2023
2023
-
[35]
VIFS: An End-to-End Variational Inference for Foley Sound Synthesis,
J. Lee, H. Nam, and Y.-H. Park, “VIFS: An End-to-End Variational Inference for Foley Sound Synthesis,” DCASE Challenge , 2023
2023
-
[36]
mixup: Beyond Empirical Risk Minimiza- tion,
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz, “mixup: Beyond Empirical Risk Minimiza- tion,” in International Conference on Learning Representations (ICLR) , 2018
2018
-
[37]
Training sound event detection on a heterogeneous dataset,
N. Turpault and R. Serizel, “Training sound event detection on a heterogeneous dataset,” inDCASE Workshop, 2020
2020
-
[38]
Analysis of weak labels for sound event tagging,
N. Turpault, R. Serizel, and E. Vincent, “Analysis of weak labels for sound event tagging,” hal- 03203692, 2021
2021
-
[39]
Dynamic Convolution: Attention Over Convolution Kernels,
Y. Chen, X. Dai, M. Liu, D. Chen, L. Yuan, and Z. Liu, “Dynamic Convolution: Attention Over Convolution Kernels,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020
2020
-
[40]
Attention is All you Need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polo- sukhin, “Attention is All you Need,” in Advances in Neural Information Processing Systems , 2017
2017
-
[42]
Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results,
A. Tarvainen and H. Valpola, “Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results,” in Advances in Neural Information Process- ing Systems , vol. 30, 2017
2017
-
[43]
Multi-Dimensional Frequency Dynamic Convolution with Con- fident Mean Teacher for Sound Event Detection,
S. Xiao, X. Zhang, and P. Zhang, “Multi-Dimensional Frequency Dynamic Convolution with Con- fident Mean Teacher for Sound Event Detection,” in ICASSP, 2023
2023
-
[44]
AST-SED: An Effective Sound Event Detection Method Based on Audio Spectrogram Transformer,
K. Li, Y. Song, L.-R. Dai, I. McLoughlin, X. Fang, and L. Liu, “AST-SED: An Effective Sound Event Detection Method Based on Audio Spectrogram Transformer,” in ICASSP, 2023. 89
2023
-
[45]
Fine-Tune the Pretrained ATST Model for Sound Event Detection,
N. Shao, X. Li, and X. Li, “Fine-Tune the Pretrained ATST Model for Sound Event Detection,” in ICASSP, 2024
2024
-
[46]
Squeeze-and-Excitation Networks,
J. Hu, L. Shen, and G. Sun, “Squeeze-and-Excitation Networks,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2018
2018
-
[47]
Learnable Sparse Filterbank for Speaker Verification,
J. Peng, R. Gu, L. Moˇ sner, O. Plchot, L. Burget, and J. ˇCernock´ y, “Learnable Sparse Filterbank for Speaker Verification,” in Proc. Interspeech, 2022
2022
-
[48]
Selective Kernel Networks,
X. Li, W. Wang, X. Hu, and J. Yang, “Selective Kernel Networks,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2019
2019
-
[49]
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” in International Conference on Learning...
2021
-
[50]
Audio Set: An Ontology and Human-Labeled Dataset for Audio Events,
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio Set: An Ontology and Human-Labeled Dataset for Audio Events,” in ICASSP, 2017
2017
-
[51]
Convolution- Augmented Transformer for Semi-Supervised Sound Event Detection,
K. Miyazaki, T. Komatsu, T. Hayashi, S. Watanabe, T. Toda, and K. Takeda, “Convolution- Augmented Transformer for Semi-Supervised Sound Event Detection,” DCASE Challenge , 2020
2020
-
[52]
Zheng USTC Team’s Submission for DCASE2021 Task4 – Semi- Supervised Sound Event Detection,
X. Zheng, H. Chen, and Y. Song, “Zheng USTC Team’s Submission for DCASE2021 Task4 – Semi- Supervised Sound Event Detection,” DCASE Challenge , 2021
2021
-
[53]
Pre-Training and Self-Training for Sound Event Detection in Domestic Environments,
J. E. and R. Haeb-Umbach, “Pre-Training and Self-Training for Sound Event Detection in Domestic Environments,” DCASE Challenge , 2022
2022
-
[54]
Semi-Supervised Sound Event Detection System for DCASE 2022 Task 4,
K. He, X. Shu, S. Jia, and Y. He, “Semi-Supervised Sound Event Detection System for DCASE 2022 Task 4,” DCASE Challenge , 2022
2022
-
[55]
Data Engineering for Noisy Student Model in Sound Event Detection,
S. Suh and D. Y. Lee, “Data Engineering for Noisy Student Model in Sound Event Detection,” DCASE Challenge , 2022
2022
-
[56]
Pretrained Models in Sound Event Detection for DCASE 2022 Challenge Task4,
S. Xiao, “Pretrained Models in Sound Event Detection for DCASE 2022 Challenge Task4,” DCASE Challenge, 2022
2022
-
[57]
Semi-Supervised Learning- Based Sound Event Detection Using Frequency Dynamic Convolution with Large Kernel Attention for DCASE Challenge 2023 Task 4,
J. W. Kim, S. W. Son, Y. Song, H. K. Kim, I. H. Song, and J. E. Lim, “Semi-Supervised Learning- Based Sound Event Detection Using Frequency Dynamic Convolution with Large Kernel Attention for DCASE Challenge 2023 Task 4,” DCASE Challenge , 2023
2023
-
[58]
How Information on Soft Labels and Hard Labels Mutually Benefits Sound Event Detection Tasks,
H. Yin, J. Bai, S. Huang, and J. Chen, “How Information on Soft Labels and Hard Labels Mutually Benefits Sound Event Detection Tasks,” DCASE Challenge , 2023
2023
-
[59]
Sound Event Detection with Weak Prediction for DCASE 2023 Challenge Task4A,
S. Xiao, J. Shen, A. Hu, X. Zhang, P. Zhang, and Y. Yan, “Sound Event Detection with Weak Prediction for DCASE 2023 Challenge Task4A,” DCASE Challenge , 2023
2023
-
[60]
CHT+NSYSU Sound Event Detection System With Multiscale Channel Attention And Multiple Consistency Training For DCASE 2021 Task 4,
Y.-W. Wang, C.-P. Chen, C.-L. Lu, and B.-C. Chan, “CHT+NSYSU Sound Event Detection System With Multiscale Channel Attention And Multiple Consistency Training For DCASE 2021 Task 4,” DCASE Challenge , 2021. 90
2021
-
[61]
Sound Event Detection Based on Self-Supervised Learning of Wav2vec 2.0,
H. Koo, H.-M. Park, J. Park, and M. Oh, “Sound Event Detection Based on Self-Supervised Learning of Wav2vec 2.0,” DCASE Challenge , 2021
2021
-
[62]
Convolution-Augmented Conformer for Sound Event Detection,
Y.-H. Chen, “Convolution-Augmented Conformer for Sound Event Detection,” DCASE Challenge , 2021
2021
-
[63]
Convolutional Network with Conformer for Semi-Supervised Sound Event Detection,
T. Na and Q. Zhang, “Convolutional Network with Conformer for Semi-Supervised Sound Event Detection,” DCASE Challenge , 2021
2021
-
[64]
Integrating Advantages of Recurrent and Transformer Struc- tures for Sound Event Detection in Multiple Scenarios,
R. Lu, W. Hu, Z. Duan, and J. Liu, “Integrating Advantages of Recurrent and Transformer Struc- tures for Sound Event Detection in Multiple Scenarios,” DCASE Challenge , 2021
2021
-
[65]
Leveraging Audio-Tagging Assisted Sound Event Detection Using Weakified Strong Labels and Frequency Dynamic Convolutions,
T. Khandelwal, R. K. Das, A. Koh, and E. S. Chng, “Leveraging Audio-Tagging Assisted Sound Event Detection Using Weakified Strong Labels and Frequency Dynamic Convolutions,” arXiv preprint arXiv:2304.12688, 2023
2023 arXiv
-
[66]
Semi-Supervised Sound Event Detection with Pre- Trained Model,
L. Xu, L. Wang, S. Bi, H. Liu, and J. Wang, “Semi-Supervised Sound Event Detection with Pre- Trained Model,” in ICASSP, 2023
2023
-
[67]
ATST: Audio Representation Learning with Teacher-Student Transformer,
X. Li and X. Li, “ATST: Audio Representation Learning with Teacher-Student Transformer,” in Proc. Interspeech, 2022
2022
-
[68]
Self-Supervised Audio Teacher-Student Transformer for Both Clip-Level and Frame-Level Tasks,
X. Li, N. Shao, and X. Li, “Self-Supervised Audio Teacher-Student Transformer for Both Clip-Level and Frame-Level Tasks,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2024
2024
-
[69]
Multi-Scale Context Aggregation by Dilated Convolutions,
F. Yu and V. Koltun, “Multi-Scale Context Aggregation by Dilated Convolutions,” in International Conference on Learning Representations (ICLR) , 2016
2016
-
[70]
Sound Event Detection with Depth- wise Separable and Dilated Convolutions,
K. Drossos, S. I. Mimilakis, S. Gharib, Y. Li, and T. Virtanen, “Sound Event Detection with Depth- wise Separable and Dilated Convolutions,” in International Joint Conference on Neural Networks , 2020, doi:10.1109/IJCNN48605.2020.9207532
2020
-
[71]
Sound Event Detection Via Dilated Convolutional Recurrent Neural Networks,
Y. Li, M. Liu, K. Drossos, and T. Virtanen, “Sound Event Detection Via Dilated Convolutional Recurrent Neural Networks,” in ICASSP, 2020, doi:10.1109/ICASSP40776.2020.9054433
2020
-
[72]
Dilated Convolution Neural Network with LeakyReLU for Envi- ronmental Sound Classification,
X. Zhang, Y. Zou, and W. Shi, “Dilated Convolution Neural Network with LeakyReLU for Envi- ronmental Sound Classification,” in International Conference on Digital Signal Processing , 2017, doi:10.1109/ICDSP.2017.8096153
2017
-
[73]
Efficient Large-Scale Audio Tagging Via Transformer-to- CNN Knowledge Distillation,
F. Schmid, K. Koutini, and G. Widmer, “Efficient Large-Scale Audio Tagging Via Transformer-to- CNN Knowledge Distillation,” in ICASSP, 2023
2023
-
[74]
Dynamic Convolutional Neural Networks as Efficient Pre-Trained Audio Models,
F. Schmid, K. Koutini, and G. Widmer, “Dynamic Convolutional Neural Networks as Efficient Pre-Trained Audio Models,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2024
2024
-
[75]
BERT: Pre-training of Deep Bidirec- tional Transformers for Language Understanding,
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirec- tional Transformers for Language Understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics , pp. 4171–4186, 2019, d...
2019 doi
-
[76]
A. V. Oppenheim and R. W. Schafer, Discrete-Time Signal Processing: International Edition , Pear- son, 3rd ed., 2010, pp. 522–525, 850–851, ISBN: 9780131988422
2010
-
[77]
L. E. Kinsler, A. R. Frey, A. B. Coppens, and J. V. Sanders, Fundamentals of Acoustics, Wiley, 4th ed., 2000, pp. 149, 210, 224, 291–296, 333–334, ISBN: 9780471847892
2000
-
[78]
J. S. Bendat and A. G. Piersol, Random Data: Analysis and Measurement Procedures , Wiley, 4th ed., 2011, pp. 8-12, 123, ISBN: 9781118210826
2011
-
[79]
D. J. Inman, Engineering Vibrations, Pearson, 4th ed., 2013, pp. 172-177
2013
-
[80]
Rabiner and R
L. Rabiner and R. Schafer, Theory and Applications of Digital Speech Processing , Pearson, 1st ed., 2010, pp. 89-123, ISBN: 0136034284
2010
-
[81]
DCASE2021 Task4 Baseline,
N. Turpault, “DCASE2021 Task4 Baseline,” GitHub, Available: https://github.com/ DCASE-REPO/DESED_task
-
[82]
DCASE 2021 Challenge Task4: Sound event detection and separation in domestic environments,
DCASE, “DCASE 2021 Challenge Task4: Sound event detection and separation in domestic environments,” Available: http://dcase.community/challenge2021/ task-sound-event-detection-and-separation-in-domestic-environments
2021
-
[83]
Threshold Independent Evaluation of Sound Event Detection Scores,
J. Ebbers, R. Haeb-Umbach, and R. Serizel, “Threshold Independent Evaluation of Sound Event Detection Scores,” in ICASSP, 2022
2022
-
[84]
MAT-SED: A Masked Audio Transformer with Masked-Reconstruction Based Pre-training for Sound Event Detection,
P. Cai, Y. Song, K. Li, H. Song, and I. McLoughlin, “MAT-SED: A Masked Audio Transformer with Masked-Reconstruction Based Pre-training for Sound Event Detection,” in Proc. Interspeech, 2024
2024
-
[85]
Sound Event Bounding Boxes,
J. Ebbers, F. G. Germain, G. Wichern, and J. L. Roux, “Sound Event Bounding Boxes,” in Proc. Interspeech, 2024
2024
-
[86]
Prototype Based Masked Audio Model for Self-Supervised Learning of Sound Event Detection,
P. Cai, Y. Song, N. Jiang, Q. Gu, and I. McLoughlin, “Prototype Based Masked Audio Model for Self-Supervised Learning of Sound Event Detection,” arXiv preprint arXiv:2409.17656 , 2024
2024 arXiv
-
[87]
Efficient Training of Audio Transformers with Patchout,
K. Koutini, J. Schl¨ uter, H. Eghbal-zadeh, and G. Widmer, “Efficient Training of Audio Transformers with Patchout,” in Proc. Interspeech, 2022
2022
-
[88]
IMPROVING AUDIO SPECTRO- GRAM TRANSFORMERS FOR SOUND EVENT DETECTION THROUGH MULTI-STAGE TRAINING,
F. Schmid, P. Primus, T. Morocutti, J. Greif, and G. Widmer, “IMPROVING AUDIO SPECTRO- GRAM TRANSFORMERS FOR SOUND EVENT DETECTION THROUGH MULTI-STAGE TRAINING,” DCASE2024 Challenge , 2024
2024
-
[89]
The Ins and Outs of Speaker Recognition: Lessons from VoxSRC 2020,
Y. Kwon, H.-S. Heo, B.-J. Lee, and J. S. Chung, “The Ins and Outs of Speaker Recognition: Lessons from VoxSRC 2020,” in ICASSP, pp. 5809-5813, 2021, doi:10.1109/ICASSP39728.2021.9413948
2020
-
[90]
In Defence of Metric Learning for Speaker Recognition,
J. S. Chung, J. Huh, S. Mun, M. Lee, H.-S. Heo, S. Choe, C. Ham, S. Jung, B.-J. Lee, and I. Han, “In Defence of Metric Learning for Speaker Recognition,” in Proc. Interspeech, 2020
2020
-
[91]
SSAST: Self-Supervised Audio Spectrogram Trans- former,
Y. Gong, C.-I. Lai, Y.-A. Chung, and J. Glass, “SSAST: Self-Supervised Audio Spectrogram Trans- former,” Proceedings of the AAAI Conference on Artificial Intelligence , 2022
2022
-
[92]
Masked Autoencoders that Listen,
P.-Y. Huang, H. Xu, J. Li, A. Baevski, M. Auli, W. Galuba, F. Metze, and C. Feichtenhofer, “Masked Autoencoders that Listen,” in Advances in Neural Information Processing Systems , 2022. 92
2022
-
[93]
BYOL for Audio: Self-Supervised Learning for General-Purpose Audio Representation,
D. Niizumi, D. Takeuchi, Y. Ohishi, N. Harada, and K. Kashino, “BYOL for Audio: Self-Supervised Learning for General-Purpose Audio Representation,” in 2021 International Joint Conference on Neural Networks (IJCNN) , 2021
2021
-
[94]
Jigsaw Clustering for Unsupervised Visual Representation Learning,
P. Chen, S. Liu, and J. Jia, “Jigsaw Clustering for Unsupervised Visual Representation Learning,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2021
2021
-
[95]
Jigsaw-ViT: Learning Jigsaw Puzzles in Vision Transformer,
Y. Chen, X. Shen, Y. Liu, Q. Tao, and J. A. K. Suykens, “Jigsaw-ViT: Learning Jigsaw Puzzles in Vision Transformer,” Pattern Recognition Letters, 2023
2023
-
[96]
Shuffle and Learn: Unsupervised Learning Using Temporal Order Verification,
I. Misra, C. L. Zitnick, and M. Hebert, “Shuffle and Learn: Unsupervised Learning Using Temporal Order Verification,” in ECCV, 2016
2016
-
[97]
TERA: Self-Supervised Learning of Transformer Encoder Representation for Speech,
A. T. Liu, S.-W. Li, and H.-Y. Lee, “TERA: Self-Supervised Learning of Transformer Encoder Representation for Speech,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2021
2021
-
[98]
Transformer-XL: Language Modeling with Longer-Term Dependency,
Z. Dai, Z. Yang, Y. Yang, W. W. Cohen, J. Carbonell, Q. V. Le, and R. Salakhutdinov, “Transformer-XL: Language Modeling with Longer-Term Dependency,” in International Confer- ence on Learning Representations (ICLR) , 2019
2019
-
[99]
You Only Hear Once: A YOLO-Like Algorithm for Audio Segmentation and Sound Event Detection,
S. Venkatesh, D. Moffat, and E. R. Miranda, “You Only Hear Once: A YOLO-Like Algorithm for Audio Segmentation and Sound Event Detection,” Applied Sciences, 2022
2022
-
[100]
Adaptive Convolutional Neural Network for Text-Independent Speaker Recognition,
S.-H. Kim and Y.-H. Park, “Adaptive Convolutional Neural Network for Text-Independent Speaker Recognition,” in Proc. Interspeech, 2021
2021
-
[101]
Temporal Dynamic Convolutional Neural Network for Text- Independent Speaker Verification and Phonemetic Analysis,
S.-H. Kim, H. Nam, and Y.-H. Park, “Temporal Dynamic Convolutional Neural Network for Text- Independent Speaker Verification and Phonemetic Analysis,” in ICASSP, 2022
2022
-
[102]
Decomposed Temporal Dynamic CNN: Efficient Time- Adaptive Network for Text-Independent Speaker Verification Explained with Speaker Activation Map,
S.-H. Kim, H. Nam, and Y.-H. Park, “Decomposed Temporal Dynamic CNN: Efficient Time- Adaptive Network for Text-Independent Speaker Verification Explained with Speaker Activation Map,” arXiv preprint arXiv:2203.15277 , 2022
2022 arXiv
-
[103]
Application of Spectro-Temporal Receptive Field on Soft Labeled Sound Event Detection,
D. Min, H. Nam, and Y.-H. Park, “Application of Spectro-Temporal Receptive Field on Soft Labeled Sound Event Detection,” DCASE Challenge , 2023
2023
-
[104]
Auditory Neural Response Inspired Sound Event Detection Based on Spectro-Temporal Receptive Field,
D. Min, H. Nam, and Y.-H. Park, “Auditory Neural Response Inspired Sound Event Detection Based on Spectro-Temporal Receptive Field,” in DCASE Workshop , 2023
2023
-
[105]
Few-Shot Bioacoustic Event Detection Utilizing Spectro-Temporal Receptive Field,
B.-Y. Ko, H. Nam, D. Min, G.-T. Lee, and Y.-H. Park, “Few-Shot Bioacoustic Event Detection Utilizing Spectro-Temporal Receptive Field,” in Proc. INTER-NOISE, 2023
2023
-
[106]
FilterAugment: An Acoustic Environmental Data Augmen- tation Method,
H. Nam, S.-H. Kim, and Y.-H. Park, “FilterAugment: An Acoustic Environmental Data Augmen- tation Method,” in ICASSP, 2022
2022
-
[107]
Frequency Dynamic Convolution: Frequency- Adaptive Pattern Recognition for Sound Event Detection,
H. Nam, S.-H. Kim, B.-Y. Ko, and Y.-H. Park, “Frequency Dynamic Convolution: Frequency- Adaptive Pattern Recognition for Sound Event Detection,” in Proc. Interspeech, 2022
2022
-
[108]
Frequency & Channel Attention for Computationally Efficient Sound Event Detection,
H. Nam, S.-H. Kim, D. Min, and Y.-H. Park, “Frequency & Channel Attention for Computationally Efficient Sound Event Detection,” in DCASE Workshop , 2023. 93
2023
-
[109]
Heavily Augmented Sound Event Detection Utilizing Weak Predictions,
H. Nam, B.-Y. Ko, G.-T. Lee, S.-H. Kim, W.-H. Jung, S.-M. Choi, and Y.-H. Park, “Heavily Augmented Sound Event Detection Utilizing Weak Predictions,” DCASE Challenge , 2021
2021
-
[110]
Frequency Dependent Sound Event Detection for DCASE 2022 Challenge Task 4,
H. Nam, S.-H. Kim, D. Min, B.-Y. Ko, S.-D. Choi, and Y.-H. Park, “Frequency Dependent Sound Event Detection for DCASE 2022 Challenge Task 4,” DCASE Challenge , 2022
2022
-
[111]
Diversifying and Expanding Frequency- Adaptive Convolution Kernels for Sound Event Detection,
H. Nam, S.-H. Kim, D. Min, J. Lee, and Y.-H. Park, “Diversifying and Expanding Frequency- Adaptive Convolution Kernels for Sound Event Detection,” in Proc. Interspeech, 2024
2024
-
[112]
Pushing the Limit of Sound Event Detection with Multi-Dilated Fre- quency Dynamic Convolution,
H. Nam and Y.-H. Park, “Pushing the Limit of Sound Event Detection with Multi-Dilated Fre- quency Dynamic Convolution,” arXiv preprint arXiv:2406.13312 , 2024
2024 arXiv
-
[113]
Self Training and Ensembling Frequency Dependent Networks with Coarse Prediction Pooling and Sound Event Bounding Boxes,
H. Nam, D. Min, I. Choi, S.-D. Choi, and Y.-H. Park, “Self Training and Ensembling Frequency Dependent Networks with Coarse Prediction Pooling and Sound Event Bounding Boxes,” DCASE Challenge, 2024
2024
-
[114]
Self Training and Ensembling Frequency Dependent Networks with Coarse Prediction Pooling and Sound Event Bounding Boxes,
H. Nam, D. Min, I. Choi, S.-D. Choi, and Y.-H. Park, “Self Training and Ensembling Frequency Dependent Networks with Coarse Prediction Pooling and Sound Event Bounding Boxes,” inDCASE Workshop, 2024
2024
-
[115]
Towards Understanding of Frequency Dependence on Sound Event Detection,
H. Nam, S.-H. Kim, D. Min, B.-Y. Ko, and Y.-H. Park, “Towards Understanding of Frequency Dependence on Sound Event Detection,” arXiv preprint arXiv:2502.07208 , 2025
2025 arXiv
-
[116]
JiTTER: Jigsaw Temporal Transformer for Event Reconstruction for Self-Supervised Sound Event Detection,
H. Nam and Y.-H. Park, “JiTTER: Jigsaw Temporal Transformer for Event Reconstruction for Self-Supervised Sound Event Detection,” arXiv preprint arXiv:2502.20857 , 2025
2025 arXiv
-
[117]
Temporal Attention Pooling for Frequency Dynamic Convolution in Sound Event Detection,
H. Nam and Y.-H. Park, “Temporal Attention Pooling for Frequency Dynamic Convolution in Sound Event Detection,” arXiv preprint arXiv:2504.12670 , 2025
2025 arXiv
-
[118]
Multi-output Classification Framework and Frequency Layer Normalization for Compound Fault Diagnosis in Motor,
W. Yi and Y.-H. Park, “Multi-output Classification Framework and Frequency Layer Normalization for Compound Fault Diagnosis in Motor,” arXiv preprint arXiv:2504.11513 , 2025
2025 arXiv
-
[119]
Multi-output Classification using a Cross-talk Archi- tecture for Compound Fault Diagnosis of Motors in Partially Labeled Condition,
W. Yi, W. Jung, K. Jang and Y.-H. Park, “Multi-output Classification using a Cross-talk Archi- tecture for Compound Fault Diagnosis of Motors in Partially Labeled Condition,” arXiv preprint arXiv:2505.24001, 2025. 94 Acknowledgments in Korean 카이스트에학부신입생으로써제일 처음교문을지나던 날이 생각납니다....
2025
-
[125]
Heavily augmented sound event detection utilizing weak predictions,
H. Nam , B. Y. Ko, G. T. Lee, S. -H. Kim, W. H. Jung, S. M. Choi and Y. -H. Park, “Heavily augmented sound event detection utilizing weak predictions,” DCASE Challenge Tech. rep. , 2021
2021
-
[126]
FilterAugment: An Acoustic Environmental Data Augmen- tation Method,
H. Nam, S. -H. Kim and Y. -H. Park, “FilterAugment: An Acoustic Environmental Data Augmen- tation Method,” International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2022
2022
-
[127]
Temporal dynamic convolutional neural network for text- independent speaker verification and phonemic analysis,
S. -H. Kim, H. Nam and Y. -H. Park, “Temporal dynamic convolutional neural network for text- independent speaker verification and phonemic analysis,” International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2022
2022
-
[128]
Frequency dependent sound event detection for DCASE 2022 Challenge Task 4,
H. Nam, S. -H. Kim, D. Min, B. Y. Ko, S. D. Choi and Y. -H. Park, “Frequency dependent sound event detection for DCASE 2022 Challenge Task 4,” DCASE Challenge Tech. rep. , 2022
2022
-
[129]
Data augmentation and squeeze-and-excitation network on multiple dimension for sound event localization and detection in real scenes,
B. Y. Ko, H. Nam , S. -H. Kim, D. Min, S. Choi and Y. -H. Park, “Data augmentation and squeeze-and-excitation network on multiple dimension for sound event localization and detection in real scenes,” DCASE Challenge Tech. rep. , 2022
2022
-
[130]
Frequency Dynamic Convolution: Frequency- Adaptive Pattern Recognition for Sound Event Detection,
H. Nam , S. -H. Kim, B. Y. Ko and Y. -H. Park, “Frequency Dynamic Convolution: Frequency- Adaptive Pattern Recognition for Sound Event Detection,” Interspeech, 2022
2022
-
[131]
Deep learning based cough detection camera using enhanced features,
G. T. Lee, H. Nam, S. -H. Kim, S. M. Choi and Y. -H. Park, “Deep learning based cough detection camera using enhanced features,” Expert Systems with Applications , 2022. 99
2022
-
[132]
VIFS: An end-to-end variational inference for foley sound synthesis,
J. Lee, H. Nam and Y. -H. Park, “VIFS: An end-to-end variational inference for foley sound synthesis,” DCASE Challenge Tech. rep. , 2023
2023
-
[133]
Application of spectro-temporal receptive field on soft labeled sound event detection,
D. Min, H. Nam and Y. -H. Park, “Application of spectro-temporal receptive field on soft labeled sound event detection,” DCASE Challenge Tech. rep. , 2023
2023
-
[134]
Analysis-based optimization of temporal dynamic convo- lutional neural network for text-independent speaker verification,
S. -H. Kim, H. Nam and Y. -H. Park, “Analysis-based optimization of temporal dynamic convo- lutional neural network for text-independent speaker verification,” IEEE Access, 2023
2023
-
[135]
Frequency & Channel Attention for Computationally Efficient Sound Event Detection,
H. Nam, S. -H. Kim, D. Min and Y. -H. Park, “Frequency & Channel Attention for Computationally Efficient Sound Event Detection,” DCASE Workshop , 2023
2023
-
[136]
Auditory Neural Response Inspired Sound Event Detection Based on Spectro-temporal Receptive Field,
D. Min, H. Nam and Y. -H. Park, “Auditory Neural Response Inspired Sound Event Detection Based on Spectro-temporal Receptive Field,” DCASE Workshop , 2023
2023
-
[137]
PRTFNet: HRTF Individualization for Accurate Spectral Cues Using a Compact PRTF,
B. Y. Ko, G. T. Lee, H. Nam and Y. -H. Park, “PRTFNet: HRTF Individualization for Accurate Spectral Cues Using a Compact PRTF,” IEEE Access, 2023
2023
-
[138]
Real-Time Sound Recognition System for Human Care Robot Considering Custom Sound Events,
S. -H. Kim, H. Nam , S. M. Choi and Y. -H. Park, “Real-Time Sound Recognition System for Human Care Robot Considering Custom Sound Events,” IEEE Access, 2024
2024
-
[139]
ChatGPT Caption Paraphrasing and FENSE- based Caption Filtering for Automated Audio Captioning,
I. Choi, H. Nam, D. Min, S. Choi and Y. -H. Park, “ChatGPT Caption Paraphrasing and FENSE- based Caption Filtering for Automated Audio Captioning,” DCASE Challenge Tech. rep. , 2024
2024
-
[140]
Diversifying and Expanding Frequency- Adaptive Convolution Kernels for Sound Event Detection,
H. Nam , S. -H. Kim, D. Min, J. Lee and Y. -H. Park, “Diversifying and Expanding Frequency- Adaptive Convolution Kernels for Sound Event Detection,” Interspeech, 2024
2024
-
[141]
Self Training and Ensembling Frequency Dependent Networks with Coarse Prediction Pooling and Sound Event Bounding Boxes,
H. Nam , D. Min, S. Choi, I. Choi and Y. -H. Park, “Self Training and Ensembling Frequency Dependent Networks with Coarse Prediction Pooling and Sound Event Bounding Boxes,” DCASE Workshop, 2024
2024
-
[142]
Coherence-based Phonemic Analysis on the Effect of Reverberation to Practical Automatic Speech Recognition,
H. Nam and Y. -H. Park, “Coherence-based Phonemic Analysis on the Effect of Reverberation to Practical Automatic Speech Recognition,” Applied Acoustics, 2025
2025
-
[143]
Towards Understanding of Frequency Dependence on Sound Event Detection,
H. Nam , S. -H. Kim, D. Min, B. Y. Ko and Y. -H. Park, “Towards Understanding of Frequency Dependence on Sound Event Detection,” arXiv, 2025
2025
-
[144]
Pushing the limit of sound event detection with multi-dilated frequency dynamic convolution,
H. Nam and Y. -H. Park, “Pushing the limit of sound event detection with multi-dilated frequency dynamic convolution,” arXiv, 2025
2025
-
[145]
JiTTER: Jigsaw Temporal Transformer for Event Reconstruction for Self-Supervised Sound Event Detection,
H. Nam and Y. -H. Park, “JiTTER: Jigsaw Temporal Transformer for Event Reconstruction for Self-Supervised Sound Event Detection,” arXiv, 2025
2025
-
[146]
Temporal Attention Pooling for Frequency Dynamic Convolution in Sound Event Detection,
H. Nam and Y. -H. Park, “Temporal Attention Pooling for Frequency Dynamic Convolution in Sound Event Detection,” arXiv, 2025
2025
-
[147]
DNN based HRIRs Identification with a Continuously Rotating Speaker Array,
B. Y. Ko, D. Min, H. Nam and Y. -H. Park, “DNN based HRIRs Identification with a Continuously Rotating Speaker Array,” arXiv, 2025. 100
2025
-
[2009]
1. – 2012. 12. NUS High School of Math and Science, Singapore (NUS High diploma)
2012
-
[2013]
9. – 2018. 2. 한국과학기술원기계공학과 (학사)
2018
-
[2017]
3. – 2024. 2. 한국과학기술원기계공학과조교 연 구 업 적
2024
-
[2018]
3. – 2020. 2. 한국과학기술원기계공학과 (석사)
2020
-
[2020]
3. – 2025. 8. 한국과학기술원기계공학과 (박사) 경 력
2025
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.