Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

Frequency Dynamic Convolutions for Sound Event Detection

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This paper claims that convolutions whose kernels adapt to the input's frequency content raise sound event detection on DESED by up to 10.98% over a baseline CRNN, and that a lighter variant matches the top score with 30% fewer parameters.

desk verdict Useful ablations around a known frequency-adaptive convolution family, but the headline PSDS1 gains are confounded with a 3-4x parameter increase and no matched-capacity static baseline. read the letter →

arxiv 2506.12785 v1 pith:YE7XLY3Y submitted 2025-06-15 eess.AS cs.SD

classification eess.AScs.SD
keywords soundeventdetectionfrequencydynamicconvolutionfrequency-adaptivekernelstemporalattentionpoolingdilatedconvolutionalrecurrentneuralnetworkDESEDdatasetPSDS
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Conventional 2D convolutions treat the frequency axis like the time axis, assuming that a pattern shifted upward or downward in pitch means the same thing. The paper argues this is wrong for audio, and sets out to show that letting convolutional kernels adapt to the input's frequency content improves sound event detection. It proposes Frequency Dynamic Convolution (FDY conv), in which the effective filter is a weighted combination of basis kernels with weights chosen per frequency bin from the input itself, plus variants: a dilated version (DFD), a hybrid static/dynamic version (PFD) that cuts parameters, a multi-branch dilated version (MDFD), and a temporal-attention-pooling version (TFD) that reaches the same top score with fewer parameters. On the DESED benchmark the best models raise the polyphonic sound detection score PSDS1 from 0.410 to 0.455, a 10.98% improvement, and class-wise analysis links each variant to a different event type (non-stationary, broad-spectrum, quasi-stationary, transient). If the gain is genuine, frequency-adaptive convolutions are a drop-in alternative to standard 2D convolutions for audio spectrogram processing.

What carries the argument

The load-bearing object is the frequency-adaptive attention weight vector $\pi(f, x) \in \mathbb{R}^K$: a small learned subnetwork reads the input feature map, aggregates over time, and outputs, for every mel-frequency bin, $K$ weights that recombine a shared bank of $K$ basis kernels into a frequency-specific effective filter. Because the basis kernels are shared across frequencies, parameter growth stays modest; because the weights depend on the input, the same layer behaves differently for a broadband noise burst and a narrow tone. Each later variant changes how the basis kernels are organized (dilation per kernel, multiple dynamic branches, a static/dynamic channel split) or how time is aggregated (Temporal Attention Pooling combining time attention, velocity attention, and average pooling), but the input-dependent, per-frequency recombination of kernels is the mechanism that carries the whole argument.

What would settle it

Train the final MDFD or TFD model with per-frequency attention weights replaced by one shared weight across all frequency bins, keeping parameter count and the training recipe identical, and compare PSDS1 on DESED's strongly labeled real test recordings; if the gain over the baseline CRNN does not shrink, the reported improvement is not due to frequency adaptivity. Replication of the 0.455 PSDS1 under the paper's exact augmentation and post-processing settings, with standard errors across seeds, would also settle whether the 10.98% figure lies outside evaluation noise.

Watch

Extended reading notes

Core claim

The central claim is that the shift invariance built into standard 2D convolution is the wrong inductive bias for the frequency axis of audio, and that replacing it with frequency-adaptive convolution measurably improves sound event detection. In FDY conv, the convolution output at frequency bin $f$ is the attention-weighted sum of $K$ basis-kernel responses, $$y(f) = \sum_{i=1}^{K} \pi_i(f, x)\, (W_i * x + b_i),$$ where the weights $\pi_i$ depend on both the frequency bin $f$ and the input $x$, so each frequency bin receives a different effective filter matched to the content there. Extending this idea, DFD conv diversifies basis kernels by dilation, PFD conv mixes a static branch with a small dynamic branch to cut parameters, MDFD conv combines several dilated dynamic branches, and TFD conv replaces temporal average pooling with attention-based pooling over time. The dissertation reports that MDFD achieves the highest PSDS1, 0.455 versus 0.410 for the baseline CRNN (a 10.98% gain), and that TFD matches 0.455 with 12.703M parameters versus 18.157M for MDFD.

Load-bearing premise

The paper assumes the performance gain is caused by frequency adaptivity, while its comparisons also add parameters, a learned attention mechanism, and hyperparameters selected on the same validation set used for the headline numbers.

Editorial extensions

If this is right

  • CRNN-style sound event detection systems can be upgraded by swapping standard 2D convolutions for FDY-type layers, with reported PSDS1 gains of 7.56% for FDY conv and up to 10.98% for MDFD conv over the baseline on DESED.
  • The parameter-lean variants matter for deployment: PFD conv keeps roughly baseline-level accuracy while cutting parameters by 54.4% relative to FDY conv, and TFD conv matches the best PSDS1 at 12.703M parameters, about 30% fewer than MDFD conv.
  • Different event classes favor different designs: FDY helps non-stationary events, DFD helps broad-spectrum events, PFD helps quasi-stationary events, and TFD helps transient events, so architecture choice can be guided by the target sound inventory.
  • Combining the MDFD-CRNN with pretrained transformer encoders (ATST-frame plus BEATs) and change-detection-based event bounding raises the true PSDS1 to 0.577 without external pretraining data or ensembling.
  • The paper's controlled comparison of 2D and 1D front-ends (PSDS1 0.410 versus 0.192) indicates that preserving the frequency axis as a spatial dimension matters more than raw model capacity in this setting.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural experiment not reported in the paper is to hold capacity constant while collapsing the per-frequency attention weights to a single shared weight; that ablation would isolate how much of the 10.98% gain is frequency adaptivity itself rather than extra parameters, the attention mechanism, or architecture selection.
  • The smooth, class-dependent clustering of attention weights along the frequency axis (the paper's PCA analysis) suggests FDY conv is learning an input-dependent filterbank, which connects this work to learnable front-end filter designs in speaker verification where the same mechanism could transfer.
  • The class-wise specialization across variants points toward a learned router or mixture-of-experts design in which the model picks the dilated, partial, or temporal-attention adapter per event type; the paper does not explore that combination.
  • Because the headline numbers come from a single validation set that also drove the architecture sweeps, multi-seed replication on a fresh test split is the prudent way to confirm the ordering of the variants before building on it.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The dissertation proposes a family of frequency-adaptive convolutions for sound event detection: Frequency Dynamic Convolution (FDY conv) and its extensions DFD, PFD, MDFD, and TFD. The core idea is to replace standard 2D convolutions with kernels that are dynamically combined using frequency-dependent attention weights, motivated by the argument that frequency-axis shift invariance is inappropriate for audio. The models are evaluated on DESED using PSDS1, with reported improvements over a CRNN baseline, parameter-efficiency claims for PFD and TFD, a class-wise analysis, and three engineering case studies. The paper claims MDFD is the best performer (PSDS1 0.455 vs 0.410 baseline) and that TFD matches it with 30% fewer parameters.

Significance. If the reported gains are real, the work is significant: it offers a systematic exploration of frequency-adaptive convolutions in a realistic SED benchmark and produces a parameter-efficient variant that matches the best accuracy. The paper also deserves credit for evaluating on the external DESED benchmark, using a standard mean-teacher pipeline, reporting parameter counts, and attempting class-wise and qualitative analyses. However, the central claim that frequency adaptivity causes the observed improvements is currently not established: the comparisons are not capacity-controlled, the headline model is selected from many configurations on the same validation set without variance estimates, and there are internal numeric inconsistencies in the baseline and parameter-efficiency numbers. These issues are load-bearing because the paper's contribution is precisely the attribution of the gains to the proposed mechanism.

major comments (4)
  1. [§4.2, Table 4.2 vs §3.2.4, Table 3.1] The baseline score is inconsistent across the manuscript: Table 3.1 reports a baseline PSDS1 of 0.396, while Table 4.2 (and Table 3.2) reports 0.410 for the same CRNN with 4.428M parameters. Correspondingly, Section 3.3.3 states that FDY conv improves over the baseline by 8.95% (0.434 vs 0.396), whereas the abstract and Table 4.2 imply a 7.56% improvement (0.441 vs 0.410). The headline 10.98% gain for MDFD conv is computed from 0.455 vs 0.410, so the choice of baseline directly changes the central quantitative claim. The authors should identify which configuration and post-processing setting each table refers to and make the numbers consistent.
  2. [Table 4.2] The main comparison is not capacity-controlled. MDFD conv uses 18.157M parameters and TFD conv uses 12.703M, versus 4.428M for the baseline CRNN, a 4.1x and 2.9x gap. No standard 2D CRNN with a matched parameter count is trained under the same mean-teacher, augmentation, and post-processing pipeline. The 7.6%-10.98% PSDS1 gains could therefore be partly or wholly due to extra model capacity rather than to frequency adaptivity. PFD conv at 5.041M/0.442 provides partial evidence, but even a static CRNN at 5M parameters is untested. The load-bearing condition for the paper's claim should be tested by training static CRNNs with matched parameter counts (for example, approximately 5M, 12M, and 18M parameters) under otherwise identical settings.
  3. [§3.3-§3.7 and §4.2-§4.3] All proposed modules are selected on the same DESED real validation set that is later used for the headline comparison. Tables 3.3-3.9 and 4.3-4.5 sweep dilation sets, branch proportions, channel widths, and pooling components, and the best configurations are then reported in Table 4.2. No repeated-seed experiments or variance estimates are provided, so differences of 0.004-0.007 PSDS1 (e.g., 0.448 vs 0.441, or 0.455 vs 0.451) are not shown to be above evaluation noise. The authors should report multiple runs (or confidence intervals) for at least the key comparisons, or otherwise demonstrate that the selected configurations are not artifacts of validation-set overfitting.
  4. [§3.5.3, §3.7.4, Table 4.2, Abstract] The parameter-efficiency claim is internally inconsistent. Section 3.5.3 states that PFD-CRNN (1/8) has 5.401M parameters and reduces parameters by 51.9% relative to FDY conv (11.061M), while Section 3.7.4 and Table 4.2 list PFD conv as 5.041M, and the abstract reports a 54.4% reduction. These numbers imply different model configurations or an arithmetic error. Since parameter efficiency is a stated contribution, this needs to be resolved with a single consistent set of model sizes and a clear definition of which proportion and channel configuration is being reported.
minor comments (5)
  1. [Abstract and Keywords] There are typos in the abstract and keywords: 'leadin to inconsistencies' should be 'leading to', 'translational equivariacne' should be 'translational equivariance', and 'temproal attention pooling' should be 'temporal attention pooling'.
  2. [§3.1, Eq. (3.1)] The text calls the property in Eq. (3.1) 'shift invariance' but the equation T(F(x)) = F(T(x)) actually defines translation equivariance. The authors should either fix the mathematical statement or adjust the terminology, since the distinction matters for the paper's motivation.
  3. [§3.5.3] The text refers to 'Table 3.7' when presenting PFD conv results, but the results are in Table 3.5. The cross-reference should be corrected.
  4. [§3.7.2, Eqs. (3.11)-(3.13)] In Eq. (3.11), the first two terms use xs,t but the third term uses xt; the subscript convention should be unified, and the dimensions of the attention weights should be stated explicitly.
  5. [§4.5, Table 4.8] The comparison with 1D CRNN is interesting but the 1D model is described as 'adapted from Wav2Vec2.0 and HuBERT' without a precise layer configuration; please provide the exact architecture (kernel sizes, strides, channels, pooling) so the result is reproducible.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: all central claims are empirical DESED benchmark results; self-citations provide provenance, not load-bearing assumptions.

full rationale

The paper's central claims are measured, not derived from the definitions of the proposed modules. Table 4.2 reports PSDS1 scores for the baseline CRNN and each frequency-adaptive variant on the DESED real validation set using the official DESED evaluation toolkit (Section 4.1.8). Each module is defined by explicit equations (FDY conv in Eqs. 3.2-3.4; DFD conv in Eqs. 3.5-3.7; PFD conv in Eq. 3.9; MDFD conv in Eq. 3.10; TAP in Eqs. 3.11-3.13), and the reported gains are benchmark outcomes, not terms defined in terms of the target metric. No parameter is fitted to PSDS1 and then renamed as a prediction: the only learned quantities are network weights trained with the mean-teacher objective (Eqs. 4.1-4.10), and the evaluation is on held-out real recordings. The paper does cite the author's own prior publications for the origins of FDY/DFD/PFD/MDFD/TFD conv and for the efficient implementation trick in Section 3.3.2, but those citations are descriptive provenance: the equations and experimental protocol are reproduced inside the thesis and the results are independently measured against the external DESED benchmark. There is no uniqueness theorem imported from the authors, no ansatz smuggled in only through a citation, and no renaming of a known empirical pattern as a derivation. The main methodological risks, such as hyperparameter selection on the same validation set (Tables 3.3-3.9, 4.3-4.5) and the parameter-count gap between the baseline and MDFD conv (4.428M vs 18.157M), are experimental-control concerns rather than circularity; they do not make any equation reduce to its own input. I therefore find no circular step.

Assumptions & free parameters 5 free parameters · 5 assumptions · 1 invented entities

The paper contributes no formal derivation; its claims are empirical. The central result depends on a handful of architecture hyperparameters chosen by hand or by validation-set search, several domain assumptions about audio representation and evaluation, and one family of new computational modules. This is typical for applied deep learning, but it means the reported percentages are conditional on those choices.

free parameters (5)
  • Number of basis kernels K = 4 (5 in early dilation experiments)
    FDY, DFD, PFD and MDFD use four basis kernels; K=5 is tested in Section 3.4.4. Performance depends on this architectural choice.
  • Dilation size set for DFD/MDFD = DFD best (1,1),(1,2),(1,3),(1,3); MDFD best (1)x5+(2,3)+(2,2,3)+(2,3,3)
    Selected by sweeping dilation configurations on the DESED validation set in Tables 3.4, 3.7 and 3.8; the headline MDFD score uses this specific set.
  • Dynamic branch proportion for PFD/MDFD = 1/8 for PFD; 5/8 for best TAP+PFD; 11/8 channel expansion for best MDFD
    The static-to-dynamic channel ratio is swept in Tables 3.5, 4.4 and 3.8, and performance is non-monotonic in this ratio.
  • TAP pooling component combination = Average + Time Attention + Velocity Attention
    The ablation in Table 3.9 selects the three-component combination; other combinations score lower.
  • Post-processing parameters = weak-prediction mask threshold (not stated); median filter window of 7 frames
    Section 4.1.7 fixes these for all models, and class-wise median filtering is deliberately excluded to keep comparisons fair.
assumptions (5)
  • domain assumption Shifting a sound along frequency changes its perceptual meaning, so 2D convolution's frequency shift invariance is inappropriate for SED.
    Section 3.1 states this as the motivation; it is not proven and it underpins every proposed module.
  • domain assumption The log-mel spectrogram with 128 mel bins, window 2048, hop 256, and per-frequency normalization is a sufficient input representation.
    Section 4.1.3 fixes the input representation; all comparisons use this representation.
  • domain assumption Mean teacher with EMA updates and consistency losses is an effective semi-supervised training framework for DESED.
    Section 4.1.2 adopts this framework from prior work; all reported gains are measured under this training scheme.
  • domain assumption PSDS1 computed by the official DESED toolkit is a valid threshold-independent measure of SED quality.
    Section 4.1.8 defines the primary metric, and all improvement percentages in the abstract and Chapter 4 are based on it.
  • standard math Class-wise F1 comparisons are statistically meaningful under ANOVA and Tukey HSD assumptions.
    Section 4.6 uses ANOVA plus Tukey HSD for class-wise analysis; the underlying independence and normality assumptions are not examined.
invented entities (1)
  • FDY conv and variants (DFD, PFD, MDFD, TFD) independent evidence
    purpose: Frequency-adaptive convolutional operators for sound event detection.
    These are new computational modules, not physical entities. They are validated on the public DESED benchmark and in the paper's case studies, so they do not suffer from the graviton problem: their behavior is falsifiable through public-benchmark experiments.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Frequency Dynamic Convolutions for Sound Event Detection." pith.science (2026). https://pith.science/paper/YE7XLY3Y

@misc{pith2026250612785,
  author       = {Pith},
  title        = {Pith review of: Frequency Dynamic Convolutions for Sound Event Detection},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YE7XLY3Y}},
  note         = {Machine review of arXiv:2506.12785}
}
abstract

Recent research in deep learning-based Sound Event Detection (SED) has primarily focused on Convolutional Recurrent Neural Networks (CRNNs) and Transformer models. However, conventional 2D convolution-based models assume shift invariance along both the temporal and frequency axes, leadin to inconsistencies when dealing with frequency-dependent characteristics of acoustic signals. To address this issue, this study proposes Frequency Dynamic Convolution (FDY conv), which dynamically adjusts convolutional kernels based on the frequency composition of the input signal to enhance SED performance. FDY conv constructs an optimal frequency response by adaptively weighting multiple basis kernels based on frequency-specific attention weights. Experimental results show that applying FDY conv to CRNNs improves performance on the DESED dataset by 7.56% compared to the baseline CRNN. However, FDY conv has limitations in that it combines basis kernels of the same shape across all frequencies, restricting its ability to capture diverse frequency-specific characteristics. Additionally, the $3\times3$ basis kernel size is insufficient to capture a broader frequency range. To overcome these limitations, this study introduces an extended family of FDY conv models. Dilated FDY conv (DFD conv) applies convolutional kernels with various dilation rates to expand the receptive field along the frequency axis and enhance frequency-specific feature representation. Experimental results show that DFD conv improves performance by 9.27% over the baseline. Partial FDY conv (PFD conv) addresses the high computational cost of FDY conv, which results from performing all convolution operations with dynamic kernels. Since FDY conv may introduce unnecessary adaptivity for quasi-stationary sound events, PFD conv integrates standard 2D convolutions with frequency-adaptive kernels to reduce computational complexity while maintaining performance. Experimental results demonstrate that PFD conv improves performance by 7.80% over the baseline while reducing the number of parameters by 54.4% compared to FDY conv. Multi-Dilated FDY conv (MDFD conv) extends DFD conv by addressing its structural limitation of applying the same dilation across all frequencies. By utilizing multiple convolutional kernels with different dilation rates, MDFD conv effectively captures diverse frequency-dependent patterns. Experimental results indicate that MDFD conv achieves the highest performance, improving the baseline CRNN performance by 10.98%. Furthermore, standard FDY conv employs Temporal Average Pooling, which assigns equal weight to all frames along the time axis, limiting its ability to effectively capture transient events. To overcome this, this study proposes TAP-FDY conv (TFD conv), which integrates Temporal Attention Pooling (TA) that focuses on salient features, Velocity Attention Pooling (VA) that emphasizes transient characteristics, and Average Pooling (AP) that captures stationary properties. TAP-FDY conv achieves the same performance as MDFD conv but reduces the number of parameters by approximately 30.01% (12.703M vs. 18.157M), achieving equivalent accuracy with lower computational complexity. Class-wise performance analysis reveals that FDY conv improves detection of non-stationary events, DFD conv is particularly effective for events with broad spectral features, and PFD conv enhances the detection of quasi-stationary events. Additionally, TFD conv (TFD-CRNN) demonstrates strong performance in detecting transient events. In the case studies, PFD conv effectively captures stable signal patterns in tank powertrain fault recognition, DFD conv recognizes wide harmonic spectral patterns on speed-varying motor fault recognition, while TFD conv outperforms other models in detecting transient signals in offshore arc detection. These results suggest that frequency-adaptive convolutions and their extended variants provide a robust alternative to conventional 2D convolutions in deep learning-based audio processing.

Figures

Figures reproduced from arXiv: 2506.12785 by the authors.

Figure 3.1
Figure 3.1. An illustration of shift invariance of (a) image and (b) 2D audio data. [PITH_FULL_IMAGE:figures/full_fig_p032_3_1.png] view at source ↗
Figure 3.2
Figure 3.2. Comparative illustration of frequency dependent convolution methods: (a) frequency-wise [PITH_FULL_IMAGE:figures/full_fig_p033_3_2.png] view at source ↗
Figure 3.3
Figure 3.3. An illustration of frequency dynamic convolution operation. [PITH_FULL_IMAGE:figures/full_fig_p035_3_3.png] view at source ↗
Figures from the paper (16 more)
Figure 3.4
Figure 3.4. Figure 3.4: Examples of PCA analysis on the attention weights used in frequency dynamic convolution. [PITH_FULL_IMAGE:figures/full_fig_p038_3_4.png]
Figure 3.5
Figure 3.5. Figure 3.5: An illustration of dilated frequency dynamic convolution operation. [PITH_FULL_IMAGE:figures/full_fig_p039_3_5.png]
Figure 3.6
Figure 3.6. Figure 3.6: Plots comparing variance of attention weights on 2nd - 6th convolution layers in FDY-CRNN [PITH_FULL_IMAGE:figures/full_fig_p044_3_6.png]
Figure 3.7
Figure 3.7. Figure 3.7: An illustration of partial frequency dynamic convolution operation. [PITH_FULL_IMAGE:figures/full_fig_p045_3_7.png]
Figure 3.8
Figure 3.8. Figure 3.8: An illustration of multi-dilated frequency dynamic convolution operation. [PITH_FULL_IMAGE:figures/full_fig_p048_3_8.png]
Figure 3.9
Figure 3.9. Figure 3.9: An illustration of TAP frequency dynamic convolution operation. [PITH_FULL_IMAGE:figures/full_fig_p053_3_9.png]
Figure 4.1
Figure 4.1. Figure 4.1: An illustration of the framework for training the SED models in this work. It applies [PITH_FULL_IMAGE:figures/full_fig_p063_4_1.png]
Figure 4.2
Figure 4.2. Figure 4.2: An illustration of SED baseline model architecture used in this work. [PITH_FULL_IMAGE:figures/full_fig_p065_4_2.png]
Figure 4.3
Figure 4.3. Figure 4.3: Comparison between 2D CRNN and 1D CRNN architectures. The 1D convolutional stack [PITH_FULL_IMAGE:figures/full_fig_p074_4_3.png]
Figure 4.4
Figure 4.4. Figure 4.4: Class-wise PSDS1 scores for different models. Frequency-adaptive convolution methods [PITH_FULL_IMAGE:figures/full_fig_p075_4_4.png]
Figure 5.1
Figure 5.1. Figure 5.1: The overall task workflow of the proposed fault diagnosis system for the K1 tank powertrain. [PITH_FULL_IMAGE:figures/full_fig_p079_5_1.png]
Figure 5.2
Figure 5.2. Figure 5.2: Visual representation of data collected from the K1 tank powertrain system. The left panel [PITH_FULL_IMAGE:figures/full_fig_p080_5_2.png]
Figure 5.3
Figure 5.3. Figure 5.3: Detailed configuration of the testbed including RPM monitoring, torque sensing, thermo [PITH_FULL_IMAGE:figures/full_fig_p084_5_3.png]
Figure 5.4
Figure 5.4. Figure 5.4: Experimental setup for compound motor fault diagnosis under variable speed conditions. [PITH_FULL_IMAGE:figures/full_fig_p085_5_4.png]
Figure 5.5
Figure 5.5. Figure 5.5: Mel spectrograms of vibration signals under three speed profiles: (top) sinusoidal (2000 [PITH_FULL_IMAGE:figures/full_fig_p086_5_5.png]
Figure 5.6
Figure 5.6. Figure 5.6: Spectrogram comparison of different acoustic signals recorded in offshore environments. The [PITH_FULL_IMAGE:figures/full_fig_p090_5_6.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Auditory Intelligence: Understanding the World Through Sound

    eess.AS 2025-08 conditional novelty 4.0 of 10

    A position paper proposing four cognitively inspired audio tasks (ASPIRE, SODA, AUX, AUGMENT) to push machine hearing beyond recognition toward explanation, reasoning, and interaction.

Reference graph

Works this paper leans on

145 extracted references · 69 canonical work pages · cited by 1 Pith paper

  1. [1]

    Virtanen, M

    T. Virtanen, M. D. Plumbley, and D. Ellis, Computational Analysis of Sound Scenes and Events , Springer Publishing Company, Incorporated, 1st ed., 2017, pp. 3-11, 71-77, ISBN: 3319634496

  2. [2]

    Sound event detection in domestic environ- ments with weakly labeled data and soundscape synthesis,

    N. Turpault, R. Serizel, A. P. Shah, and J. Salamon, “Sound event detection in domestic environ- ments with weakly labeled data and soundscape synthesis,” in DCASE Workshop , 2019

  3. [3]

    Convolutional Recurrent Neural Networks for Polyphonic Sound Event Detection,

    E. C ¸ akır, G. Parascandolo, T. Heittola, H. Huttunen, and T. Virtanen, “Convolutional Recurrent Neural Networks for Polyphonic Sound Event Detection,” IEEE/ACM Transac- tions on Audio, Speech, and Language Processing , vol. 25, no. 6, pp. 1291-1303, 2017, doi:10.1109/TASLP.2017.2690575

  4. [4]

    Metrics for Polyphonic Sound Event Detection,

    A. Mesaros, T. Heittola, and T. Virtanen, “Metrics for Polyphonic Sound Event Detection,” Applied Sciences, vol. 6, no. 6, article 162, 2016, doi:10.3390/app6060162

  5. [6]

    Scaper: A library for soundscape synthesis and augmentation

    Salamon, J., MacConnell, D., Cartwright, M., Li, P. and Bello, J. “Scaper: A library for soundscape synthesis and augmentation”, 2017 IEEE Workshop On Applications Of Signal Processing To Audio And Acoustics (WASPAA), 2017

  6. [7]

    Study on Frequency Dependent Convolution Methods for Sound Event Detection,

    H. Nam, S.-H. Kim, B.-Y. Ko, D. Min, and Y.-H. Park, “Study on Frequency Dependent Convolution Methods for Sound Event Detection,” in Proc. INTER-NOISE, 2024

  7. [8]

    SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition,

    D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition,” in Proc. Interspeech, 2019

  8. [9]

    Conformer: Convolution-augmented Transformer for Speech Recognition,

    A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented Transformer for Speech Recognition,” in Proc. Interspeech, 2020

Show all 145 references
  1. [10]

    Coherence-Based Phonemic Analysis on the Effect of Reverberation to Practical Automatic Speech Recognition,

    H. Nam and Y.-H. Park, “Coherence-Based Phonemic Analysis on the Effect of Reverberation to Practical Automatic Speech Recognition,” Applied Acoustics, vol. 227, p. 110233, 2025

  2. [11]

    wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,

    A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,” in Advances in Neural Information Processing Systems, 2020

  3. [12]

    Hu- BERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,

    W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hu- BERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2021

  4. [13]

    Attentive Statistics Pooling for Deep Speaker Embed- ding,

    K. Okabe, T. Koshinaka, and K. Shinoda, “Attentive Statistics Pooling for Deep Speaker Embed- ding,” in Proc. Interspeech, 2018. 87

  5. [14]

    Exploring the Encoding Layer and Loss Function in End-to-End Speaker and Language Recognition System,

    W. Cai, J. Chen, and M. Li, “Exploring the Encoding Layer and Loss Function in End-to-End Speaker and Language Recognition System,” in Proc. Interspeech, 2018

  6. [15]

    Analysis-Based Optimization of Temporal Dynamic Convo- lutional Neural Network for Text-Independent Speaker Verification,

    S.-H. Kim, H. Nam, and Y.-H. Park, “Analysis-Based Optimization of Temporal Dynamic Convo- lutional Neural Network for Text-Independent Speaker Verification,” IEEE Access, vol. 11, 2023

  7. [16]

    Integrating Frequency Translational Invariance in TDNNs and Frequency Positional Information in 2D ResNets to Enhance Speaker Verification,

    J. Thienpondt, B. Desplanques, and K. Demuynck, “Integrating Frequency Translational Invariance in TDNNs and Frequency Positional Information in 2D ResNets to Enhance Speaker Verification,” in Proc. Interspeech, 2021

  8. [17]

    Convolution-Based Channel-Frequency Attention for Text-Independent Speaker Verification,

    J. Li, Y. Tian, and T. Lee, “Convolution-Based Channel-Frequency Attention for Text-Independent Speaker Verification,” in ICASSP, 2023, doi:10.1109/ICASSP49357.2023.10095415

  9. [18]

    PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition,

    Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumbley, “PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2020

  10. [19]

    Deep Learning Based Cough Detection Camera Using Enhanced Features,

    G.-T. Lee, H. Nam, S.-H. Kim, S.-M. Choi, Y. Kim, and Y.-H. Park, “Deep Learning Based Cough Detection Camera Using Enhanced Features,” Expert Systems with Applications , vol. 206, 2022, doi:10.1016/j.eswa.2022.117811

  11. [20]

    Real-Time Sound Recognition System for Human Care Robot Considering Custom Sound Events,

    S.-H. Kim, H. Nam, S.-M. Choi, and Y.-H. Park, “Real-Time Sound Recognition System for Human Care Robot Considering Custom Sound Events,” IEEE Access, vol. 12, 2024

  12. [21]

    AST: Audio Spectrogram Transformer,

    Y. Gong, Y.-A. Chung, and J. Glass, “AST: Audio Spectrogram Transformer,” in Proc. Interspeech, 2021

  13. [22]

    BEATs: Audio Pre-Training with Acoustic Tokenizers,

    S. Chen, Y. Wu, C. Wang, S. Liu, D. Tompkins, Z. Chen, W. Che, X. Yu, and F. Wei, “BEATs: Audio Pre-Training with Acoustic Tokenizers,” in ICML, 2023

  14. [23]

    Overview and Evaluation of Sound Event Localization and Detection in DCASE 2019,

    A. Politis, A. Mesaros, S. Adavanne, T. Heittola, and T. Virtanen, “Overview and Evaluation of Sound Event Localization and Detection in DCASE 2019,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, 2020

  15. [24]

    STARSS22: A Dataset of Spatial Recordings of Real Scenes with Spatiotemporal Annotations of Sound Events,

    A. Politis, K. Shimada, P. Sudarsanam, S. Adavanne, D. Krause, Y. Koyama, N. Takahashi, S. Takahashi, Y. Mitsufuji, and T. Virtanen, “STARSS22: A Dataset of Spatial Recordings of Real Scenes with Spatiotemporal Annotations of Sound Events,” in DCASE Workshop , 2022

  16. [25]

    Data Augmentation and Squeeze-and-Excitation Network on Multiple Dimension for Sound Event Localization and Detection in Real Scenes,

    B.-Y. Ko, H. Nam, S.-H. Kim, D. Min, S.-D. Choi, and Y.-H. Park, “Data Augmentation and Squeeze-and-Excitation Network on Multiple Dimension for Sound Event Localization and Detection in Real Scenes,” DCASE Challenge , 2022

  17. [26]

    Automated Audio Captioning with Recurrent Neural Networks,

    K. Drossos, S. Adavanne, and T. Virtanen, “Automated Audio Captioning with Recurrent Neural Networks,” in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics , 2017

  18. [27]

    Clotho: An Audio Captioning Dataset,

    K. Drossos, S. Lipping, and T. Virtanen, “Clotho: An Audio Captioning Dataset,” in ICASSP, 2020

  19. [28]

    ChatGPT Caption Paraphrasing and FENSE-Based Caption Filtering for Automated Audio Captioning,

    I. Choi, H. Nam, D. Min, S.-D. Choi, and Y.-H. Park, “ChatGPT Caption Paraphrasing and FENSE-Based Caption Filtering for Automated Audio Captioning,” DCASE Challenge , 2024. 88

  20. [29]

    Mind the Domain Gap: A Systematic Analysis on Bioacoustic Sound Event Detection,

    J. Liang, I. Nolasco, B. Ghani, H. Phan, E. Benetos, and D. Stowell, “Mind the Domain Gap: A Systematic Analysis on Bioacoustic Sound Event Detection,” arXiv preprint arXiv:2403.18638 , 2024

  21. [30]

    Few-Shot Bioacoustic Event Detection Utilizing Spectro- Temporal Receptive Field,

    D. Min, H. Nam, and Y.-H. Park, “Few-Shot Bioacoustic Event Detection Utilizing Spectro- Temporal Receptive Field,” in Proc. INTER-NOISE, 2024

  22. [31]

    PRTFNet: HRTF Individualization for Accurate Spectral Cues Using a Compact PRTF,

    B.-Y. Ko, G.-T. Lee, H. Nam, and Y.-H. Park, “PRTFNet: HRTF Individualization for Accurate Spectral Cues Using a Compact PRTF,” IEEE Access, vol. 11, 2023

  23. [32]

    FilterAugment: An Acoustic Environmental Data Augmentation Method,

    B.-Y. Ko, Y.-H. Park, G.-T. Lee, and H. Nam, “FilterAugment: An Acoustic Environmental Data Augmentation Method,” in International Congress on Acoustics (ICA) , 2022

  24. [33]

    AudioLDM: Text-to-Audio Generation with Latent Diffusion Models,

    H. Liu, Z. Chen, Y. Yuan, X. Mei, X. Liu, D. Mandic, W. Wang, and M. D. Plumbley, “AudioLDM: Text-to-Audio Generation with Latent Diffusion Models,” in ICML, 2023

  25. [34]

    AudioGen: Textually Guided Audio Generation,

    F. Kreuk, G. Synnaeve, A. Polyak, U. Singer, A. D´ efossez, J. Copet, D. Parikh, Y. Taigman, and Y. Adi, “AudioGen: Textually Guided Audio Generation,” in International Conference on Learning Representations (ICLR), 2023

  26. [35]

    VIFS: An End-to-End Variational Inference for Foley Sound Synthesis,

    J. Lee, H. Nam, and Y.-H. Park, “VIFS: An End-to-End Variational Inference for Foley Sound Synthesis,” DCASE Challenge , 2023

  27. [36]

    mixup: Beyond Empirical Risk Minimiza- tion,

    H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz, “mixup: Beyond Empirical Risk Minimiza- tion,” in International Conference on Learning Representations (ICLR) , 2018

  28. [37]

    Training sound event detection on a heterogeneous dataset,

    N. Turpault and R. Serizel, “Training sound event detection on a heterogeneous dataset,” inDCASE Workshop, 2020

  29. [38]

    Analysis of weak labels for sound event tagging,

    N. Turpault, R. Serizel, and E. Vincent, “Analysis of weak labels for sound event tagging,” hal- 03203692, 2021

  30. [39]

    Dynamic Convolution: Attention Over Convolution Kernels,

    Y. Chen, X. Dai, M. Liu, D. Chen, L. Yuan, and Z. Liu, “Dynamic Convolution: Attention Over Convolution Kernels,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020

  31. [40]

    Attention is All you Need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polo- sukhin, “Attention is All you Need,” in Advances in Neural Information Processing Systems , 2017

  32. [42]

    Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results,

    A. Tarvainen and H. Valpola, “Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results,” in Advances in Neural Information Process- ing Systems , vol. 30, 2017

  33. [43]

    Multi-Dimensional Frequency Dynamic Convolution with Con- fident Mean Teacher for Sound Event Detection,

    S. Xiao, X. Zhang, and P. Zhang, “Multi-Dimensional Frequency Dynamic Convolution with Con- fident Mean Teacher for Sound Event Detection,” in ICASSP, 2023

  34. [44]

    AST-SED: An Effective Sound Event Detection Method Based on Audio Spectrogram Transformer,

    K. Li, Y. Song, L.-R. Dai, I. McLoughlin, X. Fang, and L. Liu, “AST-SED: An Effective Sound Event Detection Method Based on Audio Spectrogram Transformer,” in ICASSP, 2023. 89

  35. [45]

    Fine-Tune the Pretrained ATST Model for Sound Event Detection,

    N. Shao, X. Li, and X. Li, “Fine-Tune the Pretrained ATST Model for Sound Event Detection,” in ICASSP, 2024

  36. [46]

    Squeeze-and-Excitation Networks,

    J. Hu, L. Shen, and G. Sun, “Squeeze-and-Excitation Networks,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2018

  37. [47]

    Learnable Sparse Filterbank for Speaker Verification,

    J. Peng, R. Gu, L. Moˇ sner, O. Plchot, L. Burget, and J. ˇCernock´ y, “Learnable Sparse Filterbank for Speaker Verification,” in Proc. Interspeech, 2022

  38. [48]

    Selective Kernel Networks,

    X. Li, W. Wang, X. Hu, and J. Yang, “Selective Kernel Networks,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2019

  39. [49]

    An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” in International Conference on Learning...

  40. [50]

    Audio Set: An Ontology and Human-Labeled Dataset for Audio Events,

    J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio Set: An Ontology and Human-Labeled Dataset for Audio Events,” in ICASSP, 2017

  41. [51]

    Convolution- Augmented Transformer for Semi-Supervised Sound Event Detection,

    K. Miyazaki, T. Komatsu, T. Hayashi, S. Watanabe, T. Toda, and K. Takeda, “Convolution- Augmented Transformer for Semi-Supervised Sound Event Detection,” DCASE Challenge , 2020

  42. [52]

    Zheng USTC Team’s Submission for DCASE2021 Task4 – Semi- Supervised Sound Event Detection,

    X. Zheng, H. Chen, and Y. Song, “Zheng USTC Team’s Submission for DCASE2021 Task4 – Semi- Supervised Sound Event Detection,” DCASE Challenge , 2021

  43. [53]

    Pre-Training and Self-Training for Sound Event Detection in Domestic Environments,

    J. E. and R. Haeb-Umbach, “Pre-Training and Self-Training for Sound Event Detection in Domestic Environments,” DCASE Challenge , 2022

  44. [54]

    Semi-Supervised Sound Event Detection System for DCASE 2022 Task 4,

    K. He, X. Shu, S. Jia, and Y. He, “Semi-Supervised Sound Event Detection System for DCASE 2022 Task 4,” DCASE Challenge , 2022

  45. [55]

    Data Engineering for Noisy Student Model in Sound Event Detection,

    S. Suh and D. Y. Lee, “Data Engineering for Noisy Student Model in Sound Event Detection,” DCASE Challenge , 2022

  46. [56]

    Pretrained Models in Sound Event Detection for DCASE 2022 Challenge Task4,

    S. Xiao, “Pretrained Models in Sound Event Detection for DCASE 2022 Challenge Task4,” DCASE Challenge, 2022

  47. [57]

    Semi-Supervised Learning- Based Sound Event Detection Using Frequency Dynamic Convolution with Large Kernel Attention for DCASE Challenge 2023 Task 4,

    J. W. Kim, S. W. Son, Y. Song, H. K. Kim, I. H. Song, and J. E. Lim, “Semi-Supervised Learning- Based Sound Event Detection Using Frequency Dynamic Convolution with Large Kernel Attention for DCASE Challenge 2023 Task 4,” DCASE Challenge , 2023

  48. [58]

    How Information on Soft Labels and Hard Labels Mutually Benefits Sound Event Detection Tasks,

    H. Yin, J. Bai, S. Huang, and J. Chen, “How Information on Soft Labels and Hard Labels Mutually Benefits Sound Event Detection Tasks,” DCASE Challenge , 2023

  49. [59]

    Sound Event Detection with Weak Prediction for DCASE 2023 Challenge Task4A,

    S. Xiao, J. Shen, A. Hu, X. Zhang, P. Zhang, and Y. Yan, “Sound Event Detection with Weak Prediction for DCASE 2023 Challenge Task4A,” DCASE Challenge , 2023

  50. [60]

    CHT+NSYSU Sound Event Detection System With Multiscale Channel Attention And Multiple Consistency Training For DCASE 2021 Task 4,

    Y.-W. Wang, C.-P. Chen, C.-L. Lu, and B.-C. Chan, “CHT+NSYSU Sound Event Detection System With Multiscale Channel Attention And Multiple Consistency Training For DCASE 2021 Task 4,” DCASE Challenge , 2021. 90

  51. [61]

    Sound Event Detection Based on Self-Supervised Learning of Wav2vec 2.0,

    H. Koo, H.-M. Park, J. Park, and M. Oh, “Sound Event Detection Based on Self-Supervised Learning of Wav2vec 2.0,” DCASE Challenge , 2021

  52. [62]

    Convolution-Augmented Conformer for Sound Event Detection,

    Y.-H. Chen, “Convolution-Augmented Conformer for Sound Event Detection,” DCASE Challenge , 2021

  53. [63]

    Convolutional Network with Conformer for Semi-Supervised Sound Event Detection,

    T. Na and Q. Zhang, “Convolutional Network with Conformer for Semi-Supervised Sound Event Detection,” DCASE Challenge , 2021

  54. [64]

    Integrating Advantages of Recurrent and Transformer Struc- tures for Sound Event Detection in Multiple Scenarios,

    R. Lu, W. Hu, Z. Duan, and J. Liu, “Integrating Advantages of Recurrent and Transformer Struc- tures for Sound Event Detection in Multiple Scenarios,” DCASE Challenge , 2021

  55. [65]

    Leveraging Audio-Tagging Assisted Sound Event Detection Using Weakified Strong Labels and Frequency Dynamic Convolutions,

    T. Khandelwal, R. K. Das, A. Koh, and E. S. Chng, “Leveraging Audio-Tagging Assisted Sound Event Detection Using Weakified Strong Labels and Frequency Dynamic Convolutions,” arXiv preprint arXiv:2304.12688, 2023

  56. [66]

    Semi-Supervised Sound Event Detection with Pre- Trained Model,

    L. Xu, L. Wang, S. Bi, H. Liu, and J. Wang, “Semi-Supervised Sound Event Detection with Pre- Trained Model,” in ICASSP, 2023

  57. [67]

    ATST: Audio Representation Learning with Teacher-Student Transformer,

    X. Li and X. Li, “ATST: Audio Representation Learning with Teacher-Student Transformer,” in Proc. Interspeech, 2022

  58. [68]

    Self-Supervised Audio Teacher-Student Transformer for Both Clip-Level and Frame-Level Tasks,

    X. Li, N. Shao, and X. Li, “Self-Supervised Audio Teacher-Student Transformer for Both Clip-Level and Frame-Level Tasks,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2024

  59. [69]

    Multi-Scale Context Aggregation by Dilated Convolutions,

    F. Yu and V. Koltun, “Multi-Scale Context Aggregation by Dilated Convolutions,” in International Conference on Learning Representations (ICLR) , 2016

  60. [70]

    Sound Event Detection with Depth- wise Separable and Dilated Convolutions,

    K. Drossos, S. I. Mimilakis, S. Gharib, Y. Li, and T. Virtanen, “Sound Event Detection with Depth- wise Separable and Dilated Convolutions,” in International Joint Conference on Neural Networks , 2020, doi:10.1109/IJCNN48605.2020.9207532

  61. [71]

    Sound Event Detection Via Dilated Convolutional Recurrent Neural Networks,

    Y. Li, M. Liu, K. Drossos, and T. Virtanen, “Sound Event Detection Via Dilated Convolutional Recurrent Neural Networks,” in ICASSP, 2020, doi:10.1109/ICASSP40776.2020.9054433

  62. [72]

    Dilated Convolution Neural Network with LeakyReLU for Envi- ronmental Sound Classification,

    X. Zhang, Y. Zou, and W. Shi, “Dilated Convolution Neural Network with LeakyReLU for Envi- ronmental Sound Classification,” in International Conference on Digital Signal Processing , 2017, doi:10.1109/ICDSP.2017.8096153

  63. [73]

    Efficient Large-Scale Audio Tagging Via Transformer-to- CNN Knowledge Distillation,

    F. Schmid, K. Koutini, and G. Widmer, “Efficient Large-Scale Audio Tagging Via Transformer-to- CNN Knowledge Distillation,” in ICASSP, 2023

  64. [74]

    Dynamic Convolutional Neural Networks as Efficient Pre-Trained Audio Models,

    F. Schmid, K. Koutini, and G. Widmer, “Dynamic Convolutional Neural Networks as Efficient Pre-Trained Audio Models,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2024

  65. [75]

    BERT: Pre-training of Deep Bidirec- tional Transformers for Language Understanding,

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirec- tional Transformers for Language Understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics , pp. 4171–4186, 2019, d...

  66. [76]

    A. V. Oppenheim and R. W. Schafer, Discrete-Time Signal Processing: International Edition , Pear- son, 3rd ed., 2010, pp. 522–525, 850–851, ISBN: 9780131988422

  67. [77]

    L. E. Kinsler, A. R. Frey, A. B. Coppens, and J. V. Sanders, Fundamentals of Acoustics, Wiley, 4th ed., 2000, pp. 149, 210, 224, 291–296, 333–334, ISBN: 9780471847892

  68. [78]

    J. S. Bendat and A. G. Piersol, Random Data: Analysis and Measurement Procedures , Wiley, 4th ed., 2011, pp. 8-12, 123, ISBN: 9781118210826

  69. [79]

    D. J. Inman, Engineering Vibrations, Pearson, 4th ed., 2013, pp. 172-177

  70. [80]

    Rabiner and R

    L. Rabiner and R. Schafer, Theory and Applications of Digital Speech Processing , Pearson, 1st ed., 2010, pp. 89-123, ISBN: 0136034284

  71. [81]

    DCASE2021 Task4 Baseline,

    N. Turpault, “DCASE2021 Task4 Baseline,” GitHub, Available: https://github.com/ DCASE-REPO/DESED_task

  72. [82]

    DCASE 2021 Challenge Task4: Sound event detection and separation in domestic environments,

    DCASE, “DCASE 2021 Challenge Task4: Sound event detection and separation in domestic environments,” Available: http://dcase.community/challenge2021/ task-sound-event-detection-and-separation-in-domestic-environments

  73. [83]

    Threshold Independent Evaluation of Sound Event Detection Scores,

    J. Ebbers, R. Haeb-Umbach, and R. Serizel, “Threshold Independent Evaluation of Sound Event Detection Scores,” in ICASSP, 2022

  74. [84]

    MAT-SED: A Masked Audio Transformer with Masked-Reconstruction Based Pre-training for Sound Event Detection,

    P. Cai, Y. Song, K. Li, H. Song, and I. McLoughlin, “MAT-SED: A Masked Audio Transformer with Masked-Reconstruction Based Pre-training for Sound Event Detection,” in Proc. Interspeech, 2024

  75. [85]

    Sound Event Bounding Boxes,

    J. Ebbers, F. G. Germain, G. Wichern, and J. L. Roux, “Sound Event Bounding Boxes,” in Proc. Interspeech, 2024

  76. [86]

    Prototype Based Masked Audio Model for Self-Supervised Learning of Sound Event Detection,

    P. Cai, Y. Song, N. Jiang, Q. Gu, and I. McLoughlin, “Prototype Based Masked Audio Model for Self-Supervised Learning of Sound Event Detection,” arXiv preprint arXiv:2409.17656 , 2024

  77. [87]

    Efficient Training of Audio Transformers with Patchout,

    K. Koutini, J. Schl¨ uter, H. Eghbal-zadeh, and G. Widmer, “Efficient Training of Audio Transformers with Patchout,” in Proc. Interspeech, 2022

  78. [88]

    IMPROVING AUDIO SPECTRO- GRAM TRANSFORMERS FOR SOUND EVENT DETECTION THROUGH MULTI-STAGE TRAINING,

    F. Schmid, P. Primus, T. Morocutti, J. Greif, and G. Widmer, “IMPROVING AUDIO SPECTRO- GRAM TRANSFORMERS FOR SOUND EVENT DETECTION THROUGH MULTI-STAGE TRAINING,” DCASE2024 Challenge , 2024

  79. [89]

    The Ins and Outs of Speaker Recognition: Lessons from VoxSRC 2020,

    Y. Kwon, H.-S. Heo, B.-J. Lee, and J. S. Chung, “The Ins and Outs of Speaker Recognition: Lessons from VoxSRC 2020,” in ICASSP, pp. 5809-5813, 2021, doi:10.1109/ICASSP39728.2021.9413948

  80. [90]

    In Defence of Metric Learning for Speaker Recognition,

    J. S. Chung, J. Huh, S. Mun, M. Lee, H.-S. Heo, S. Choe, C. Ham, S. Jung, B.-J. Lee, and I. Han, “In Defence of Metric Learning for Speaker Recognition,” in Proc. Interspeech, 2020

  81. [91]

    SSAST: Self-Supervised Audio Spectrogram Trans- former,

    Y. Gong, C.-I. Lai, Y.-A. Chung, and J. Glass, “SSAST: Self-Supervised Audio Spectrogram Trans- former,” Proceedings of the AAAI Conference on Artificial Intelligence , 2022

  82. [92]

    Masked Autoencoders that Listen,

    P.-Y. Huang, H. Xu, J. Li, A. Baevski, M. Auli, W. Galuba, F. Metze, and C. Feichtenhofer, “Masked Autoencoders that Listen,” in Advances in Neural Information Processing Systems , 2022. 92

  83. [93]

    BYOL for Audio: Self-Supervised Learning for General-Purpose Audio Representation,

    D. Niizumi, D. Takeuchi, Y. Ohishi, N. Harada, and K. Kashino, “BYOL for Audio: Self-Supervised Learning for General-Purpose Audio Representation,” in 2021 International Joint Conference on Neural Networks (IJCNN) , 2021

  84. [94]

    Jigsaw Clustering for Unsupervised Visual Representation Learning,

    P. Chen, S. Liu, and J. Jia, “Jigsaw Clustering for Unsupervised Visual Representation Learning,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2021

  85. [95]

    Jigsaw-ViT: Learning Jigsaw Puzzles in Vision Transformer,

    Y. Chen, X. Shen, Y. Liu, Q. Tao, and J. A. K. Suykens, “Jigsaw-ViT: Learning Jigsaw Puzzles in Vision Transformer,” Pattern Recognition Letters, 2023

  86. [96]

    Shuffle and Learn: Unsupervised Learning Using Temporal Order Verification,

    I. Misra, C. L. Zitnick, and M. Hebert, “Shuffle and Learn: Unsupervised Learning Using Temporal Order Verification,” in ECCV, 2016

  87. [97]

    TERA: Self-Supervised Learning of Transformer Encoder Representation for Speech,

    A. T. Liu, S.-W. Li, and H.-Y. Lee, “TERA: Self-Supervised Learning of Transformer Encoder Representation for Speech,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2021

  88. [98]

    Transformer-XL: Language Modeling with Longer-Term Dependency,

    Z. Dai, Z. Yang, Y. Yang, W. W. Cohen, J. Carbonell, Q. V. Le, and R. Salakhutdinov, “Transformer-XL: Language Modeling with Longer-Term Dependency,” in International Confer- ence on Learning Representations (ICLR) , 2019

  89. [99]

    You Only Hear Once: A YOLO-Like Algorithm for Audio Segmentation and Sound Event Detection,

    S. Venkatesh, D. Moffat, and E. R. Miranda, “You Only Hear Once: A YOLO-Like Algorithm for Audio Segmentation and Sound Event Detection,” Applied Sciences, 2022

  90. [100]

    Adaptive Convolutional Neural Network for Text-Independent Speaker Recognition,

    S.-H. Kim and Y.-H. Park, “Adaptive Convolutional Neural Network for Text-Independent Speaker Recognition,” in Proc. Interspeech, 2021

  91. [101]

    Temporal Dynamic Convolutional Neural Network for Text- Independent Speaker Verification and Phonemetic Analysis,

    S.-H. Kim, H. Nam, and Y.-H. Park, “Temporal Dynamic Convolutional Neural Network for Text- Independent Speaker Verification and Phonemetic Analysis,” in ICASSP, 2022

  92. [102]

    Decomposed Temporal Dynamic CNN: Efficient Time- Adaptive Network for Text-Independent Speaker Verification Explained with Speaker Activation Map,

    S.-H. Kim, H. Nam, and Y.-H. Park, “Decomposed Temporal Dynamic CNN: Efficient Time- Adaptive Network for Text-Independent Speaker Verification Explained with Speaker Activation Map,” arXiv preprint arXiv:2203.15277 , 2022

  93. [103]

    Application of Spectro-Temporal Receptive Field on Soft Labeled Sound Event Detection,

    D. Min, H. Nam, and Y.-H. Park, “Application of Spectro-Temporal Receptive Field on Soft Labeled Sound Event Detection,” DCASE Challenge , 2023

  94. [104]

    Auditory Neural Response Inspired Sound Event Detection Based on Spectro-Temporal Receptive Field,

    D. Min, H. Nam, and Y.-H. Park, “Auditory Neural Response Inspired Sound Event Detection Based on Spectro-Temporal Receptive Field,” in DCASE Workshop , 2023

  95. [105]

    Few-Shot Bioacoustic Event Detection Utilizing Spectro-Temporal Receptive Field,

    B.-Y. Ko, H. Nam, D. Min, G.-T. Lee, and Y.-H. Park, “Few-Shot Bioacoustic Event Detection Utilizing Spectro-Temporal Receptive Field,” in Proc. INTER-NOISE, 2023

  96. [106]

    FilterAugment: An Acoustic Environmental Data Augmen- tation Method,

    H. Nam, S.-H. Kim, and Y.-H. Park, “FilterAugment: An Acoustic Environmental Data Augmen- tation Method,” in ICASSP, 2022

  97. [107]

    Frequency Dynamic Convolution: Frequency- Adaptive Pattern Recognition for Sound Event Detection,

    H. Nam, S.-H. Kim, B.-Y. Ko, and Y.-H. Park, “Frequency Dynamic Convolution: Frequency- Adaptive Pattern Recognition for Sound Event Detection,” in Proc. Interspeech, 2022

  98. [108]

    Frequency & Channel Attention for Computationally Efficient Sound Event Detection,

    H. Nam, S.-H. Kim, D. Min, and Y.-H. Park, “Frequency & Channel Attention for Computationally Efficient Sound Event Detection,” in DCASE Workshop , 2023. 93

  99. [109]

    Heavily Augmented Sound Event Detection Utilizing Weak Predictions,

    H. Nam, B.-Y. Ko, G.-T. Lee, S.-H. Kim, W.-H. Jung, S.-M. Choi, and Y.-H. Park, “Heavily Augmented Sound Event Detection Utilizing Weak Predictions,” DCASE Challenge , 2021

  100. [110]

    Frequency Dependent Sound Event Detection for DCASE 2022 Challenge Task 4,

    H. Nam, S.-H. Kim, D. Min, B.-Y. Ko, S.-D. Choi, and Y.-H. Park, “Frequency Dependent Sound Event Detection for DCASE 2022 Challenge Task 4,” DCASE Challenge , 2022

  101. [111]

    Diversifying and Expanding Frequency- Adaptive Convolution Kernels for Sound Event Detection,

    H. Nam, S.-H. Kim, D. Min, J. Lee, and Y.-H. Park, “Diversifying and Expanding Frequency- Adaptive Convolution Kernels for Sound Event Detection,” in Proc. Interspeech, 2024

  102. [112]

    Pushing the Limit of Sound Event Detection with Multi-Dilated Fre- quency Dynamic Convolution,

    H. Nam and Y.-H. Park, “Pushing the Limit of Sound Event Detection with Multi-Dilated Fre- quency Dynamic Convolution,” arXiv preprint arXiv:2406.13312 , 2024

  103. [113]

    Self Training and Ensembling Frequency Dependent Networks with Coarse Prediction Pooling and Sound Event Bounding Boxes,

    H. Nam, D. Min, I. Choi, S.-D. Choi, and Y.-H. Park, “Self Training and Ensembling Frequency Dependent Networks with Coarse Prediction Pooling and Sound Event Bounding Boxes,” DCASE Challenge, 2024

  104. [114]

    Self Training and Ensembling Frequency Dependent Networks with Coarse Prediction Pooling and Sound Event Bounding Boxes,

    H. Nam, D. Min, I. Choi, S.-D. Choi, and Y.-H. Park, “Self Training and Ensembling Frequency Dependent Networks with Coarse Prediction Pooling and Sound Event Bounding Boxes,” inDCASE Workshop, 2024

  105. [115]

    Towards Understanding of Frequency Dependence on Sound Event Detection,

    H. Nam, S.-H. Kim, D. Min, B.-Y. Ko, and Y.-H. Park, “Towards Understanding of Frequency Dependence on Sound Event Detection,” arXiv preprint arXiv:2502.07208 , 2025

  106. [116]

    JiTTER: Jigsaw Temporal Transformer for Event Reconstruction for Self-Supervised Sound Event Detection,

    H. Nam and Y.-H. Park, “JiTTER: Jigsaw Temporal Transformer for Event Reconstruction for Self-Supervised Sound Event Detection,” arXiv preprint arXiv:2502.20857 , 2025

  107. [117]

    Temporal Attention Pooling for Frequency Dynamic Convolution in Sound Event Detection,

    H. Nam and Y.-H. Park, “Temporal Attention Pooling for Frequency Dynamic Convolution in Sound Event Detection,” arXiv preprint arXiv:2504.12670 , 2025

  108. [118]

    Multi-output Classification Framework and Frequency Layer Normalization for Compound Fault Diagnosis in Motor,

    W. Yi and Y.-H. Park, “Multi-output Classification Framework and Frequency Layer Normalization for Compound Fault Diagnosis in Motor,” arXiv preprint arXiv:2504.11513 , 2025

  109. [119]

    Multi-output Classification using a Cross-talk Archi- tecture for Compound Fault Diagnosis of Motors in Partially Labeled Condition,

    W. Yi, W. Jung, K. Jang and Y.-H. Park, “Multi-output Classification using a Cross-talk Archi- tecture for Compound Fault Diagnosis of Motors in Partially Labeled Condition,” arXiv preprint arXiv:2505.24001, 2025. 94 Acknowledgments in Korean 카이스트에학부신입생으로써제일 처음교문을지나던 날이 생각납니다....

  110. [125]

    Heavily augmented sound event detection utilizing weak predictions,

    H. Nam , B. Y. Ko, G. T. Lee, S. -H. Kim, W. H. Jung, S. M. Choi and Y. -H. Park, “Heavily augmented sound event detection utilizing weak predictions,” DCASE Challenge Tech. rep. , 2021

  111. [126]

    FilterAugment: An Acoustic Environmental Data Augmen- tation Method,

    H. Nam, S. -H. Kim and Y. -H. Park, “FilterAugment: An Acoustic Environmental Data Augmen- tation Method,” International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2022

  112. [127]

    Temporal dynamic convolutional neural network for text- independent speaker verification and phonemic analysis,

    S. -H. Kim, H. Nam and Y. -H. Park, “Temporal dynamic convolutional neural network for text- independent speaker verification and phonemic analysis,” International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2022

  113. [128]

    Frequency dependent sound event detection for DCASE 2022 Challenge Task 4,

    H. Nam, S. -H. Kim, D. Min, B. Y. Ko, S. D. Choi and Y. -H. Park, “Frequency dependent sound event detection for DCASE 2022 Challenge Task 4,” DCASE Challenge Tech. rep. , 2022

  114. [129]

    Data augmentation and squeeze-and-excitation network on multiple dimension for sound event localization and detection in real scenes,

    B. Y. Ko, H. Nam , S. -H. Kim, D. Min, S. Choi and Y. -H. Park, “Data augmentation and squeeze-and-excitation network on multiple dimension for sound event localization and detection in real scenes,” DCASE Challenge Tech. rep. , 2022

  115. [130]

    Frequency Dynamic Convolution: Frequency- Adaptive Pattern Recognition for Sound Event Detection,

    H. Nam , S. -H. Kim, B. Y. Ko and Y. -H. Park, “Frequency Dynamic Convolution: Frequency- Adaptive Pattern Recognition for Sound Event Detection,” Interspeech, 2022

  116. [131]

    Deep learning based cough detection camera using enhanced features,

    G. T. Lee, H. Nam, S. -H. Kim, S. M. Choi and Y. -H. Park, “Deep learning based cough detection camera using enhanced features,” Expert Systems with Applications , 2022. 99

  117. [132]

    VIFS: An end-to-end variational inference for foley sound synthesis,

    J. Lee, H. Nam and Y. -H. Park, “VIFS: An end-to-end variational inference for foley sound synthesis,” DCASE Challenge Tech. rep. , 2023

  118. [133]

    Application of spectro-temporal receptive field on soft labeled sound event detection,

    D. Min, H. Nam and Y. -H. Park, “Application of spectro-temporal receptive field on soft labeled sound event detection,” DCASE Challenge Tech. rep. , 2023

  119. [134]

    Analysis-based optimization of temporal dynamic convo- lutional neural network for text-independent speaker verification,

    S. -H. Kim, H. Nam and Y. -H. Park, “Analysis-based optimization of temporal dynamic convo- lutional neural network for text-independent speaker verification,” IEEE Access, 2023

  120. [135]

    Frequency & Channel Attention for Computationally Efficient Sound Event Detection,

    H. Nam, S. -H. Kim, D. Min and Y. -H. Park, “Frequency & Channel Attention for Computationally Efficient Sound Event Detection,” DCASE Workshop , 2023

  121. [136]

    Auditory Neural Response Inspired Sound Event Detection Based on Spectro-temporal Receptive Field,

    D. Min, H. Nam and Y. -H. Park, “Auditory Neural Response Inspired Sound Event Detection Based on Spectro-temporal Receptive Field,” DCASE Workshop , 2023

  122. [137]

    PRTFNet: HRTF Individualization for Accurate Spectral Cues Using a Compact PRTF,

    B. Y. Ko, G. T. Lee, H. Nam and Y. -H. Park, “PRTFNet: HRTF Individualization for Accurate Spectral Cues Using a Compact PRTF,” IEEE Access, 2023

  123. [138]

    Real-Time Sound Recognition System for Human Care Robot Considering Custom Sound Events,

    S. -H. Kim, H. Nam , S. M. Choi and Y. -H. Park, “Real-Time Sound Recognition System for Human Care Robot Considering Custom Sound Events,” IEEE Access, 2024

  124. [139]

    ChatGPT Caption Paraphrasing and FENSE- based Caption Filtering for Automated Audio Captioning,

    I. Choi, H. Nam, D. Min, S. Choi and Y. -H. Park, “ChatGPT Caption Paraphrasing and FENSE- based Caption Filtering for Automated Audio Captioning,” DCASE Challenge Tech. rep. , 2024

  125. [140]

    Diversifying and Expanding Frequency- Adaptive Convolution Kernels for Sound Event Detection,

    H. Nam , S. -H. Kim, D. Min, J. Lee and Y. -H. Park, “Diversifying and Expanding Frequency- Adaptive Convolution Kernels for Sound Event Detection,” Interspeech, 2024

  126. [141]

    Self Training and Ensembling Frequency Dependent Networks with Coarse Prediction Pooling and Sound Event Bounding Boxes,

    H. Nam , D. Min, S. Choi, I. Choi and Y. -H. Park, “Self Training and Ensembling Frequency Dependent Networks with Coarse Prediction Pooling and Sound Event Bounding Boxes,” DCASE Workshop, 2024

  127. [142]

    Coherence-based Phonemic Analysis on the Effect of Reverberation to Practical Automatic Speech Recognition,

    H. Nam and Y. -H. Park, “Coherence-based Phonemic Analysis on the Effect of Reverberation to Practical Automatic Speech Recognition,” Applied Acoustics, 2025

  128. [143]

    Towards Understanding of Frequency Dependence on Sound Event Detection,

    H. Nam , S. -H. Kim, D. Min, B. Y. Ko and Y. -H. Park, “Towards Understanding of Frequency Dependence on Sound Event Detection,” arXiv, 2025

  129. [144]

    Pushing the limit of sound event detection with multi-dilated frequency dynamic convolution,

    H. Nam and Y. -H. Park, “Pushing the limit of sound event detection with multi-dilated frequency dynamic convolution,” arXiv, 2025

  130. [145]

    JiTTER: Jigsaw Temporal Transformer for Event Reconstruction for Self-Supervised Sound Event Detection,

    H. Nam and Y. -H. Park, “JiTTER: Jigsaw Temporal Transformer for Event Reconstruction for Self-Supervised Sound Event Detection,” arXiv, 2025

  131. [146]

    Temporal Attention Pooling for Frequency Dynamic Convolution in Sound Event Detection,

    H. Nam and Y. -H. Park, “Temporal Attention Pooling for Frequency Dynamic Convolution in Sound Event Detection,” arXiv, 2025

  132. [147]

    DNN based HRIRs Identification with a Continuously Rotating Speaker Array,

    B. Y. Ko, D. Min, H. Nam and Y. -H. Park, “DNN based HRIRs Identification with a Continuously Rotating Speaker Array,” arXiv, 2025. 100

  133. [2009]

    1. – 2012. 12. NUS High School of Math and Science, Singapore (NUS High diploma)

  134. [2013]

    9. – 2018. 2. 한국과학기술원기계공학과 (학사)

  135. [2017]

    3. – 2024. 2. 한국과학기술원기계공학과조교 연 구 업 적

  136. [2018]

    3. – 2020. 2. 한국과학기술원기계공학과 (석사)

  137. [2020]

    3. – 2025. 8. 한국과학기술원기계공학과 (박사) 경 력

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.