REVIEW 3 major objections 5 minor 51 references
Auditory Intelligence: Understanding the World Through Sound
T0 review · 3 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read Machine hearing should be reframed from surface recognition into a layered, situated process of perception, reasoning, and interaction, instantiated by four new task paradigms.
desk verdict An honest, well-scoped position paper that repackages existing tasks into a layered roadmap; the real risk is whether the highest cognitive layers (explanation and intent) can be annotated reliably enough to train and evaluate. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing structure is a three-layer cognitive model of auditory intelligence—perceptual recognition, contextual reasoning, generative interaction—operationalized as a task pipeline that converts audio into structured text. ASPIRE first textualizes spectrograms into spectro-temporal descriptors such as “broadband noise from 1.5–4 kHz for 0.8 s”; SODA organizes those descriptions into an open event→context→scene hierarchy; AUX adds causal/explanatory sentences; AUGMENT fills semantic slots (trigger, event, motivation, goal). This stack is the mechanism that would let recognition, explanation, and interaction share one representational scaffold, and it is what shifts evaluation from clo
What would settle it
Annotate a sample of DESED foreground clips with the ASPIRE pattern vocabulary and with SODA and AUGMENT slots using two independent experts, then train the proposed hierarchy with and without the event→context→scene structure on a DESED-derived corpus. If expert agreement on pattern descriptions or on motivation/goal slots is near chance (e.g., Cohen’s kappa below 0.6), or if the hierarchical model does not beat a flat multi-task baseline on unseen scenes, the roadmap’s core empirical bet is settled against it.
Extended reading notes
Core claim
The central claim is that sound, for an intelligent agent, is not a raw signal to classify but a cognitive medium encoding events, intentions, and contexts. The author argues that machine hearing should therefore be built as a layered stack: low-level spectro-temporal evidence (ASPIRE) feeds structured event–context–scene descriptions (SODA), which support causal and explanatory captions (AUX), and ultimately goal-driven interpretation with trigger, event, motivation, and goal slots (AUGMENT). Generation and audio–text–image grounding complete the loop by turning perception and reasoning into action. Within this view, existing tasks such as SED, ASC, and AAC remain useful but underspecified,
Load-bearing premise
The load-bearing premise is that human annotators can label spectro-temporal patterns, event–context–scene hierarchies, and motivation/goal slots with usable agreement, and that models can learn them well enough to beat flat recognition baselines; the paper itself signals the risk by noting that even Clotho captions leak inferred content despite instructions to describe only what is heard.
Editorial extensions
If this is right
- Recognition benchmarks like SED, ASC, and AAC would be reframed as components of a larger stack, with evaluation newly focused on structure fidelity across event, context, and scene, and on consistency between levels.
- Datasets would need layered annotations—spectro-temporal pattern descriptions, open hierarchical scene structures, and a clean separation of observed facts from inferred explanations—including a re-annotation of Clotho.
- Models trained under SODA and AUX should generalize to unseen classes through text prototypes and should be able to justify decisions by citing explicit acoustic evidence such as timing, loudness, and frequency ranges.
- AUGMENT would give assistive systems and robots access to the motivations and goals behind everyday sounds, enabling intent-aware interaction and safety diagnostics.
- Audio generation would be conditioned on structured perception outputs, producing sound effects, speech, and music that stay consistent with inferred scene, affect, and goals, and audio–text–image grounding would extend the same structure to cross-modal retrieval.
Reading between the lines
- Inference: if ASPIRE’s spectro-temporal description is learned first, downstream SODA and AUX models may need far less paired audio–text data, because the intermediate description supplies compositional structure; the paper suggests transferability but does not quantify this benefit.
- Inference: the proposed observation/explanation split for Clotho implies a direct grounding probe the paper does not design: train an explanation model with and without audio input at test time; the performance gap measures how much of the explanation is acoustically grounded rather than carried by language priors.
- Inference: the same event→context→scene and trigger/event/motivation/goal structure could be applied to speech and music understanding—for instance, attributing speaker intent from prosody—which would make the framework a candidate unified benchmark for all non-speech auditory cognition, not just environmental sound.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper is a position paper arguing that current machine-listening research (SED, ASC, AAC, AQA) is limited to surface-level recognition and should be reframed as layered, situated auditory intelligence encompassing perception, contextual reasoning, and generative interaction. To instantiate this view, it proposes four task paradigms: ASPIRE (spectro-temporal pattern parsing), SODA (hierarchical event-to-context-to-scene description), AUX (causal/explanatory audio captioning), and AUGMENT (trigger/event/motivation/goal slot filling), plus extensions to generative audio and audio–text–image grounding. No model or quantitative evaluation is presented; the paper explicitly positions itself as a research agenda and invites collaboration. It is candid about open challenges (Section V) and about the conditional nature of the proposals, but the paper's strongest claims depend on the tractability of the proposed annotation and learning tasks, especially at the higher cognitive layers.
Significance. If the proposed task formulations prove feasible and learnable, they would provide a valuable organizing framework and shift benchmarking from isolated recognition to structured, explainable, and interaction-oriented evaluation. The paper’s strengths are its clarity, the concreteness of the named paradigms (ASPIRE, SODA, AUX, AUGMENT), and its explicit acknowledgment of absent quantitative support and of open problems (Section V and the footnote in Section I). The central risk, which I think is real, is that the highest layers (AUX and especially AUGMENT) require human annotations of causes, motivations, and goals from non-speech audio, and the paper offers no evidence that such annotations can be made reliably. The claim that the roadmap leads to 'human-aligned auditory intelligence' therefore rests on an empirical bet that the paper does not yet support.
major comments (3)
- [Section IV-D and Table I] AUGMENT defines semantic slots for trigger, event, motivation, and goal. The last two slots require annotators to infer intentions and goals that are not directly observable in the acoustic signal. The paper provides no inter-annotator agreement data or pilot protocol; Table I lists 'inter-annotator agreement' only as an evaluation sketch. This matters because the paper's own Section IV-C notes that even Clotho captions, collected under an explicit 'describe only what is heard' instruction, leak inference—demonstrating how difficult the observation/explanation split is in practice. Since AUGMENT is the layer that supports 'intent-aware interaction' and the 'human-aligned' claim, the strongest version of the roadmap depends on this feasibility. I request a small-scale annotation feasibility study (even on a few examples) or an explicit downgrade of AUGMENT from a planned paradigm to a lon
- [Section IV-B, step 2] The SODA roadmap asserts that hierarchy-based learning (event→context→scene) 'outperforms single-task and flat multi-task baselines,' but gives no evidence or reference for this expectation. The claimed generalization and data-efficiency benefits of SODA hinge on this premise. A position paper is not required to contain full experiments, but to make the roadmap argumentative rather than assertive, it should cite relevant multi-task or hierarchical representation-learning results, or present a concrete pilot comparison plan with a minimal dataset. As written, this is a testable hypothesis, not a supported component of the roadmap—please mark it as such.
- [Section IV-A] ASPIRE is introduced as the foundational layer that textualizes spectrograms into low-level pattern descriptions. It assumes (a) expert annotators can reliably produce such descriptions with the proposed primitive vocabulary, and (b) an acoustic model and language model can map these descriptions to event concepts. No annotation protocol, vocabulary evaluation, or agreement evidence is provided. The explainability and transferability benefits claimed for ASPIRE rest on this assumption. Adding a pilot annotation study on a small set of DESED foreground events, or citing perceptual studies of spectrogram reading, would materially strengthen the proposal.
minor comments (5)
- [Abstract] The phrase 'those structure auditory understanding' should read 'that structure auditory understanding'; the relative clause is grammatically incorrect.
- [Figure 1 caption] The caption notes that 'SODA, AUX, and AUGMENT can also be developed without ASPIRE,' but Section IV-A describes ASPIRE as foundational to the other paradigms. Please reconcile this tension, as the architecture's dependence is currently ambiguous.
- [Section IV-D] The acronym AUGMENT is derived from 'Goal, Motivation, Event, and Trigger,' but the output order is listed as 'trigger, event, motivation, goal.' Align the acronym derivation and the slot order to avoid confusion.
- [Section V] The list of open challenges could explicitly mention the annotation feasibility of AUGMENT and AUX, since that is the riskiest component of the proposed roadmap.
- [General] The first-person 'I' and the footnote about the author's career transition are transparent but unusual for an archival paper; consider moving such material to an acknowledgment or a less prominent note.
Circularity Check
No significant circularity: the paper is a position paper that makes no derivation claim; self-citations are background motivation, not load-bearing.
full rationale
This is a conceptual position paper, not a technical derivation. It explicitly states: 'This work does not introduce a new model; instead, it offers a research lens and organizing scaffold for future datasets, benchmarks, and systems.' There are no equations, no fitted parameters, and no empirical predictions that reduce to inputs. The four proposed paradigms (ASPIRE, SODA, AUX, AUGMENT) are defined as new task formulations and research proposals; they are not derived from prior results. Self-citations (roughly 20 of 51 references) are used to support the background claim that SED/SER recognition is 'mature and yield[s] diminishing returns' (Section IV), but this claim is not load-bearing for the proposed roadmap: the roadmap would stand even if recognition were immature. The paper also explicitly postpones validation, proposing experiments to test whether hierarchy-aware training outperforms flat baselines and whether inter-annotator agreement can be reached for AUGMENT slots. These are acknowledged as open empirical bets, not as already-established results. The only oblique circular-adjacent point, that Clotho captions already leak inference (Section IV-C), is used as a motivation for the observation/explanation split, not as a result that presupposes the conclusion. No self-definitional, fitted-input, or self-citation-chain circularity is present.
Assumptions & free parameters
assumptions (4)
- domain assumption Spectrograms are partially readable by trained listeners and spectro-temporal pattern descriptions can be annotated reliably and learned as an intermediate representation.
- domain assumption Hierarchical event->context->scene learning outperforms flat single-task or flat multi-task recognition.
- domain assumption Causal explanations (AUX) and trigger/event/motivation/goal slots (AUGMENT) can be annotated with usable inter-annotator agreement for everyday non-speech audio.
- domain assumption LLMs can map compositional spectro-temporal pattern descriptions to event concepts, enabling generalization beyond the seed classes.
invented entities (4)
-
ASPIRE (Acoustic Structure-based Parsing and Interpretation for REcognition)
-
SODA (Structured Open Description of Acoustics) and the planned HEAR dataset
-
AUX (Acoustic Understanding and eXplanation)
-
AUGMENT (Acoustic Understanding via Goal, Motivation, Event, and Trigger)
Cite this review
Pith. "Pith review of Auditory Intelligence: Understanding the World Through Sound." pith.science (2026). https://pith.science/paper/K3LIWOLP
@misc{pith2026250807829,
author = {Pith},
title = {Pith review of: Auditory Intelligence: Understanding the World Through Sound},
year = {2026},
howpublished = {\url{https://pith.science/paper/K3LIWOLP}},
note = {Machine review of arXiv:2508.07829}
}
read the original abstract
Recent progress in auditory intelligence has yielded high-performing systems for sound event detection (SED), acoustic scene classification (ASC), automated audio captioning (AAC), and audio question answering (AQA). Yet these tasks remain largely constrained to surface-level recognition-capturing what happened but not why, what it implies, or how it unfolds in context. I propose a conceptual reframing of auditory intelligence as a layered, situated process that encompasses perception, reasoning, and interaction. To instantiate this view, I introduce four cognitively inspired task paradigms-ASPIRE, SODA, AUX, and AUGMENT-those structure auditory understanding across time-frequency pattern captioning, hierarchical event/scene description, causal explanation, and goal-driven interpretation, respectively. Together, these paradigms provide a roadmap toward more generalizable, explainable, and human-aligned auditory intelligence, and are intended to catalyze a broader discussion of what it means for machines to understand sound.
Figures
Reference graph
Works this paper leans on
-
[1]
OpenAI, “Gpt-4 technical report,” arXiv preprint arXiv:2303.08774 , 2024
arXiv 2024
-
[2]
Llama 2: Open foundation and fine-tuned chat models,
MetaGenAI, “Llama 2: Open foundation and fine-tuned chat models,” arXiv preprint arXiv:2307.09288 , 2023
arXiv 2023
-
[3]
DeepSeek-AI, “Deepseek-v3 technical report,” arXiv preprint arXiv:2412.19437, 2025
arXiv 2025
-
[4]
T. Virtanen, M. D. Plumbley, and D. Ellis, Computational Analysis of Sound Scenes and Events , 1st ed. Springer Publishing Company, Incorporated, 2017, pp. 3–11, 71–77
work page 2017
-
[5]
Sound event detection in domestic environments with weakly labeled data and soundscape synthesis,
N. Turpault, R. Serizel, A. P. Shah, and J. Salamon, “Sound event detection in domestic environments with weakly labeled data and soundscape synthesis,” in DCASE Workshop, 2019
work page 2019
-
[6]
Convolutional recurrent neural networks for polyphonic sound event detection,
E. C ¸ akır, G. Parascandolo, T. Heittola, H. Huttunen, and T. Virtanen, “Convolutional recurrent neural networks for polyphonic sound event detection,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 25, no. 6, pp. 1291–1303, 2017
work page 2017
-
[7]
Metrics for polyphonic sound event detection,
A. Mesaros, T. Heittola, and T. Virtanen, “Metrics for polyphonic sound event detection,” Applied Sciences , vol. 6, no. 6, 2016
work page 2016
-
[8]
A framework for the robust evaluation of sound event detection,
C ¸ . Bilen, G. Ferroni, F. Tuveri, J. Azcarreta, and S. Krstulovi ´c, “A framework for the robust evaluation of sound event detection,” in ICASSP, 2020, pp. 61–65
work page 2020
Show all 51 references
-
[9]
Towards understanding of frequency dependence on sound event detection,
H. Nam, S.-H. Kim, D. Min, B.-Y . Ko, and Y .-H. Park, “Towards understanding of frequency dependence on sound event detection,” arXiv preprint arXiv:2502.07208, 2025
2025 arXiv
-
[10]
Jitter: Jigsaw temporal transformer for event reconstruction for self-supervised sound event detection,
H. Nam and Y .-H. Park, “Jitter: Jigsaw temporal transformer for event reconstruction for self-supervised sound event detection,” arXiv preprint arXiv:2502.20857, 2025
2025 arXiv
-
[11]
SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition,
D. S. Park, W. Chan, Y . Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V . Le, “SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition,” in Proc. Interspeech, 2019
2019
-
[12]
Conformer: Convolution- augmented Transformer for Speech Recognition,
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y . Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y . Wu, and R. Pang, “Conformer: Convolution- augmented Transformer for Speech Recognition,” in Proc. Interspeech, 2020
2020
-
[13]
Coherence-based phonemic analysis on the ef- fect of reverberation to practical automatic speech recognition,
H. Nam and Y .-H. Park, “Coherence-based phonemic analysis on the ef- fect of reverberation to practical automatic speech recognition,” Applied Acoustics, vol. 227, p. 110233, 2025
2025
-
[14]
wav2vec 2.0: A framework for self-supervised learning of speech representations,
A. Baevski, Y . Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in Advances in Neural Information Processing Systems , 2020
2020
-
[15]
Hubert: Self-supervised speech representation learning by masked prediction of hidden units,
W.-N. Hsu, B. Bolte, Y .-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2021
2021
-
[16]
Attentive statistics pooling for deep speaker embedding,
K. Okabe, T. Koshinaka, and K. Shinoda, “Attentive statistics pooling for deep speaker embedding,” in Proc. Interspeech, 2018
2018
-
[17]
Exploring the encoding layer and loss function in end-to-end speaker and language recognition system,
W. Cai, J. Chen, and M. Li, “Exploring the encoding layer and loss function in end-to-end speaker and language recognition system,” in Proc. Interspeech, 2018
2018
-
[18]
Analysis-based optimization of temporal dynamic convolutional neural network for text-independent speaker verification,
S.-H. Kim, H. Nam, and Y .-H. Park, “Analysis-based optimization of temporal dynamic convolutional neural network for text-independent speaker verification,” IEEE Access , vol. 11, 2023
2023
-
[19]
Integrating fre- quency translational invariance in tdnns and frequency positional in- formation in 2d resnets to enhance speaker verification,
J. Thienpondt, B. Desplanques, and K. Demuynck, “Integrating fre- quency translational invariance in tdnns and frequency positional in- formation in 2d resnets to enhance speaker verification,” in Proc. Interspeech, 2021
2021
-
[20]
Convolution-based channel-frequency atten- tion for text-independent speaker verification,
J. Li, Y . Tian, and T. Lee, “Convolution-based channel-frequency atten- tion for text-independent speaker verification,” in ICASSP, 2023
2023
-
[21]
Panns: Large-scale pretrained audio neural networks for audio pattern recognition,
Q. Kong, Y . Cao, T. Iqbal, Y . Wang, W. Wang, and M. D. Plumbley, “Panns: Large-scale pretrained audio neural networks for audio pattern recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2020
2020
-
[22]
Deep learning based cough detection camera using enhanced features,
G.-T. Lee, H. Nam, S.-H. Kim, S.-M. Choi, Y . Kim, and Y .-H. Park, “Deep learning based cough detection camera using enhanced features,” Expert Systems with Applications , vol. 206, 2022
2022
-
[23]
Real-time sound recognition system for human care robot considering custom sound events,
S.-H. Kim, H. Nam, S.-M. Choi, and Y .-H. Park, “Real-time sound recognition system for human care robot considering custom sound events,” IEEE Access , vol. 12, 2024
2024
-
[24]
Ast: Audio spectrogram trans- former,
Y . Gong, Y .-A. Chung, and J. Glass, “Ast: Audio spectrogram trans- former,” in Proc. Interspeech, 2021
2021
-
[25]
Beats: Audio pre-training with acoustic tokenizers,
S. Chen, Y . Wu, C. Wang, S. Liu, D. Tompkins, Z. Chen, W. Che, X. Yu, and F. Wei, “Beats: Audio pre-training with acoustic tokenizers,” in ICML, 2023
2023
-
[26]
Overview and evaluation of sound event localization and detection in dcase 2019,
A. Politis, A. Mesaros, S. Adavanne, T. Heittola, and T. Virtanen, “Overview and evaluation of sound event localization and detection in dcase 2019,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, 2020
2019
-
[27]
STARSS22: A dataset of spatial recordings of real scenes with spa- tiotemporal annotations of sound events,
A. Politis, K. Shimada, P. Sudarsanam, S. Adavanne, D. Krause, Y . Koyama, N. Takahashi, S. Takahashi, Y . Mitsufuji, and T. Virtanen, “STARSS22: A dataset of spatial recordings of real scenes with spa- tiotemporal annotations of sound events,” in DCASE Workshop, 2022
2022
-
[28]
Data augmentation and squeeze-and-excitation network on multiple dimension for sound event localization and detection in real scenes,
B.-Y . Ko, H. Nam, S.-H. Kim, D. Min, S.-D. Choi, and Y .-H. Park, “Data augmentation and squeeze-and-excitation network on multiple dimension for sound event localization and detection in real scenes,” DCASE Challenge, Tech. Rep., 2022
2022
-
[29]
Binaural sound event localization and detection based on hrtf cues for humanoid robots,
G.-T. Lee, H. Nam, and Y .-H. Park, “Binaural sound event localization and detection based on hrtf cues for humanoid robots,” arXiv preprint arXiv:2507.20530, 2025
2025 arXiv
-
[30]
Automated audio captioning with recurrent neural networks,
K. Drossos, S. Adavanne, and T. Virtanen, “Automated audio captioning with recurrent neural networks,” in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics , 2017
2017
-
[31]
Clotho: an audio captioning dataset,
K. Drossos, S. Lipping, and T. Virtanen, “Clotho: an audio captioning dataset,” in ICASSP, 2020
2020
-
[32]
Chatgpt caption paraphrasing and fense-based caption filtering for automated audio captioning,
I. Choi, H. Nam, D. Min, S.-D. Choi, and Y .-H. Park, “Chatgpt caption paraphrasing and fense-based caption filtering for automated audio captioning,” DCASE Challenge, Tech. Rep., 2024
2024
-
[33]
Mind the Domain Gap: a Systematic Analysis on Bioacoustic Sound Event Detection,
J. Liang, I. Nolasco, B. Ghani, H. Phan, E. Benetos, and D. Stowell, “Mind the Domain Gap: a Systematic Analysis on Bioacoustic Sound Event Detection,” arXiv preprint arXiv:2403.18638 , 2024
2024 arXiv
-
[34]
Few-shot bioacoustic event detection utilizing spectro-temporal receptive field,
D. Min, H. Nam, and Y .-H. Park, “Few-shot bioacoustic event detection utilizing spectro-temporal receptive field,” in Proc. INTER-NOISE, 2024
2024
-
[35]
Prtfnet: Hrtf individual- ization for accurate spectral cues using a compact prtf,
B.-Y . Ko, G.-T. Lee, H. Nam, and Y .-H. Park, “Prtfnet: Hrtf individual- ization for accurate spectral cues using a compact prtf,” IEEE Access , vol. 11, 2023
2023
-
[36]
Filteraugment: An acoustic environmental data augmentation method,
B.-Y . Ko, Y .-H. Park, G.-T. Lee, and H. Nam, “Filteraugment: An acoustic environmental data augmentation method,” in International Congress on Acoustics (ICA) , 2022
2022
-
[37]
Dnn based hrirs iden- tification with a continuously rotating speaker array,
B.-Y . Ko, D. Min, H. Nam, and Y .-H. Park, “Dnn based hrirs iden- tification with a continuously rotating speaker array,” arXiv preprint arXiv:2504.14817, 2025
2025 arXiv
-
[38]
AudioLDM: Text-to-audio generation with latent diffusion models,
H. Liu, Z. Chen, Y . Yuan, X. Mei, X. Liu, D. Mandic, W. Wang, and M. D. Plumbley, “AudioLDM: Text-to-audio generation with latent diffusion models,” in ICML, 2023
2023
-
[39]
Audiogen: Textually guided audio generation,
F. Kreuk, G. Synnaeve, A. Polyak, U. Singer, A. D ´efossez, J. Copet, D. Parikh, Y . Taigman, and Y . Adi, “Audiogen: Textually guided audio generation,” in International Conference on Learning Representations (ICLR), 2023
2023
-
[40]
Vifs: An end-to-end variational inference for foley sound synthesis,
J. Lee, H. Nam, and Y .-H. Park, “Vifs: An end-to-end variational inference for foley sound synthesis,” DCASE Challenge, Tech. Rep., 2023
2023
-
[41]
Heavily augmented sound event detection utilizing weak predictions,
H. Nam, B.-Y . Ko, G.-T. Lee, S.-H. Kim, W.-H. Jung, S.-M. Choi, and Y .-H. Park, “Heavily augmented sound event detection utilizing weak predictions,” DCASE Challenge, Tech. Rep., 2021
2021
-
[42]
Filteraugment: An acoustic environmental data augmentation method,
H. Nam, S.-H. Kim, and Y .-H. Park, “Filteraugment: An acoustic environmental data augmentation method,” in ICASSP, 2022
2022
-
[43]
Frequency Dynamic Convolution: Frequency-Adaptive Pattern Recognition for Sound Event Detection,
H. Nam, S.-H. Kim, B.-Y . Ko, and Y .-H. Park, “Frequency Dynamic Convolution: Frequency-Adaptive Pattern Recognition for Sound Event Detection,” in Proc. Interspeech, 2022
2022
-
[44]
Self training and ensembling frequency dependent networks with coarse prediction pooling and sound event bounding boxes,
H. Nam, D. Min, I. Choi, S.-D. Choi, and Y .-H. Park, “Self training and ensembling frequency dependent networks with coarse prediction pooling and sound event bounding boxes,” DCASE Challenge, Tech. Rep., 2024
2024
-
[45]
Diversifying and expanding frequency-adaptive convolution kernels for sound event detection,
H. Nam, S.-H. Kim, D. Min, J. Lee, and Y .-H. Park, “Diversifying and expanding frequency-adaptive convolution kernels for sound event detection,” in Proc. Interspeech, 2024
2024
-
[46]
Pushing the limit of sound event detec- tion with multi-dilated frequency dynamic convolution,
H. Nam and Y .-H. Park, “Pushing the limit of sound event detec- tion with multi-dilated frequency dynamic convolution,” arXiv preprint arXiv:2406.13312, 2024
2024 arXiv
-
[47]
Temporal attention pooling for frequency dynamic convolution in sound event detection,
——, “Temporal attention pooling for frequency dynamic convolution in sound event detection,” arXiv preprint arXiv:2504.12670 , 2025
2025 arXiv
-
[48]
Frequency dynamic convolutions for sound event detection,
H. Nam, “Frequency dynamic convolutions for sound event detection,” arXiv preprint arXiv:2506.12785 , 2025
2025 arXiv
-
[49]
Flam: Frame-wise language-audio modeling,
Y . Wu, C. Tsirigotis, K. Chen, C.-Z. A. Huang, A. Courville, O. Nieto, P. Seetharaman, and J. Salamon, “Flam: Frame-wise language-audio modeling,” arXiv preprint arXiv:2505.05335 , 2025
2025 arXiv
-
[50]
AudioCaps: Generating captions for audios in the wild,
C. D. Kim, B. Kim, H. Lee, and G. Kim, “AudioCaps: Generating captions for audios in the wild,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, V olume 1 (Long and Short Papers), 2019
2019
-
[51]
Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,
X. Mei, C. Meng, H. Liu, Q. Kong, T. Ko, C. Zhao, M. D. Plumbley, Y . Zou, and W. Wang, “Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2024
2024
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.