REVIEW 10 cited by
FSD50K: An Open Dataset of Human-Labeled Sound Events
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
FSD50K: An Open Dataset of Human-Labeled Sound Events
read the original abstract
Most existing datasets for sound event recognition (SER) are relatively small and/or domain-specific, with the exception of AudioSet, based on over 2M tracks from YouTube videos and encompassing over 500 sound classes. However, AudioSet is not an open dataset as its official release consists of pre-computed audio features. Downloading the original audio tracks can be problematic due to YouTube videos gradually disappearing and usage rights issues. To provide an alternative benchmark dataset and thus foster SER research, we introduce FSD50K, an open dataset containing over 51k audio clips totalling over 100h of audio manually labeled using 200 classes drawn from the AudioSet Ontology. The audio clips are licensed under Creative Commons licenses, making the dataset freely distributable (including waveforms). We provide a detailed description of the FSD50K creation process, tailored to the particularities of Freesound data, including challenges encountered and solutions adopted. We include a comprehensive dataset characterization along with discussion of limitations and key factors to allow its audio-informed usage. Finally, we conduct sound event classification experiments to provide baseline systems as well as insight on the main factors to consider when splitting Freesound audio data for SER. Our goal is to develop a dataset to be widely adopted by the community as a new open benchmark for SER research.
Forward citations
Cited by 10 Pith papers
-
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio
Trained connectors and audio-only gated adapters integrate audio into a frozen vision-language embedding space, preserving base outputs bit-exactly and yielding emergent audio-image retrieval.
-
The World is Not Mono: Enabling Spatial Understanding in Large Audio-Language Models
TWNM framework equips audio-language models with spatial scene analysis via FOA simulation and metadata-grounded training, reaching 70.8% accuracy on a new ASA benchmark.
-
Fine-grained Soundscape Control for Augmented Hearing
Aurchestra enables real-time, on-device per-class sound extraction and volume control for up to five simultaneous sound classes on hearables.
-
Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification
A binarized prototypical probing method that pools per-class evidence from patch tokens substantially outperforms [cls]-token and attentive probes on multi-label audio classification benchmarks.
-
CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents
CodecSep performs prompt-driven universal sound separation directly in neural audio codec latents by combining a frozen DAC backbone with a lightweight FiLM-conditioned Transformer masker driven by CLAP embeddings, yi...
-
CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents
A FiLM-conditioned transformer masker on DAC codec latents performs text-guided sound separation with claimed efficiency, but the main comparison against AudioSep is confounded by asymmetric input processing.
-
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models
An audio agent trained with trajectory-based SFT and multi-turn GRPO improves tool-use and reasoning on a new AI-generated audio agent benchmark, including tasks with unseen tools and workflows.
-
Inside the Latent Flow: Causal Deciphering of Attention Dynamics in Audio Separation Foundation Models
Causal probing of attention in audio separation transformers identifies dual pathways and asynchronous convergence, enabling a training-free Layer-Selective Attention Caching method that reduces self-attention computa...
-
DPDFNet: Boosting DeepFilterNet2 via Dual-Path RNN
DPDFNet inserts dual-path RNN blocks into DeepFilterNet2's encoder, adds an over-attenuation loss and long-context fine-tuning, and reports superior causal speech enhancement on a 12-language low-SNR test set.
-
Low-latency Assistive Audio Enhancement for Neurodivergent People
Among DSP and ML audio enhancement approaches evaluated on trigger-sound mixtures, Dynamic Range Compression (DRC) attenuates distressing sounds most effectively in both objective metrics and a neurodivergent listening test.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.