REVIEW 4 cited by
General-purpose Tagging of Freesound Audio with AudioSet Labels: Task Description, Dataset, and Baseline
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper describes Task 2 of the DCASE 2018 Challenge, titled "General-purpose audio tagging of Freesound content with AudioSet labels". This task was hosted on the Kaggle platform as "Freesound General-Purpose Audio Tagging Challenge". The goal of the task is to build an audio tagging system that can recognize the category of an audio clip from a subset of 41 diverse categories drawn from the AudioSet Ontology. We present the task, the dataset prepared for the competition, and a baseline system.
Forward citations
Cited by 4 Pith papers
-
Exploring Efficient Waveform Diffusion Models for Foley Sound Generation
Dual-path attention over time-frequency representations lets a 3.26M-parameter waveform diffusion model match the quality of 50M+ parameter Foley generators.
-
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance
X-ARES evaluates 13 audio encoders on 22 speech, sound, and music tasks using linear probing and nearest-neighbor classifiers, revealing strong domain-dependent performance differences.
-
WQ-Fusion: Dynamic Gated Attention for Cross-Domain Audio Representation
WQ-Fusion combines Whisper and Qwen encoders with gated attention to reach 0.836 on the Interspeech 2026 Audio Encoder Capability Challenge, outperforming single-encoder baselines.
-
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
Transferring I-JEPA's masked latent prediction to mel-spectrograms yields competitive audio representations on music and environmental sound tasks with a small fraction of the training data.
Discussion (0). Continue with ORCID to comment.