Pith. sign in

REVIEW 2 cited by

DCASE 2024 Task 4: Sound Event Detection with Heterogeneous Data and Missing Labels

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.08056 v1 pith:EZEMHEUA submitted 2024-06-12 eess.AS cs.SD

classification eess.AScs.SD
keywords labelstrainingdifferentsystemdatadetectionmissingsound
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The Detection and Classification of Acoustic Scenes and Events Challenge Task 4 aims to advance sound event detection (SED) systems in domestic environments by leveraging training data with different supervision uncertainty. Participants are challenged in exploring how to best use training data from different domains and with varying annotation granularity (strong/weak temporal resolution, soft/hard labels), to obtain a robust SED system that can generalize across different scenarios. Crucially, annotation across available training datasets can be inconsistent and hence sound labels of one dataset may be present but not annotated in the other one and vice-versa. As such, systems will have to cope with potentially missing target labels during training. Moreover, as an additional novelty, systems will also be evaluated on labels with different granularity in order to assess their robustness for different applications. To lower the entry barrier for participants, we developed an updated baseline system with several caveats to address these aforementioned problems. Results with our baseline system indicate that this research direction is promising and is possible to obtain a stronger SED system by using diverse domain training data with missing labels compared to training a SED system for each domain separately.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Semantics to Readout: Mechanistic Understanding of Audio Tokens after Fine-Tuning for Temporal Audio Grounding

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Fine-tuning for temporal audio grounding mostly improves the decoder's ability to read pre-existing event evidence in audio tokens and align it with predicted timestamps, rather than creating that evidence from scratch.

  2. FLAM: Frame-Wise Language-Audio Modeling

    cs.SD 2025-05 conditional novelty 6.0 of 10

    FLAM trains an audio-language model with a frame-level contrastive goal and per-text logit adjustment, enabling open-vocabulary temporal localization of sound events while preserving global retrieval.

Pith tools