Pith. sign in

REVIEW 2 major objections 4 minor 25 references

RealDESED is a 5,710-clip benchmark of real home recordings with multi-annotator temporal labels; its transformer baseline reaches a macro-averaged PSDS1 of 0.731, making it a realistic alternative to synthetic domestic SED benchmarks.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review

2026-08-01 20:04 UTC pith:LHRU4T7W

load-bearing objection RealDESED is a solid new real-home SED benchmark; the 'real-world' claim is slightly overstated but the resource is valuable and honestly reported. the 2 major comments →

arxiv 2607.16736 v1 pith:LHRU4T7W submitted 2026-07-18 eess.AS cs.AIcs.SD

RealDESED: A Real-World Domestic Sound Event Detection Benchmark

classification eess.AS cs.AIcs.SD
keywords sound event detectionreal-world datasetdomestic environmentmulti-annotator labelstemporal annotationbenchmarktransformer baselinePSDS
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper introduces RealDESED, a benchmark of 5,710 domestic recordings made by 652 people in their own homes, with temporally precise labels for 15 everyday sound classes. Unlike popular SED datasets that rely on synthetic soundscapes or short web-crawled clips, each RealDESED recording is 15–35 seconds long, captured on consumer devices, and annotated by multiple independent annotators, with validation and test labels reviewed for quality. The paper argues that this combination makes RealDESED a more realistic testbed for developing sound event detection systems that work in actual homes. As evidence, it trains a transformer baseline and reports a macro-averaged PSDS1 of 0.731 on the test set, and it shows that annotation aggregation, post-processing, and long-form inference all noticeably affect that number. If the dataset is as natural as claimed, it gives the field a public benchmark that better predicts deployment performance than synthetic or fixed-length alternatives.

Core claim

The paper's central claim is that a benchmark built from naturally occurring home recordings, with multiple independent annotators and a reviewed evaluation split, is a viable and more realistic alternative to existing domestic SED benchmarks. On RealDESED, a transformer pre-trained on strongly labeled audio and fine-tuned on the new data reaches 0.731 macro-averaged PSDS1 on the test set. The paper further shows that soft labels weighted by annotator quality outperform simple majority or union aggregation, that temporal post-processing with sound event bounding boxes is essential, and that long-form inference with overlapping windows and triangular weighting improves scores. It also documen

What carries the argument

The central object is the dataset itself: 5,710 recordings (about 38 hours) from 652 collectors, each 15–35 seconds long, covering 15 domestic sound classes, with an average of 2.45 independent annotators per file and reviewed labels for the full validation and test splits. Two mechanisms carry the benchmark's value: multi-annotator aggregation, where soft frame-level targets are formed by averaging per-annotator labels with weights derived from pairwise Dice agreement, turning annotator disagreement into a training signal rather than noise; and temporal post-processing with sound event bounding boxes (cSEBBs), which converts raw frame predictions into event-level detections. The dataset als

Load-bearing premise

The core premise is that the course-collected recordings are natural domestic scenes and not noticeably staged or simplified; no independent audit verifies that the homes, background conditions, and event sequences are representative of real everyday life.

What would settle it

Audit a random sample of RealDESED test recordings with an independent, opportunistic home-recording campaign matched by environment and time of day, and compare background sound levels, non-target event rates, and event co-occurrence statistics; if the benchmark recordings are systematically quieter or more event-dense than the independent sample, the real-world claim and the 0.731 PSDS1 would not transfer to deployment.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • If RealDESED is representative of real homes, SED systems trained on it should transfer to deployment better than those trained on synthetic mixes or fixed 10-second web clips.
  • The finding that annotator-quality-weighted soft labels match or beat human review suggests future dataset builders can reduce expensive review effort by collecting more cheap annotations and weighting them automatically.
  • The metadata-driven performance gaps across devices, placements, and environments imply that device-robust and domain-adaptive SED should be evaluated on RealDESED to see whether proposed gains actually generalize.
  • The strong effect of long-form inference with overlapping windows indicates that scores on short clips may understate performance on longer, real-world recordings.
  • Reviewed validation and test labels, combined with multi-annotator coverage, make RealDESED a reliable common ground for comparing SED systems in the domestic domain.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Going beyond the paper: a natural next test is to use RealDESED's metadata to train a device-robust model, for example via adversarial domain adaptation or device dropout; the paper shows gaps exist but does not attempt to close them.
  • Going beyond the paper: the weighted-soft-label aggregation recipe is likely to transfer to other audio tasks with noisy or subjective labels, not just sound event detection.
  • Going beyond the paper: the real-world claim could be stress-tested by comparing ambient noise levels and event co-occurrence rates in RealDESED against independently collected passive home recordings; the paper does not report such a comparison.
  • Going beyond the paper: because collectors knew they were recording for a benchmark, there may be a self-selection bias toward quiet or staged scenes, so the reported PSDS1 should be treated as an upper bound for fully opportunistic deployment until independent validation.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. RealDESED is a new domestic sound event detection benchmark consisting of 5,710 audio recordings collected by 652 participants in their homes, with multi-annotator strong labels for 15 domestic sound classes, a reviewed validation/test split, and rich metadata (device, placement, environment, scene descriptions). The paper describes the data collection, annotation, review, and split procedures, reports dataset statistics including annotation agreement and overlap analysis, and establishes an ATST-F baseline. It further compares annotation aggregation strategies, post-processing methods (median filter vs. cSEBBs), long-form inference settings, and metadata-dependent performance. The strongest reported result is a macro-averaged PSDS1 of 0.731 on the test set using quality-weighted soft labels, cSEBBs post-processing, and a 5-s hop long-form aggregation.

Significance. If the recordings are indeed representative of natural domestic soundscapes, RealDESED would be a valuable community resource: it is large, strongly annotated, multi-annotator, and released with protocols and baselines, thereby addressing a real gap left by synthetic benchmarks such as DESED and by fixed-length web-crawled datasets. The experimental methodology is generally careful: results are averaged over three runs with standard deviations, aggregation and post-processing are compared systematically, and the validation/test labels receive an additional review pass. The main uncertainty is conceptual rather than computational: the 'real-world deployment' contribution rests on the authenticity of the collected scenes, which is self-reported and protocol-driven. The paper should either provide quantitative evidence of naturalness or explicitly narrow its claims; with that clarification, the benchmark has clear value for the SED community.

major comments (2)
  1. [§2.1, Abstract, §5] The central claim that RealDESED comprises 'real-world domestic recordings' and provides a path to 'real-world deployment' is only as strong as the authenticity of the collection protocol. Section 2.1 states that each participant was 'asked to collect eight to ten realistic domestic sound scenes', that the protocol 'encouraged balanced target-class coverage', and that 'recordings containing speech were not permitted'. This is an instructed and explicitly curated setup, not passive observation of everyday home activity. Speech is a ubiquitous component of real domestic soundscapes, and its exclusion, together with the instructions to produce balanced multi-event scenes, likely shifts the event distribution, background conditions, and silence patterns away from natural homes. The manuscript provides no independent audit (e.g., comparisons with unconstrained household recordings, background
  2. [§3.3] The reported improvement in annotation agreement after review (Jaccard 0.875 to 0.994; temporal IoU 0.694 to 0.862) is presented as evidence of labeling quality, but the protocol described in §3.3 is not a clean before/after comparison. Because 'approximately 300 low-quality reviews were identified based on remaining disagreement between annotations and reassigned for review', the final numbers may reflect selection or replacement of the hardest annotations rather than the effect of correction alone. The manuscript should specify exactly how the final validation/test annotations are compiled from the reviewed versions and report agreement statistics computed on identical file sets before and after the correction/reassignment step. Without this, the 'reviewed labels' claim is difficult to interpret quantitatively.
minor comments (4)
  1. [Table 2] The 'Micro' row contains a formatting error: '0.4040.7870.751' should read '0.404 0.787 0.751' (or similar). Please fix the table layout.
  2. [§1] The claim 'first large real-world domestic SED benchmark' is unsupported as stated. AudioSet Strong is real-world and includes domestic content; DESED also contains real web-crawled evaluation clips. Please qualify the novelty claim (e.g., 'first large home-collected, multi-annotator domestic SED benchmark') or add a systematic comparison with prior real-world domestic datasets.
  3. [§4.3–4.4] The cSEBBs hyperparameters and the long-form inference settings (triangular floor, hop size, aggregation) are selected on the validation set, but the selected cSEBBs hyperparameter values are not reported. For reproducibility, list them (or point to the released configuration) along with the hyperparameter search ranges.
  4. [§2.4] The split by collector ID assumes each collector records in a single domestic environment. If collectors could record in multiple rooms or homes, this assumption is not guaranteed. Please state whether the environment metadata was used to verify that all recordings from one collector share the same environment.

Circularity Check

0 steps flagged

No significant circularity: RealDESED is an independently collected benchmark with a standard held-out evaluation; no prediction reduces to its inputs.

full rationale

The paper's central contribution is a new dataset and a benchmark evaluation, not a derivation from first principles. The dataset was collected by 652 participants according to a stated protocol (Section 2.1), independently annotated by 645 annotators (Section 2.2), and reviewed before use in the validation/test splits (Section 2.3). The baseline performance (PSDS1-M = 0.731) is obtained by fine-tuning an external transformer model (ATST-F) on the RealDESED training split and evaluating on the held-out test split using standard metrics (PSDS1/PSDS2 from [24]). No equation in the paper defines the target metric in terms of fitted parameters or renames an input as a prediction. The only self-citations are [13] and [16] in the introduction and baseline setup; these are not load-bearing for the central benchmark claim. In particular, the use of a pretrained model from the authors' prior work is an external, published result and does not by construction force the reported test score. The potential concern that recordings may be staged or unrepresentative of real homes is a data-quality/representativeness issue, not a circularity issue: the paper does not derive its conclusions from an assumption that already contains the conclusion. The paper is self-contained as a benchmark paper, and no circular step was found.

Axiom & Free-Parameter Ledger

4 free parameters · 6 axioms · 0 invented entities

The central product is a dataset; the free parameters belong to the baseline pipeline rather than to the dataset claim. The main loaded assumptions are annotation quality and recording authenticity, both explicitly discussed in Sections 2 and 3.

free parameters (4)
  • Annotator quality exponent α = 16
    Selected on validation set (Section 4.2, Eq 2) to maximize PSDS1-M; affects all soft-label results.
  • Median filter window size = 360 ms
    Post-processing setting chosen on validation (Section 4.1); used for all baseline runs.
  • cSEBBs hyperparameters = not specified (optimized on validation)
    Optimized on validation set (Section 4.3); required to reach PSDS1-M 0.731.
  • Long-form inference hop size and triangular filter floor = 5 s; 0.3
    Selected on validation set (Section 4.4) for the final pipeline.
axioms (6)
  • domain assumption The review process in Section 2.3 yields ground-truth labels that are accurate enough for benchmarking.
    The dataset's validity depends on reviewers catching annotation errors; the evidence is internal agreement metrics (Jaccard 0.994, temporal IoU 0.862), not external validation.
  • domain assumption The collected recordings are natural domestic scenes rather than posed recordings.
    Collection was guided by instructions (Section 2.1), but authenticity is self-reported by 652 course participants; no independent audit.
  • domain assumption The 15 target classes are mutually exclusive and cover the practical domestic sound space described.
    Class definitions are given in Section 2; some acoustic overlap (e.g., door and window) may exist, but the task assumes separability.
  • standard math PSDS1/PSDS2 are appropriate metrics for this benchmark.
    The paper adopts the field-standard threshold-independent metrics from [24]; the choice is conventional but not derived in the paper.
  • domain assumption Ten-second crops with frame-level labels at 25 Hz are a sufficient representation for training/evaluation on longer recordings.
    Section 4.1 and 4.4; the baseline relies on this sliding-window approach for all 15–35 s recordings.
  • domain assumption User-provided metadata (device, placement, environment) is accurate enough for the performance grouping in Section 4.5.
    Normalization was applied (161→65 environments, 523→342 devices), but errors in source strings could misclassify recordings.

reviewed 2026-08-01 · how reviews work

0 comments
Cite this review

Pith. "Pith review of RealDESED: A Real-World Domestic Sound Event Detection Benchmark." pith.science (2026). https://pith.science/paper/LHRU4T7W

@misc{pith2026260716736,
  author       = {Pith},
  title        = {Pith review of: RealDESED: A Real-World Domestic Sound Event Detection Benchmark},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LHRU4T7W}},
  note         = {Machine review of arXiv:2607.16736}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

This paper presents RealDESED, a real-world domestic sound event detection (SED) benchmark comprising 5,710 audio recordings collected by 652 participants in their homes. Each recording is between 15 and 35 seconds long and contains temporally precise annotations for 15 common domestic sound classes. In contrast to existing SED datasets, which typically rely on simulated soundscapes or broad web-crawled audio, RealDESED consists exclusively of recordings captured in natural domestic environments, reflecting realistic variability in recording devices, device placement, acoustic conditions, background sounds, and naturally occurring event co-occurrences. A distinguishing characteristic of the dataset is its multi-annotator labeling scheme, where each recording is independently annotated by multiple annotators, while the validation and test sets undergo an additional review process to ensure high annotation quality and reliable benchmarking. Furthermore, the dataset provides rich metadata, including recording device, device placement, environment labels, and textual scene descriptions. We establish a strong transformer-based baseline and investigate annotation aggregation strategies, post-processing methods, long-form inference, and the impact of recording metadata on model performance. Our baseline achieves a macro-averaged PSDS1 score of 0.731 on the test set. We believe RealDESED provides a valuable benchmark for developing and evaluating robust SED systems under realistic domestic conditions, helping to bridge the gap between current research benchmarks and real-world deployment.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

25 extracted references · 1 linked inside Pith

  1. [1]

    Audio analysis for surveillance applications,

    R. Radhakrishnan, A. Divakaran, and A. Smaragdis, “Audio analysis for surveillance applications,” inProc. WASPAA. IEEE, 2005, pp. 158–161

  2. [2]

    Monitoring activities of daily living in smart homes: Understanding human behavior,

    C. Debes, A. Merentitis, S. Sukhanov, M. E. Niessen, N. Frangiadakis, and A. Bauer, “Monitoring activities of daily living in smart homes: Understanding human behavior,”IEEE Signal Process. Mag., vol. 33, no. 2, pp. 81–94, 2016

  3. [3]

    homesound: Real-time audio event detection based on high performance computing for behaviour and surveillance remote monitoring,

    R. M. Alsina-Pag `es, J. Navarro, F. Al ´ıas, and M. Herv ´as, “homesound: Real-time audio event detection based on high performance computing for behaviour and surveillance remote monitoring,”Sensors, vol. 17, no. 4, p. 854, 2017

  4. [4]

    A method for automatic fall detection of elderly people using floor vibrations and sound - proof of concept on human mimicking doll falls,

    Y . Zigel, D. Litvak, and I. Gannot, “A method for automatic fall detection of elderly people using floor vibrations and sound - proof of concept on human mimicking doll falls,”IEEE Trans. Biomed. Eng., vol. 56, no. 12, pp. 2858–2867, 2009

  5. [5]

    Audio set: An ontology and human-labeled dataset for audio events,

    J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” inProc. ICASSP, 2017, pp. 776–780

  6. [6]

    The benefit of temporally-strong labels in audio event classification,

    S. Hershey, D. P. W. Ellis, E. Fonseca, A. Jansen, C. Liu, R. C. Moore, and M. Plakal, “The benefit of temporally-strong labels in audio event classification,” inProc. ICASSP, 2021, pp. 366–370

  7. [7]

    Sound event detection in domestic environments with weakly labeled data and soundscape synthesis,

    N. Turpault, R. Serizel, J. Salamon, and A. P. Shah, “Sound event detection in domestic environments with weakly labeled data and soundscape synthesis,” inProc. DCASE, 2019, pp. 253–257

  8. [8]

    Sound event detection in synthetic domestic environments,

    R. Serizel, N. Turpault, A. P. Shah, and J. Salamon, “Sound event detection in synthetic domestic environments,” inProc. ICASSP, 2020, pp. 86–90

  9. [9]

    Scaper: A library for soundscape synthesis and augmentation,

    J. Salamon, D. MacConnell, M. Cartwright, P. Li, and J. P. Bello, “Scaper: A library for soundscape synthesis and augmentation,” inProc. WASPAA, 2017, pp. 344–348

  10. [10]

    Strong labeling of sound events using crowdsourced weak labels and annotator competence estimation,

    I. Mart ´ın-Morat´o and A. Mesaros, “Strong labeling of sound events using crowdsourced weak labels and annotator competence estimation,” IEEE/ACM Trans. Audio, Speech, Lang. Process., vol. 31, pp. 902–914, 2023

  11. [11]

    Audio set classification with attention model: A probabilistic perspective,

    Q. Kong, Y . Xu, W. Wang, and M. D. Plumbley, “Audio set classification with attention model: A probabilistic perspective,” inProc. ICASSP, 2018, pp. 316–320

  12. [12]

    Large-scale weakly supervised audio classification using gated convolutional neural network,

    Y . Xu, Q. Kong, W. Wang, and M. D. Plumbley, “Large-scale weakly supervised audio classification using gated convolutional neural network,” inProc. ICASSP, 2018, pp. 121–125

  13. [13]

    Multi- iteration multi-stage fine-tuning of transformers for sound event detection with heterogeneous datasets,

    F. Schmid, P. Primus, T. Morocutti, J. Greif, and G. Widmer, “Multi- iteration multi-stage fine-tuning of transformers for sound event detection with heterogeneous datasets,” inProc. DCASE, 2024, pp. 141–145

  14. [14]

    Self training and ensembling frequency dependent networks with coarse prediction pooling and sound event bounding boxes,

    H. Nam, D. Min, S. Choi, I. Choi, and Y .-H. Park, “Self training and ensembling frequency dependent networks with coarse prediction pooling and sound event bounding boxes,” inProc. DCASE, 2024, pp. 96–100

  15. [15]

    Dcase 2024 task 4: Sound event detection with heterogeneous data and missing labels,

    S. Cornell, J. Ebbers, C. Douwes, I. Mart´ın-Morat´o, M. Harju, A. Mesaros, and R. Serizel, “Dcase 2024 task 4: Sound event detection with heterogeneous data and missing labels,” inProc. DCASE, 2024, pp. 31–35

  16. [16]

    Effective pre-training of audio transformers for sound event detection,

    F. Schmid, T. Morocutti, F. Foscarin, J. Schl ¨uter, P. Primus, and G. Widmer, “Effective pre-training of audio transformers for sound event detection,” inProc. ICASSP, 2025, pp. 1–5

  17. [17]

    MAT-SED: A masked audio transformer with masked-reconstruction based pre-training for sound event detection,

    P. Cai, Y . Song, K. Li, H. Song, and I. McLoughlin, “MAT-SED: A masked audio transformer with masked-reconstruction based pre-training for sound event detection,” inProc. Interspeech, 2024

  18. [18]

    Prototype based masked audio model for self-supervised learning of sound event detection,

    P. Cai, Y . Song, N. Jiang, Q. Gu, and I. McLoughlin, “Prototype based masked audio model for self-supervised learning of sound event detection,” inProc. ICASSP, 2025, pp. 1–5

  19. [19]

    Jitter: Jigsaw temporal transformer for event reconstruction for self-supervised sound event detection,

    H. Nam and Y . Park, “Jitter: Jigsaw temporal transformer for event reconstruction for self-supervised sound event detection,”CoRR, vol. abs/2502.20857, 2025

  20. [20]

    Self-supervised audio teacher-student transformer for both clip-level and frame-level tasks,

    X. Li, N. Shao, and X. Li, “Self-supervised audio teacher-student transformer for both clip-level and frame-level tasks,”IEEE/ACM Trans. Audio, Speech, Lang. Process., vol. 32, pp. 1336–1351, 2024

  21. [21]

    Label Studio: Data labeling software,

    M. Tkachenko, M. Malyuk, A. Holmanyuk, and N. Liubimov, “Label Studio: Data labeling software,” 2020-2025, open source software available from https://github.com/HumanSignal/label-studio

  22. [22]

    Decoupled weight decay regularization,

    I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in Proc. ICLR, 2019

  23. [23]

    mixup: Beyond empirical risk minimization,

    H. Zhang, M. Ciss ´e, Y . N. Dauphin, and D. Lopez-Paz, “mixup: Beyond empirical risk minimization,” inProc. ICLR, 2018

  24. [24]

    Threshold independent evaluation of sound event detection scores,

    J. Ebbers, R. Haeb-Umbach, and R. Serizel, “Threshold independent evaluation of sound event detection scores,” inProc. ICASSP, 2022, pp. 1021–1025

  25. [25]

    Sound event bounding boxes,

    J. Ebbers, F. G. Germain, G. Wichern, and J. L. Roux, “Sound event bounding boxes,” inProc. Interspeech, 2024

This paper was first reviewed by deepseek-v4-flash on August 1, 2026.