Pith. sign in

REVIEW 1 cited by

Joint speech and overlap detection: a benchmark over multiple audio setup and speech domains

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.13012 v1 pith:AK4366YV submitted 2023-07-24 cs.SD cs.AIcs.NEeess.ASeess.SP

classification cs.SDcs.AIcs.NEeess.ASeess.SP
keywords speechsystemsaudiobenchmarkdetectiondomainsjointmultiple
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Voice activity and overlapped speech detection (respectively VAD and OSD) are key pre-processing tasks for speaker diarization. The final segmentation performance highly relies on the robustness of these sub-tasks. Recent studies have shown VAD and OSD can be trained jointly using a multi-class classification model. However, these works are often restricted to a specific speech domain, lacking information about the generalization capacities of the systems. This paper proposes a complete and new benchmark of different VAD and OSD models, on multiple audio setups (single/multi-channel) and speech domains (e.g. media, meeting...). Our 2/3-class systems, which combine a Temporal Convolutional Network with speech representations adapted to the setup, outperform state-of-the-art results. We show that the joint training of these two tasks offers similar performances in terms of F1-score to two dedicated VAD and OSD systems while reducing the training cost. This unique architecture can also be used for single and multichannel speech processing.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Robust Overlapping Speech Detection: A Speaker-Aware Progressive Approach Using WavLM

    cs.SD 2025-05 conditional novelty 6.0 of 10

    A speaker-aware progressive OSD model using WavLM, Campplus, and VAD-gated masking reports 82.76% F1 on AMI, above the listed prior best of 79.21%.

Pith tools