Pith. sign in

REVIEW 5 cited by

The Third DIHARD Diarization Challenge

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.01477 v3 pith:2BFAVEEB submitted 2020-12-02 eess.AS cs.SD

classification eess.AScs.SD
keywords diarizationspeechconditionsdiharddomainsspeakeractivityconversational
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

DIHARD III was the third in a series of speaker diarization challenges intended to improve the robustness of diarization systems to variability in recording equipment, noise conditions, and conversational domain. Speaker diarization was evaluated under two speech activity conditions (diarization from a reference speech activity vs. diarization from scratch) and 11 diverse domains. The domains span a range of recording conditions and interaction types, including read audio-books, meeting speech, clinical interviews, web videos, and, for the first time, conversational telephone speech. A total of 30 organizations (forming 21teams) from industry and academia submitted 499 valid system outputs. The evaluation results indicate that speaker diarization has improved markedly since DIHARD I, particularly for two-party interactions, but that for many domains (e.g., web video) the problem remains far from solved.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SDBench: A Comprehensive Benchmark Suite for Speaker Diarization

    cs.SD 2025-07 conditional novelty 6.0 of 10

    SDBench provides a reproducible 13-dataset benchmark for speaker diarization, and its companion SpeakerKit achieves a claimed 9.6x speedup over Pyannote v3.1 with comparable DER.

  2. HPP-Voice: A Large-Scale Evaluation of Speech Embeddings for Multi-Phenotypic Classification

    eess.AS 2025-05 conditional novelty 6.0 of 10

    A 30-second counting task, embedded with speaker-identification models, predicts male sleep apnea (AUC 0.64) and shows gender- and condition-specific model rankings across a new 7,188-recording clinical speech benchmark.

  3. M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset

    eess.AS 2025-06 reject novelty 5.0 of 10

    Release of M3SD, a 770+ hour pseudo-labeled multi-scenario, multi-language audio-visual speaker diarization dataset, built from YouTube and Bilibili videos without manual annotation.

  4. The tttAI System for the TSA-ASR Task of the SmartGlasses Challenge 2026

    eess.AS 2026-07 conditional novelty 4.0 of 10

    A cascaded smart-glasses TSA-ASR system with a dominant-speaker overlap fallback achieved 7.10% tcpCER on two-person dialogues and 34.04% on multi-party meetings, ranking second on the meeting track.

  5. Speaker Diarization with Overlapping Community Detection Using Graph Attention Networks and Label Propagation Algorithm

    cs.SD 2025-06 conditional novelty 4.0 of 10

    Graph attention refinement plus overlapping label propagation yields a reported 15.94% DER on DIHARD-III without oracle VAD, though internal configuration inconsistencies and missing error bars make the SOTA claim pro...

Pith tools