Pith. sign in

REVIEW 2 cited by

Spot the conversation: speaker diarisation in the wild

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.01216 v3 pith:V56YWO7Y submitted 2020-07-02 cs.SD cs.CVeess.ASeess.IV

classification cs.SDcs.CVeess.ASeess.IV
keywords speakerdiarisationvideosdatasetmethodwildaudio-visualcollected
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The goal of this paper is speaker diarisation of videos collected 'in the wild'. We make three key contributions. First, we propose an automatic audio-visual diarisation method for YouTube videos. Our method consists of active speaker detection using audio-visual methods and speaker verification using self-enrolled speaker models. Second, we integrate our method into a semi-automatic dataset creation pipeline which significantly reduces the number of hours required to annotate videos with diarisation labels. Finally, we use this pipeline to create a large-scale diarisation dataset called VoxConverse, collected from 'in the wild' videos, which we will release publicly to the research community. Our dataset consists of overlapping speech, a large and diverse speaker pool, and challenging background conditions.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion

    cs.SD 2025-06 conditional novelty 5.0 of 10

    FusionVAD shows that simple addition or concatenation of MFCC and pre-trained model features outperforms cross-attention fusion for voice activity detection, with the best model beating Pyannote by 2.04 average DER.

  2. The tttAI System for the TSA-ASR Task of the SmartGlasses Challenge 2026

    eess.AS 2026-07 conditional novelty 4.0 of 10

    A cascaded smart-glasses TSA-ASR system with a dominant-speaker overlap fallback achieved 7.10% tcpCER on two-person dialogues and 34.04% on multi-party meetings, ranking second on the meeting track.

Pith tools