Pith. sign in

REVIEW

The DKU-MSXF Diarization System for the VoxCeleb Speaker Recognition Challenge 2023

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.07595 v2 pith:33FIRXID submitted 2023-08-15 eess.AS

classification eess.AS
keywords diarizationchallengedetectionactivityclustering-baseddku-msxfrecognitionspeaker
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper describes the DKU-MSXF submission to track 4 of the VoxCeleb Speaker Recognition Challenge 2023 (VoxSRC-23). Our system pipeline contains voice activity detection, clustering-based diarization, overlapped speech detection, and target-speaker voice activity detection, where each procedure has a fused output from 3 sub-models. Finally, we fuse different clustering-based and TSVAD-based diarization systems using DOVER-Lap and achieve the 4.30% diarization error rate (DER), which ranks first place on track 4 of the challenge leaderboard.

Discussion (0). Continue with ORCID to comment.

Pith tools