Pith. sign in

REVIEW 2 cited by

Target Speech Diarization with Multimodal Prompts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.07198 v1 pith:RRNK7DI2 submitted 2024-06-11 eess.AS cs.MM

classification eess.AScs.MM
keywords speechdiarizationpromptstargetmm-tsdframeworkspeakeraccording
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Traditional speaker diarization seeks to detect ``who spoke when'' according to speaker characteristics. Extending to target speech diarization, we detect ``when target event occurs'' according to the semantic characteristics of speech. We propose a novel Multimodal Target Speech Diarization (MM-TSD) framework, which accommodates diverse and multi-modal prompts to specify target events in a flexible and user-friendly manner, including semantic language description, pre-enrolled speech, pre-registered face image, and audio-language logical prompts. We further propose a voice-face aligner module to project human voice and face representation into a shared space. We develop a multi-modal dataset based on VoxCeleb2 for MM-TSD training and evaluation. Additionally, we conduct comparative analysis and ablation studies for each category of prompts to validate the efficacy of each component in the proposed framework. Furthermore, our framework demonstrates versatility in performing various signal processing tasks, including speaker diarization and overlap speech detection, using task-specific prompts. MM-TSD achieves robust and comparable performance as a unified system compared to specialized models. Moreover, MM-TSD shows capability to handle complex conversations for real-world dataset.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sequence-to-Sequence Neural Diarization with Automatic Speaker Detection and Representation

    eess.AS 2024-11 conditional novelty 7.0 of 10

    A single sequence-to-sequence network with detection and representation decoders achieves state-of-the-art online and offline speaker diarization on DIHARD-II and DIHARD-III.

  2. Beyond Speaker Identity: Text Guided Target Speech Extraction

    eess.AS 2025-01 conditional novelty 6.0 of 10

    StyleTSE extracts target speech from mixtures using natural-language speaking style descriptions, optionally combined with reference audio, trained on the new TextrolMix dataset.

Pith tools