Pith. sign in

REVIEW 2 cited by

Streaming Speaker-Attributed ASR with Token-Level Speaker Embeddings

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.16685 v2 pith:BV75GKRG submitted 2022-03-30 eess.AS cs.CLcs.SD

classification eess.AScs.CLcs.SD
keywords modelspeakerproposedspeechstreamingembeddingevenjoint
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This paper presents a streaming speaker-attributed automatic speech recognition (SA-ASR) model that can recognize ``who spoke what'' with low latency even when multiple people are speaking simultaneously. Our model is based on token-level serialized output training (t-SOT) which was recently proposed to transcribe multi-talker speech in a streaming fashion. To further recognize speaker identities, we propose an encoder-decoder based speaker embedding extractor that can estimate a speaker representation for each recognized token not only from non-overlapping speech but also from overlapping speech. The proposed speaker embedding, named t-vector, is extracted synchronously with the t-SOT ASR model, enabling joint execution of speaker identification (SID) or speaker diarization (SD) with the multi-talker transcription with low latency. We evaluate the proposed model for a joint task of ASR and SID/SD by using LibriSpeechMix and LibriCSS corpora. The proposed model achieves substantially better accuracy than a prior streaming model and shows comparable or sometimes even superior results to the state-of-the-art offline SA-ASR model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization

    cs.CL 2025-07 reject novelty 4.0 of 10

    Timestamp alignment between ASR transcripts and speaker diarization is claimed to improve speech emotion recognition, but the experiment conflates alignment with fine-tuning of the feature extractors.

  2. Streaming Speaker Change Detection and Gender Classification for Transducer-Based Multi-Talker Speech Translation

    cs.SD 2025-02 conditional novelty 4.0 of 10

    Cosine similarity between t-vector speaker embeddings in a streaming transducer speech translation model detects speaker changes (F1 up to 0.68) and classifies gender (0.989 accuracy).

Pith tools