Pith. sign in

REVIEW 3 cited by

Serialized Output Training for End-to-End Overlapped Speech Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.12687 v2 pith:AVK4FVMX submitted 2020-03-28 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords speakersoutputoverlappedspeechtrainingmodelsmultiplenumber
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

This paper proposes serialized output training (SOT), a novel framework for multi-speaker overlapped speech recognition based on an attention-based encoder-decoder approach. Instead of having multiple output layers as with the permutation invariant training (PIT), SOT uses a model with only one output layer that generates the transcriptions of multiple speakers one after another. The attention and decoder modules take care of producing multiple transcriptions from overlapped speech. SOT has two advantages over PIT: (1) no limitation in the maximum number of speakers, and (2) an ability to model the dependencies among outputs for different speakers. We also propose a simple trick that allows SOT to be executed in $O(S)$, where $S$ is the number of the speakers in the training sample, by using the start times of the constituent source utterances. Experimental results on LibriSpeech corpus show that the SOT models can transcribe overlapped speech with variable numbers of speakers significantly better than PIT-based models. We also show that the SOT models can accurately count the number of speakers in the input audio.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition

    eess.AS 2025-06 conditional novelty 5.0 of 10

    A speech large language model trained on beamformed multi-channel audio performs directional speech recognition and source localization across 12 discrete angles on simulated smart glasses data.

  2. SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition

    cs.SD 2025-06 conditional novelty 5.0 of 10

    Conditioning an SOT multi-talker ASR decoder on EEND-EDA speaker embeddings and activity information lowers WER on Libri2Mix and Libri3Mix, provided the diarization branch is accurate.

  3. Joint ASR and Speaker Role Tagging with Serialized Output Training

    eess.AS 2025-06 conditional novelty 5.0 of 10

    Fine-tuning Whisper with serialized output training and role-specific tokens produces role-aware transcripts in one pass, cutting multi-talker word error rate by 10 to 40 percent versus a WavLM CTC baseline.

Pith tools