Pith. sign in

REVIEW 2 cited by

Meeting Recognition with Continuous Speech Separation and Transcription-Supported Diarization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.16482 v2 pith:HWYA4NPJ submitted 2023-09-28 eess.AS cs.SD

Meeting Recognition with Continuous Speech Separation and Transcription-Supported Diarization

classification eess.AS cs.SD
keywords diarizationrecognitionseparationspeakerspeechcontinuouserrormeeting
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We propose a modular pipeline for the single-channel separation, recognition, and diarization of meeting-style recordings and evaluate it on the Libri-CSS dataset. Using a Continuous Speech Separation (CSS) system with a TF-GridNet separation architecture, followed by a speaker-agnostic speech recognizer, we achieve state-of-the-art recognition performance in terms of Optimal Reference Combination Word Error Rate (ORC WER). Then, a d-vector-based diarization module is employed to extract speaker embeddings from the enhanced signals and to assign the CSS outputs to the correct speaker. Here, we propose a syntactically informed diarization using sentence- and word-level boundaries of the ASR module to support speaker turn detection. This results in a state-of-the-art Concatenated minimum-Permutation Word Error Rate (cpWER) for the full meeting recognition pipeline.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Diarization-Guided Qwen-ASR Adaptation for Multilingual Two-Speaker Conversational Speech

    cs.CL 2026-07 conditional novelty 4.0

    A Qwen3-ASR-based two-speaker, 21-language transcription system cuts its official error metric from 30.53 to 23.70 on the MLC-SLM 2026 dev set; supervised fine-tuning delivers most of the gain.

  2. Diarization-Guided Qwen-ASR Adaptation for Multilingual Two-Speaker Conversational Speech

    cs.CL 2026-07 conditional novelty 3.5

    Diarization-guided full SFT, synthetic-speech LoRA, and GRPO RL adapt Qwen3-ASR-1.7B to 23.70 average tcpMER on the MLC-SLM 2026 Task 1 development set.