Pith. sign in

REVIEW 2 cited by

Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.05796 v1 pith:AG7HIORX submitted 2025-06-06 eess.AS

Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models

classification eess.AS
keywords multi-speakerspeechtranscriptionautomaticconversationaldiarizationdiarization-awarelanguage
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Multi-speaker automatic speech recognition (MS-ASR) faces significant challenges in transcribing overlapped speech, a task critical for applications like meeting transcription and conversational analysis. While serialized output training (SOT)-style methods serve as common solutions, they often discard absolute timing information, limiting their utility in time-sensitive scenarios. Leveraging recent advances in large language models (LLMs) for conversational audio processing, we propose a novel diarization-aware multi-speaker ASR system that integrates speaker diarization with LLM-based transcription. Our framework processes structured diarization inputs alongside frame-level speaker and semantic embeddings, enabling the LLM to generate segment-level transcriptions. Experiments demonstrate that the system achieves robust performance in multilingual dyadic conversations and excels in complex, high-overlap multi-speaker meeting scenarios. This work highlights the potential of LLMs as unified back-ends for joint speaker-aware segmentation and transcription.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models

    eess.AS 2026-04 unverdicted novelty 6.0

    DM-ASR reformulates multi-speaker ASR as multi-turn dialogue generation conditioned on diarization results, achieving competitive benchmark performance with relatively small models and limited data.

  2. Speech Encoder Fusion for LLM-based Automatic Speech Recognition

    eess.AS 2026-06 unverdicted novelty 4.0

    Fusing multiple parallel pre-trained speech encoders into LLM-based ASR yields consistent performance gains across mono- and multilingual and diarized settings with limited added cost.