Pith. sign in

MOSS transcribe diarize: Accurate transcription with speaker diarization,

4 Pith papers cite this work. Polarity classification is still indexing.

4 Pith papers citing it

years

2026 4

verdicts

UNVERDICTED 4

representative citing papers

MOSS-Audio Technical Report

cs.SD · 2026-06-01 · unverdicted · novelty 4.0

MOSS-Audio is an audio-language model using a 12.5 Hz encoder, DeepStack cross-layer injection, time markers, and an event-preserving annotation pipeline for unified audio understanding.

citing papers explorer

Showing 4 of 4 citing papers.

  • Grounding Spoken LLMs in Multi-Speaker Audio via Diarization Conditioning eess.AS · 2026-06-16 · unverdicted · none · ref 11 · internal anchor

    Dixtral uses diarization conditioning on a Whisper-based encoder within Voxtral to outperform baselines on multi-speaker transcription and match or exceed on QA tasks.

  • DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models eess.AS · 2026-04-24 · unverdicted · none · ref 62 · internal anchor

    DM-ASR reformulates multi-speaker ASR as multi-turn dialogue generation conditioned on diarization results, achieving competitive benchmark performance with relatively small models and limited data.

  • Balancing ASR and diarization in end-to-end LLMs for multi-talker speech recognition eess.AS · 2026-06-11 · unverdicted · none · ref 21 · internal anchor

    LLM-based multi-talker ASR with dual-encoder, feature interleaving, length-aware speaker loss, and adaptive ASR threshold achieves 18% and 24% relative gains over baselines on AliMeeting and Aishell4.

  • MOSS-Audio Technical Report cs.SD · 2026-06-01 · unverdicted · none · ref 23 · internal anchor

    MOSS-Audio is an audio-language model using a 12.5 Hz encoder, DeepStack cross-layer injection, time markers, and an event-preserving annotation pipeline for unified audio understanding.