A hybrid pipeline (LLM to text, duration predictor, phoneme-conditioned speech language model) generates spoken dialogue with improved semantic coherence while preserving turn-taking naturalism.
Conversational short-phrase speaker diarization via self-adjusting speech segmentation and embedding ex- traction,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
SLIDE: Integrating Speech Language Model with LLM for Spontaneous Spoken Dialogue Generation
A hybrid pipeline (LLM to text, duration predictor, phoneme-conditioned speech language model) generates spoken dialogue with improved semantic coherence while preserving turn-taking naturalism.