REVIEW 5 cited by
E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This study focuses on emotion-sensitive spoken dialogue in human-machine speech interaction. With the advancement of Large Language Models (LLMs), dialogue systems can handle multimodal data, including audio. Recent models have enhanced the understanding of complex audio signals through the integration of various audio events. However, they are unable to generate appropriate responses based on emotional speech. To address this, we introduce the Emotional chat Model (E-chat), a novel spoken dialogue system capable of comprehending and responding to emotions conveyed from speech. This model leverages an emotion embedding extracted by a speech encoder, combined with LLMs, enabling it to respond according to different emotional contexts. Additionally, we introduce the E-chat200 dataset, designed explicitly for emotion-sensitive spoken dialogue. In various evaluation metrics, E-chat consistently outperforms baseline model, demonstrating its potential in emotional comprehension and human-machine interaction.
Forward citations
Cited by 5 Pith papers
-
Dual Information Speech Language Models for Emotional Conversations
A dual-adapter design with equivalence replacement regularization lets frozen LLMs perceive both paralinguistic and linguistic information from speech for emotional conversation.
-
OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios
ShareChatX and OmniChat show that large-scale synthetic spoken dialogue data improves multi-turn response quality and emotion prediction, setting a new state of the art on DailyTalk.
-
X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment
X3-OPD improves audio-grounded reasoning by training the audio student on its own rollouts with token-level teacher feedback, using a three-tier paired text-audio corpus.
-
Leveraging Chain of Thought towards Empathetic Spoken Dialogue without Corresponding Question-Answering Data
Listen, Perceive, Express (LPE) uses two-stage ASR/SER training plus chain-of-thought prompting to let a frozen LLM generate empathetic responses from speech without question-answer fine-tuning data.
-
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
A single aligned multimodal model, MinMo, achieves strong or state-of-the-art results on speech recognition, translation, emotion recognition, voice generation, and full-duplex dialogue while preserving the underlying...
Discussion (0). Continue with ORCID to comment.