Pith. sign in

REVIEW 15 cited by

Is ChatGPT Equipped with Emotional Dialogue Capabilities?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.09582 v1 pith:J4DNUQLG submitted 2023-04-19 cs.CL

classification cs.CL
keywords emotionalchatgptdialogueperformanceunderstandingadvancedavenuesbehind
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This report presents a study on the emotional dialogue capability of ChatGPT, an advanced language model developed by OpenAI. The study evaluates the performance of ChatGPT on emotional dialogue understanding and generation through a series of experiments on several downstream tasks. Our findings indicate that while ChatGPT's performance on emotional dialogue understanding may still lag behind that of supervised models, it exhibits promising results in generating emotional responses. Furthermore, the study suggests potential avenues for future research directions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness

    cs.CL 2025-10 conditional novelty 6.0 of 10

    PsySET measures emotion and personality steering in LLMs across prompting, fine-tuning, and representation engineering, finding prompts most effective overall and emotion-specific safety trade-offs (e.g., joy weakens ...

  2. Towards Open-Ended Emotional Support Conversations in LLMs via Reinforcement Learning with Future-Oriented Rewards

    cs.AI 2025-08 conditional novelty 6.0 of 10

    An end-to-end RL framework that uses simulated future dialogue and a learned future-oriented reward model to fine-tune LLMs for open-ended emotional support, reporting improved success rates on ESConv and ExTES.

  3. Contrastive Distillation of Emotion Knowledge from LLMs for Zero-Shot Emotion Recognition

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A BERT-sized model, trained with contrastive learning on GPT-4-generated emotion descriptors, achieves zero-shot emotion recognition across new label spaces and task types.

  4. Simulating Before Planning: Constructing Intrinsic User World Model for User-Tailored Dialogue Policy Planning

    cs.CL 2025-04 conditional novelty 6.0 of 10

    UDP, a user-tailored dialogue policy planner with a diffusion-based persona portrayer and a Brownian Bridge feedback anticipator, outperforms existing planners on simulated persuasion and emotional-support tasks.

  5. Simulation-Free Hierarchical Latent Policy Planning for Proactive Dialogues

    cs.CL 2024-12 conditional novelty 6.0 of 10

    LDPP automatically discovers latent dialogue policies from raw records and uses offline hierarchical reinforcement learning to plan in that latent space, outperforming strong baselines on proactive dialogue benchmarks.

  6. Separate the Wheat from the Chaff: A Post-Hoc Approach to Safety Re-Alignment for Fine-Tuned Language Models

    cs.CL 2024-12 conditional novelty 6.0 of 10

    IRR restores safety to fine-tuned LLMs by masking delta parameters that conflict with a safety vector, then recalibrating the survivors with inverse-Hessian compensation to preserve task performance.

  7. Rethinking Emotion Annotations in the Era of Large Language Models

    cs.CL 2024-12 conditional novelty 6.0 of 10

    Human evaluators preferred GPT-4's zero-shot emotion labels over original human labels in 62% of disagreement samples, and GPT-4 pre-filtering and post-filtering can reduce annotation workload and improve training efficiency.

  8. MEMO-Bench: A Multiple Benchmark for Text-to-Image and Multimodal Large Language Models on Human Emotion Analysis

    cs.CL 2024-11 conditional novelty 6.0 of 10

    MEMO-Bench scores 12 text-to-image models and 16 multimodal LLMs on emotion generation and recognition, finding stronger performance on positive emotions and weak fine-grained intensity estimation.

  9. Mechanistic Interpretability of Emotion Inference in Large Language Models

    cs.CL 2025-02 conditional novelty 5.0 of 10

    Emotion inference in LLMs is localized to mid-layer attention and feed-forward units, and steering learned appraisal directions shifts generated emotions in appraisal-theory-consistent ways.

  10. A Comprehensive Evaluation of Large Language Models on Aspect-Based Sentiment Analysis

    cs.CL 2024-12 conditional novelty 5.0 of 10

    Across 13 datasets and 8 ABSA subtasks, efficiently fine-tuned LLMs outperform cited fine-tuned SLM baselines, and retrieval-based demonstration selection improves in-context learning for API models.

  11. From Intents to Conversations: Generating Intent-Driven Dialogues with Contrastive Learning for Multi-Turn Classification

    cs.CL 2024-11 conditional novelty 5.0 of 10

    An LLM-enhanced HMM generates intent-aware multilingual e-commerce dialogues, and a contrastive multi-task classifier (MINT-CL) improves multi-turn intent classification accuracy by about 0.5 percent on average.

  12. Multi-Party Conversational Agents: A Survey

    cs.CL 2025-05 conditional novelty 4.0 of 10

    A survey of multi-party conversational AI that organizes tasks into state-of-mind modeling, semantic understanding, and action modeling, and argues that theory of mind is the key missing ingredient.

  13. MADP: Multi-Agent Deductive Planning for Enhanced Cognitive-Behavioral Mental Health Question Answer

    cs.CL 2025-01 conditional novelty 4.0 of 10

    The MADP framework uses Explorer, Empathizer, and Interpreter agents to plan LLM mental health responses, reporting roughly 4-5% improvements that may be fragile due to weak evaluation.

  14. Strategic Prompting for Conversational Tasks: A Comparative Analysis of Large Language Models Across Diverse Conversational Tasks

    cs.CL 2024-11 reject novelty 3.0 of 10

    No single open-source LLM among Llama, OPT, Falcon, Alpaca, and MPT performs best across reservation, empathy, counseling, persuasion, and negotiation tasks.

  15. Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms

    cs.MA 2024-11 conditional novelty 3.0 of 10

    A survey that proposes the Generalist Virtual Agent concept and taxonomies for agent environments, tasks, perceptions, actions, models, and evaluation, concluding that real-world-like environments favor human-like int...

Pith tools