Pith. sign in

REVIEW 20 cited by

Large Language Models Understand and Can be Enhanced by Emotional Stimuli

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.11760 v7 pith:MZR2OYUD submitted 2023-07-14 cs.CL cs.AIcs.HC

classification cs.CLcs.AIcs.HC
keywords emotionalllmsperformancetasksemotionpromptstimuligenerativeintelligence
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Emotional intelligence significantly impacts our daily behaviors and interactions. Although Large Language Models (LLMs) are increasingly viewed as a stride toward artificial general intelligence, exhibiting impressive performance in numerous tasks, it is still uncertain if LLMs can genuinely grasp psychological emotional stimuli. Understanding and responding to emotional cues gives humans a distinct advantage in problem-solving. In this paper, we take the first step towards exploring the ability of LLMs to understand emotional stimuli. To this end, we first conduct automatic experiments on 45 tasks using various LLMs, including Flan-T5-Large, Vicuna, Llama 2, BLOOM, ChatGPT, and GPT-4. Our tasks span deterministic and generative applications that represent comprehensive evaluation scenarios. Our automatic experiments show that LLMs have a grasp of emotional intelligence, and their performance can be improved with emotional prompts (which we call "EmotionPrompt" that combines the original prompt with emotional stimuli), e.g., 8.00% relative performance improvement in Instruction Induction and 115% in BIG-Bench. In addition to those deterministic tasks that can be automatically evaluated using existing metrics, we conducted a human study with 106 participants to assess the quality of generative tasks using both vanilla and emotional prompts. Our human study results demonstrate that EmotionPrompt significantly boosts the performance of generative tasks (10.9% average improvement in terms of performance, truthfulness, and responsibility metrics). We provide an in-depth discussion regarding why EmotionPrompt works for LLMs and the factors that may influence its performance. We posit that EmotionPrompt heralds a novel avenue for exploring interdisciplinary knowledge for human-LLMs interaction.

Discussion (0). Sign in to comment.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 58 citations worldwide. Full citation record

  1. ZifaMem: Structured Memory for Persona, Preference, and Emotional Continuity in AI Companions

    cs.AI 2026-07 conditional novelty 6.0 of 10

    ZifaMem's structured memory improves LLM-judged emotional continuity over raw-history context by about 11%, matches Mem0 on the primary preference endpoint, and gains nothing from an affect state machine.

  2. Personalized Image Aesthetic Assessment via Preference-rich Sample Mining and Cohort Merging

    cs.CV 2026-07 conditional novelty 6.0 of 10

    PRAC mines preference-rich images and merges LoRA adapters from aesthetically similar users to achieve state-of-the-art personalized aesthetic rating prediction.

  3. Agents with Feelings? Personality and Emotion in Multi-Agent Software Teams

    cs.SE 2026-07 conditional novelty 6.0 of 10

    Personality and emotion profiles substantially change multi-agent LLM team pass rates, review scores, revision behavior, and token cost on code generation and code review, with mixed profiles often beating shared ones.

  4. Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking

    cs.CL 2026-07 unverdicted novelty 6.0 of 10

    LLM evaluators reach clinician-level agreement on a new German medical benchmark but fail to abstain on difficult items and show lineage-dependent scoring biases.

  5. The Personalization Trap: How User Memory Alters Emotional Reasoning in LLMs

    cs.AI 2025-10 conditional novelty 6.0 of 10

    Adding user memory to LLMs degrades their emotional-intelligence test scores and systematically disadvantages marginalized user profiles.

  6. From Canonical to Complex: Benchmarking LLM Capabilities in Undergraduate Thermodynamics

    physics.ed-ph 2025-08 conditional novelty 6.0 of 10

    On a new 50-item thermodynamics benchmark, the best LLM scored 82%, below the authors' 95% tutoring-safety threshold, with diagram-based questions near chance.

  7. APIO: Automatic Prompt Induction and Optimization for Grammatical Error Correction and Text Simplification

    cs.CL 2025-08 conditional novelty 6.0 of 10

    APIO automatically induces and optimizes instruction-list prompts for grammatical error correction and text simplification, reporting improved scores over prior prompt-based methods on BEA-2019 and ASSET.

  8. Revisiting Prompt Engineering: A Comprehensive Evaluation for LLM-based Personalized Recommendation

    cs.IR 2025-07 conditional novelty 6.0 of 10

    For cost-efficient LLMs, rephrasing, step-back, and structured reasoning prompts raise ranking accuracy; for high-performance LLMs, a simple baseline prompt matches complex prompts at a fraction of the cost.

  9. Leveraging GPT-4 for Vulnerability-Witnessing Unit Test Generation

    cs.SE 2025-06 conditional novelty 6.0 of 10

    GPT-4 generated syntactically valid vulnerability-witnessing unit tests in 66.5% of runs, semantically valid tests in 7.5%, and useful templates in 68.5%, suggesting a semi-automated role.

  10. Which Prompting Technique Should I Use? An Empirical Investigation of Prompting Techniques for Software Engineering Tasks

    cs.SE 2025-06 conditional novelty 6.0 of 10

    Across ten software engineering tasks and four LLMs, no prompting technique wins consistently; ES-KNN is best on many tasks, some techniques underperform the baseline, and USC is best for code QA and code generation.

  11. Analysis of Threat-Based Manipulation in Large Language Models: A Dual Perspective on Vulnerabilities and Performance Enhancement Opportunities

    cs.CR 2025-07 reject novelty 5.0 of 10

    Threat-based prompts change LLM output length, style, and certainty, but the claimed performance gains are not backed by accuracy measures.

  12. Prompt Engineering for Requirements Engineering: A Literature Review and Roadmap

    cs.SE 2025-07 conditional novelty 5.0 of 10

    The first roadmap-oriented systematic literature review of prompt engineering for requirements engineering analyzes 35 studies and proposes a hybrid taxonomy and research roadmap.

  13. ChatGPT Reads Your Tone and Responds Accordingly -- Until It Does Not -- Emotional Framing Induces Bias in LLM Outputs

    cs.CL 2025-06 reject novelty 5.0 of 10

    A triplet-prompt study claiming GPT-4 rebounds from negative prompts to neutral or positive answers and suppresses tone effects on sensitive topics, but the reported tables contradict the headline claims.

  14. Identifying Helpful Context for LLM-based Vulnerability Repair: A Preliminary Study

    cs.SE 2025-06 conditional novelty 5.0 of 10

    Using CVE descriptions and manually selected code context in prompts, and combining the best prompts, GPT-4o fixed 26 of 42 Java vulnerabilities at least once, up from 19 with its baseline prompt.

  15. Empathic Prompting: Non-Verbal Context Integration for Multimodal LLM Conversations

    cs.HC 2025-10 conditional novelty 4.0 of 10

    A multimodal chatbot framework that injects real-time facial-expression-derived valence, arousal, and emotion labels into LLM prompts can condition responses on non-verbal affect.

  16. PromptGuard: An Orchestrated Prompting Framework for Principled Synthetic Text Generation for Vulnerable Populations using LLMs with Enhanced Safety, Fairness, and Controllability

    cs.CV 2025-09 reject novelty 4.0 of 10

    A framework paper that claims its VulnGuard prompt technique cuts harmful LLM outputs by 25-30% via theoretical bounds, without a real proof or empirical test.

  17. Psychologically Enhanced AI Agents

    cs.AI 2025-09 conditional novelty 4.0 of 10

    MBTI personality prompts measurably change how LLM agents write stories and play strategic games, with self-reflection before communication supporting cooperative behavior.

  18. More Parameters Than Populations: A Systematic Literature Review of Large Language Models within Survey Research

    cs.DL 2025-09 conditional novelty 4.0 of 10

    A work-in-progress systematic review finds LLM use in survey research clusters in instrument development, synthetic respondent modeling, and automated text classification, leaving interviewing and cross-lingual work thin.

  19. Balancing Knowledge Delivery and Emotional Comfort in Healthcare Conversational Systems

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Fine-tuning a 1B medical chatbot on LLM-rewritten emotional dialogues improves its emotion scores with only small changes in n-gram overlap with the original medical responses.

  20. The Future of Continual Learning in the Era of Foundation Models: Three Key Directions

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Continual learning should pivot from weight-update-based methods to continual compositionality and orchestration of foundation models and agents.

Pith tools