REVIEW 20 cited by
Large Language Models Understand and Can be Enhanced by Emotional Stimuli
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Emotional intelligence significantly impacts our daily behaviors and interactions. Although Large Language Models (LLMs) are increasingly viewed as a stride toward artificial general intelligence, exhibiting impressive performance in numerous tasks, it is still uncertain if LLMs can genuinely grasp psychological emotional stimuli. Understanding and responding to emotional cues gives humans a distinct advantage in problem-solving. In this paper, we take the first step towards exploring the ability of LLMs to understand emotional stimuli. To this end, we first conduct automatic experiments on 45 tasks using various LLMs, including Flan-T5-Large, Vicuna, Llama 2, BLOOM, ChatGPT, and GPT-4. Our tasks span deterministic and generative applications that represent comprehensive evaluation scenarios. Our automatic experiments show that LLMs have a grasp of emotional intelligence, and their performance can be improved with emotional prompts (which we call "EmotionPrompt" that combines the original prompt with emotional stimuli), e.g., 8.00% relative performance improvement in Instruction Induction and 115% in BIG-Bench. In addition to those deterministic tasks that can be automatically evaluated using existing metrics, we conducted a human study with 106 participants to assess the quality of generative tasks using both vanilla and emotional prompts. Our human study results demonstrate that EmotionPrompt significantly boosts the performance of generative tasks (10.9% average improvement in terms of performance, truthfulness, and responsibility metrics). We provide an in-depth discussion regarding why EmotionPrompt works for LLMs and the factors that may influence its performance. We posit that EmotionPrompt heralds a novel avenue for exploring interdisciplinary knowledge for human-LLMs interaction.
Forward citations
Cited by 20 Pith papers
-
ZifaMem: Structured Memory for Persona, Preference, and Emotional Continuity in AI Companions
ZifaMem's structured memory improves LLM-judged emotional continuity over raw-history context by about 11%, matches Mem0 on the primary preference endpoint, and gains nothing from an affect state machine.
-
Personalized Image Aesthetic Assessment via Preference-rich Sample Mining and Cohort Merging
PRAC mines preference-rich images and merges LoRA adapters from aesthetically similar users to achieve state-of-the-art personalized aesthetic rating prediction.
-
Agents with Feelings? Personality and Emotion in Multi-Agent Software Teams
Personality and emotion profiles substantially change multi-agent LLM team pass rates, review scores, revision behavior, and token cost on code generation and code review, with mixed profiles often beating shared ones.
-
Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking
LLM evaluators reach clinician-level agreement on a new German medical benchmark but fail to abstain on difficult items and show lineage-dependent scoring biases.
-
The Personalization Trap: How User Memory Alters Emotional Reasoning in LLMs
Adding user memory to LLMs degrades their emotional-intelligence test scores and systematically disadvantages marginalized user profiles.
-
From Canonical to Complex: Benchmarking LLM Capabilities in Undergraduate Thermodynamics
On a new 50-item thermodynamics benchmark, the best LLM scored 82%, below the authors' 95% tutoring-safety threshold, with diagram-based questions near chance.
-
APIO: Automatic Prompt Induction and Optimization for Grammatical Error Correction and Text Simplification
APIO automatically induces and optimizes instruction-list prompts for grammatical error correction and text simplification, reporting improved scores over prior prompt-based methods on BEA-2019 and ASSET.
-
Revisiting Prompt Engineering: A Comprehensive Evaluation for LLM-based Personalized Recommendation
For cost-efficient LLMs, rephrasing, step-back, and structured reasoning prompts raise ranking accuracy; for high-performance LLMs, a simple baseline prompt matches complex prompts at a fraction of the cost.
-
Leveraging GPT-4 for Vulnerability-Witnessing Unit Test Generation
GPT-4 generated syntactically valid vulnerability-witnessing unit tests in 66.5% of runs, semantically valid tests in 7.5%, and useful templates in 68.5%, suggesting a semi-automated role.
-
Which Prompting Technique Should I Use? An Empirical Investigation of Prompting Techniques for Software Engineering Tasks
Across ten software engineering tasks and four LLMs, no prompting technique wins consistently; ES-KNN is best on many tasks, some techniques underperform the baseline, and USC is best for code QA and code generation.
-
Analysis of Threat-Based Manipulation in Large Language Models: A Dual Perspective on Vulnerabilities and Performance Enhancement Opportunities
Threat-based prompts change LLM output length, style, and certainty, but the claimed performance gains are not backed by accuracy measures.
-
Prompt Engineering for Requirements Engineering: A Literature Review and Roadmap
The first roadmap-oriented systematic literature review of prompt engineering for requirements engineering analyzes 35 studies and proposes a hybrid taxonomy and research roadmap.
-
ChatGPT Reads Your Tone and Responds Accordingly -- Until It Does Not -- Emotional Framing Induces Bias in LLM Outputs
A triplet-prompt study claiming GPT-4 rebounds from negative prompts to neutral or positive answers and suppresses tone effects on sensitive topics, but the reported tables contradict the headline claims.
-
Identifying Helpful Context for LLM-based Vulnerability Repair: A Preliminary Study
Using CVE descriptions and manually selected code context in prompts, and combining the best prompts, GPT-4o fixed 26 of 42 Java vulnerabilities at least once, up from 19 with its baseline prompt.
-
Empathic Prompting: Non-Verbal Context Integration for Multimodal LLM Conversations
A multimodal chatbot framework that injects real-time facial-expression-derived valence, arousal, and emotion labels into LLM prompts can condition responses on non-verbal affect.
-
PromptGuard: An Orchestrated Prompting Framework for Principled Synthetic Text Generation for Vulnerable Populations using LLMs with Enhanced Safety, Fairness, and Controllability
A framework paper that claims its VulnGuard prompt technique cuts harmful LLM outputs by 25-30% via theoretical bounds, without a real proof or empirical test.
-
Psychologically Enhanced AI Agents
MBTI personality prompts measurably change how LLM agents write stories and play strategic games, with self-reflection before communication supporting cooperative behavior.
-
More Parameters Than Populations: A Systematic Literature Review of Large Language Models within Survey Research
A work-in-progress systematic review finds LLM use in survey research clusters in instrument development, synthetic respondent modeling, and automated text classification, leaving interviewing and cross-lingual work thin.
-
Balancing Knowledge Delivery and Emotional Comfort in Healthcare Conversational Systems
Fine-tuning a 1B medical chatbot on LLM-rewritten emotional dialogues improves its emotion scores with only small changes in n-gram overlap with the original medical responses.
-
The Future of Continual Learning in the Era of Foundation Models: Three Key Directions
Continual learning should pivot from weight-update-based methods to continual compositionality and orchestration of foundation models and agents.
Discussion (0). Sign in to comment.