REVIEW 27 cited by
GoEmotions: A Dataset of Fine-Grained Emotions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Understanding emotion expressed in language has a wide range of applications, from building empathetic chatbots to detecting harmful online behavior. Advancement in this area can be improved using large-scale datasets with a fine-grained typology, adaptable to multiple downstream tasks. We introduce GoEmotions, the largest manually annotated dataset of 58k English Reddit comments, labeled for 27 emotion categories or Neutral. We demonstrate the high quality of the annotations via Principal Preserved Component Analysis. We conduct transfer learning experiments with existing emotion benchmarks to show that our dataset generalizes well to other domains and different emotion taxonomies. Our BERT-based model achieves an average F1-score of .46 across our proposed taxonomy, leaving much room for improvement.
Forward citations
Cited by 27 Pith papers
-
Video2Reaction: Mapping Video to Audience Reaction Distribution in the Wild
A new dataset and benchmark maps movie clips to distributions of audience emotional reactions derived from YouTube comments, showing that finetuned vision-language models can predict these distributions from video alone.
-
E3AD: An Emotion-Aware Vision-Language-Action Model for Human-Centric End-to-End Autonomous Driving
An emotion-aware vision-language-action driving model estimates VAD emotion from commands and uses it to improve grounding and waypoint planning.
-
Backtranslation and paraphrasing in the LLM era? Comparing data augmentation methods for emotion classification
Backtranslation and paraphrasing produce competitive or better classification gains than zero-shot and few-shot generation when augmenting a low-resource emotion dataset.
-
Moodifier: MLLM-Enhanced Emotion-Driven Image Editing
A training-free editing pipeline that uses a large emotion-annotated dataset and a fine-tuned CLIP model to modify only the parts of an image that convey a target emotion.
-
Abstract Counterfactuals for Language Model Agents
Counterfactuals for LM agents computed over a high-level abstraction of the action, instead of its tokens, preserve the observed action's meaning across counterfactual contexts far more often than token-level counterfactuals.
-
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models
SocialMaze is a six-task benchmark that claims to evaluate LLM social reasoning along deep reasoning, dynamic interaction, and information uncertainty dimensions.
-
EmoMeta: A Multimodal Dataset for Fine-grained Emotion Classification in Chinese Metaphors
Introduces EmoMeta, a publicly available Chinese multimodal dataset of 5,000 metaphorical advertisements with fine-grained emotion labels across ten categories.
-
PicPersona-TOD : A Dataset for Personalizing Utterance Style in Task-Oriented Dialogue with Image Persona
Introduces PicPersona-TOD, the first task-oriented dialogue dataset with user images as persona, together with a multimodal NLG baseline called Pictor.
-
AI for the Open-World: the Learning Principles
Open-world AI requires rich features, disentangled representations, and inference-time learning; the thesis presents techniques and large-scale experiments supporting these principles.
-
Synthetic Audio Helps for Cognitive State Tasks
Adding zero-shot synthetic audio from a text-to-speech system to text-only models improves cognitive-state prediction on seven tasks, though gains are small.
-
Once More, With Feeling: Measuring Emotion of Acting Performances in Contemporary American Film
A computational pipeline aligns movie audio with script text and uses speech emotion models to show that acted emotions in American film track narrative structure, release year, genre, and dialogue function.
-
FLARE: Few-shot Learning-based Adaptive Reflective Engine
FLARE, an error-driven prompt optimizer that combines reflective rewriting with a small few-shot reference set, reports higher scores than GEPA on all tested GPT-5 task-model pairs.
-
MVRS: The Multimodal Virtual Reality Stimuli-based Emotion Recognition Dataset
A new VR-based emotion dataset with synchronized eye tracking, body motion, EMG, and GSR from 13 participants, evaluated with classifiers but with questionable validation.
-
AI in Mental Health: Emotional and Sentiment Analysis of Large Language Models' Responses to Depression, Anxiety, and Stress Queries
Eight LLMs show measurably different emotional tones in mental-health answers: anxiety prompts produced near-saturated fear scores, depression prompts the most sadness, and stress prompts the most optimism.
-
MAPLE: Many-Shot Adaptive Pseudo-Labeling for In-Context Learning
MAPLE uses graph-influence scores to select and pseudo-label the most useful unlabeled examples, then adaptively chooses demonstrations per query, improving many-shot in-context learning with few human labels.
-
Co-Designing a Chatbot for Culturally Competent Clinical Communication: Experience and Reflections
A co-designed GPT-4o chatbot with ACT-based feedback gave 29 medical student consultations a low-pressure practice environment for culturally competent communication, but the pilot lacked a formal design and external ...
-
AI with Emotions: Exploring Emotional Expressions in Large Language Models
LLMs given numerical arousal and valence coordinates produce text that sentiment analysis places in the same region of Russell's circumplex, demonstrating limited but real control over emotional tone.
-
AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues
AV-EmoDialog uses speech and face encoders with a large language model to generate emotion-aware dialogue responses from audio-visual input, reporting better emotional alignment than the compared baselines.
-
Investigating Algorithmic Bias in YouTube Shorts
YouTube Shorts recommendations from political seeds drift to entertainment and positive-emotion content within the first few steps, with the drift unchanged by simulated watch-time.
-
Crowd-SFT: Crowdsourcing for LLM Alignment
A competitive multi-group fine-tuning framework with point rewards correlated to Shapley values reduced simulated model distance by up to 55% and tracked user contributions reasonably in vector-space experiments.
-
The Super Emotion Dataset
The paper presents SuperEmotion, an aggregated NLP dataset of roughly 555k samples (abstract says 519k) remapped to Shaver's six emotions plus neutral, with no experimental validation.
-
Examining gender and cultural influences on customer emotions
Across 129,297 reviews, gender and culture interact in emotion scores, with larger female-male gaps among Western than Eastern reviewers, though the models explain under 1.2% of variance.
-
Performance Evaluation of Emotion Classification in Japanese Using RoBERTa and DeBERTa
DeBERTa-v3-large achieves the highest mean F1 (0.662) among compared models for binary detection of eight Plutchik emotions in Japanese WRIME posts, though the paper's stated accuracy advantage is not supported by its...
-
Examining the sentiment and emotional differences in product and service reviews: The moderating role of culture
Service reviews show stronger, more varied emotions than product reviews, and the gap is larger for Eastern than Western consumers, while gender has little moderating effect.
-
Hype and Adoption of Generative Artificial Intelligence Applications
The paper interprets sentiment and emotion curves from tweets as evidence that generative AI adoption follows the Gartner Hype Cycle and Kübler-Ross Change Curve.
-
AIMA at SemEval-2024 Task 10: History-Based Emotion Recognition in Hindi-English Code-Mixed Conversations
An ensemble of four context-aware models with a Hinglish-to-English translation pipeline achieves weighted F1 0.4080 on SemEval 2024 Task 10 subtask 1, marginally above its strongest single model.
-
Analyzing Emotions in Bangla Social Media Comments Using Machine Learning and LIME
On the EmoNoBa Bangla emotion dataset, a boosted decision tree achieves 0.7860 macro F1, the best among classical models tested but below transformer-based baselines.
Discussion (0). Continue with ORCID to comment.