Evaluator-specific demonstrations with retrospective reasoning improve LLM simulation of individual preference judges by up to 9.9 points over a non-personalized base judge, while interface telemetry often degrades accuracy.
Language Resources and Evaluation , volume=
4 Pith papers cite this work. Polarity classification is still indexing.
years
2026 4representative citing papers
LLMs can be statistically superior to humans at estimating group-level judgments on subjective tasks because of their low variance and decoupled representation-processing biases.
Analyses of labeled social media sentences and interpretations show 30% divergence in ethos and pathos, greater variability for charged content, and predictive power for audience attitudes toward the author.
Disagreement in health-literacy annotations is driven by conceptual task difficulty rather than annotator differences, with social effects varying or reversing by agreement level, making perspectivist modeling necessary.
citing papers explorer
-
PERSONAJUDGE: Simulating Individual Human Preference Judgments with Evaluator-Specific Demonstration Data
Evaluator-specific demonstrations with retrospective reasoning improve LLM simulation of individual preference judges by up to 9.9 points over a non-personalized base judge, while interface telemetry often degrades accuracy.
-
From Fallback to Frontline: When Can LLMs be Superior Annotators of Human Perspectives?
LLMs can be statistically superior to humans at estimating group-level judgments on subjective tasks because of their low variance and decoupled representation-processing biases.
-
How Ethos and Pathos Appeals Resonate in Reader Interpretations of Social Media Messages
Analyses of labeled social media sentences and interpretations show 30% divergence in ethos and pathos, greater variability for charged content, and predictive power for audience attitudes toward the author.
-
Structured Disagreement in Health-Literacy Annotation: Epistemic Stability, Conceptual Difficulty, and Agreement-Stratified Inference
Disagreement in health-literacy annotations is driven by conceptual task difficulty rather than annotator differences, with social effects varying or reversing by agreement level, making perspectivist modeling necessary.