Evaluator-specific demonstrations with retrospective reasoning improve LLM simulation of individual preference judges by up to 9.9 points over a non-personalized base judge, while interface telemetry often degrades accuracy.
Transactions of the Association for Computational Linguistics , volume=
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3representative citing papers
LLMs can be statistically superior to humans at estimating group-level judgments on subjective tasks because of their low variance and decoupled representation-processing biases.
Unsupervised embedding proxies for social constructs are non-identified mixtures of target and confounders, so the Construct Validity Protocol plus Counterfactual Neutralization are required to turn them into defensible measures.
citing papers explorer
-
PERSONAJUDGE: Simulating Individual Human Preference Judgments with Evaluator-Specific Demonstration Data
Evaluator-specific demonstrations with retrospective reasoning improve LLM simulation of individual preference judges by up to 9.9 points over a non-personalized base judge, while interface telemetry often degrades accuracy.
-
From Fallback to Frontline: When Can LLMs be Superior Annotators of Human Perspectives?
LLMs can be statistically superior to humans at estimating group-level judgments on subjective tasks because of their low variance and decoupled representation-processing biases.
-
The Proxy Presumption: From Semantic Embeddings to Valid Social Measures
Unsupervised embedding proxies for social constructs are non-identified mixtures of target and confounders, so the Construct Validity Protocol plus Counterfactual Neutralization are required to turn them into defensible measures.