Pith. sign in

hub

LLM Evaluators Recognize and Favor Their Own Generations

50 Pith papers cite this work. Polarity classification is still indexing.

50 Pith papers citing it
abstract

Self-evaluation using large language models (LLMs) has proven valuable not only in benchmarking but also methods like reward modeling, constitutional AI, and self-refinement. But new biases are introduced due to the same LLM acting as both the evaluator and the evaluatee. One such bias is self-preference, where an LLM evaluator scores its own outputs higher than others' while human annotators consider them of equal quality. But do LLMs actually recognize their own outputs when they give those texts higher scores, or is it just a coincidence? In this paper, we investigate if self-recognition capability contributes to self-preference. We discover that, out of the box, LLMs such as GPT-4 and Llama 2 have non-trivial accuracy at distinguishing themselves from other LLMs and humans. By fine-tuning LLMs, we discover a linear correlation between self-recognition capability and the strength of self-preference bias; using controlled experiments, we show that the causal explanation resists straightforward confounders. We discuss how self-recognition can interfere with unbiased evaluations and AI safety more generally.

hub tools

citation-role summary

background 3

citation-polarity summary

roles

background 3

polarities

background 3

representative citing papers

Noncrossing Duality and the Geometry of Positive Tropical Linear Spaces

math.CO · 2026-04-28 · unverdicted · novelty 7.0

Establishes noncrossing duality linking positive tropical Grassmannian fan structure to noncrossing fans with a bijection to noncrossing tableaux, and realizes bounded complexes of tropical linear spaces as subdifferentials whose diameter is set by the planar kinematics weight.

AGC-Bench: Measuring Artificial General Creativity

cs.CL · 2026-07-01 · unverdicted · novelty 6.0 · 2 refs

AGC-Bench introduces a multi-domain creativity benchmark for LLMs, recovers a general 'c' factor explaining 81.5% of variance, and finds humans still outperform top models on matched tasks.

AMEL: Accumulated Message Effects on LLM Judgments

cs.AI · 2026-05-21 · unverdicted · novelty 6.0 · 2 refs

LLMs exhibit an accumulated message effect where conversation history polarity biases subsequent judgments, stronger for high-entropy items, independent of context length, and with a negativity bias.

citing papers explorer

Showing 50 of 50 citing papers.