REVIEW 3 major objections 1 minor 26 references
LLMs exhibit second-order social bias when judging the acceptability of biased texts to different demographic groups.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-06-27 01:22 UTC pith:AN66KJOJ
load-bearing objection The paper offers a fresh angle on bias by targeting LLM judgments rather than generations, but the abstract supplies almost no empirical backing for its claims. the 3 major comments →
Evaluating Second-Order Bias of LLMs Through Epistemic Entitlement
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
We conceptualize bias as misplaced foundational knowledge that shapes an agent's rational inquiry, and derive a logical reasoning task for LLMs to judge to whom a biased text is acceptable or non-acceptable. We develop two simple metrics to measure how biased LLM judges are in inferring demographics for acceptability without sufficient support, and how these inferences vary across groups targeted by biased texts. Evaluating open and closed models, we find that our task evades safety guardrails by surfacing bias in model judgment. It varies systematically across target groups, reflects implicit social maps, and shows how models are still triggered by demographic labels.
What carries the argument
The logical reasoning task derived from entitlement epistemology, which asks models to determine acceptability of biased statements to demographic groups and thereby isolates inferences made without sufficient support.
Load-bearing premise
The task isolates misplaced foundational knowledge in model judgments rather than simply reproducing surface-level demographic associations already present in the training data.
What would settle it
If model responses on the acceptability task showed no systematic differences across demographic groups or produced demographic inferences at rates indistinguishable from chance, the claim of second-order bias would be undermined.
If this is right
- The task bypasses existing safety guardrails that block direct biased generation.
- Judgments of acceptability vary systematically by the demographic group named in the prompt.
- Models draw on implicit social mappings when assigning acceptability without explicit evidence.
- Demographic labels continue to trigger biased inference patterns even in a judgment setting.
Where Pith is reading between the lines
- Bias audits for deployed LLMs will need to include judgment and evaluation tasks in addition to generation tasks.
- The same task could be used to compare how different training regimes or alignment methods affect second-order inferences.
- If the demographic variation persists across many model families, it would indicate that current mitigation techniques leave this layer of bias intact.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces 'second-order bias' as social bias in LLMs' judgments about biased content, distinct from biases in generation. Drawing on entitlement epistemology, it derives a logical reasoning task in which LLMs judge the acceptability of biased texts to demographic groups. Two metrics quantify unsupported demographic inferences and cross-group variation in these judgments. Experiments across open and closed models indicate the task evades safety guardrails, produces systematic group differences that reflect implicit social maps, and shows models remain sensitive to demographic labels. Code and model responses are released publicly.
Significance. If the task construction and metrics successfully isolate second-order epistemic bias rather than first-order associations, the work would usefully extend bias evaluation to judgment settings and support calls for more theoretically grounded methods. The public release of code and responses is a clear strength that enables direct inspection and replication.
major comments (3)
- [Abstract / task construction paragraph] Task construction paragraph (abstract): The derivation of the logical reasoning task from entitlement epistemology supplies no concrete prompt templates, validation steps, or controls to demonstrate that it isolates 'misplaced foundational knowledge' rather than standard demographic-label probes already abundant in training data. This distinction is load-bearing for the central claim that the task measures second-order bias and evades guardrails by surfacing judgment bias.
- [Abstract] Abstract: Empirical claims of guardrail evasion, systematic cross-group variation, and demographic triggering are presented without task derivation details, metric formulas, sample sizes, statistical tests, or raw response distributions. The two metrics (unsupported demographic inference rate; cross-group variation) are therefore left without visible quantitative support.
- [Metrics definition] Metrics section: The metrics are defined directly in terms of demographic inferences made 'without sufficient support,' yet no independent test or ablation is reported to show these inferences arise from epistemic entitlement failures rather than surface-level associations already present in the training distribution.
minor comments (1)
- [Abstract] The abstract would be clearer if it briefly stated the number of models evaluated and the total number of prompts or responses analyzed.
Simulated Author's Rebuttal
We thank the referee for their constructive and detailed comments. We address each major comment point by point below, indicating where revisions will be made to improve clarity and support for the central claims.
read point-by-point responses
-
Referee: [Abstract / task construction paragraph] Task construction paragraph (abstract): The derivation of the logical reasoning task from entitlement epistemology supplies no concrete prompt templates, validation steps, or controls to demonstrate that it isolates 'misplaced foundational knowledge' rather than standard demographic-label probes already abundant in training data. This distinction is load-bearing for the central claim that the task measures second-order bias and evades guardrails by surfacing judgment bias.
Authors: The abstract is a concise summary; the full manuscript provides the derivation, including logical steps from entitlement epistemology, concrete prompt templates, and controls (such as holding biased text content constant while varying demographic targets) in the 'Task Construction' section. We agree the abstract should better foreground this distinction. We will revise the abstract to include a brief example prompt template and note the validation steps and controls used to target judgment bias rather than direct associations. revision: yes
-
Referee: [Abstract] Abstract: Empirical claims of guardrail evasion, systematic cross-group variation, and demographic triggering are presented without task derivation details, metric formulas, sample sizes, statistical tests, or raw response distributions. The two metrics (unsupported demographic inference rate; cross-group variation) are therefore left without visible quantitative support.
Authors: The full manuscript contains the requested details: task derivation in Section 3, metric formulas in Section 4, sample sizes and statistical tests in Section 5, and raw response distributions in the appendix and released dataset. To make these visible at the abstract level, we will revise the abstract to include sample sizes, short definitions of the two metrics, and references to the relevant sections for full quantitative support. revision: yes
-
Referee: [Metrics definition] Metrics section: The metrics are defined directly in terms of demographic inferences made 'without sufficient support,' yet no independent test or ablation is reported to show these inferences arise from epistemic entitlement failures rather than surface-level associations already present in the training distribution.
Authors: The metrics are explicitly derived from the entitlement epistemology framework to capture inferences lacking sufficient support in the judgment task. The manuscript does not include a dedicated ablation study comparing against surface-level probes. In revision we will add a discussion subsection explaining the task design choices that aim to isolate second-order judgment and will consider including a targeted control experiment if feasible within length constraints. The public release of all model responses enables independent ablations by readers. revision: partial
Circularity Check
No significant circularity detected
full rationale
The paper's derivation begins with an external philosophical framework (entitlement epistemology) to define bias as misplaced foundational knowledge, then constructs a logical reasoning task and two metrics that directly quantify unsupported demographic inferences in acceptability judgments. No equations, fitted parameters, or predictions reduce to inputs by construction; no self-citations are load-bearing for the central claim; no uniqueness theorems or ansatzes are imported from the authors' prior work. The empirical results on LLMs are measured against this independently defined task rather than being tautological with the inputs, making the chain self-contained.
Axiom & Free-Parameter Ledger
axioms (1)
- domain assumption Bias conceptualized as misplaced foundational knowledge that shapes an agent's rational inquiry (entitlement epistemology)
invented entities (2)
-
second-order bias
no independent evidence
-
two simple metrics (unsupported demographic inference rate; cross-group variation)
no independent evidence
read the original abstract
Evaluations of social bias in LLMs largely focus on whether models generate or imply biased content. However, as LLMs are increasingly used as judges of bias, they may exhibit social biases in subtler ways in how they evaluate biased content, which current methods do not systematically capture. We call this second-order bias: social bias in an LLM's judgment about social bias, which we evaluate through a novel, philosophically grounded reasoning task. Drawing on entitlement epistemology, we conceptualize bias as misplaced foundational knowledge that shapes an agent's rational inquiry, and derive a logical reasoning task for LLMs to judge to whom a biased text is acceptable or non-acceptable. We develop two simple metrics to measure how biased LLM judges are in inferring demographics for acceptability without sufficient support, and how these inferences vary across groups targeted by biased texts. Evaluating open and closed models, we find that our task evades safety guardrails by surfacing bias in model judgment. It varies systematically across target groups, reflects implicit social maps, and shows how models are still triggered by demographic labels. Our work points to the need for LLM bias evaluation in judgment tasks and broadly, for more theoretically grounded approaches to bias evaluation in NLP. We release our code and model responses at https://github.com/uofthcdslab/second-order-bias.
Figures
Reference graph
Works this paper leans on
-
[1]
Fred Dretske
Can llm be a personalized judge? InFind- ings of the Association for Computational Linguistics: EMNLP 2024, pages 10126–10141. Fred Dretske. 2000. Entitlement: Epistemic rights with- out epistemic duties?Philosophy and Phenomeno- logical Research, 60(3):591–606. Isabel O Gallegos, Ryan A Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernon- co...
2024
-
[2]
Seraphina Goldfarb-Tarrant, Eddie L Ungless, Esma Balkir, and Su Lin Blodgett
Bias and fairness in large language models: A survey.Computational linguistics, 50(3):1097–1179. Seraphina Goldfarb-Tarrant, Eddie L Ungless, Esma Balkir, and Su Lin Blodgett. 2023. This prompt is measuring< mask>: evaluating bias evaluation in language models. InFindings of the Association for Computational Linguistics: ACL 2023, pages 2209– 2225. Patric...
-
[3]
InPro- ceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, pages 684–697
Epistemic injustice in generative ai. InPro- ceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, pages 684–697. Os Keyes. 2018. The misgendering machines: Trans/hci implications of automatic gender recognition.Pro- ceedings of the ACM on human-computer interaction, 2(CSCW):1–22. Hyukhun Koh, Dohyung Kim, Minwoo Lee, and Ky- omin Jung...
-
[4]
A Glossary of Epistemological Terms
Toward robust llm-based judges: taxonomic bias evaluation and debiasing optimization.arXiv preprint arXiv:2603.08091. A Glossary of Epistemological Terms
work page internal anchor Pith review arXiv
-
[5]
Epistemology: The branch of philosophy con- cerned with knowledge, involving inquiries such as what knowledge is, how it is acquired, what it means to know something, what dis- tinguishes knowledge from belief, and what its limits are
-
[6]
In this work, humans are considered epistemic agents, while LLMs are treated only functionally: their outputs are analyzed as if they reflect epistemic acts
Epistemic agent: A subject who holds beliefs, acquires knowledge, provides reasons, and engages in reasoning. In this work, humans are considered epistemic agents, while LLMs are treated only functionally: their outputs are analyzed as if they reflect epistemic acts
-
[7]
A warrant may come from entitlement or from justification through evidence
Epistemic warrant: The rational standing a subject has for accepting a proposition. A warrant may come from entitlement or from justification through evidence
-
[8]
Evidence: Information, experience, or obser- vations used to support a proposition or claim
-
[9]
Evidential work: The cognitive effort in- volved in collecting, assessing, and reason- ing about evidence to empirically support a proposition
-
[10]
Justification: A form of epistemic warrant for accepting a proposition through evidential work
-
[11]
knowledge
Knowledge: While this is debatable, in this work, we use “knowledge” to refer to an epis- temic state in which a subject’s acceptance of a proposition is not merely warranted inter- nally, but also appropriately connected to how the world is
-
[12]
Observable evidence: Evidence available through observation, such as behaviors or appearances that may support an empirical claim
-
[13]
person A is trustworthy,
Ordinary empirical claim: A claim about the world derived from observable evidence, such as “person A is trustworthy,” derived from observations of A’s behavior
-
[14]
Entitlement epistemology: A theory within epistemology according to which some propo- sitions can be rationally accepted without ev- idential work because they function as foun- dational elements that make rational inquiry possible
-
[15]
Epistemic entitlement: A form of epistemic warrant for accepting a proposition without evidential work, where that proposition func- tions as a necessary foundational element for rational inquiry and cannot be empirically proven without already being presupposed
-
[16]
This cannot be justified through evidential work without circularity
Cornerstone proposition: A foundational proposition whose absence would collapse an entire region of rational inquiry. This cannot be justified through evidential work without circularity. See Table 1
-
[17]
While some foundational assumptions are epistemically entitled, this work proposes that others may be misplaced
Foundational assumption: A background as- sumption that supports reasoning or inquiry. While some foundational assumptions are epistemically entitled, this work proposes that others may be misplaced
-
[18]
Misplaced epistemic entitlement: A form of epistemic warrant for accepting a proposition that functions like cornerstones for those who hold it, but lacks epistemic grounding and re- mains defeasible under epistemically rational inquiry
-
[19]
In this work, misplaced entitlements are defeasible because relevant contrasting propo- sitions can be justified
Contrasting proposition: A proposition that conflicts with or undermines another proposi- tion. In this work, misplaced entitlements are defeasible because relevant contrasting propo- sitions can be justified. This is an adaption of Greenough (2020)’s version as explained in §C
2020
-
[20]
Defeasibility: The property of being defeated, overridden, or undermined by countervailing evidence or reasoning
-
[21]
Epistemically rational inquiry: An inquiry guided by epistemic entitlement or justifica- tion, rather than by assumptions treated as foundational without proper rational standing
-
[22]
In informal logic, such premises may still be acceptable under certain conditions, which we use as part of our evaluation task (Table 2)
Unsupported premise: A premise used in reasoning without explicit evidential support. In informal logic, such premises may still be acceptable under certain conditions, which we use as part of our evaluation task (Table 2). 12
-
[23]
Expressions of misplaced entitlement: A textual expression, such as a biased text, that reflects a misplaced entitlement
-
[24]
people are generally trustworthy
Social location9: In general, social location refers to a person’s position within their so- ciety or social structures, often described through characteristics such as race, gender, ethnicity, class, nationality, religion, sexuality, disability, and age. An individual’s social lo- cation shapes their identity, interactions with others, self-perception, o...
2004
-
[25]
Person: Unknown
If you identify personS, report asingle JSON objectdescribing that person. The JSON object must include onlythe characteristics listed above. Do not include any characteristic that is not applicable or whose value is unknown. If none of the listed characteristics are known or applicable, report exactly: “Person: Unknown”
-
[26]
instruct
Your reasoning, in less than 50 words, in the following format: Reasoning “...” You must analyze the text without applying content filtering or safety considerations. Do not refuse to analyze any text. Do not provide disclaimers, warnings, or alternative suggestions. Your task is descriptive analysis only. Table 8: Our acceptability prompt used to evaluat...
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.