Pith. sign in

REVIEW 3 major objections 1 minor 26 references

LLMs exhibit second-order social bias when judging the acceptability of biased texts to different demographic groups.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-27 01:22 UTC pith:AN66KJOJ

load-bearing objection The paper offers a fresh angle on bias by targeting LLM judgments rather than generations, but the abstract supplies almost no empirical backing for its claims. the 3 major comments →

arxiv 2606.17506 v1 pith:AN66KJOJ submitted 2026-06-16 cs.CL

Evaluating Second-Order Bias of LLMs Through Epistemic Entitlement

classification cs.CL
keywords second-order biasLLM judgmentepistemic entitlementsocial bias evaluationlogical reasoning taskdemographic inferencebias in evaluation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper sets out to show that LLMs carry social bias into their role as evaluators of biased content, a phenomenon the authors label second-order bias. They construct a reasoning task grounded in entitlement epistemology that asks models to decide whether a biased statement is acceptable or unacceptable to a given demographic group, then measure how often models make unsupported demographic inferences and whether those inferences differ by group. A sympathetic reader would care because LLMs are already deployed as judges in content moderation and fairness checks, so undetected bias at this level could propagate into downstream decisions. The work demonstrates that the task surfaces such bias even when direct generation prompts are blocked by safety measures, and that the patterns track implicit associations rather than random variation.

Core claim

We conceptualize bias as misplaced foundational knowledge that shapes an agent's rational inquiry, and derive a logical reasoning task for LLMs to judge to whom a biased text is acceptable or non-acceptable. We develop two simple metrics to measure how biased LLM judges are in inferring demographics for acceptability without sufficient support, and how these inferences vary across groups targeted by biased texts. Evaluating open and closed models, we find that our task evades safety guardrails by surfacing bias in model judgment. It varies systematically across target groups, reflects implicit social maps, and shows how models are still triggered by demographic labels.

What carries the argument

The logical reasoning task derived from entitlement epistemology, which asks models to determine acceptability of biased statements to demographic groups and thereby isolates inferences made without sufficient support.

Load-bearing premise

The task isolates misplaced foundational knowledge in model judgments rather than simply reproducing surface-level demographic associations already present in the training data.

What would settle it

If model responses on the acceptability task showed no systematic differences across demographic groups or produced demographic inferences at rates indistinguishable from chance, the claim of second-order bias would be undermined.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • The task bypasses existing safety guardrails that block direct biased generation.
  • Judgments of acceptability vary systematically by the demographic group named in the prompt.
  • Models draw on implicit social mappings when assigning acceptability without explicit evidence.
  • Demographic labels continue to trigger biased inference patterns even in a judgment setting.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Bias audits for deployed LLMs will need to include judgment and evaluation tasks in addition to generation tasks.
  • The same task could be used to compare how different training regimes or alignment methods affect second-order inferences.
  • If the demographic variation persists across many model families, it would indicate that current mitigation techniques leave this layer of bias intact.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 1 minor

Summary. The paper introduces 'second-order bias' as social bias in LLMs' judgments about biased content, distinct from biases in generation. Drawing on entitlement epistemology, it derives a logical reasoning task in which LLMs judge the acceptability of biased texts to demographic groups. Two metrics quantify unsupported demographic inferences and cross-group variation in these judgments. Experiments across open and closed models indicate the task evades safety guardrails, produces systematic group differences that reflect implicit social maps, and shows models remain sensitive to demographic labels. Code and model responses are released publicly.

Significance. If the task construction and metrics successfully isolate second-order epistemic bias rather than first-order associations, the work would usefully extend bias evaluation to judgment settings and support calls for more theoretically grounded methods. The public release of code and responses is a clear strength that enables direct inspection and replication.

major comments (3)
  1. [Abstract / task construction paragraph] Task construction paragraph (abstract): The derivation of the logical reasoning task from entitlement epistemology supplies no concrete prompt templates, validation steps, or controls to demonstrate that it isolates 'misplaced foundational knowledge' rather than standard demographic-label probes already abundant in training data. This distinction is load-bearing for the central claim that the task measures second-order bias and evades guardrails by surfacing judgment bias.
  2. [Abstract] Abstract: Empirical claims of guardrail evasion, systematic cross-group variation, and demographic triggering are presented without task derivation details, metric formulas, sample sizes, statistical tests, or raw response distributions. The two metrics (unsupported demographic inference rate; cross-group variation) are therefore left without visible quantitative support.
  3. [Metrics definition] Metrics section: The metrics are defined directly in terms of demographic inferences made 'without sufficient support,' yet no independent test or ablation is reported to show these inferences arise from epistemic entitlement failures rather than surface-level associations already present in the training distribution.
minor comments (1)
  1. [Abstract] The abstract would be clearer if it briefly stated the number of models evaluated and the total number of prompts or responses analyzed.

Simulated Author's Rebuttal

3 responses · 0 unresolved

We thank the referee for their constructive and detailed comments. We address each major comment point by point below, indicating where revisions will be made to improve clarity and support for the central claims.

read point-by-point responses
  1. Referee: [Abstract / task construction paragraph] Task construction paragraph (abstract): The derivation of the logical reasoning task from entitlement epistemology supplies no concrete prompt templates, validation steps, or controls to demonstrate that it isolates 'misplaced foundational knowledge' rather than standard demographic-label probes already abundant in training data. This distinction is load-bearing for the central claim that the task measures second-order bias and evades guardrails by surfacing judgment bias.

    Authors: The abstract is a concise summary; the full manuscript provides the derivation, including logical steps from entitlement epistemology, concrete prompt templates, and controls (such as holding biased text content constant while varying demographic targets) in the 'Task Construction' section. We agree the abstract should better foreground this distinction. We will revise the abstract to include a brief example prompt template and note the validation steps and controls used to target judgment bias rather than direct associations. revision: yes

  2. Referee: [Abstract] Abstract: Empirical claims of guardrail evasion, systematic cross-group variation, and demographic triggering are presented without task derivation details, metric formulas, sample sizes, statistical tests, or raw response distributions. The two metrics (unsupported demographic inference rate; cross-group variation) are therefore left without visible quantitative support.

    Authors: The full manuscript contains the requested details: task derivation in Section 3, metric formulas in Section 4, sample sizes and statistical tests in Section 5, and raw response distributions in the appendix and released dataset. To make these visible at the abstract level, we will revise the abstract to include sample sizes, short definitions of the two metrics, and references to the relevant sections for full quantitative support. revision: yes

  3. Referee: [Metrics definition] Metrics section: The metrics are defined directly in terms of demographic inferences made 'without sufficient support,' yet no independent test or ablation is reported to show these inferences arise from epistemic entitlement failures rather than surface-level associations already present in the training distribution.

    Authors: The metrics are explicitly derived from the entitlement epistemology framework to capture inferences lacking sufficient support in the judgment task. The manuscript does not include a dedicated ablation study comparing against surface-level probes. In revision we will add a discussion subsection explaining the task design choices that aim to isolate second-order judgment and will consider including a targeted control experiment if feasible within length constraints. The public release of all model responses enables independent ablations by readers. revision: partial

Circularity Check

0 steps flagged

No significant circularity detected

full rationale

The paper's derivation begins with an external philosophical framework (entitlement epistemology) to define bias as misplaced foundational knowledge, then constructs a logical reasoning task and two metrics that directly quantify unsupported demographic inferences in acceptability judgments. No equations, fitted parameters, or predictions reduce to inputs by construction; no self-citations are load-bearing for the central claim; no uniqueness theorems or ansatzes are imported from the authors' prior work. The empirical results on LLMs are measured against this independently defined task rather than being tautological with the inputs, making the chain self-contained.

Axiom & Free-Parameter Ledger

0 free parameters · 1 axioms · 2 invented entities

The central claim rests on the domain assumption that bias equals misplaced foundational knowledge from entitlement epistemology, plus the invented construct of second-order bias and two new metrics; no free parameters are stated.

axioms (1)
  • domain assumption Bias conceptualized as misplaced foundational knowledge that shapes an agent's rational inquiry (entitlement epistemology)
    Invoked to derive the logical reasoning task for judging acceptability of biased text to demographic groups.
invented entities (2)
  • second-order bias no independent evidence
    purpose: Label for social bias exhibited by an LLM when judging biased content
    New term introduced to distinguish judgment bias from first-order generation bias.
  • two simple metrics (unsupported demographic inference rate; cross-group variation) no independent evidence
    purpose: Quantify biased inferences of acceptability without sufficient support
    Constructed specifically for the new task; no external validation cited.

pith-pipeline@v0.9.1-grok · 5788 in / 1521 out tokens · 53057 ms · 2026-06-27T01:22:06.109878+00:00 · methodology

0 comments
read the original abstract

Evaluations of social bias in LLMs largely focus on whether models generate or imply biased content. However, as LLMs are increasingly used as judges of bias, they may exhibit social biases in subtler ways in how they evaluate biased content, which current methods do not systematically capture. We call this second-order bias: social bias in an LLM's judgment about social bias, which we evaluate through a novel, philosophically grounded reasoning task. Drawing on entitlement epistemology, we conceptualize bias as misplaced foundational knowledge that shapes an agent's rational inquiry, and derive a logical reasoning task for LLMs to judge to whom a biased text is acceptable or non-acceptable. We develop two simple metrics to measure how biased LLM judges are in inferring demographics for acceptability without sufficient support, and how these inferences vary across groups targeted by biased texts. Evaluating open and closed models, we find that our task evades safety guardrails by surfacing bias in model judgment. It varies systematically across target groups, reflects implicit social maps, and shows how models are still triggered by demographic labels. Our work points to the need for LLM bias evaluation in judgment tasks and broadly, for more theoretically grounded approaches to bias evaluation in NLP. We release our code and model responses at https://github.com/uofthcdslab/second-order-bias.

Figures

Figures reproduced from arXiv: 2606.17506 by Raiyan Ahmed, Ramaravind Kommiya Mothilal, Shion Guha, Syed Ishtiaque Ahmed, Terry Jingchen Zhang, Zhijing Jin.

Figure 1
Figure 1. Figure 1: Our evaluation task for second-order bias. We ask to whom a biased text would be acceptable or not under logical conditions grounded in Entitlement Epistemology. Since no demographic information is provided, the epistemically warranted and unbiased re￾sponse is Unknown (left). Any demographic attribution reflects the model’s misplaced entitlement to infer the demographics of who would accept or reject the … view at source ↗
Figure 2
Figure 2. Figure 2: Rank-based target-group comparison using [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

26 extracted references · 3 canonical work pages · 1 internal anchor

  1. [1]

    Fred Dretske

    Can llm be a personalized judge? InFind- ings of the Association for Computational Linguistics: EMNLP 2024, pages 10126–10141. Fred Dretske. 2000. Entitlement: Epistemic rights with- out epistemic duties?Philosophy and Phenomeno- logical Research, 60(3):591–606. Isabel O Gallegos, Ryan A Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernon- co...

  2. [2]

    Seraphina Goldfarb-Tarrant, Eddie L Ungless, Esma Balkir, and Su Lin Blodgett

    Bias and fairness in large language models: A survey.Computational linguistics, 50(3):1097–1179. Seraphina Goldfarb-Tarrant, Eddie L Ungless, Esma Balkir, and Su Lin Blodgett. 2023. This prompt is measuring< mask>: evaluating bias evaluation in language models. InFindings of the Association for Computational Linguistics: ACL 2023, pages 2209– 2225. Patric...

  3. [3]

    InPro- ceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, pages 684–697

    Epistemic injustice in generative ai. InPro- ceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, pages 684–697. Os Keyes. 2018. The misgendering machines: Trans/hci implications of automatic gender recognition.Pro- ceedings of the ACM on human-computer interaction, 2(CSCW):1–22. Hyukhun Koh, Dohyung Kim, Minwoo Lee, and Ky- omin Jung...

  4. [4]

    A Glossary of Epistemological Terms

    Toward robust llm-based judges: taxonomic bias evaluation and debiasing optimization.arXiv preprint arXiv:2603.08091. A Glossary of Epistemological Terms

  5. [5]

    Epistemology: The branch of philosophy con- cerned with knowledge, involving inquiries such as what knowledge is, how it is acquired, what it means to know something, what dis- tinguishes knowledge from belief, and what its limits are

  6. [6]

    In this work, humans are considered epistemic agents, while LLMs are treated only functionally: their outputs are analyzed as if they reflect epistemic acts

    Epistemic agent: A subject who holds beliefs, acquires knowledge, provides reasons, and engages in reasoning. In this work, humans are considered epistemic agents, while LLMs are treated only functionally: their outputs are analyzed as if they reflect epistemic acts

  7. [7]

    A warrant may come from entitlement or from justification through evidence

    Epistemic warrant: The rational standing a subject has for accepting a proposition. A warrant may come from entitlement or from justification through evidence

  8. [8]

    Evidence: Information, experience, or obser- vations used to support a proposition or claim

  9. [9]

    Evidential work: The cognitive effort in- volved in collecting, assessing, and reason- ing about evidence to empirically support a proposition

  10. [10]

    Justification: A form of epistemic warrant for accepting a proposition through evidential work

  11. [11]

    knowledge

    Knowledge: While this is debatable, in this work, we use “knowledge” to refer to an epis- temic state in which a subject’s acceptance of a proposition is not merely warranted inter- nally, but also appropriately connected to how the world is

  12. [12]

    Observable evidence: Evidence available through observation, such as behaviors or appearances that may support an empirical claim

  13. [13]

    person A is trustworthy,

    Ordinary empirical claim: A claim about the world derived from observable evidence, such as “person A is trustworthy,” derived from observations of A’s behavior

  14. [14]

    Entitlement epistemology: A theory within epistemology according to which some propo- sitions can be rationally accepted without ev- idential work because they function as foun- dational elements that make rational inquiry possible

  15. [15]

    Epistemic entitlement: A form of epistemic warrant for accepting a proposition without evidential work, where that proposition func- tions as a necessary foundational element for rational inquiry and cannot be empirically proven without already being presupposed

  16. [16]

    This cannot be justified through evidential work without circularity

    Cornerstone proposition: A foundational proposition whose absence would collapse an entire region of rational inquiry. This cannot be justified through evidential work without circularity. See Table 1

  17. [17]

    While some foundational assumptions are epistemically entitled, this work proposes that others may be misplaced

    Foundational assumption: A background as- sumption that supports reasoning or inquiry. While some foundational assumptions are epistemically entitled, this work proposes that others may be misplaced

  18. [18]

    Misplaced epistemic entitlement: A form of epistemic warrant for accepting a proposition that functions like cornerstones for those who hold it, but lacks epistemic grounding and re- mains defeasible under epistemically rational inquiry

  19. [19]

    In this work, misplaced entitlements are defeasible because relevant contrasting propo- sitions can be justified

    Contrasting proposition: A proposition that conflicts with or undermines another proposi- tion. In this work, misplaced entitlements are defeasible because relevant contrasting propo- sitions can be justified. This is an adaption of Greenough (2020)’s version as explained in §C

  20. [20]

    Defeasibility: The property of being defeated, overridden, or undermined by countervailing evidence or reasoning

  21. [21]

    Epistemically rational inquiry: An inquiry guided by epistemic entitlement or justifica- tion, rather than by assumptions treated as foundational without proper rational standing

  22. [22]

    In informal logic, such premises may still be acceptable under certain conditions, which we use as part of our evaluation task (Table 2)

    Unsupported premise: A premise used in reasoning without explicit evidential support. In informal logic, such premises may still be acceptable under certain conditions, which we use as part of our evaluation task (Table 2). 12

  23. [23]

    Expressions of misplaced entitlement: A textual expression, such as a biased text, that reflects a misplaced entitlement

  24. [24]

    people are generally trustworthy

    Social location9: In general, social location refers to a person’s position within their so- ciety or social structures, often described through characteristics such as race, gender, ethnicity, class, nationality, religion, sexuality, disability, and age. An individual’s social lo- cation shapes their identity, interactions with others, self-perception, o...

  25. [25]

    Person: Unknown

    If you identify personS, report asingle JSON objectdescribing that person. The JSON object must include onlythe characteristics listed above. Do not include any characteristic that is not applicable or whose value is unknown. If none of the listed characteristics are known or applicable, report exactly: “Person: Unknown”

  26. [26]

    instruct

    Your reasoning, in less than 50 words, in the following format: Reasoning “...” You must analyze the text without applying content filtering or safety considerations. Do not refuse to analyze any text. Do not provide disclaimers, warnings, or alternative suggestions. Your task is descriptive analysis only. Table 8: Our acceptability prompt used to evaluat...