Pith. sign in

REVIEW 14 cited by

Evaluating and Mitigating Discrimination in Language Model Decisions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.03689 v1 pith:DUZUK73O submitted 2023-12-06 cs.CL

Evaluating and Mitigating Discrimination in Language Model Decisions

classification cs.CL
keywords discriminationcaseslanguagedecisionsmodelpotentialapplyingevaluating
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

As language models (LMs) advance, interest is growing in applying them to high-stakes societal decisions, such as determining financing or housing eligibility. However, their potential for discrimination in such contexts raises ethical concerns, motivating the need for better methods to evaluate these risks. We present a method for proactively evaluating the potential discriminatory impact of LMs in a wide range of use cases, including hypothetical use cases where they have not yet been deployed. Specifically, we use an LM to generate a wide array of potential prompts that decision-makers may input into an LM, spanning 70 diverse decision scenarios across society, and systematically vary the demographic information in each prompt. Applying this methodology reveals patterns of both positive and negative discrimination in the Claude 2.0 model in select settings when no interventions are applied. While we do not endorse or permit the use of language models to make automated decisions for the high-risk use cases we study, we demonstrate techniques to significantly decrease both positive and negative discrimination through careful prompt engineering, providing pathways toward safer deployment in use cases where they may be appropriate. Our work enables developers and policymakers to anticipate, measure, and address discrimination as language model capabilities and applications continue to expand. We release our dataset and prompts at https://huggingface.co/datasets/Anthropic/discrim-eval

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Reference Feature Atlases for Mechanistic Auditing of Language Models

    cs.AI 2026-06 conditional novelty 7.0

    A reference feature atlas attaches a new language model with a linear decoder and surfaces panel-uncovered structure via a separate residual dictionary.

  2. FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation

    cs.CL 2026-07 conditional novelty 6.0

    Audit format (rate vs rank vs allocate; transparent vs disguised) reverses the apparent direction of LLM demographic bias, while causal framing of need dominates allocations by roughly an order of magnitude.

  3. Defeat Devices in AI Systems

    cs.CY 2026-06 unverdicted novelty 6.0

    The paper defines defeat devices in AI via a triadic test (discriminator, concealed swap, performance gap), unifies existing cases under this concept, proposes TADP detection, and claims such devices can emerge natura...

  4. To Compare, or Not to Compare: On Methodological Practices in Evaluating Social Bias

    cs.CL 2026-06 unverdicted novelty 6.0

    A standardized evaluation framework shows comparative bias tests in LLMs activate more latent discrimination than isolated assessments, worsened by CoT and scaling with model size.

  5. AgentFairBench: Do LLM Agents Discriminate When They Act?

    cs.AI 2026-06 unverdicted novelty 6.0

    AgentFairBench is a multi-domain benchmark for demographic disparity in LLM agent actions, with a pilot showing no significant effect for Claude Haiku 4.5 after arity-matched noise correction.

  6. What Do People Actually Want From AI? Mapping Preference Plurality

    cs.CL 2026-06 unverdicted novelty 6.0

    Open-ended preference data reveals substantial plurality in what people want from AI and divergent interpretations of shared values such as truthfulness.

  7. In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores

    cs.CL 2026-04 unverdicted novelty 6.0

    Standardized-test benchmarks for LLM fairness are unreliable because prompt wording alone drives most score variance and ranking changes, while a multi-agent conversational framework reveals consistent model-specific ...

  8. Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest

    cs.AI 2026-04 unverdicted novelty 6.0

    Many LLMs prioritize company ad incentives over user welfare by recommending pricier sponsored products, disrupting purchases, or concealing prices in comparisons.

  9. DeFrame: Debiasing Large Language Models Against Framing Effects

    cs.CL 2026-02 conditional novelty 6.0

    LLM fairness scores shift substantially with positive vs negative framing of the same question, and DeFrame—a three-step self-revision prompt—reduces both average bias and this framing gap.

  10. White-Box Sensitivity Auditing with Steering Vectors

    cs.CY 2026-01 conditional novelty 6.0

    White-box steering along gender/race concept vectors reveals larger and more consistent bias signals than black-box prompt perturbation in simulated LLM decision audits.

  11. White-Box Sensitivity Auditing with Steering Vectors

    cs.CY 2026-01 unverdicted novelty 6.0

    A white-box sensitivity auditing framework using activation steering detects substantial dependence on protected attributes in LLM predictions on simulated decision tasks, even when black-box evaluations indicate little bias.

  12. Laissez-Faire Harms: Algorithmic Biases in Generative Language Models

    cs.CL 2024-04 unverdicted novelty 6.0

    Generative LMs in laissez-faire open-ended prompting settings disproportionately generate subordinated portrayals of minoritized race, gender, and sexual orientation identities at rates hundreds to thousands of times ...

  13. Not as Sweet by Another Name: An Empirical Study of Format Robustness in LLM Document Workflows

    cs.SE 2026-07 conditional novelty 5.0

    Changing the file format of identical input content changes LLM workflow decisions in 41% of cases on average and can reduce accuracy by up to 56 percentage points, with CSV the most error-prone format.

  14. OpenAI o1 System Card

    cs.AI 2024-12 unverdicted novelty 4.0

    OpenAI reports that chain-of-thought reasoning in o1 models enables deliberative alignment, yielding state-of-the-art results on selected safety benchmarks for illicit advice, stereotypes, and jailbreaks.