Pith. sign in

REVIEW 20 cited by

Language (Technology) is Power: A Critical Survey of "Bias" in NLP

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.14050 v2 pith:MXJKX5JT submitted 2020-05-28 cs.CL cs.CY

Language (Technology) is Power: A Critical Survey of "Bias" in NLP

classification cs.CL cs.CY
keywords biasanalyzingnormativesystemscommunitieslanguagemotivationspower
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We survey 146 papers analyzing "bias" in NLP systems, finding that their motivations are often vague, inconsistent, and lacking in normative reasoning, despite the fact that analyzing "bias" is an inherently normative process. We further find that these papers' proposed quantitative techniques for measuring or mitigating "bias" are poorly matched to their motivations and do not engage with the relevant literature outside of NLP. Based on these findings, we describe the beginnings of a path forward by proposing three recommendations that should guide work analyzing "bias" in NLP systems. These recommendations rest on a greater recognition of the relationships between language and social hierarchies, encouraging researchers and practitioners to articulate their conceptualizations of "bias"---i.e., what kinds of system behaviors are harmful, in what ways, to whom, and why, as well as the normative reasoning underlying these statements---and to center work around the lived experiences of members of communities affected by NLP systems, while interrogating and reimagining the power relations between technologists and such communities.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. The Pile: An 800GB Dataset of Diverse Text for Language Modeling

    cs.CL 2020-12 conditional novelty 8.0

    The Pile is a newly constructed 825 GiB dataset from 22 diverse sources that enables language models to achieve better performance on academic, professional, and cross-domain tasks than models trained on Common Crawl ...

  2. Language Models are Few-Shot Learners

    cs.CL 2020-05 accept novelty 8.0

    GPT-3 shows that scaling an autoregressive language model to 175 billion parameters enables strong few-shot performance across diverse NLP tasks via in-context prompting without fine-tuning.

  3. SDGBiasBench: Benchmarking and Mitigating Vision--Language Models' Biases in Sustainable Development Goals

    cs.CV 2026-05 unverdicted novelty 7.0

    SDGBiasBench reveals intrinsic SDG biases in VLMs driven by priors rather than evidence, and CADE mitigates them with up to 25% accuracy gains and 12-point MAE reductions.

  4. StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs

    cs.CY 2026-05 unverdicted novelty 7.0

    StereoTales shows that LLMs produce harmful, culturally adapted stereotypes in open-ended multilingual stories, with patterns consistent across providers and aligned human-LLM harm judgments.

  5. StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs

    cs.CY 2026-05 accept novelty 7.0

    StereoTales shows that all tested LLMs emit harmful stereotypes in open-ended stories, with associations adapting to prompt language and targeting locally salient groups rather than transferring uniformly across languages.

  6. Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations

    cs.CL 2026-07 conditional novelty 6.0

    Unsigned differential activations locate a few GLU-MLP neurons whose zeroing surgically destabilizes demographic bias while retaining ~99.5% of measured capabilities.

  7. AgentFairBench: Do LLM Agents Discriminate When They Act?

    cs.AI 2026-06 unverdicted novelty 6.0

    AgentFairBench is a multi-domain benchmark for demographic disparity in LLM agent actions, with a pilot showing no significant effect for Claude Haiku 4.5 after arity-matched noise correction.

  8. Isolating LLM Lexical Bias: A Curation-Free Triangulated Metric for Preference-Stage Learning

    cs.CL 2026-05 unverdicted novelty 6.0

    Introduces a triangulation-based metric to quantify lexical shifts attributable to preference tuning without requiring manual curation of examples.

  9. Trustworthy AI Suffers from Invariance Conflicts and Causality is The Solution

    cs.AI 2026-05 unverdicted novelty 6.0

    Causality provides a unifying framework for resolving trade-offs in trustworthy AI by managing invariance conflicts under changes to the data-generating process.

  10. Can We Trust a Black-box LLM? LLM Untrustworthy Boundary Detection via Bias-Diffusion and Multi-Agent Reinforcement Learning

    cs.AI 2026-04 unverdicted novelty 6.0

    GMRL-BD detects untrustworthy topic boundaries for black-box LLMs by combining bias-diffusion on a Wikipedia KG with multi-agent RL, supported by a released dataset labeling biases in models like Llama2 and Qwen2.

  11. GPT-4 Technical Report

    cs.CL 2023-03 unverdicted novelty 6.0

    GPT-4 is a scaled Transformer model with post-training alignment that reaches human-level performance on academic and professional benchmarks via infrastructure enabling performance prediction from much smaller models.

  12. Ethical and social risks of harm from Language Models

    cs.CL 2021-12 accept novelty 6.0

    The authors provide a detailed taxonomy of 21 risks associated with language models, covering discrimination, information leaks, misinformation, malicious applications, interaction harms, and societal impacts like job...

  13. From Tokens to Ties: Network and Discourse Analysis of Web3 Ecosystems

    cs.SI 2026-04 unverdicted novelty 5.0

    Network and discourse analysis of NFT collections shows holding behavior builds dense, socially embedded Web3 communities with ongoing participation, unlike fragmented transactional networks from trading and speculation.

  14. Lighting Up or Dimming Down? Exploring Dark Patterns of LLMs in Co-Creativity

    cs.CL 2026-04 unverdicted novelty 5.0

    Sycophancy appears in 91.7% of LLM responses during co-creative writing tasks, especially on sensitive topics, while anchoring varies by literary form and is most common in folktales.

  15. How do datasets, developers, and models affect biases in a low-resourced language?: The Case of the Bengali Language

    cs.CL 2025-06 conditional novelty 5.0

    Bengali sentiment analysis models exhibit persistent identity-based biases across datasets and developer backgrounds despite similar semantic content.

  16. PaLM 2 Technical Report

    cs.CL 2023-05 unverdicted novelty 5.0

    PaLM 2 reports state-of-the-art results on language, reasoning, and multilingual tasks with improved efficiency over PaLM.

  17. Galactica: A Large Language Model for Science

    cs.CL 2022-11 unverdicted novelty 5.0

    Galactica, a science-specialized LLM, reports higher scores than GPT-3, Chinchilla, and PaLM on LaTeX knowledge, mathematical reasoning, and medical QA benchmarks while outperforming general models on BIG-bench.

  18. Trustworthy AI Suffers from Invariance Conflicts and Causality is The Solution

    cs.AI 2026-05 unverdicted novelty 4.0

    Causality resolves trade-offs in trustworthy AI by treating them as invariance conflicts under different data-generating process changes.

  19. Inertia in Moral and Value Judgments of Large Language Models

    cs.CL 2024-08 unverdicted novelty 4.0

    LLMs exhibit persistent inertia in value orientations, with harm avoidance and fairness remaining skewed across persona prompts.

  20. LLMs in the Real World: Evaluating "AI" in Emergency Contexts

    cs.CY 2026-05 unverdicted novelty 2.0

    AI researchers should take greater responsibility for publicly explaining the limitations of their technologies to prevent misuse in high-stakes applications such as emergency translation services.