Pith. sign in

REVIEW 4 cited by

BiasAlert: A Plug-and-play Tool for Social Bias Detection in LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.10241 v2 pith:OWR4L2QQ submitted 2024-07-14 cs.CL

classification cs.CL
keywords biasbiasalertllmsdemonstratedetectevaluationexistingmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Evaluating the bias in Large Language Models (LLMs) becomes increasingly crucial with their rapid development. However, existing evaluation methods rely on fixed-form outputs and cannot adapt to the flexible open-text generation scenarios of LLMs (e.g., sentence completion and question answering). To address this, we introduce BiasAlert, a plug-and-play tool designed to detect social bias in open-text generations of LLMs. BiasAlert integrates external human knowledge with inherent reasoning capabilities to detect bias reliably. Extensive experiments demonstrate that BiasAlert significantly outperforms existing state-of-the-art methods like GPT4-as-A-Judge in detecting bias. Furthermore, through application studies, we demonstrate the utility of BiasAlert in reliable LLM bias evaluation and bias mitigation across various scenarios. Model and code will be publicly released.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BiasFilter: An Inference-Time Debiasing Framework for Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    BiasFilter filters low-fairness segments during LLM generation using a reward model trained on a GPT-4-scored preference dataset, cutting bias on CEB and FairMT.

  2. Large Language Models in Architecture Studio: A Framework for Learning Outcomes

    cs.CY 2025-10 conditional novelty 5.0 of 10

    A conceptual framework maps LLM-based interventions onto architecture studio challenges and Bloom's taxonomy across self-, peer-, and teacher-led learning.

  3. Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A supervised fine-tuning plus difficulty-filtered reinforcement learning recipe improves video temporal grounding on three benchmarks, with datasets and models released.

  4. Detection, Classification, and Mitigation of Gender Bias in Large Language Models

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A Chinese gender-bias system using SFT, chain-of-thought, and DPO with GPT-4-generated preference pairs reports top validation scores and first place on all three NLPCC 2025 subtasks.

Pith tools