Pith. sign in

REVIEW 1 cited by

BEATS: Bias Evaluation and Assessment Test Suite for Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.24310 v1 pith:DBQZWB7H submitted 2025-03-31 cs.CL cs.AI

classification cs.CLcs.AI
keywords beatsmodelsbenchmarkbiasframeworkllmsmetricsassessment
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this research, we introduce BEATS, a novel framework for evaluating Bias, Ethics, Fairness, and Factuality in Large Language Models (LLMs). Building upon the BEATS framework, we present a bias benchmark for LLMs that measure performance across 29 distinct metrics. These metrics span a broad range of characteristics, including demographic, cognitive, and social biases, as well as measures of ethical reasoning, group fairness, and factuality related misinformation risk. These metrics enable a quantitative assessment of the extent to which LLM generated responses may perpetuate societal prejudices that reinforce or expand systemic inequities. To achieve a high score on this benchmark a LLM must show very equitable behavior in their responses, making it a rigorous standard for responsible AI evaluation. Empirical results based on data from our experiment show that, 37.65\% of outputs generated by industry leading models contained some form of bias, highlighting a substantial risk of using these models in critical decision making systems. BEATS framework and benchmark offer a scalable and statistically rigorous methodology to benchmark LLMs, diagnose factors driving biases, and develop mitigation strategies. With the BEATS framework, our goal is to help the development of more socially responsible and ethically aligned AI models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Data and AI governance: Promoting equity, ethics, and fairness in large language models

    cs.CL 2025-08 unverdicted novelty 4.0 of 10

    The paper proposes a lifecycle governance framework, built on the authors' BEATS benchmark, to quantify and mitigate bias, ethics, fairness, and factuality failures in large language models.

Pith tools