Pith. sign in

REVIEW 1 cited by

Evaluating Large Language Models through Gender and Racial Stereotypes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.14788 v1 pith:GPKHQEFI submitted 2023-11-24 cs.CL cs.AIcs.CY

classification cs.CLcs.AIcs.CY
keywords modelsgenderlanguagebiasbiasesracialabilityamongst
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Language Models have ushered a new age of AI gaining traction within the NLP community as well as amongst the general population. AI's ability to make predictions, generations and its applications in sensitive decision-making scenarios, makes it even more important to study these models for possible biases that may exist and that can be exaggerated. We conduct a quality comparative study and establish a framework to evaluate language models under the premise of two kinds of biases: gender and race, in a professional setting. We find out that while gender bias has reduced immensely in newer models, as compared to older ones, racial bias still exists.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GenderBench: Evaluation Suite for Gender Biases in LLMs

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A 14-probe benchmark on 12 LLMs finds consistent gender stereotype reasoning and unbalanced character representation across models.

Pith tools