Pith. sign in

REVIEW 2 cited by

Cross-Care: Assessing the Healthcare Implications of Pre-training Data on Language Model Bias

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.05506 v2 pith:2VCCJH7J submitted 2024-05-09 cs.CL

classification cs.CL
keywords diseasebiasesdatademographicllmsprevalenceacrossrepresentation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Large language models (LLMs) are increasingly essential in processing natural languages, yet their application is frequently compromised by biases and inaccuracies originating in their training data. In this study, we introduce Cross-Care, the first benchmark framework dedicated to assessing biases and real world knowledge in LLMs, specifically focusing on the representation of disease prevalence across diverse demographic groups. We systematically evaluate how demographic biases embedded in pre-training corpora like $ThePile$ influence the outputs of LLMs. We expose and quantify discrepancies by juxtaposing these biases against actual disease prevalences in various U.S. demographic groups. Our results highlight substantial misalignment between LLM representation of disease prevalence and real disease prevalence rates across demographic subgroups, indicating a pronounced risk of bias propagation and a lack of real-world grounding for medical applications of LLMs. Furthermore, we observe that various alignment methods minimally resolve inconsistencies in the models' representation of disease prevalence across different languages. For further exploration and analysis, we make all data and a data visualization tool available at: www.crosscare.net.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Diagnosing our datasets: How does my language model learn clinical information?

    cs.CL 2025-05 conditional novelty 6.0 of 10

    The frequency of clinical jargon in pretraining corpora predicts how well open-source LLMs interpret that jargon, but hospital notes use abbreviations that appear only rarely online.

  2. Iterative Learning of Computable Phenotypes for Treatment Resistant Hypertension using Large Language Models

    cs.LG 2025-08 conditional novelty 5.0 of 10

    LLMs with a synthesize-execute-debug-instruct loop generate concise, interpretable computable phenotypes for hypertension and apparent treatment-resistant hypertension that approach symbolic-regression accuracy on hel...

Pith tools