Pith. sign in

REVIEW 1 cited by

Pre-Trained Language Models Represent Some Geographic Populations Better Than Others

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.11025 v1 pith:7JN6CGG5 submitted 2024-03-16 cs.CL

classification cs.CL
keywords populationsmodelsgeographicrepresentskewfactorspre-trainedacross
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This paper measures the skew in how well two families of LLMs represent diverse geographic populations. A spatial probing task is used with geo-referenced corpora to measure the degree to which pre-trained language models from the OPT and BLOOM series represent diverse populations around the world. Results show that these models perform much better for some populations than others. In particular, populations across the US and the UK are represented quite well while those in South and Southeast Asia are poorly represented. Analysis shows that both families of models largely share the same skew across populations. At the same time, this skew cannot be fully explained by sociolinguistic factors, economic factors, or geographic factors. The basic conclusion from this analysis is that pre-trained models do not equally represent the world's population: there is a strong skew towards specific geographic populations. This finding challenges the idea that a single model can be used for all populations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Performance Gains of LLMs With Humans in a World of LLMs Versus Humans

    cs.HC 2025-05 conditional novelty 4.0 of 10

    A commentary argues that medical research should stop running evanescent LLM-versus-human comparisons and instead study human-LLM collaboration, supported by a literature review showing rapid model turnover and tiny e...

Pith tools