Pith. sign in

REVIEW 17 cited by

Large Language Models are Geographically Biased

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.02680 v2 pith:YU3YSHEM submitted 2024-02-05 cs.CL cs.AIcs.CYcs.LG

classification cs.CLcs.AIcs.CYcs.LG
keywords llmsbiaseslanguagemodelsacrossbiasbiasedgeographic
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Large Language Models (LLMs) inherently carry the biases contained in their training corpora, which can lead to the perpetuation of societal harm. As the impact of these foundation models grows, understanding and evaluating their biases becomes crucial to achieving fairness and accuracy. We propose to study what LLMs know about the world we live in through the lens of geography. This approach is particularly powerful as there is ground truth for the numerous aspects of human life that are meaningfully projected onto geographic space such as culture, race, language, politics, and religion. We show various problematic geographic biases, which we define as systemic errors in geospatial predictions. Initially, we demonstrate that LLMs are capable of making accurate zero-shot geospatial predictions in the form of ratings that show strong monotonic correlation with ground truth (Spearman's $\rho$ of up to 0.89). We then show that LLMs exhibit common biases across a range of objective and subjective topics. In particular, LLMs are clearly biased against locations with lower socioeconomic conditions (e.g. most of Africa) on a variety of sensitive subjective topics such as attractiveness, morality, and intelligence (Spearman's $\rho$ of up to 0.70). Finally, we introduce a bias score to quantify this and find that there is significant variation in the magnitude of bias across existing LLMs. Code is available on the project website: https://rohinmanvi.github.io/GeoLLM

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mapping the City Through the Lens of Language Models

    cs.CL 2026-08 conditional novelty 7.0 of 10

    Ten open-weight language models, rating anonymized 40-indicator city profiles, systematically favor larger, faster-growing, infrastructure-rich, less sparse urban forms.

  2. Representational Equality in Cross-country Value Simulation: A Systematic Analysis of Large Language Models

    cs.CY 2026-08 conditional novelty 6.0 of 10

    LLM-based value simulation is systematically more accurate for wealthy, high-governance, individualist countries, and common interventions rarely fix the imbalance.

  3. Benchmarking Open-Weight Foundation Models for Global AI Technical Governance

    cs.CY 2026-04 conditional novelty 6.0 of 10

    Open-weight frontier models fabricate ~72% of AI-governance numeric answers, almost never refuse, and show inverted North/South accuracy driven largely by a proportional ±10% scoring rule and sparse high-value indicators.

  4. Generating the Modal Worker: A Cross-Model Audit of Race and Gender in LLM-Generated Personas Across 41 Occupations

    cs.HC 2025-10 conditional novelty 6.0 of 10

    Four major LLMs generate occupational personas whose race and gender distributions deviate from U.S. workforce data in shared, patterned ways: White and Black workers are underrepresented while Hispanic and Asian work...

  5. CAMS: A CityGPT-Powered Agentic Framework for Urban Human Mobility Simulation

    cs.CL 2025-06 conditional novelty 6.0 of 10

    CAMS, an agentic framework built on CityGPT, generates urban mobility trajectories from user profiles and claims superior JSD-based performance over prior simulators.

  6. Around the World in 24 Hours: Probing LLM Knowledge of Time and Place

    cs.CL 2025-06 conditional novelty 6.0 of 10

    On GeoTemp, the best open model answers only 56% of two-city time questions and 33% when an hour shift is added, despite near-perfect scores on pure time arithmetic.

  7. Personalisation or Prejudice? Addressing Geographic Bias in Hate Speech Detection using Debias Tuning in Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Country and language personas degrade LLM hate speech detection F1 scores, and a custom reweighted fine-tuning loss reduces the degradation for Llama and Nemo, but less for Phi.

  8. Breaking Down Bias: On The Limits of Generalizable Pruning Strategies

    cs.CL 2025-02 conditional novelty 6.0 of 10

    Pruning-based bias removal in Llama-3-8B reduces racial bias mainly in the context used to choose what to prune, and transfers poorly across contexts.

  9. ValuesRAG: Enhancing Cultural Alignment Through Retrieval-Augmented Contextual Learning

    cs.CL 2025-01 conditional novelty 6.0 of 10

    ValuesRAG retrieves and reranks value summaries of demographically similar WVS respondents and uses them as in-context evidence, beating four baselines on six regional survey datasets.

  10. On the missing data layer and a potential solution

    cs.AI 2026-08 unverdicted novelty 5.0 of 10

    A position paper proposing DataHub, a task/domain/language organized catalog to close Latin America's AI data discovery and supply gap.

  11. Mitigating Geospatial Knowledge Hallucination in Large Language Models: Benchmarking and Dynamic Factuality Aligning

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A new benchmark called GEOHALUBENCH measures how often LLMs invent, omit, or confuse real-world places and relations, and a dynamic-beta KTO method reduces these errors on the benchmark.

  12. The Democratic Paradox in Large Language Models' Underestimation of Press Freedom

    cs.CY 2025-06 conditional novelty 5.0 of 10

    Six LLMs systematically underrate press freedom in 180 countries, penalize freer countries most, and five give their home countries favorable treatment.

  13. TravelAgent: Generative Agents in the Built Environment

    cs.AI 2024-12 conditional novelty 5.0 of 10

    A new LLM-driven agent platform navigates virtual urban environments with multimodal sensory inputs, achieving a 76% self-reported task completion rate across 100 simulations.

  14. On the missing benchmarks layer and a potential solution

    cs.AI 2026-08 unverdicted novelty 4.0 of 10

    An open regional EvalsHub, starting with LatamBoard, is proposed as the answer to Latin America's missing benchmark layer, but no benchmark or evaluation is presented.

  15. The Rise of AI in Weather and Climate Information and its Impact on Global Inequality

    physics.ao-ph 2026-03 conditional novelty 4.0 of 10

    AI weather and climate tools inherit Northern-controlled data and compute, risking worse forecasts and maladaptation for the Global South rather than democratizing climate information.

  16. Musical ethnocentrism in Large Language Models

    cs.CL 2025-01 conditional novelty 4.0 of 10

    LLMs like ChatGPT and Mixtral show a strong Western bias when asked to name top musical contributors and to rate the musical cultures of countries.

  17. Fantastic Biases (What are They) and Where to Find Them

    cs.CL 2024-11 conditional novelty 3.0 of 10

    A survey that defines bias broadly, catalogs commonly discussed AI and NLP biases, and reviews methods to detect and mitigate them.

Pith tools