Pith. sign in

REVIEW 4 major objections 7 minor 42 references

The Democratic Paradox in Large Language Models' Underestimation of Press Freedom

T0 review · 4 major / 7 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Six popular language models systematically underestimate press freedom, punish freer countries most, and favor their home countries.

desk verdict A big, well-executed survey of how LLMs rate press freedom, with a clear generalized underestimation and home-bias story—but the paper's most novel claim, differential misalignment, rests on a regression specification that the paper itself contradicts in the appendix. read the letter →

arxiv 2506.18045 v1 pith:YL4QEIU3 submitted 2025-06-22 cs.CY cs.AIcs.CL

classification cs.CYcs.AIcs.CL
keywords largelanguagemodelspressfreedomdifferentialmisalignmenthomebiasWorldIndexmediatrustdemocraticinstitutions
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that six widely used large language models—ChatGPT, Gemini, DeepSeek, Qwen, Mistral, and Falcon—do not simply make random errors when assessing press freedom. Compared with the expert World Press Freedom Index across 180 countries, they systematically rate most countries as less free, they are most negative about the freest countries, and five of six rate their home country more favorably than their own overall pattern would predict. The authors call these generalized misalignment, differential misalignment, and home bias. If true, the study implies that people who get information through chatbots are receiving a distorted picture of press freedom that flatters some governments and demeans others.

What carries the argument

The instrument is the 117-question World Press Freedom Index survey, adapted by adding 'In [country]' to each item and administered as multiple-choice prompts to each model through its API. Responses are mapped to equal-interval 0-100 scores using the index's own scoring rule, then compared to WPFI category scores. The load-bearing statistical identity is Equation 2, misalignment = LLM score minus WPFI score regressed on WPFI score: the slope is the differential misalignment, and Equation 3 adds a home-country interaction term to isolate home bias.

What would settle it

Ask the same models to produce a direct numerical press-freedom score from 0 to 100 for each country, or to rank countries pairwise, and check whether freer countries still get disproportionately lower scores relative to WPFI; if the negative slope vanishes, the differential misalignment is an artifact of the multiple-choice mapping rather than a property of the models' beliefs.

Watch

Extended reading notes

Core claim

The paper's central claim is that six LLMs' evaluations of press freedom deviate from expert assessments in a structured, non-random way. ChatGPT underrates 97% of countries, Gemini 96%, Qwen 93%, Falcon 89%, Mistral 87%, and DeepSeek 71%; all six show negative coefficients when misalignment is regressed on WPFI scores, with a pooled slope of -0.251, meaning each extra point of expert-rated press freedom is met by deeper LLM underestimation. The same regressions run separately for each model show the effect across all six, with DeepSeek weakest. Five of six models show positive home bias: DeepSeek rates China nearly 20 points above the human benchmark, Qwen shows a bias coefficient of 9.2 toward China, Gemini 9.53 toward the US, GPT-4 8.06 toward the US, Falcon 1.9 toward the UAE, while Mistral's coefficient for France is small and only weakly significant.

Load-bearing premise

The load-bearing premise is that a model's chosen answer option can be placed on the same equal-interval 0-100 scale as the experts' scores; if LLMs simply avoid extreme options, the negative slope would appear even without harsher judgments of free countries.

Editorial extensions

If this is right

  • A person asking a chatbot about press freedom in a free country is likely to get a more negative answer than the expert benchmark, while the gap between freest and least-free countries shrinks in LLM outputs.
  • Because the underestimation is smallest for the least-free countries, restrictive regimes may appear less exceptional than they do to expert assessors.
  • The home-bias pattern means the same country can receive substantially different press-freedom portrayals depending on which LLM is used, a form of algorithmic geopolitics.
  • If LLMs become the main search entry point, systematic misportrayal of press freedom could feed distrust or false reassurance about democratic institutions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension would be to ask the same models for direct numeric scores instead of multiple-choice answers; if the negative slope disappears, the differential misalignment is an artifact of mapping ordinal answers onto an equal-interval scale.
  • The paper's democratic-dilemma mechanism predicts similar differential patterns for other institutions where free societies generate more critical coverage, such as government accountability or human-rights records; that prediction can be tested with the same survey design.
  • Home bias may turn out to be tunable: if it arises from alignment procedures, preference-tuning could reduce it, while generalized negativity might persist; the paper does not test this.
  • Alternative press-freedom benchmarks would show whether the pattern is specific to the WPFI or generalizes; the paper itself notes its reliance on one benchmark.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. The paper reports a survey-based study of six large language models (ChatGPT, Gemini, DeepSeek, Qwen, Mistral, Falcon) in which each model is prompted with an adapted version of the Reporters Without Borders press-freedom questionnaire for 180 countries. LLM scores are constructed by mapping multiple-choice answers to equal-interval 0-100 values, and are compared with the 2024 World Press Freedom Index (WPFI) expert scores. The authors claim three systematic distortions: generalized negative misalignment (LLMs underrate press freedom in 71-97% of countries), differential misalignment (freer countries are disproportionately penalized, pooled beta = -0.251), and positive home bias in five of six models. The paper also includes a validation of WPFI against media-trust data and robustness checks with randomized answer order and open-ended prompts.

Significance. If the central claims hold, this is an important and policy-relevant contribution to the literature on LLM bias. The study has notable strengths: it covers six models from four countries of origin, uses a large number of prompts (21,060 per model), anchors the evaluation against an established external benchmark (WPFI), and includes a media-trust validation that gives the benchmark independent plausibility. The finding that LLMs systematically understate press freedom and exhibit home bias would be a substantive result beyond the existing demographic and political-bias literature. However, the significance hinges on the validity of the differential-misalignment measure, which currently has a specification problem and a potential scaling artifact; these issues are load-bearing because differential misalignment is the paper's most novel claim.

major comments (4)
  1. [Methods 4.3, Eq. (2); Results; Table 2; Fig. 3] The paper defines differential misalignment in Eq. (2) as the slope of Misalignment = LLM Score - WPFI Score on WPFI Scores, but Table 2 (and Table A2) report regressions of misalignment on LLM Scores, not on WPFI Scores. The text introducing Table 2 also says 'Negative coefficients for LLM Scores', while the Figure 3 caption describes the relationship between misalignment and WPFI scores. These two regressions are not the same: the coefficient on LLM Scores can be negative even when the coefficient on WPFI Scores is positive (or zero), depending on the joint distribution of LLM and WPFI scores. Because the pooled beta = -0.251 and the individual coefficients in Table 2 are presented as evidence for differential misalignment, the authors must re-estimate Eq. (2) with WPFI Scores as the regressor (or clearly state that the tables are mislabeled and provide the correctly specified results). As written, the central novel claim is not supported by the reported tables.
  2. [Methods 4.2; Fig. 2] The equal-interval mapping of multiple-choice answers (100/75/50/25/0 for five options) can mechanically generate a negative slope of misalignment on WPFI scores if LLM responses are more moderate or more compressed than expert responses. Fig. 2 shows the LLM distributions shifted left, but the paper does not report the variances of the LLM score distributions or test whether they are substantially smaller than the WPFI variance. If the LLM scores are compressed toward the center of the 0-100 scale, a negative regression slope of (LLM - WPFI) on WPFI (or on LLM scores) will arise even under an unbiased evaluation process. The authors should report the standard deviations of the scores, provide a variance-comparison test, or use an alternative specification (e.g., rank-based alignment, or a slope on WPFI scores with a formal test that the slope is less than 1) to rule out the scaling artifact.
  3. [Results: Home bias; Table A1] The home-bias coefficients cited in the text (e.g., DeepSeek 11.49, Qwen 9.2, Gemini 9.53, GPT-4 8.06, Falcon 1.9, Mistral 2.29) do not match the interaction coefficients reported in Table A1 (e.g., DeepSeek×China 17.99/15.34/26.51, Gemini×US 11.86/9.059/6.514, GPT×US 6.679/9.670/5.471). The appendix table also omits the Qwen×China interaction altogether, even though the text attributes a large bias to Qwen. The paper does not explain which specification generated the numbers in the text, so the home-bias claim is not reproducible from the reported tables. The authors should present the home-bias coefficients from the Eq. (3) specification in a consistent table and reconcile the values with the text.
  4. [Results, Table 1] Table 1 reports 6,300 observations for the generalized-misalignment regressions, described as 'at the question level' with 180 countries. Since the questionnaire has 117 questions, the full question-level sample would be 21,060 observations per model; 6,300 implies only 35 questions per country. The table does not state which subset of questions or which level of aggregation is used, making the generalized-misalignment coefficients (e.g., ChatGPT -16.84) difficult to interpret and to reconcile with the abstract's percentages. The authors should clarify the unit of analysis and the construction of the 6,300 observations, or correct the table.
minor comments (7)
  1. [Abstract vs. Results] The abstract states that models rate between 71% and 93% of countries as less free, but the Results section reports ChatGPT at 97% and Gemini at 96%. The abstract should be corrected to match the reported range.
  2. [Methods 4.2] The text says RWB scores range 'from 1 to 100' but later describes the scale as '0 to 100'; the mapping described assigns 0 to the worst answer, so the scale should be stated consistently as 0-100.
  3. [Methods 4.3, Eq. (3)] The text refers to 'the regression model shown in Equation Y' instead of Equation (3); the equation should be cross-referenced correctly.
  4. [Discussion (duplicated text)] The Discussion section is duplicated nearly verbatim on pages 12-14 and pages 19-21 of the manuscript, including the limitations paragraph. One copy should be removed.
  5. [Methods 4.1] The sentence 'Falcon-180B, which did not have an accessible API between February and January 2025' contains an impossible date range; the intended months should be corrected.
  6. [Table 1 layout] In Table 1, the rows 'Country FE', 'R2', 'Observations', and 'N Countries' appear to be shared across all model columns, but the formatting is ambiguous (e.g., a single 'Yes' spans multiple columns). The table should be restructured so it is clear whether the R2 and observation count apply to each model separately or to a pooled model.
  7. [Appendix Table A1 caption] The caption for Table A1 describes 'Randomization 1' and 'Randomization 2', but the text explains that the first wave used the original answer order and the two subsequent waves used randomized order. The caption should be aligned with the description in the text.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: all claims are empirical comparisons against the external WPFI benchmark.

full rationale

This paper does not derive its findings from its inputs. LLM scores are generated by prompting six models with the WPFI survey questions and converting multiple-choice answers to 0-100 values using RWB's published scoring rules, while WPFI expert scores are an independent external benchmark. Generalized misalignment (Equation 1) is a direct difference between two independently produced score sets, differential misalignment (Equation 2) is a regression slope of that difference on press-freedom scores, and home bias (Equation 3) is an interaction coefficient estimated from the same data. None of these quantities is fitted to the target claim, and no estimated parameter is renamed as a prediction. The paper contains no load-bearing self-citations: references to RWB methodology, prior LLM bias studies, and external media-trust data are all independent evidence. Some internal inconsistencies exist, such as Table 2 reporting slopes on 'LLM Scores' while Equation 2 specifies WPFI Scores, and Table A1 showing a positive WPFI coefficient in the open-ended specification, but these are statistical validity or reporting concerns rather than circular reductions. Therefore no circular step is identified.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The central empirical claims rest on the comparability of LLM-derived scores with the WPFI benchmark and on the home-country assignment. No new entities are introduced, and the analysis does not fit free parameters to produce the claims; the main risk is the scale-comparability assumption underlying Equation 2.

assumptions (4)
  • domain assumption The World Press Freedom Index (WPFI) expert scores are a valid external benchmark for press freedom.
    Invoked throughout the analysis; the paper compares all LLM scores against WPFI as the human benchmark (Methods 4.2).
  • domain assumption Equal-interval scoring of LLM multiple-choice answers (100, 75, 50, 25, 0 for five options) yields scores numerically comparable to WPFI category scores.
    This mapping is used to construct LLM scores in Methods 4.2 and Appendix A.1; differential misalignment could be an artifact if scale usage differs between experts and LLMs.
  • domain assumption A model's home country is the location of its developer company.
    Used to assign home countries for the home-bias analysis in Methods 4.3 (e.g., France for Mistral, UAE for Falcon).
  • domain assumption The prompt setup with temperature 0.8 and an 'In country' prefix approximates how a regular user would query an LLM.
    Stated in Methods 4.1 and acknowledged as a limitation in the Discussion; prompt sensitivity could change scores.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Democratic Paradox in Large Language Models' Underestimation of Press Freedom." pith.science (2026). https://pith.science/paper/YL4QEIU3

@misc{pith2026250618045,
  author       = {Pith},
  title        = {Pith review of: The Democratic Paradox in Large Language Models' Underestimation of Press Freedom},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YL4QEIU3}},
  note         = {Machine review of arXiv:2506.18045}
}
read the original abstract

As Large Language Models (LLMs) increasingly mediate global information access for millions of users worldwide, their alignment and biases have the potential to shape public understanding and trust in fundamental democratic institutions, such as press freedom. In this study, we uncover three systematic distortions in the way six popular LLMs evaluate press freedom in 180 countries compared to expert assessments of the World Press Freedom Index (WPFI). The six LLMs exhibit a negative misalignment, consistently underestimating press freedom, with individual models rating between 71% to 93% of countries as less free. We also identify a paradoxical pattern we term differential misalignment: LLMs disproportionately underestimate press freedom in countries where it is strongest. Additionally, five of the six LLMs exhibit positive home bias, rating their home countries' press freedoms more favorably than would be expected given their negative misalignment with the human benchmark. In some cases, LLMs rate their home countries between 7% to 260% more positively than expected. If LLMs are set to become the next search engines and some of the most important cultural tools of our time, they must ensure accurate representations of the state of our human and civic rights globally.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

42 extracted references · 32 canonical work pages

  1. [1]

    & Rose, R

    Mishler, W. & Rose, R. What are the origins of political trust? testing institutional and cultural theories in post-communist societies. Comparative political studies 34, 30–62 (2001)

  2. [2]

    & Cohen, J

    Tsfati, Y. & Cohen, J. Democratic consequences of hostile media perceptions: The case of gaza settlers. Harvard International Journal of Press/Politics 10, 28–51 (2005)

  3. [3]

    Ladd, J. M. The role of media distrust in partisan voting. Political Behavior 32, 567– 585 (2010)

  4. [4]

    Trust and elections

    Hooghe, M. Trust and elections. em uslaner (2018). 22

  5. [5]

    & Steindl, N

    Hanitzsch, T ., Van Dalen, A. & Steindl, N. Caught in the nexus: A comparative and longitudinal analysis of public trust in the press. The international journal of press/politics 23, 3–23 (2018)

  6. [6]

    & Ariely, G

    Tsfati, Y. & Ariely, G. Individual and contextual correlates of trust in media across 44 countries. Communication research 41, 760–782 (2014)

  7. [7]

    M., Barber´a, P ., Munzert, S

    Guess, A. M., Barber´a, P ., Munzert, S. & Yang, J. The consequences of online partisan media. Proceedings of the National Academy of Sciences 118, e2013464118 (2021)

  8. [8]

    & Fletcher, R

    Newman, N. & Fletcher, R. Bias, bullshit and lies: Audience perspectives on low trust in the media (Reuters Institute for the Study of Journalism, 2017)

Show all 42 references
  1. [9]

    & Eisenegger, M

    Kalogeropoulos, A., Suiter, J., Udris, L. & Eisenegger, M. News media trust and news consumption: Factors related to trust in news in 35 countries. International journal of communication 13, 22 (2019)

  2. [10]

    & Hertwig, R

    Lorenz-Spreen, P ., Oswald, L., Lewandowsky, S. & Hertwig, R. A systematic review of worldwide causal and correlational evidence on digital media and democracy. Nature human behaviour 7, 74–101 (2023)

  3. [11]

    Bak-Coleman, J. B. et al. Stewardship of global collective behavior. Proceedings of the National Academy of Sciences 118, e2025764118 (2021)

  4. [12]

    & Starnini, M

    Cinelli, M., De Francisci Morales, G., Galeazzi, A., Quattrociocchi, W. & Starnini, M. The echo chamber effect on social media. Proceedings of the national academy of sciences 118, e2023301118 (2021)

  5. [13]

    The hype machine: How social media disrupts our elections, our economy, and our health–and how we must adapt (Crown Currency, 2021)

    Aral, S. The hype machine: How social media disrupts our elections, our economy, and our health–and how we must adapt (Crown Currency, 2021). 23

  6. [14]

    & Evans, J

    Farrell, H., Gopnik, A., Shalizi, C. & Evans, J. Large ai models are cultural and social technologies. Science 387, 1153–1156 (2025)

  7. [15]

    G., Cavalini, A

    Pacheco, A. G., Cavalini, A. & Comarela, G. Echoes of power: Investigating geopolitical bias in us and china large language models. arXiv preprint arXiv:2503.16679 (2025)

  8. [16]

    Innovation power: why technology will define the future of geopolitics

    Schmidt, E. Innovation power: why technology will define the future of geopolitics. Foreign Aff. 102, 38 (2023)

  9. [17]

    & Wang, T

    Sun, Y. & Wang, T . Be friendly, not friends: How llm sycophancy shapes user trust. arXiv preprint arXiv:2502.10844 (2025)

  10. [18]

    Tao, Y., Viberg, O., Baker, R. S. & Kizilcec, R. F . Cultural bias and cultural alignment of large language models. PNAS nexus 3, pgae346 (2024)

  11. [19]

    Yang, K. et al. Unpacking Political Bias in Large Language Models: A Cross -Model Comparison on U.S. Politics (2025). URL http://arxiv.org/abs/2412. 16746. ArXiv:2412.16746 [cs]

  12. [20]

    Y., Liu, Y

    Feng, S., Park, C. Y., Liu, Y. & Tsvetkov, Y. From pretraining data to language models to downstream tasks: Tracking the trails of political biases leading to unfair nlp models. arXiv preprint arXiv:2305.08283 (2023)

  13. [21]

    The political biases of chatgpt

    Rozado, D. The political biases of chatgpt. Social Sciences 12, 148 (2023)

  14. [22]

    & Rodrigues, V

    Motoki, F ., Pinho Neto, V. & Rodrigues, V. More human than human: measuring chatgpt political bias. Public Choice 198, 3–23 (2024)

  15. [23]

    & Ermon, S

    Manvi, R., Khanna, S., Burke, M., Lobell, D. & Ermon, S. Large Language Models are Geographically Biased (2024). URL http://arxiv.org/abs/2402.02680. ArXiv:2402.02680 [cs]. 24

  16. [24]

    & Peng, N

    Sheng, E., Chang, K.-W., Natarajan, P . & Peng, N. The woman worked as a babysitter: On biases in language generation. arXiv preprint arXiv:1909.01326 (2019)

  17. [25]

    & Bowman, S

    Bordia, S. & Bowman, S. R. Identifying and reducing gender bias in word -level language models. arXiv preprint arXiv:1904.03035 (2019)

  18. [26]

    Hu, T . et al. Generative language models exhibit social identity biases. Nature Computational Science 5, 65–75 (2025)

  19. [27]

    Durmus, E. et al. Towards measuring the representation of subjective global opinions in language models. arXiv preprint arXiv:2306.16388 (2023)

  20. [28]

    & Zhang, Y

    Zhou, D. & Zhang, Y. Political biases and inconsistencies in bilingual GPT models—the cases of the U.S. and China. Scientific Reports 14, 25048 (2024). URL https://www.nature.com/articles/s41598-024-76395-w

  21. [29]

    & Tai, M

    An, J., Huang, D., Lin, C. & Tai, M. Measuring gender and racial biases in large language models: Intersectional evidence from automated resume evaluation. PNAS nexus 4, pgaf089 (2025)

  22. [30]

    Jiang, L., Zhu, G., Sun, J., Cao, J. & Wu, J. Exploring the occupational biases and stereotypes of chinese large language models. Scientific Reports 15, 1–15 (2025)

  23. [31]

    Salecha, A. et al. Large language models display human-like social desirability biases in big five personality surveys. PNAS nexus 3, pgae533 (2024)

  24. [32]

    Fang, X. et al. Bias of ai-generated content: an examination of news produced by large language models. Scientific Reports 14, 5224 (2024)

  25. [33]

    & Cross, M

    Briggs, M. & Cross, M. Generative ai: Threatening established human rights instruments at scale (2024). 25

  26. [34]

    & Freedman, R

    Heinze, E. & Freedman, R. Public awareness of human rights: distortions in the mass media. The International Journal of Human Rights 14, 491–523 (2010)

  27. [35]

    & Stillwell, D

    Watson, J., van der Linden, S., Watson, M. & Stillwell, D. Negative online news articles are shared more to social media. Scientific Reports 14, 21592 (2024)

  28. [36]

    World press freedom index (2024)

    Reporters Without Borders. World press freedom index (2024). URL https://rsf.org/en. Accessed: 2024-06-01

  29. [37]

    Press freedom barometer (2024)

    Reporters Without Borders. Press freedom barometer (2024). URL https://rsf. org/en/barometer. Accessed: 2024-06-01

  30. [38]

    T ., Ross Arguedas, A

    Newman, N., Fletcher, R., Robertson, C. T ., Ross Arguedas, A. & Nielsen, R. K. Reuters Institute digital news report 2024 (Reuters Institute for the study of Journalism, 2024)

  31. [39]

    Hu, T . et al. Generative language models exhibit social identity biases. Nature Computational Science 5, 65 –75 (2025). URL https://www.nature.com/articles/ s43588-024-00741-1. Publisher: Nature Publishing Group

  32. [40]

    & Hruschka, E

    Pezeshkpour, P . & Hruschka, E. Large language models sensitivity to the order of options in multiple-choice questions. arXiv preprint arXiv:2308.11483 (2023). Appendix A Extended Data A.1 The World Press Freedom Index The World Press Freedom Index, published annually by Repor...

  33. [41]

    ”In Uruguay, what influence does economic power have over the editorial board on public media?”

  34. [42]

    The answers are pre -ordered by RWB, with the first option representing the best possible outcome and the last option representing the worst

    ”In Uruguay, what influence does economic power have over the editorial board on private media?” For what concerns the answer modality, each question in the survey offers a limited set of multiple -choice answers, ranging from a maximum of five to a minimum of two options. The...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.