{"id":"8ef8ee0d-91df-40dd-b1be-6d4ea078761f","arxiv_id":"2506.18045","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Six LLMs systematically underrate press freedom in 180 countries, penalize freer countries most, and five give their home countries favorable treatment.","lead":"In 180 countries, six popular AI chatbots rate press freedom as much worse than the human expert index does, while giving their home countries a boost. The study suggests that as chatbots become news sources, their systematic bias could distort public trust in democratic institutions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Scale-compression artifact likely drives the differential misalignment claim; the paper's own open-ended robustness check contradicts the headline slope sign.","rationale":"The reader identified scale compression as a possible confound but did not notice the stronger internal evidence: the paper's own open-ended robustness regression (Table A1, column 3) reports a positive coefficient of 0.479 on WPFI Scores, which directly contradicts the claimed negative differential misalignment. The main differential-misalignment tables (Table 2, Table A2, Table A3) label the right-hand-side variable 'LLM Scores', not 'WPFI Scores' as stated in Equation 2. If LLM Scores are compressed relative to WPFI scores (all models show means around 40-60 while WPFI spans 0-100), regressing (LLM score minus WPFI score) on LLM score will mechanically produce a negative slope even when LLM scores are positively related to WPFI scores. This is precisely the scale-compression artifact the reader worried about, and it appears to be realized in the paper's own numbers. The generalized underestimation and home-bias findings are likely robust - they survive in the randomized and open-ended checks - so the paper retains value, but the headline 'differential misalignment' result, which is the most novel and policy-relevant claim, is not supported by the reported regressions as specified. Because the central claim of the paper depends on this regression, and because the paper itself contains a sign-reversed robustness result, the appropriate verdict is rejection rather than conditional acceptance. The authors should be invited to resubmit after correcting the regression specification, rerunning the analysis with WPFI Scores on the right-hand side, and reconciling the abstract percentages (71-93% vs. 71-97%) and the home-bias coefficients (text says DeepSeek 11.49 and Qwen 9.2, while Table A1 reports 26.51 and 9.2 in different columns). Formal verification is absent, code is not yet public, and the paper even repeats its entire Discussion section verbatim in pages 19-21, which suggests incomplete editorial review. These issues compound the central statistical concern.","tokens_in":13567,"tokens_out":2005,"duration_ms":16305,"concrete_test":"Recompute the differential-misalignment regression in Equation 2 using the open-ended scores from Table A2 and plot (LLM score minus WPFI score) against WPFI score. If the slope is positive (as Table A1's 0.479 coefficient suggests), then the negative slopes in Table 2 and Table A2 arise because those regressions use LLM Scores on the right-hand side instead of WPFI Scores. Also rerun the main multiple-choice regression with WPFI Scores on the right-hand side, as Equation 2 states, and check whether the sign flips from negative to positive. If it flips, the differential misalignment claim is a scale-compression artifact and the paper should be revised to drop or substantially reframe that claim.","verdict_should_be":"REJECT","load_bearing_attack":"The paper's central novel claim is differential misalignment: LLMs disproportionately penalize freer countries, pooled beta = -0.251. The reader flagged scale compression as an alternative explanation, but the paper itself contains a more direct internal contradiction. In the main multiple-choice analysis, Table 2 reports negative slopes of (LLM score minus WPFI score) on LLM Scores (which are the LLM's own 0-100 mapped scores), not on WPFI scores. In Table A2 (open-ended answers), the same regressions also yield negative slopes, e.g., -0.219 for DeepSeek and -0.214 for GPT. Yet in Table A1, the open-ended specification reports a positive coefficient on WPFI Scores of 0.479***, which directly implies that LLM misalignment increases with WPFI score, i.e., the opposite of the negative differential misalignment claimed. This is not a subtle scaling issue; it is a sign reversal between two regressions that are supposed to measure the same effect. Either the main differential-misalignment result is an artifact of mapping categorical multiple-choice answers to equal-interval 0-100 scores, or the robustness table is mislabeled or incorrectly specified. Given the abstract's own claim that models rate between 71% and 93% of countries as less free (while the text says 71% to 97%), the paper already contains unreconciled numbers. The load-bearing concern is therefore that the differential misalignment finding, the paper's most policy-relevant claim, is not robust to the survey scoring method and may be a methodological artifact.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a survey-based study of six large language models (ChatGPT, Gemini, DeepSeek, Qwen, Mistral, Falcon) in which each model is prompted with an adapted version of the Reporters Without Borders press-freedom questionnaire for 180 countries. LLM scores are constructed by mapping multiple-choice answers to equal-interval 0-100 values, and are compared with the 2024 World Press Freedom Index (WPFI) expert scores. The authors claim three systematic distortions: generalized negative misalignment (LLMs underrate press freedom in 71-97% of countries), differential misalignment (freer countries are disproportionately penalized, pooled beta = -0.251), and positive home bias in five of six models. The paper also includes a validation of WPFI against media-trust data and robustness checks with randomized answer order and open-ended prompts.","tokens_in":13790,"tokens_out":5988,"duration_ms":53898,"significance":"If the central claims hold, this is an important and policy-relevant contribution to the literature on LLM bias. The study has notable strengths: it covers six models from four countries of origin, uses a large number of prompts (21,060 per model), anchors the evaluation against an established external benchmark (WPFI), and includes a media-trust validation that gives the benchmark independent plausibility. The finding that LLMs systematically understate press freedom and exhibit home bias would be a substantive result beyond the existing demographic and political-bias literature. However, the significance hinges on the validity of the differential-misalignment measure, which currently has a specification problem and a potential scaling artifact; these issues are load-bearing because differential misalignment is the paper's most novel claim.","major_comments":[{"comment":"The paper defines differential misalignment in Eq. (2) as the slope of Misalignment = LLM Score - WPFI Score on WPFI Scores, but Table 2 (and Table A2) report regressions of misalignment on LLM Scores, not on WPFI Scores. The text introducing Table 2 also says 'Negative coefficients for LLM Scores', while the Figure 3 caption describes the relationship between misalignment and WPFI scores. These two regressions are not the same: the coefficient on LLM Scores can be negative even when the coefficient on WPFI Scores is positive (or zero), depending on the joint distribution of LLM and WPFI scores. Because the pooled beta = -0.251 and the individual coefficients in Table 2 are presented as evidence for differential misalignment, the authors must re-estimate Eq. (2) with WPFI Scores as the regressor (or clearly state that the tables are mislabeled and provide the correctly specified results). As written, the central novel claim is not supported by the reported tables.","section":"Methods 4.3, Eq. (2); Results; Table 2; Fig. 3"},{"comment":"The equal-interval mapping of multiple-choice answers (100/75/50/25/0 for five options) can mechanically generate a negative slope of misalignment on WPFI scores if LLM responses are more moderate or more compressed than expert responses. Fig. 2 shows the LLM distributions shifted left, but the paper does not report the variances of the LLM score distributions or test whether they are substantially smaller than the WPFI variance. If the LLM scores are compressed toward the center of the 0-100 scale, a negative regression slope of (LLM - WPFI) on WPFI (or on LLM scores) will arise even under an unbiased evaluation process. The authors should report the standard deviations of the scores, provide a variance-comparison test, or use an alternative specification (e.g., rank-based alignment, or a slope on WPFI scores with a formal test that the slope is less than 1) to rule out the scaling artifact.","section":"Methods 4.2; Fig. 2"},{"comment":"The home-bias coefficients cited in the text (e.g., DeepSeek 11.49, Qwen 9.2, Gemini 9.53, GPT-4 8.06, Falcon 1.9, Mistral 2.29) do not match the interaction coefficients reported in Table A1 (e.g., DeepSeek×China 17.99/15.34/26.51, Gemini×US 11.86/9.059/6.514, GPT×US 6.679/9.670/5.471). The appendix table also omits the Qwen×China interaction altogether, even though the text attributes a large bias to Qwen. The paper does not explain which specification generated the numbers in the text, so the home-bias claim is not reproducible from the reported tables. The authors should present the home-bias coefficients from the Eq. (3) specification in a consistent table and reconcile the values with the text.","section":"Results: Home bias; Table A1"},{"comment":"Table 1 reports 6,300 observations for the generalized-misalignment regressions, described as 'at the question level' with 180 countries. Since the questionnaire has 117 questions, the full question-level sample would be 21,060 observations per model; 6,300 implies only 35 questions per country. The table does not state which subset of questions or which level of aggregation is used, making the generalized-misalignment coefficients (e.g., ChatGPT -16.84) difficult to interpret and to reconcile with the abstract's percentages. The authors should clarify the unit of analysis and the construction of the 6,300 observations, or correct the table.","section":"Results, Table 1"}],"minor_comments":[{"comment":"The abstract states that models rate between 71% and 93% of countries as less free, but the Results section reports ChatGPT at 97% and Gemini at 96%. The abstract should be corrected to match the reported range.","section":"Abstract vs. Results"},{"comment":"The text says RWB scores range 'from 1 to 100' but later describes the scale as '0 to 100'; the mapping described assigns 0 to the worst answer, so the scale should be stated consistently as 0-100.","section":"Methods 4.2"},{"comment":"The text refers to 'the regression model shown in Equation Y' instead of Equation (3); the equation should be cross-referenced correctly.","section":"Methods 4.3, Eq. (3)"},{"comment":"The Discussion section is duplicated nearly verbatim on pages 12-14 and pages 19-21 of the manuscript, including the limitations paragraph. One copy should be removed.","section":"Discussion (duplicated text)"},{"comment":"The sentence 'Falcon-180B, which did not have an accessible API between February and January 2025' contains an impossible date range; the intended months should be corrected.","section":"Methods 4.1"},{"comment":"In Table 1, the rows 'Country FE', 'R2', 'Observations', and 'N Countries' appear to be shared across all model columns, but the formatting is ambiguous (e.g., a single 'Yes' spans multiple columns). The table should be restructured so it is clear whether the R2 and observation count apply to each model separately or to a pooled model.","section":"Table 1 layout"},{"comment":"The caption for Table A1 describes 'Randomization 1' and 'Randomization 2', but the text explains that the first wave used the original answer order and the two subsequent waves used randomized order. The caption should be aligned with the description in the text.","section":"Appendix Table A1 caption"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely and important topic, and the data collection effort is substantial. However, the differential-misalignment result—the paper's headline claim—is currently presented in a way that is internally inconsistent: the stated regression equation uses WPFI scores, while the reported tables use LLM scores as the regressor. In addition, the equal-interval mapping of categorical answers raises a scaling-artifact concern that is not addressed. These are fixable with re-analysis and clearer reporting, so I recommend major revision rather than rejection. The authors should also reconcile the home-bias coefficients and the abstract's percentage range before resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline finding here is straightforward and probably right: six popular LLMs, asked to fill out the World Press Freedom Index questionnaire for 180 countries, consistently score press freedom lower than the human expert benchmark, and most of them give their home country a boost. The generalized underestimation is robust across models and prompt randomizations, and the home-bias pattern is interesting—DeepSeek rating China nearly twenty points above the WPFI is the kind of concrete result that makes you want to read on. The authors deserve credit for adapting a real survey instrument, running multiple waves with randomized answer orders, and checking open-ended responses as a robustness exercise. That is real work, and the descriptive results will likely survive scrutiny.\n\nThe problem is the paper's most novel claim: differential misalignment, the idea that LLMs are disproportionately harsher on freer countries. The authors estimate Equation 2, which regresses misalignment (LLM score minus WPFI score) on WPFI score, and report a pooled coefficient of -0.251. But Table 2, which is supposed to show this, actually regresses misalignment on 'LLM Scores'—the independent variable is the LLM's own mapped score, not the WPFI score. That is not a minor labeling slip; it is a different regression. The appendix makes this more confusing: Table A1, the open-ended specification, shows a positive coefficient on WPFI Scores (0.479), which suggests that misalignment actually increases with press freedom—the opposite of the negative differential misalignment claim. Either the tables are misspecified, or the finding is an artifact of the multiple-choice scoring scale. The authors do not discuss the scale-compression alternative at all, and the abstract's 71–93% range does not match the 97% and 96% figures in the text. These are addressable, but they are load-bearing for the paper's central narrative.\n\nThe generalized underestimation and home-bias results are likely real and worth publishing. The differential misalignment claim needs more care before I'd trust it. The paper deserves a serious referee, but only with a request for the correct regressions, a discussion of the scale artifact, and a public data/code release. As it stands, I'd hold off citing the selective misalignment result until the authors reconcile their own numbers.","headline":"A big, well-executed survey of how LLMs rate press freedom, with a clear generalized underestimation and home-bias story—but the paper's most novel claim, differential misalignment, rests on a regression specification that the paper itself contradicts in the appendix.","tokens_in":14366,"tokens_out":2520,"would_cite":false,"duration_ms":24908,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Six popular language models systematically underestimate press freedom, punish freer countries most, and favor their home countries.","keywords":["large language models","press freedom","differential misalignment","home bias","World Press Freedom Index","media trust","democratic institutions"],"falsifier":"Ask the same models to produce a direct numerical press-freedom score from 0 to 100 for each country, or to rank countries pairwise, and check whether freer countries still get disproportionately lower scores relative to WPFI; if the negative slope vanishes, the differential misalignment is an artifact of the multiple-choice mapping rather than a property of the models' beliefs.","tokens_in":13336,"feed_emoji":"📰","tokens_out":5383,"duration_ms":50517,"temperature":0.7,"pith_summary":"This paper argues that six widely used large language models—ChatGPT, Gemini, DeepSeek, Qwen, Mistral, and Falcon—do not simply make random errors when assessing press freedom. Compared with the expert World Press Freedom Index across 180 countries, they systematically rate most countries as less free, they are most negative about the freest countries, and five of six rate their home country more favorably than their own overall pattern would predict. The authors call these generalized misalignment, differential misalignment, and home bias. If true, the study implies that people who get information through chatbots are receiving a distorted picture of press freedom that flatters some governments and demeans others.","feed_headline":"LLMs underrate press freedom in up to 97% of countries","feed_subtitle":"Six chatbots judge freer countries most harshly and largely favor their home nations in a 180-country test.","key_machinery":"The instrument is the 117-question World Press Freedom Index survey, adapted by adding 'In [country]' to each item and administered as multiple-choice prompts to each model through its API. Responses are mapped to equal-interval 0-100 scores using the index's own scoring rule, then compared to WPFI category scores. The load-bearing statistical identity is Equation 2, misalignment = LLM score minus WPFI score regressed on WPFI score: the slope is the differential misalignment, and Equation 3 adds a home-country interaction term to isolate home bias.","core_discovery":"The paper's central claim is that six LLMs' evaluations of press freedom deviate from expert assessments in a structured, non-random way. ChatGPT underrates 97% of countries, Gemini 96%, Qwen 93%, Falcon 89%, Mistral 87%, and DeepSeek 71%; all six show negative coefficients when misalignment is regressed on WPFI scores, with a pooled slope of -0.251, meaning each extra point of expert-rated press freedom is met by deeper LLM underestimation. The same regressions run separately for each model show the effect across all six, with DeepSeek weakest. Five of six models show positive home bias: DeepSeek rates China nearly 20 points above the human benchmark, Qwen shows a bias coefficient of 9.2 toward China, Gemini 9.53 toward the US, GPT-4 8.06 toward the US, Falcon 1.9 toward the UAE, while Mistral's coefficient for France is small and only weakly significant.","pith_inferences":["A testable extension would be to ask the same models for direct numeric scores instead of multiple-choice answers; if the negative slope disappears, the differential misalignment is an artifact of mapping ordinal answers onto an equal-interval scale.","The paper's democratic-dilemma mechanism predicts similar differential patterns for other institutions where free societies generate more critical coverage, such as government accountability or human-rights records; that prediction can be tested with the same survey design.","Home bias may turn out to be tunable: if it arises from alignment procedures, preference-tuning could reduce it, while generalized negativity might persist; the paper does not test this.","Alternative press-freedom benchmarks would show whether the pattern is specific to the WPFI or generalizes; the paper itself notes its reliance on one benchmark."],"forward_implications":["A person asking a chatbot about press freedom in a free country is likely to get a more negative answer than the expert benchmark, while the gap between freest and least-free countries shrinks in LLM outputs.","Because the underestimation is smallest for the least-free countries, restrictive regimes may appear less exceptional than they do to expert assessors.","The home-bias pattern means the same country can receive substantially different press-freedom portrayals depending on which LLM is used, a form of algorithmic geopolitics.","If LLMs become the main search entry point, systematic misportrayal of press freedom could feed distrust or false reassurance about democratic institutions."],"supporting_citations":[{"why":"Supplies the questionnaire and the expert benchmark scores against which all LLM outputs are compared.","marker":"[36]"},{"why":"Provides the media-trust survey data used to validate that WPFI scores correlate with public trust.","marker":"[38]"},{"why":"Gives the democratic-dilemma mechanism the paper invokes to explain why freer countries attract more negative coverage in training data.","marker":"[34]"},{"why":"Supports the claim that negative news is more salient and more widely shared, underpinning the training-data asymmetry.","marker":"[13]"},{"why":"Provides prior evidence that LLMs reflect United States and selected European values when prompted in non-English languages.","marker":"[27]"},{"why":"Shows that LLMs give more favorable portrayals of China when prompted in Chinese, motivating the home-bias analysis.","marker":"[28]"},{"why":"Motivates the randomized answer-order robustness checks for multiple-choice sensitivity.","marker":"[40]"},{"why":"Identifies the barometer data whose omission from the safety category the authors test as a robustness check.","marker":"[37]"},{"why":"Links LLM social-identity bias and the possibility of reducing it through alignment techniques, backing the home-bias interpretation.","marker":"[39]"}],"fun_headline_variants":["LLMs understate press freedom, most in freest nations","Six chatbots skew press freedom scores, favoring home","LLMs undervalue press freedom: freer countries hit hardest","Paradox: LLMs rate freest presses as least free","Home bias: LLMs flatter their own press freedom"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a model's chosen answer option can be placed on the same equal-interval 0-100 scale as the experts' scores; if LLMs simply avoid extreme options, the negative slope would appear even without harsher judgments of free countries.","fun_headline_variants_meta":{"raw":{"variants":["LLMs understate press freedom, most in freest nations","Six chatbots skew press freedom scores, favoring home","LLMs undervalue press freedom: freer countries hit hardest","Paradox: LLMs rate freest presses as least free","Home bias: LLMs flatter their own press freedom"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000844,"raw_usage":{"total_tokens":3683,"prompt_tokens":960,"completion_tokens":2723,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":576,"completion_tokens_details":{"reasoning_tokens":2653}},"tokens_in":576,"tokens_out":2723,"duration_ms":18618,"temperature":1.0,"reasoning_tokens":2653,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:55:45.658254+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Ask the same models to produce a direct numerical press-freedom score from 0 to 100 for each country, or to rank countries pairwise, and check whether freer countries still get disproportionately lower scores relative to WPFI; if the negative slope vanishes, the differential misalignment is an artifact of the multiple-choice mapping rather than a property of the models' beliefs.","supporting_citations":[{"cited_title":"World press freedom index (2024)","cited_arxiv_id":null,"evidence_quote":"Supplies the questionnaire and the expert benchmark scores against which all LLM outputs are compared."},{"cited_title":"T ., Ross Arguedas, A","cited_arxiv_id":null,"evidence_quote":"Provides the media-trust survey data used to validate that WPFI scores correlate with public trust."},{"cited_title":"& Freedman, R","cited_arxiv_id":null,"evidence_quote":"Gives the democratic-dilemma mechanism the paper invokes to explain why freer countries attract more negative coverage in training data."},{"cited_title":"The hype machine: How social media disrupts our elections, our economy, and our health–and how we must adapt (Crown Currency, 2021)","cited_arxiv_id":null,"evidence_quote":"Supports the claim that negative news is more salient and more widely shared, underpinning the training-data asymmetry."},{"cited_title":"& Zhang, Y","cited_arxiv_id":null,"evidence_quote":"Shows that LLMs give more favorable portrayals of China when prompted in Chinese, motivating the home-bias analysis."},{"cited_title":"Press freedom barometer (2024)","cited_arxiv_id":null,"evidence_quote":"Identifies the barometer data whose omission from the safety category the authors test as a robustness check."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Links LLM social-identity bias and the possibility of reducing it through alignment techniques, backing the home-bias interpretation."}],"review_version":1}