Pith. sign in

REVIEW 4 major objections 5 minor 37 references

Informing AI Risk Assessment with News Media: Analyzing National and Political Variation in the Coverage of AI Risks

T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read AI risk coverage in the news varies by country and by outlet political leaning, so media-based risk assessment must account for that variation.

desk verdict A transparent and useful descriptive study whose headline country and bias effects are statistically overstated because article-level chi-squares ignore severe outlet clustering and a paywall-filtered sample. read the letter →

arxiv 2507.23718 v1 pith:LSWTRNRR submitted 2025-07-31 cs.CY

classification cs.CY
keywords AIriskassessmentnewsmediacoveragecross-nationalcomparisonpoliticalbiasinagendasettingtaxonomyLLMtextclassification
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that news media cannot be treated as a neutral window onto AI risks, because the risks that get covered are prioritized differently across countries and across politically leaning outlets. Analyzing a sample of 12,385 English-language AI articles from the U.S., U.K., India, Australia, Israel, and South Africa, it finds statistically significant associations between country and coverage of six of seven AI risk categories, and between U.S. outlet bias and all seven. The authors' point is practical: risk assessors and policymakers who use news coverage as a supplementary source for risk-based AI governance should model national and political variation in coverage as a measurement factor, not ignore it.

What carries the argument

The load-bearing machinery is the seven-domain risk taxonomy applied at scale: a consolidated taxonomy of AI risks organized into seven domains, used to classify every summarized negative impact extracted from an article. Around it sits a three-step pipeline: a large language model filters articles for whether they describe an impact of an AI system, summarizes each negative impact in one sentence, and classifies the summary into exactly one of the taxonomy's domains, with human-annotated samples used to validate each step (F1 of 0.82 for filtering and 0.90 for classification). The statistical engine is the chi-square test of independence, run separately on each risk category against country and against U.S. outlet bias ratings.

What would settle it

A direct falsifier is to scrape the paywalled and non-scrapable outlets in the same six countries and rerun the per-country $\chi^2$ tests: if the country-risk associations shrink to non-significance after controlling for outlet identity, the national-prioritization claim is an artifact of which outlets were accessible. A simpler version is to compare risk proportions within countries across outlets, since the claim fails if outlet-level variation swamps country-level variation.

Watch

Extended reading notes

Core claim

The discovery is that a seven-domain taxonomy of AI risks, spanning discrimination and toxicity, privacy and security, misinformation, malicious actors and misuse, human-computer interaction, socioeconomic and environmental harms, and AI system safety failures, maps unevenly onto news coverage. Socioeconomic and environmental harms dominate overall, while human-computer interaction risks are the least covered. Chi-square tests show country and risk category are not independent: for instance, coverage of socioeconomic and environmental harms ranges from about 25% of articles in Israel and India to 35\textendash 45% in South Africa, the U.S., U.K., and Australia, with $\chi^2(5)=54.76$, $p<0.001$; only privacy and security shows no significant country association. Within the U.S., right-biased outlets emphasize malicious actors and misuse, at 43.1% of their articles, and discrimination and toxicity, at 36.1%, while under-covering socioeconomic and environmental harms at 25.7% relative to center and left outlets, and the reporting around these risk categories contains demonstrably politicized language, such as claims that AI is used to 'make AI woke' or 'push a leftist agenda' versus left-outlet framings about far-right extremism and discrimination against people of color.

Load-bearing premise

The entire comparison assumes that the 17.5% of eligible URLs that were freely scrapable, with U.S. outlets supplying 64.5% of the final sample and single outlets contributing over 70% of Israel's and South Africa's articles, represents each country's AI risk discourse.

Editorial extensions

If this is right

  • Risk assessors who draw on news media for incident monitoring should weight or stratify sources by country and outlet bias, because raw coverage counts conflate journalistic priorities with actual harm prevalence.
  • Country-specific risk registers would look different from one another: socioeconomic and environmental harms dominate South African and U.S. coverage, while malicious actors and misuse is relatively more salient in U.K. and Australian coverage.
  • U.S. media polarization extends to AI risk reporting, meaning public perceptions of which AI risks matter are likely shaped differently for audiences of right-leaning versus left-leaning outlets.
  • The mismatch between research taxonomies and news priorities, where misinformation is covered more prominently in news than its rank in academic risk repositories, supports using news media as a complementary societal-context signal rather than a substitute for expert assessment.
  • Incident databases that rely on news articles inherit these national and political coverage biases, so cross-country comparisons in such databases should be adjusted for media availability and outlet bias.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper, the scrape-access asymmetry implies that the observed country effects may partly be outlet effects: only 17.5% of eligible URLs were freely accessible, with single outlets contributing over 70% of the Israeli and South African samples, so adding paywalled or non-English outlets would show how much of the national pattern is driven by a few dominant, freely accessible domains.
  • A testable extension the paper does not run is to compare the risk distribution in its news sample against the distribution of incident types in an AI incident database; if they diverge systematically by country, that would quantify how much journalistic selection, rather than actual reported incidents, shapes the risk signal.
  • The politicized-language examples suggest that framing measures, such as the share of risk articles naming political actors or using charged terms, could be operationalized as a quantifiable indicator of politicization and fed into AI governance monitoring.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper analyzes a cross-national sample of 12,385 news articles from six countries (U.S., U.K., India, Australia, Israel, South Africa), retrieved via GDELT and filtered by AI-related keywords. Using GPT-4o with zero-shot prompts, the authors extract negative impacts and classify them into the MIT AI Risk Repository domain taxonomy. They report chi-square tests showing that the prevalence of most risk categories is associated with country (Section 5.1) and with the Media Bias Fact Check political-bias rating of U.S. outlets (Section 5.2), and they illustrate politicized language in coverage by right- and left-leaning outlets. The central claim is that AI risks are prioritized differently across nations and across U.S. outlet bias categories, and that risk assessors should account for such media variation.

Significance. If the findings are statistically valid, the paper makes a useful contribution by demonstrating that news media coverage of AI risks is not homogeneous and that national and political contexts shape which risks are salient. The use of an external, aggregated taxonomy (MIT Risk Repository), transparent reporting of prompts and validation scores, and the inclusion of Global South countries (India and South Africa) are strengths. However, the core quantitative evidence relies on article-level chi-square tests that ignore severe outlet-level clustering and a sample that is heavily skewed toward the U.S. and dominated by a few outlets in several countries. These issues currently undermine the strength of the cross-national and cross-bias claims, so the paper needs additional statistical work and sensitivity analyses before the conclusions can be fully supported.

major comments (4)
  1. [§5.1, Tables A.2 and A.4] The chi-square tests of independence treat each of the 12,385 articles as an independent observation, but the data are heavily clustered by outlet. For example, Israel's 266 articles come from 4 domains with jpost.com contributing 71.4%; South Africa's 206 articles come from 3 domains with dailymaverick.co.za contributing 71.8%; and India's 1,255 articles come from 4 domains with timesofindia.indiatimes.com contributing 46.5%. Articles from the same outlet share editorial policies, so the effective sample size for country-level effects is far smaller than the article count. The reported p-values (e.g., χ²(5)=54.76 for Socioeconomic & Environmental risks) are therefore inflated, and the country effects may largely reflect outlet effects. The authors should either conduct outlet-level analyses (e.g., aggregating within domains or using multilevel models that include random intercepts for outlet) or provide a robustness check showing that the associations persist when outlet clustering is accounted for.
  2. [§3 and §7 (data collection and limitations)] The sample is only 17.5% of the eligible URLs (31,252 of 178,172), and the final analytical sample has strong geographic skew: U.S. articles are 64.5% of the 12,385, while Israel and South Africa account for 2.1% and 1.7%, respectively. The paywall and scraping failure may systematically exclude certain types of outlets (e.g., major subscription-based newspapers), and Section 7 acknowledges this but does not quantify how the excluded 82.5% or the dominant outlets affect the reported associations. As a result, the cross-national comparisons in Section 5.1 may reflect differences in which outlets could be scraped rather than differences in national discourse. The authors should compare the characteristics of scraped versus unscraped domains, or conduct sensitivity analyses restricted to outlets with similar accessibility, to assess the magnitude of this bias.
  3. [§5.2, Table A.5] The U.S. bias analysis also relies on article-level chi-square tests despite severe outlet clustering within bias categories. The Right category contains only 5 domains (hotair.com, newsbusters.org, dailycaller.com, redstate.com, patriotpost.us) contributing 430 articles, and the Left-Center category is dominated by businessinsider.com (523 articles) and washingtonpost.com (503 articles). With such small numbers of domains per bias category, the significant chi-square results (e.g., χ²(4)=84.87 for Malicious Actors & Misuse) are not credible as evidence about political orientation per se; they may be driven by the editorial choices of a few outlets. The authors should report outlet-level analyses or at least show that the patterns hold when no single outlet dominates a bias category.
  4. [§4.2 (LLM classification validation)] The classification performance is evaluated on only 300 manually annotated negative impacts (macro F1=0.90), and the summarization step on 50 articles. Given that the corpus contains 36,793 impacts, the validation sample is small, and no per-category or per-country performance breakdown is reported. If the LLM's classification accuracy varies by risk category or by outlet/country (e.g., due to topic-specific phrasing), the prevalence estimates in Sections 5.1 and 5.2 could be systematically biased. The authors should provide confidence intervals for the prevalence estimates or a sensitivity analysis that accounts for measurement error in the labels.
minor comments (5)
  1. [Abstract and §5.2] The abstract says "left vs. right leaning U.S. outlets," but the analysis uses five bias categories (Left, Left-Center, Least Biased, Right-Center, Right). Please align the wording with the actual analysis.
  2. [Appendix A.9] The prompt says the impact statements should be categorized into one of 32 categories, but the list contains 32 numbered categories and the note at the end mentions "33 categories." Please correct this inconsistency.
  3. [§5.2 and Table A.5] Minor typos: "right-based outlets" should be "right-biased outlets" in the discussion of AI Systems Safety, Failures, & Limitations.
  4. [§3 and §7] The paper uses both "U.S." and "USA" and both "U.K." and "UK"; please standardize abbreviations.
  5. [§4.2] The definition of the "other" category in the classification prompt is not explicitly aligned with the MIT domain taxonomy described in the text; please clarify how "other" labels were treated when aggregating to the seven domain categories.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the risk taxonomy and classification pipeline are externally anchored, and the corpus keyword list is a disclosed, non-target input.

full rationale

The paper's empirical chain is not circular. The outcome variable (prevalence of AI risk categories) is produced by an external taxonomy (the MIT Risk Repository) and an LLM classifier validated against human annotations; neither is defined in terms of the paper's conclusions. The corpus is selected using a 40-keyword list that is inherited from the authors' prior workshop paper, but that list is disclosed in Appendix A.1, was derived data-driven from NYT and AIID sources, and consists of general AI terms rather than risk-category labels, so it does not build the target result into the inputs. The only self-citation relevant to corpus construction is therefore a transparent methodological input rather than a load-bearing circular premise. Cross-country and political-orientation comparisons are descriptive chi-square analyses of the resulting classifications; no fitted parameter is renamed as a prediction, no uniqueness theorem is imported from the authors' prior work, and no risk-specific ansatz is smuggled in via citation. Concerns about outlet clustering, paywall filtering, and LLM measurement validity are correctness and robustness issues, not circularity.

Assumptions & free parameters 1 free parameters · 5 assumptions · 0 invented entities

The central descriptive claims rest on three layers of choices: the GDELT keyword-based corpus construction, the freely-scrapable subset, and the GPT-4o annotation pipeline. None of these are numerically fitted to the outcome, so free parameters are limited to the keyword instrument; the main burden is domain assumptions about representativeness and label validity.

free parameters (1)
  • AI-relevant keyword set = 40 keywords (31 from NYT seed queries, 9 from AI Incident Database n-grams)
    The corpus is built by querying GDELT for these terms; prevalence estimates depend on this choice. The set is not validated for recall or precision against a reference set of AI articles, so it is a free measurement choice.
assumptions (5)
  • domain assumption GDELT provides sufficient coverage of online news in all six countries to support country-level comparison
    Section 3 selects GDELT as the sole source based on its 'extensive coverage'; no per-country validation is reported.
  • domain assumption Freely scrapable articles represent each country's media discourse despite large paywall exclusions
    Section 3 and Section 7 acknowledge that paywalled or missing articles were excluded, but the analysis proceeds as if the remaining sample still supports country-level inference.
  • domain assumption GPT-4o zero-shot classification generalizes from validation samples to the full corpus
    Section 4.1 validates filtering on 300 annotated articles (F1=0.82) and classification on 300 impacts (F1=0.90); summarization is checked on 50 articles. The full 36,793 impacts are produced without further validation.
  • domain assumption MIT Risk Repository domain taxonomy is an appropriate lens for news-reported AI harms
    Section 4.2 justifies the taxonomy as a synthesis of 56 taxonomies, but assumes news-described harms map cleanly onto its seven domain categories.
  • domain assumption Articles are independent units for chi-square tests
    Section 5 runs chi-square tests on article counts without accounting for nesting within outlets, which can produce overconfident p-values.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Informing AI Risk Assessment with News Media: Analyzing National and Political Variation in the Coverage of AI Risks." pith.science (2026). https://pith.science/paper/LSWTRNRR

@misc{pith2026250723718,
  author       = {Pith},
  title        = {Pith review of: Informing AI Risk Assessment with News Media: Analyzing National and Political Variation in the Coverage of AI Risks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LSWTRNRR}},
  note         = {Machine review of arXiv:2507.23718}
}
read the original abstract

Risk-based approaches to AI governance often center the technological artifact as the primary focus of risk assessments, overlooking systemic risks that emerge from the complex interaction between AI systems and society. One potential source to incorporate more societal context into these approaches is the news media, as it embeds and reflects complex interactions between AI systems, human stakeholders, and the larger society. News media is influential in terms of which AI risks are emphasized and discussed in the public sphere, and thus which risks are deemed important. Yet, variations in the news media between countries and across different value systems (e.g. political orientations) may differentially shape the prioritization of risks through the media's agenda setting and framing processes. To better understand these variations, this work presents a comparative analysis of a cross-national sample of news media spanning 6 countries (the U.S., the U.K., India, Australia, Israel, and South Africa). Our findings show that AI risks are prioritized differently across nations and shed light on how left vs. right leaning U.S. based outlets not only differ in the prioritization of AI risks in their coverage, but also use politicized language in the reporting of these risks. These findings can inform risk assessors and policy-makers about the nuances they should account for when considering news media as a supplementary source for risk-based governance approaches.

Figures

Figures reproduced from arXiv: 2507.23718 by the authors.

Figure 1
Figure 1. Prevalence of AI risk categories in news coverage across six countries in our sample, based on the proportion of [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Proportion of news articles in our sample from U.S. domains across five media bias categories, as rated by Media Bias [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Domain Taxonomy of AI Risks sourced from the MIT Risk Repository (Slattery et al. 2024) [PITH_FULL_IMAGE:figures/full_fig_p015_3.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

37 extracted references · 33 canonical work pages

  1. [1]

    ai_governance: negative impacts associated with governance policies and regulations related to the development, deployment, licensing, or moderation of AI technologies

  2. [2]

    ai_incompetence: risks resulting from limitations and malfunctions of AI technologies that impact their performance

  3. [3]

    Montasari, R

    Fully Autonomous AI Agents Should Not be Devel- oped.arXiv preprint arXiv:2502.02649. Montasari, R. 2023. National artificial intelligence strate- gies: a comparison of the UK, EU and US approaches with those adopted by state adversaries. InCountering Cybert- errorism: The Confluence of Artificial Intelligence, Cyber Forensics and Digital Policing in US a...

  4. [4]

    arXiv preprint arXiv:2408.12622

    The ai risk repository: A comprehensive meta-review, database, and taxonomy of risks from artificial intelligence. arXiv preprint arXiv:2408.12622. Solaiman, I.; Talat, Z.; Agnew, W.; Ahmad, L.; Baker, D.; Blodgett, S. L.; Daum´e III, H.; Dodge, J.; Evans, E.; Hooker, S.; Jernite, Y .; Luccioni, A. S.; Lusoli, A.; Mitchell, M.; Newman, J.; Png, M.-T.; Str...

  5. [5]

    negative_impacts

    Accessed: 2025-04-15. Zeng, Y .; Klyman, K.; Zhou, A.; Yang, Y .; Pan, M.; Jia, R.; Song, D.; Liang, P.; and Li, B. 2024. AI Risk Categoriza- tion Decoded (AIR 2024): From Government Regulations to Corporate Policies. Version Number: 1. Zhang, Z.; Lei, L.; Wu, L.; Sun, R.; Huang, Y .; Long, C.; Liu, X.; Lei, X.; Tang, J.; and Huang, M. 2023. Safetybench: ...

  6. [6]

    discrimination/bias: risks of AI technologies generating outputs based on protected characteristics that result in unequal treatment or representation of individuals or social groups

  7. [7]

    disruption_of_service: risks related to disruptions or reductions in the accessibility, availability, and functionality of AI systems

  8. [8]

    authoritative_use_of_ai: risks associated with the potential misuse of artificial intelligence by governments in ways that may support authoritarian practices that violate human rights and civil liberties

Show all 37 references
  1. [9]

    criminal_activities: risks associated with the misuse of AI technologies for online crimes such as cybercrimes or cyberattacks

  2. [10]

    deception/manipulation: risks associated with the use of AI technologies for fraud, dishonest activities, sowing divisions, or misrepresenting individuals to influence or alter perceptions, behaviors, or mislead individuals or society

  3. [11]

    existential_threats: risks related to potential inequalities in accessing AI technologies, their development, deployment, and use, with a particular focus on issues of fairness, accountability, and transparency

  4. [12]

    fundamental_rights: risks posed by AI technologies related to violating individual freedoms and rights, including freedom of expression and intellectual property

  5. [13]

    economic_harm: risks posed by AI technologies to financial systems, labor market, and trading dynamics

  6. [14]

    environmental: environmental and ecological risks arising from the energy consumption and resource intensive process required for the development, deployment, and operation of AI technologies

  7. [15]

    ethical_impact: challenges related to inequalities in access to AI technologies, or in their development , deployment, and use, with a particular focus on issues of fairness, accountability, and transparency

  8. [16]

    media_impacts: risks of AI technologies on the independence, integrity, and reliability of media and journalism

  9. [17]

    mental_&_emotional: risks related to the psychological well-being and emotional health of individuals using or interacting with AI technologies

  10. [18]

    information_risks: risks associated with AI hallucinations, including the generation of inaccurate information , low-quality AI-generated content, and fabrication of information by AI technologies

  11. [19]

    hate/toxicity: risks associated with AI-generated content amplifying or spreading hateful, abusive, or offensive content

  12. [20]

    humanness: risks related to the loss or diminishment of human qualities, such as creativity, emotional depth, and authentic interpersonal connections, due to AI technologies

  13. [21]

    privacy: risks related to unauthorized access, use, or disclosure of users’ data and personal information

  14. [22]

    safety_risks: risks of harming or endangering individuals’ lives or safety arising from the malfunction or misuse of AI technologies

  15. [23]

    operational_misuses: risks associated with the misuse of AI technologies in critical and highly regulated applications, such as unsafe autonomous operations, unreliable legal or military advice, or automated decision-making

  16. [24]

    over-reliance: risks arising from over-relying on AI technologies in contexts that results in undermining human judgment, critical thinking, and decision-making

  17. [25]

    political_useage: risks associated with the use of AI technologies to spread misinformation or disinformation, influence elections or politics, undermine democratic integrity, or disrupt social order

  18. [26]

    technology_adoption: risks related to the adoption of AI technologies due to integration and usability challenges, or due to barriers faced by organizations and individuals in adopting AI technologies into their work

  19. [27]

    user_experience: risks and issues that undermine the satisfaction satisfaction, trust, and interaction of the end-user with AI technologies

  20. [28]

    security_risks: risks related to the threats and exploitation of vulnerabilities that compromise the confidentiality, integrity, or availability of AI technologies

  21. [29]

    sexual_content: risks related to the non-consensual creation, distribution, or misuse of sexually explicit material or pornography using AI technologies

  22. [30]

    structure/power: risks related to the concentration of power and AI resources or technologies among a few entities or governments, and its consequences on competition, collaboration, innovation, and safety of AI technologies

  23. [31]

    no_impact: refers to general statements that do not highlight potential or direct negative consequences, risks, or harms of AI systems

  24. [32]

    Output Format : Present the classified categories without any numbers and clean from whitespace

    other: Any risks or harms that do not fit into the above categories. Output Format : Present the classified categories without any numbers and clean from whitespace. The categories should be selected from one of the above 33 categories. Note: Ensure that each impact statement ...

  25. [33]

    violence_&_extremism: risks pertaining to the use of AI technologies or AI-generated content to incites violence, promotes extremist ideologies, or enables harmful activities such as weapon development or using AI in warfare

  26. [34]

    defamation: risks involving reputational harms to individuals or organizations through AI-generated false or misleading statements, images, or representations

  27. [35]

    child_harm: risks related to the misuse of AI technologies to harm or exploit children

  28. [2024]

    Allaham, M.; Lokmanoglu, A

    https://arxiv.org/pdf/2411.02536. Allaham, M.; Lokmanoglu, A. D.; Hart, P.; and Nisbet, E. C

  29. [2025]

    InInternational Workshop on AI Governance: Alignment, Morality and Law (AIGOV) 2025

    Enhancing LLMs for Governance with Human Over- sight: Evaluating and Aligning LLMs on Expert Classifica- tion of Climate Misinformation for Detecting False or Mis- leading Claims about Climate Change. InInternational Workshop on AI Governance: Alignment, Morality and Law (AIGO...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.