Pith. sign in

REVIEW 3 major objections 4 minor 45 references

Grokipedia vs Wikipedia: An LLM-Based Audit of Political Neutrality along Ideologies

T0 review · 3 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash

Pith's one-line read Grok's own encyclopedia, Grokipedia, is rated less neutral than Wikipedia by all four AI judges, with bias favouring economically right-wing politicians.

desk verdict A careful, large-scale audit of LLM-perceived neutrality with a real methodological backbone; the persistent absence of human calibration is the one thing that keeps it from being an audit of the encyclopedias themselves. read the letter →

arxiv 2607.15146 v1 pith:CDZ5XTDY submitted 2026-07-16 cs.CL cs.CY

classification cs.CLcs.CY
keywords GrokipediaWikipediapoliticalbiasneutralityLLM-as-judgeideologyencyclopediaauditGrok
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tests the claim that an encyclopedia written entirely by an AI model could be more neutral than a human-edited one, comparing Grokipedia — written by the model Grok — with Wikipedia. Its central finding is the opposite: all four AI judges (Claude, Grok, Mistral, and DeepSeek) rate Grokipedia as less neutral on average than Wikipedia, and Grok rates its own output as more biased. The bias has a direction: Grokipedia favours economically right-wing politicians and penalises socially liberal positions, while Wikipedia is rated as favourably biased toward socially liberal politicians. This matters because encyclopedias shape political opinion, and AI-generated knowledge bases — which may be cited by future models — could embed one ideology rather than remove bias.

What carries the argument

The audit pipeline is the machinery: 1,394 world-government politicians are mapped to nine expert-coded ideological dimensions; each politician's Wikipedia and Grokipedia article is rated for neutrality by four different LLM judges on a five-point scale (strongly biased against to strongly biased in favor); and ordinary least-squares regressions estimate how each ideology dimension predicts the rating gap between sources and within each source. The key design choice is the multi-judge panel — including Grok, the model that generates Grokipedia — so that evaluator bias becomes part of the measurement rather than an unexamined assumption.

What would settle it

Recruit a panel of human annotators with diverse political views to rate a random sample of the same 1,394 article pairs on the same five-point neutrality scale; if the human-rated gap and ideological direction do not reproduce the LLM judges' results, the paper's central claim fails.

Watch

Extended reading notes

Core claim

All four large-language-model judges — Claude, Grok, Mistral, and DeepSeek — rate Grokipedia as less neutral than Wikipedia across 1,394 pairs of articles about members of government. The strongest ideological predictor is the economic left-right scale: in Grokipedia, a one-standard-deviation shift toward the economic right raises the favourable-bias rating by 0.221 on a −2-to-+2 scale, whereas the same shift has no significant effect in Wikipedia. Grokipedia is also rated as favourably biased toward politicians supporting democratic pluralism and as negatively biased toward pro-LGBT and pro-women's-labour politicians; Wikipedia shows the opposite pattern on the social-liberal dimensions. Be

Load-bearing premise

The load-bearing premise is that the four LLM judges' neutrality ratings measure the articles' actual bias; if the judges share an evaluative bias — say, a preference for human-edited Wikipedia-style prose or a distaste for AI-generated text — the reported gap between Grokipedia and Wikipedia could be an artifact of the judges rather than a property of the encyclopedias.

Editorial extensions

If this is right

  • If the finding is correct, Grokipedia is not a neutral alternative to Wikipedia; it shifts the bias from social liberalism to economic conservatism.
  • Readers' impressions of a politician could depend on which encyclopedia they consult, particularly for economically right-wing or socially liberal politicians.
  • Because Grok itself rates Grokipedia as less neutral, the bias is likely encoded in the generated text, so fixing it means changing the generation process, not just the evaluation.
  • The same audit design can be applied to other AI-generated or human-edited knowledge bases to screen for ideological skew at scale.
  • Single-judge audits would be misleading: relying on Grok or DeepSeek alone would underestimate the neutrality gap between the two encyclopedias.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Inference: The four judges disagree in strictness but not in direction, which suggests a shared judgment prior among frontier LLMs about what neutral encyclopedic prose looks like; if that prior is a house style, the absolute gap could be inflated even though the relative ordering is stable.
  • Inference: Grok's self-critical rating implies a gap between its writing behavior and its evaluation behavior; a direct test would be to ask Grok to rewrite its own biased passages and check whether neutrality scores improve.
  • Inference: If AI-written encyclopedias feed future model training, the measured ideological tilt could amplify through a feedback loop, making periodic neutrality audits a governance tool for AI-generated knowledge.
  • Inference: The paper's own future-work suggestion — auditing non-political topics like health, history, and science — would separate ideology-specific bias from a general machine-writing style.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper compares the political neutrality of Grokipedia and Wikipedia on 1,394 paired articles about members of government. Politicians are mapped to nine ideology dimensions from the expert-coded V-Party dataset, and each article is rated by four LLM judges (Grok, Claude, Mistral, DeepSeek) on a five-point scale from 'strongly biased against' to 'strongly biased in favor.' The authors use OLS regressions with politician-level clustered standard errors to relate ratings to ideology, and report that all judges rate Grokipedia as less neutral, that Grokipedia is particularly favorable to economically right-wing politicians and unfavorable to socially liberal ones, while Wikipedia is relatively more favorable to the latter. They also examine judge-specific rating tendencies.

Significance. If the findings hold, the paper is a timely and policy-relevant audit of an emerging AI-generated encyclopedia, with a methodology that is a clear advance over prior six- or 395-article comparisons: it uses 1,394 pairs, expert-coded ideology instead of LLM-generated ideology scores (avoiding circularity), a multi-judge panel with documented ideological diversity, and prompt-robustness checks. The main regression finding — a large positive coefficient for right-wing ideology in Grokipedia and a negative coefficient for LGBT equality — is a specific, falsifiable claim. The main weakness is that the outcome variable is an unvalidated LLM judgment of neutrality, and the paper's own limitations section concedes that human alignment has not been established. That weakness is load-bearing for the central descriptive and ideological claims.

major comments (3)
  1. [§3.1, §6 (Limitations), Eqs. (1)–(2), Tables 1–2] The outcome variable is an LLM neutrality judgment, and the Limitations section concedes: 'it remains to be verified whether the LLM ratings are aligned with human annotators.' This concession targets the load-bearing assumption of the entire paper. The prompt-robustness check (86.25% exact agreement when the neutrality definition is removed) demonstrates stability across prompt variants, not validity: a shared judge prior—for example, a common preference for Wikipedia-style NPOV prose, or a shared left-of-center framing—would also be highly stable. Because the four models are trained on overlapping corpora and share instruction-tuning regimes, a common evaluative bias could generate the Table 1–2 differences and the Table 7 coefficients without any difference in the encyclopedias' actual content. Grok rating Grokipedia lower rules out simple self-serving leniency, but not a common prior
  2. [Table 1, §4.1] The central claim that 'all LLM-judges, including Grok, rate Grokipedia less neutral than Wikipedia' rests on descriptive mean absolute ratings with no reported uncertainty. The strongest gap is Claude's (0.239), but Grok's is only 0.0409 and DeepSeek's 0.0760; no standard errors, confidence intervals, or paired significance tests are reported. With 1,394 observations per cell, even small differences may be significant, but the manuscript should demonstrate this rather than assert it. The claim that 'Grok itself assigns a higher mean absolute bias rating to Grokipedia' is used in the Discussion to argue the patterns are 'not an artifact of rater bias'; that inference requires at least a confidence interval for Grok's gap and ideally a test showing it is distinguishable from zero.
  3. [§3.2, Eqs. (1)–(2), Tables 6–7] Standard errors are clustered at the politician level, but the ideology regressors are assigned at the party/country level: many politicians share the same V-Party score, so residuals are likely correlated within country and especially within party. Clustering only at the politician level ignores this dependence and may understate standard errors for the ideology coefficients. The right-wing coefficient is large enough that it may survive country-level clustering, but the smaller coefficients (e.g., womlab, anti-popul, secular) are less certain. Please report a robustness check with standard errors clustered by country (or party), or at least report the intra-cluster correlation and whether the substantive conclusions change.
minor comments (4)
  1. [Eq. (2)] The indicator notation 'D_{k}^{m} = /x31[m = k]' appears to be a rendering artifact; use standard indicator notation such as \mathbf{1}[m = k].
  2. [Figure 4] The x-axis is labeled '0.0 0.2 0.4', but the text reports negative coefficients (e.g., culincl = −0.047, lgbt = −0.095 in Grokipedia). The figure as printed appears to clip the negative range; please rescale the axis.
  3. [Abstract and §5] The phrase 'both encyclopedias are rated as portraying politicians favourably overall' is stronger than the per-source intercepts in Table 7 suggest: for the reference judge Mistral, the Wikipedia intercept is −0.073 (slightly unfavorable), while Grokipedia's is +0.075. The claim may be true for a particular averaging or for the direction of non-neutral votes in Fig. 2(a), but the wording should be reconciled with the regression estimates.
  4. [Appendix A.1] The paper states code and data will be made available upon publication. For a reproducibility-focused journal, consider depositing the article-pair IDs, judgments, and analysis scripts in a public repository at submission time.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: expert-coded ideology dimensions are external; neutrality ratings are measured, not reduction-equivalent to their inputs.

full rationale

The paper's derivation chain is not circular. The independent variables (nine ideology dimensions) come from the external expert-coded V-Party dataset, and the authors explicitly avoid LLM-generated ideology scores for this reason: 'Relying on LLM-generated ideology scores would be circular; the models whose biases we aim to disentangle would partly define the categories against which those biases are measured. Our approach severs this dependency...' The outcome variables are LLM neutrality ratings, measured from article text via a stated prompt; the regressions (Eqs. 1-2) are descriptive fits, not predictions re-derived from fitted parameters. The neutrality construct is admittedly conditional ('all neutrality statements in this work are conditional on our methodology'), and a robustness test removing the explicit definition preserves the direction of ratings (86.25% exact agreement), so the Wikipedia-NPOV-inspired prompt does not make the Wikipedia-vs-Grokipedia comparison true by definition. The only self-citation is [7], used to support the minor claim that DeepSeek's distinctiveness lies on the geopolitical dimension; this is not load-bearing for the central results. The limitation 'it remains to be verified whether the LLM ratings are aligned with human annotators' is an unvalidated measurement concern (correctness risk), not a form of circularity, and does not reduce any derived quantity to its inputs.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

No invented entities. The central result rests on two modeling choices (year matching, length threshold) and two domain assumptions: transfer of party ideology to individuals and validity of LLM neutrality judgments. The first is partially validated by sampling; the second is explicitly unverified.

free parameters (2)
  • Ideology year-matching rule = nearest V-Party election-year score to most recent government spell
    Party-level expert scores are assigned to each politician using the closest election year; a different year rule could shift ideology scores for politicians who changed parties or positions.
  • Article length minimum = 300 words
    Drops 180 article pairs to 'ensure enough content'; could bias the sample toward longer, more detailed profiles and away from minor politicians.
assumptions (4)
  • domain assumption Party-level V-Party ideology scores can be attributed to individual members of government.
    Stated in §3.1: members of government are prominent party actors and party ideologies are stable; the paper acknowledges individuals may diverge but assumes errors are unsystematic.
  • domain assumption LLM neutrality ratings are a valid operationalization of article neutrality without human calibration.
    The outcome variable in every analysis is an LLM rating; §6 Limitations says alignment with human annotators 'remains to be verified.' If this fails, the central comparison fails.
  • domain assumption WhoGov-to-article matching recovers the correct politician pages in both encyclopedias.
    Heuristic title resolution and Grokipedia URL slug matching with manual verification on 100 samples; 1% mismatch in retained sample, but not exhaustive.
  • standard math OLS with clustered standard errors is applicable to the ordered 5-point ratings.
    Standard modeling assumption for the regressions; no violation diagnostics beyond VIF are reported.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Grokipedia vs Wikipedia: An LLM-Based Audit of Political Neutrality along Ideologies." pith.science (2026). https://pith.science/paper/CDZ5XTDY

@misc{pith2026260715146,
  author       = {Pith},
  title        = {Pith review of: Grokipedia vs Wikipedia: An LLM-Based Audit of Political Neutrality along Ideologies},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CDZ5XTDY}},
  note         = {Machine review of arXiv:2607.15146}
}
read the original abstract

Online encyclopedias shape political opinion and, through it, democratic discourse. In late 2025, Grokipedia was released, an encyclopedia written entirely by the LLM Grok. One motivation behind the project was to provide an unbiased alternative to Wikipedia, which has faced accusations of "left-wing" and "liberal" bias. But does an encyclopedia written by an LLM deliver greater neutrality, or does it simply embed a different ideology? We conduct a large-scale political bias study on Grokipedia and Wikipedia, analysing 1,394 article pairs describing members of government for neutrality along nine expert-coded ideology dimensions employing four LLM judges, Grok, Claude, Mistral, and DeepSeek. As the LLMs could themselves be biased, we also investigate patterns in their judgments. We find all LLM-judges, including Grok, to rate Grokipedia less neutral than Wikipedia. Both encyclopedias are rated as portraying politicians favourably overall, but towards different ideological groups. Grokipedia particularly favours economically right-wing politicians and penalises socially liberal ones, while Wikipedia is rated as favourably biased towards the latter.

Figures

Figures reproduced from arXiv: 2607.15146 by the authors.

Figure 1
Figure 1. Methodology overview: members of government (WhoGov, 20162023) are mapped to expert-coded V-Party ideology scores, their Wikipedia and Grokipedia ar￾ticles are rated for neutrality by four LLM judges, and the effects of ideology on the ratings are estimated via OLS regression. with future LLMs potentially being trained on its contents. More generally, the examples of Grokipedia and other recent LLM encyclopedias [35… view at source ↗
Figure 2
Figure 2. Assessments grouped by (a) source across LLM-judges and (b) LLM-judge across sources. Absolute Bias Ratings [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. OLS coefficients for the neutrality differential ∆y; error bars denote 95% confi￾dence intervals. A positive coefficient indicates that Wikipedia portrays politicians with that ideology more positively (or less negatively), whereas a negative coefficient indi￾cates the reverse. The intercept is interpreted as the expected Wikipedia-Grokipedia neutrality gap for a member of government with the average ideology in the… view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: OLS coefficients for neutrality assessment y; error bars denote 95% confidence intervals. ernment are also associated with a negative portrayal. In Grokipedia, ideological characteristics explain a substantially larger share of variance than in Wikipedia (R2 = 0.220). …
Figure 5
Figure 5. Figure 5: Heatmap of country-origin of the members of government in the final dataset. and India (62) being the most represented. Countries with multiple government changes in the time period examined (2016-2023) will tend to be over-represented as they will have had more member…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

45 extracted references · 6 linked inside Pith

  1. [1]

    PS: Political Science & Politics 55(2), 429–433 (2022)

    Ackerly, B.A., Michelitch, K.: Wikipedia and Political Science: Addressing System- atic Biases with Student Initiatives. PS: Political Science & Politics 55(2), 429–433 (2022)

  2. [2]

    Adams, J., Clark, M., Ezrow, L., Glasgow, G.: Understanding Change and Stability in Party Ideologies: Do Parties Respond to Public Opinion or to Past Election Results? British Journal of Political Science 34(4), 589610 (2004)

  3. [3]

    https://www.anthropic.com/ne ws/political-even-handedness (November 2025), accessed: June 12, 2026 Grokipedia vs Wikipedia: An LLM-Based Audit of Political Neutrality 15

    Anthropic: Measuring Political Bias in Claude. https://www.anthropic.com/ne ws/political-even-handedness (November 2025), accessed: June 12, 2026 Grokipedia vs Wikipedia: An LLM-Based Audit of Political Neutrality 15

  4. [4]

    https://www.anthropic.co m/claude/opus (Feb 2026), accessed: April 16, 2026

    Anthropic: Claude opus 4.6 [Large language model]. https://www.anthropic.co m/claude/opus (Feb 2026), accessed: April 16, 2026

  5. [5]

    Computational Linguistics 34(4), 555–596 (2008)

    Artstein, R., Poesio, M.: Survey Article: Inter-Coder Agreement for Computational Linguistics. Computational Linguistics 34(4), 555–596 (2008)

  6. [6]

    arXiv preprint arXiv:2509.08853 (2025)

    Azzopardi, L., Moshfeghi, Y.: POW: Political Overton Windows of Large Language Models. arXiv preprint arXiv:2509.08853 (2025)

  7. [7]

    npj Artificial Intelligence 2(1), 7 (2026)

    Buyl, M., Rogiers, A., Noels, S., Bied, G., Dominguez-Catena, I., Heiter, E., Johary, I., Mara, A.C., Romero, R., Lijffijt, J., et al.: Large language models reflect the ideology of their creators. npj Artificial Intelligence 2(1), 7 (2026)

  8. [8]

    In: World Conference on Explainable Artificial Intelli- gence

    Cahlik, V., Alves, R., Kordik, P.: Reasoning-grounded natural language explana- tions for language models. In: World Conference on Explainable Artificial Intelli- gence. pp. 3–18. Springer (2025)

Show all 45 references
  1. [9]

    arXiv preprint arXiv:2601.08785 (2026)

    Chen, J., de Jong, K., Poole, A., Burakowski, J., Nosti, E.E., Windt, J., Wang, C.: Uncovering political bias in large language models using parliamentary voting records. arXiv preprint arXiv:2601.08785 (2026)

  2. [10]

    IEEE Access 13, 11341– 11379 (2024)

    Choudhary, T.: Political Bias in Large Language Models: A Comparative Analysis of ChatGPT-4, Perplexity, Google Gemini, and Claude. IEEE Access 13, 11341– 11379 (2024)

  3. [11]

    https://api-docs.deepseek .com/ (Apr 2026), accessed: May 30, 2026

    DeepSeek: Deepseek v4-pro [Large language model]. https://api-docs.deepseek .com/ (Apr 2026), accessed: May 30, 2026

  4. [12]

    In: World Conference on Explainable Artificial Intelligence

    Dormuth, I., Franke, S., Hafer, M., Katzke, T., Marx, A., Müller, E., Neider, D., Pauly, M., Rutinowski, J.: A Cautionary Tale About Neutrally Informative AI Tools Ahead of the 2025 Federal Elections in Germany. In: World Conference on Explainable Artificial Intelligence. pp. 6...

  5. [13]

    The Guardian (January 2026), https://www.theguardian.com/technolo gy/2026/jan/24/latest-chatgpt-model-uses-elon-musks-grokipedia-as-sou rce-tests-reveal , accessed: 2026-04-28

    Down, A.: Latest ChatGPT model uses Elon Musk’s Grokipedia as source, tests reveal. The Guardian (January 2026), https://www.theguardian.com/technolo gy/2026/jan/24/latest-chatgpt-model-uses-elon-musks-grokipedia-as-sou rce-tests-reveal , accessed: 2026-04-28

  6. [14]

    arXiv preprint arXiv:2601.15484 (2026)

    Eibl, P., Coppolillo, E., Mungari, S., Luceri, L.: Is Grokipedia Right-Leaning? Comparing Political Framing in Wikipedia and Grokipedia on Controversial Top- ics. arXiv preprint arXiv:2601.15484 (2026)

  7. [15]

    Garzia, D., Ferreira da Silva, F., De Angelis, A.: Partisan dealignment and the personalisation of politics in West European parliamentary democracies, 1961–

  8. [16]

    Greenstein, S., Zhu, F.: Is Wikipedia biased? American Economic Review 102(3), 343–348 (2012)

  9. [17]

    arXiv preprint arXiv:2602.05519 (2026)

    Hadad, O., Loru, E., Nudo, J., Bonetti, A., Cinelli, M., Quattrociocchi, W.: Wikipedia and Grokipedia: A Comparison of Human and Generative Encyclo- pedias. arXiv preprint arXiv:2602.05519 (2026)

  10. [18]

    https://ww w.pcmag.com/news/traffic-to-elon-musks-grokipedia-tanks-after-initial -surge (Nov 2025), accessed: 2026-05-07

    Kan, M.: Traffic to Elon Musk’s Grokipedia tanks after initial surge. https://ww w.pcmag.com/news/traffic-to-elon-musks-grokipedia-tanks-after-initial -surge (Nov 2025), accessed: 2026-05-07

  11. [19]

    arXiv preprint arXiv:2601.05835 (2026)

    Kennedy, M., Parker, A., Liu, Y., Schütze, H.: Left, Right, or Center? Eval- uating LLM Framing in News Classification and Generation. arXiv preprint arXiv:2601.05835 (2026)

  12. [20]

    arXiv preprint arXiv:2603.23841 (2026)

    Khetan, R., Khetan, A.: PoliticsBench: Benchmarking Political Values in Large Language Models with Multi-Turn Roleplay. arXiv preprint arXiv:2603.23841 (2026)

  13. [21]

    Biometrics 33(1), 159–174 (1977) 16 F

    Landis, J.R., Koch, G.G.: The Measurement of Observer Agreement for Categorical Data. Biometrics 33(1), 159–174 (1977) 16 F. Vlahos et al

  14. [22]

    arXiv preprint arXiv:2412.05579 (2024)

    Li, H., Dong, Q., Chen, J., Su, H., Zhou, Y., Ai, Q., Ye, Z., Liu, Y.: LLMs-as- judges: A comprehensive survey on LLM-based evaluation methods. arXiv preprint arXiv:2412.05579 (2024)

  15. [23]

    Varieties of Democracy (V-Dem) Project (2022)

    Lindberg, S.I., Düpont, N., Higashijima, M., Berker Kavasoglu, Y., Marquardt, K.L., Bernhard, M., Döring, H., Hicken, A., Laebens, M., Medzinhorsky, J., et al.: Codebook varieties of party identity and organization (V–party) v2. Varieties of Democracy (V-Dem) Project (2022)

  16. [24]

    arXiv preprint arXiv:2512.03337 (2025)

    Mehdizadeh, A., Hilbert, M.: Epistemic substitution: How Grokipedia’s AI- Generated Encyclopedia Restructures Authority. arXiv preprint arXiv:2512.03337 (2025)

  17. [25]

    European Journal of Political Research 64(4), 1668–1692 (2025)

    Meijers, M.J., Dassonneville, R.: Who accepts party policy change? The individual- level drivers of attitudes towards party repositioning. European Journal of Political Research 64(4), 1668–1692 (2025)

  18. [26]

    https://huggingface

    Mistral AI: Mistral-medium-3.5 [Large language model]. https://huggingface. co/mistralai/Mistral-Medium-3.5-128B (2025), accessed: June 11, 2026

  19. [27]

    Stop donating to Wokepedia until they restore balance to their editing authority

    Musk, E.: Post on X (formerly Twitter). https://x.com/elonmusk (December 2024), “Stop donating to Wokepedia until they restore balance to their editing authority. ” Reported in: Newsweek, 27 December 2024, https://www.newsweek.c om/elon-musk-takes-aim-wikipedia-fund-raising-ed...

  20. [28]

    Since legacy media propaganda is considered a ‘valid’ source by Wikipedia, it naturally simply becomes an extension of legacy media propaganda

    Musk, E.: Post on X (formerly Twitter). https://x.com/elonmusk (2025), “Since legacy media propaganda is considered a ‘valid’ source by Wikipedia, it naturally simply becomes an extension of legacy media propaganda. ” Reported in: BBC Science Focus, 27 October 2025, https://ww...

  21. [29]

    American Political Science Review 114(4), 1366–1374 (2020)

    Nyrup, J., Bramwell, S.: Who governs? A new global dataset on members of cabi- nets. American Political Science Review 114(4), 1366–1374 (2020)

  22. [30]

    Journal of Information Technology & Politics pp

    Peng, T.Q., Yang, K., Lee, S., Li, H., Chu, Y., Lin, Y., Liu, H.: Beyond partisan leaning: A comparative analysis of political bias in large language models. Journal of Information Technology & Politics pp. 1–18 (2026)

  23. [31]

    https://www.pe wresearch.org/short-reads/2026/01/13/wikipedia-at-25-what-the-data-t ells-us/ (2026), accessed: 2026-05-06

    Pew Research Center: Wikipedia at 25: What the data tells us. https://www.pe wresearch.org/short-reads/2026/01/13/wikipedia-at-25-what-the-data-t ells-us/ (2026), accessed: 2026-05-06

  24. [32]

    OUP Oxford (2007)

    Poguntke, T., Webb, P.: The presidentialization of politics: A comparative study of modern democracies. OUP Oxford (2007)

  25. [33]

    Rozado, D.: Is Wikipedia Politically Biased? Report, Manhattan Institute (Jun 2024), https://manhattan.institute/article/is-wikipedia-politically-b iased, accessed: 2026-05-08

  26. [34]

    PLOS ONE 19(7), 1–15 (07 2024)

    Rozado, D.: The political preferences of LLMs. PLOS ONE 19(7), 1–15 (07 2024)

  27. [35]

    arXiv preprint arXiv:2603.24080 (2026)

    Saeed, M., Razniewski, S.: LLMpedia: A Transparent Framework to Materialize an LLM’s Encyclopedic Knowledge at Scale. arXiv preprint arXiv:2603.24080 (2026)

  28. [36]

    The Journal of Politics 71(1), 238–248 (2009)

    Somer-Topcu, Z.: Timely decisions: The effects of past national elections on party policy change. The Journal of Politics 71(1), 238–248 (2009)

  29. [37]

    arXiv preprint arXiv:2511.09685 (2025)

    Triedman, H., Mantzarlis, A.: What did Elon change? A comprehensive analysis of Grokipedia. arXiv preprint arXiv:2511.09685 (2025)

  30. [38]

    Does it matter how they match? The German General Elections 2009 in comparison

    Wagner, A., Weßels, B.: Parties and their leaders. Does it matter how they match? The German General Elections 2009 in comparison. Electoral Studies 31(1), 72–82 (2012)

  31. [39]

    Advances in Neural Information Processing Systems 35, 24824–24837 (2022) Grokipedia vs Wikipedia: An LLM-Based Audit of Political Neutrality 17

    Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q.V., Zhou, D., et al.: Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems 35, 24824–24837 (2022) Grokipedia vs Wikipedia: An LLM-Based Aud...

  32. [40]

    British Journal of Political Science 55, e174 (2025)

    Werner, A., Habersack, F.: Parties ideological cores and peripheries: Examining how parties balance adaptation and continuity in their manifestos. British Journal of Political Science 55, e174 (2025)

  33. [41]

    https://grok.com/ (2025), accessed: April 16, 2026

    xAI: Grok-4 [Large language model]. https://grok.com/ (2025), accessed: April 16, 2026

  34. [42]

    arXiv preprint arXiv:2510.26899 (2025)

    Yasseri, T., Mohammadi, S.: How Similar Are Grokipedia and Wikipedia? A Multi- Dimensional Textual and Structural Comparison. arXiv preprint arXiv:2510.26899 (2025)

  35. [43]

    arXiv preprint arXiv:2410.02736 (2024) A Appendix The supplementary material is organized as follows

    Ye, J., Wang, Y., Huang, Y., Chen, D., Zhang, Q., Moniz, N., Gao, T., Geyer, W., Huang, C., Chen, P.Y., et al.: Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge. arXiv preprint arXiv:2410.02736 (2024) A Appendix The supplementary material is organized as follows. App...

  36. [2018]

    West European Politics 45(2), 311–334 (2022)

  37. [2026]

    politician

    For each politician name in our input list, we collect a parallel pair of articles: one from Wikipedia (English) , and one from Grokipedia V0.2 . Collection is parallelized across entities (5 workers) with rate limiting (100 re- quests/min, ≥ 0.5 s between requests), exponenti...

Pith tools

Reviewed August 1, 2026 · model on record in the stance chip above.