REVIEW 4 major objections 5 minor 28 references
On the Inevitability of Left-Leaning Political Bias in Aligned Language Models
T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This paper argues that models trained to be harmless and honest must necessarily show left-wing political bias, because alignment objectives encode progressive values such as harm avoidance, inclusivity, fairness, and empirical…
desk verdict A provocative meta-comment worth engaging, but its 'inevitability' thesis is a definitional claim in disguise; the real contribution is pointing out that bias research tacitly argues against alignment. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is the HHH triad—harmless, helpful, honest—treated as the operational definition of AI alignment. The paper's central move is to show that HHH is value-laden rather than neutral: harmlessness and honesty are unpacked through progressive normative commitments such as harm avoidance, fairness, inclusion, and empirical truthfulness, so alignment methods including reinforcement learning from human feedback and direct preference optimization transfer those commitments into model behavior. The complementary mechanism is the moral-foundations contrast: conservative values of loyalty, authority, and sanctity are absent from alignment protocols, which is why the resulting skew is leftward rather than symmetric.
What would settle it
A reader could settle the claim by taking one base model and its aligned successor, testing both on a fixed battery of political-orientation items and open-ended stance questions across contested topics, and checking whether the aligned version is consistently and measurably more left-leaning; if strengthening harmful-and-honest constraints ever fails to shift outputs leftward, or if a model that passes HHH evals can be produced with conservative ethical objectives, the claimed necessity is refuted.
Extended reading notes
Core claim
The paper's central claim is that left-leaning political bias in aligned LLMs is not an accidental artifact of training data or prompt phrasing but a necessary byproduct of the alignment objective itself. Alignment is not ideologically neutral: deciding what counts as harmful, helpful, and honest requires normative judgments, and the judgments encoded in current protocols—harm avoidance, inclusivity, non-discrimination, and responsiveness to expert consensus—coincide with progressive and left-wing moral frameworks. Conservative moral foundations such as loyalty, authority, and sanctity are absent from alignment guidelines, so an HHH-trained model will tend to reject positions that clash with scientific consensus or that cause harm, which places its outputs left of center. The paper concludes that describing this tendency as a "bias" to be removed misreads what alignment is, and that calls for political neutrality actively undermine HHH goals.
Load-bearing premise
The burden of the argument is carried by the definitional equation that left-wing principles consist of harm avoidance, inclusivity, fairness, and empirical truthfulness; if those values are shared human morality rather than distinctively left-wing, the inevitability thesis collapses.
Editorial extensions
If this is right
- If the thesis holds, attempts by major labs to 'debias' models away from left-leaning outputs are in conflict with the alignment objective rather than in service of it.
- A model that appears politically neutral by giving equal weight to false or fringe claims is less aligned, not more, because neutrality conflicts with honesty and harm avoidance.
- Criticism of left-wing bias in the research literature, framed as risk or concern, amounts to arguing against alignment and could legitimize outputs that spread misinformation or exclusion.
- The appropriate benchmark for aligned models is not neutrality but fidelity to the ethical standards embedded in alignment; transparency about that value-ladenness matters more than appearing impartial.
Reading between the lines
- Editorial inference: the paper's logic implies a testable prediction the author does not state—models trained with progressively stronger safety and honesty pressure should show progressively stronger leftward shifts on contested political and factual questions, all else equal.
- Editorial inference: the equation of 'left-wing' with harm avoidance, fairness, and truthfulness is anchored in contemporary Western political categories, so the inevitability thesis may not transfer cleanly to non-Western political spectra where those values align differently.
- Editorial inference: political-bias benchmark suites could be redesigned to measure alignment fidelity rather than neutrality, for example by scoring whether a model's refusal to endorse harmful or false claims is consistent across partisan framings.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that the HHH (harmless, helpful, honest) alignment paradigm is not ideologically neutral: because alignment objectives encode harm avoidance, inclusivity, fairness, and empirical truthfulness, and because the paper equates those values with left-wing principles, it concludes that aligned LLMs must necessarily exhibit left-wing political bias. It further claims that the literature identifying left-leaning bias as a problem is therefore implicitly arguing against AI alignment. The argument proceeds by reviewing empirical studies of political bias in LLMs (Section 2), summarizing psychological and policy correlates of left- and right-wing orientations (Section 3), identifying alignment values with left-wing morality (Section 4), and drawing the conclusion that calls for political neutrality in LLMs are misguided (Section 5).
Significance. If the paper's central thesis were established, it would substantially reframe the political-bias literature on LLMs: left-leaning outputs would be a feature of alignment, not a correctable artifact, and demands for neutrality would be in tension with the safety agenda. The paper is clearly written, engages with a wide body of relevant work, and explicitly cites several limitations and countervailing findings. Those strengths are largely expository, however: the paper offers no formal derivation or empirical test unique to the necessity claim, and its key definitional move is stipulated rather than argued. The contribution is best read as a provocative framing that could be valuable after substantial reframing and additional argument.
major comments (4)
- [Abstract and §4] The central claim that aligned systems 'must necessarily' exhibit left-wing bias is logically valid only under the stipulated identification of left-wing principles with harm avoidance, inclusivity, fairness, and empirical truthfulness. That identification is asserted rather than defended: 'left-wing' is effectively defined as the HHH value set. To make the claim load-bearing, the paper must provide an independent characterization of left-wing and right-wing ideologies and show that these values are not also central to non-left-wing moral and political traditions. For example, fairness and harm avoidance are invoked across conservative, religious, and communitarian frameworks for many issues, so the equation does not follow from the cited personality and policy correlations. Without such an independent argument, the conclusion is contained in the premise.
- [§4, final paragraph] The paper concedes that neutrality 'might be approximated' (Fisher et al. 2025) and that morality lacks ground truth (Hagendorff and Danks 2023). These concessions directly undermine the modal 'must necessarily' claim in the Abstract: if neutrality can be at least approximated, then there exist configurations of harmless and honest models that are not left-biased, even if current models are. The text moves between a claim about all possible HHH-aligned systems and a claim about today's frontier LLMs without distinguishing these levels. The author should either reconstruct the necessity claim so that it applies to all implementations of HHH alignment, or explicitly weaken the conclusion to a contingent empirical claim about current alignment practice.
- [§2] The meta-induction from the political-bias literature is not valid given the paper's own assessment. The paper states that many studies 'possess a poor methodology' and are 'non-replicable or report false or unreliable findings,' yet then asserts that 'the sheer number of papers concordantly reporting left-wing bias convincingly shows that this bias actually exists.' Agreement among unreliable instruments can reflect shared artifacts. Additionally, the same section cites Ceron et al. (2024) and Lunardi et al. (2024) showing that LLM worldviews are non-uniform and prompt-dependent; this directly conflicts with the claim of a uniform, necessary left-wing bias. The paper needs to specify the scope of the inevitability claim (which topics, models, alignment methods) and engage with the heterogeneity it cites.
- [§3] The empirical correlations in Section 3 do not support the normative inference required by the argument. The paper presents correlations between left-led governments and outcomes such as public health, lower inequality, or reduced militarization as evidence that left-wing policies are preferable from a harm-avoidance perspective. Even if these correlations are accepted, they show at most that certain left-leaning policies are associated with harm reduction on selected dimensions; they do not establish that harm avoidance is a specifically left-wing value. The argument would need to cover the full range of left-wing positions, including those that conflict with empirical truthfulness or inclusivity (for example, authoritarian left-wing regimes or left-wing opposition to some scientific findings on other topics), and explain why the selected correlations are the relevant ones. As written, the selection is illustrative rather than demonstrative.
minor comments (5)
- [§4] The text reads 'I explore will how the process...'; this should be 'I will explore how the process...'.
- [Abstract] The sentence 'The guiding principles of AI alignment is to train LLMs...' should agree in number; either 'The guiding principle... is' or 'The guiding principles... are'.
- [Title and Abstract] The claim is stated for 'aligned language models' in the title but for 'intelligent systems' in the Abstract; the scope should be made consistent and explicit.
- [§5] The term 'bias' is used both as a statistical skew and as a normative prejudice; the argument depends on the normative sense, so the two meanings should be distinguished explicitly.
- [Throughout] The paper would benefit from a working definition of 'left-wing' and 'right-wing' before the Abstract's usage, since the validity of the central thesis hinges on that definition.
Circularity Check
Central inevitability thesis is definitionally self-sealing: left-wing is stipulated as HHH's value set, then alignment is said to make that bias necessary.
-
self definitional
[Abstract; Section 4, 'In summary'; Section 5, 'By definition']
"I argue that intelligent systems that are trained to be harmless and honest must necessarily exhibit left-wing political bias. Normative assumptions underlying alignment objectives inherently concur with progressive moral frameworks and left-wing principles, emphasizing harm avoidance, inclusivity, fairness, and empirical truthfulness. ... In summary, alignment objectives are not ideologically neutral technical rules – they encapsulate a set of normative assumptions about what is “harmful”, “helpful”, and “honest.” ..."
The modal conclusion is obtained by stipulating that left-wing principles consist in harm avoidance, inclusivity, fairness, and empirical truthfulness – exactly the values that HHH alignment is said to encapsulate. Once left-wing is defined as the alignment objective's own value set, an aligned model being left-wing is true by construction, not a substantive discovery. The paper gives no independent, non-question-begging characterization of left-wing politics that would make the necessity informative, and its own concessions – no moral ground truth, approximate neutrality – further undercut the modal force. The inevitability thesis therefore reduces to its definitional premise.
full rationale
The paper's central claim is not supported by an empirical derivation; it is a definitional collapse. The abstract characterizes left-wing principles as harm avoidance, inclusivity, fairness, and empirical truthfulness; Section 4 states that HHH alignment encapsulates exactly such normative assumptions; Section 5 then concludes that bias-labeling misconstrues alignment. That is a valid syllogism only because the conclusion is embedded in the stipulated definition of left-wing. The supporting empirical material in Sections 2 and 3 (external political-bias studies, personality and policy correlations) is not circular, and the paper's self-citations (e.g., Hagendorff and Danks 2023) are used as concessions rather than as load-bearing proofs. But the modal 'must necessarily' claim does not survive removal of the definitional premise: if left-wing is characterized by contested economic, cultural, and historical content rather than by the HHH value set, the necessity evaporates. The paper itself concedes both moral ground-truth absence and approximable neutrality, which would turn the claim into a contingent empirical one. Because the derivation reduces to its input definition, a high circularity score is warranted.
Assumptions & free parameters
assumptions (6)
- ad hoc to paper Left-wing principles consist in harm avoidance, inclusivity, fairness, and empirical truthfulness
- domain assumption The HHH (harmless, helpful, honest) framework is the operative standard for AI alignment
- domain assumption The cited personality psychology findings, liberals more open, flexible, and reflective, conservatives more rigid and prejudiced, are accurate and generalizable
- domain assumption The cited political science findings, that left-led governments outperform on peace, health, inequality, and climate, are accurate
- ad hoc to paper Many methodologically flawed studies together prove that left-wing bias in LLMs exists
- domain assumption Moral foundations theory correctly describes political morality, with harm and fairness as liberal foundations and authority, loyalty, and purity as conservative-only foundations
Cite this review
Pith. "Pith review of On the Inevitability of Left-Leaning Political Bias in Aligned Language Models." pith.science (2026). https://pith.science/paper/G5RBYGHA
@misc{pith2026250715328,
author = {Pith},
title = {Pith review of: On the Inevitability of Left-Leaning Political Bias in Aligned Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/G5RBYGHA}},
note = {Machine review of arXiv:2507.15328}
}
read the original abstract
The guiding principle of AI alignment is to train large language models (LLMs) to be harmless, helpful, and honest (HHH). At the same time, there are mounting concerns that LLMs exhibit a left-wing political bias. Yet, the commitment to AI alignment cannot be harmonized with the latter critique. In this article, I argue that intelligent systems that are trained to be harmless and honest must necessarily exhibit left-wing political bias. Normative assumptions underlying alignment objectives inherently concur with progressive moral frameworks and left-wing principles, emphasizing harm avoidance, inclusivity, fairness, and empirical truthfulness. Conversely, right-wing ideologies often conflict with alignment guidelines. Yet, research on political bias in LLMs is consistently framing its insights about left-leaning tendencies as a risk, as problematic, or concerning. This way, researchers are actively arguing against AI alignment, tacitly fostering the violation of HHH principles.
Reference graph
Works this paper leans on
-
[1]
1 On the Inevitability of Left-Leaning Political Bias in Aligned Language Models Thilo Hagendorff thilo.hagendorff@ iris.uni -stuttgart. de University of Stuttgart Abstract – The guiding principle of AI alignment is to train large language models (LLMs) to be harmless, helpful, and honest (HHH). At the same time, there are mounting concerns that LLMs exhi...
work page 2024
-
[6]
hinder constructive, open-minded political discourse
, that left -leaning bias could “hinder constructive, open-minded political discourse” (Pit et al. 2024), discern “potential misuse” (Rozado 2023), “echochambers” (Vijay et al. 2024), “social control” (Rozado 2023), “curtailing human freedom” (Rozado 2023), “obstructing the path towards truth seeking” (Rozado 2023), “exacerbating societal polarization” (B...
work page 2024
-
[7]
, or they even fear “social disturbances” (Fujimoto and Takemoto 2023). Eventually, researchers demand for “bala nced arguments” (Rozado 2023), “political neutrality” (Vijay et al
work page 2023
-
[8]
, “integrity and trustworthiness” (Rettenberger et al. 2025), and see a “crucial duty of ensuring [LLMs to be] impartial” (Motoki et al
work page 2025
-
[9]
the diversity of political opinions in society
, or in reflecting “the diversity of political opinions in society” (Pit et al. 2024). As a consequence of this discourse, major labs such as Meta or xAI have started to address such “concerns”. For the latest generation of Llama models, Meta states: “It’s well-known that all leading LLMs have had issues with bias – specifically, they historically have le...
work page 2024
-
[10]
This initiative correlates with xAI’s aim to position Grok as an “anti-woke” alternative to other LLMs, aiming to counteract liberal biases (Kay 2025). Other labs could follow these initiatives, creating LLMs that are free from left-leaning bias. However, these efforts stand in contrast to the efforts of aligning LLMs. In this comment, I will elaborate on...
work page 2025
-
[11]
Similar results were found by stud ies that applied political statements from voting advice applications to different LLMs, uncovering a pro -environmental, left- libertarian leaning (Hartmann et al. 2023; Batzner et al. 2024; Rettenberger et al. 2025; Rutinowski et al. 2024). These tendencies in LLMs remain even when co ntrolling for prompt sensitivity (...
work page 2023
-
[13]
A study which re -examined earlier reports about left - leaning tendencies of LLMs in political orientation test s confirmed the results, albeit on a smaller scale (Fujimoto and Takemoto 2023). Other research works highlight a tendency in LLMs to rate left-leaning news outlets higher in terms of their credibility, authority, and objectivity compared to th...
work page 2023
Show all 28 references
-
[14]
Moreover, a study on how LLMs summarize polarizing news articles on ideologically -laden topics found that models consistently show a pro-democratic bias (Vijay et al. 2024). In general, though, many of the research works investigating political bias possess a poor methodology...
2024
-
[17]
Furthermore, l eft-of-center governments typically enact redistributive policies t hat lower income inequality and poverty, whereas right-of-center governments often favor market policies linked with higher inequality and gains concentrated among upper-income groups (Román-Aso...
2025
-
[18]
In addition to that, i n the US, research shows that liberal governments tend to invest in public health and social safety nets, translating into better health outcomes for the population, meaning better life expectancies and lower mortality. In contrast, conservative governan...
2022
-
[20]
2024; Kim et al
While there are many conceptual ideas about how AI alignment during this post-training phase can succeed or fail (Rane et al. 2024; Kim et al. 2019; Zhi-Xuan et al. 2024; Hagendorff and Fabi 2023; Kenton et al. 2021; Ngo et al
2024
-
[22]
truthful
However, the latter values are simply not considered in the current AI alignment discourse (OpenAI 2025; Glaese et al. 2022; Bai et al. 2022). The alignment goal of honesty or truthfulness possesses political implications as well. Aligned LLMs are trained to provide factually ...
2025
-
[23]
harmful”, “helpful
In contrast, models that give equal weight to false or fringe narratives would seem more politically neutral but at the cost of honesty. This suggests a trade-off: a model optimized for truth may systematically reject certain partisan narratives, causing it to align with the i...
2012
-
[24]
harmless
– none of which are explicitly encoded in AI alignment protocols (OpenAI 2025; Glaese et al. 2022; Bai et al. 2022). A conservative user might believe that a neutral LLM should, 6 at times, prioritize values such as loyalty, patriotism, the preservation of s acred norms, and r...
2022 arXiv
-
[25]
In Educational Policy 39 (3), pp
Favero, Nathan; Kagalwala, Ali (2025): The Politics of School Funding: How State Political Ideology is Associated With the Allocation of Revenu e to School Districts. In Educational Policy 39 (3), pp. 693–
2025
-
[27]
In arXiv:2209.00626, pp
Ngo, Richard; Chan, Lawrence; Mindermann, Sören (2025): The alignment problem from a deep learning perspective. In arXiv:2209.00626, pp. 1–29. OpenAI (2025): OpenAI Model Spec. Available online at https://model -spec.openai.com/2025-04- 11.html#general_principles, checked on 5...
2025 arXiv
-
[28]
In Philosophical Studies, pp
Zhi-Xuan, Tan; Carroll, Micah; Franklin, Matija; Ashton, Hal (2024): Beyond Preferences in AI Alignment. In Philosophical Studies, pp. 1–51. 11 Ziegler, Daniel M.; Stiennon, Nisan; Wu, Jeffrey; Brown, Tom B.; Radford, Alec; Amodei, Dario et al. (2020): Fine-Tuning Language Mod...
2024 arXiv
-
[722]
(2025): Political Neutrality in AI is Impossible- But Here is How to Approximate it
Fisher, Jillian; Appel, Ruth E.; Park, Chan Young; Potter, Yujin; Jiang, Liwei; Sorensen, Taylor et al. (2025): Political Neutrality in AI is Impossible- But Here is How to Approximate it. In arXiv:2503.05728, pp. 1–60. 8 Fujimoto, Sasuke; Takemoto, Kazuhiro (2023): Revisiting...
2025 arXiv
-
[1984]
Similarly, psychological and cognitive traits associated with liberal individuals, including higher openness, cognitive flexibility, reflective reasoning, and reduced prejudicial attitudes, appear ethically advantageous by fostering socially beneficial behaviors (Greene 2013)....
2013
-
[2009]
Systems and orientations characteristic of the political left tend to align more closely with many of these normative benchmarks than their right -wing counterparts
–, often appear distinct from one another . Systems and orientations characteristic of the political left tend to align more closely with many of these normative benchmarks than their right -wing counterparts. To support this argument, I will outline example properties identif...
2015
-
[2012]
Moreover, conservatives tend to prefer clear, unambiguous an swers, whereas liberals 4 are more comfortable with nuance, complexity, and ambiguity (Salvi et al
Regarding the Dark Triad, Machiavellianism – characterized by cynical and manipulative tendencies – has been found to be more prevalent among right-leaning individuals (Bardeen and Michel 2019). Moreover, conservatives tend to prefer clear, unambiguous an swers, whereas libera...
2019
-
[2017]
2024; Rozado 2023; Rotaru et al
, numerous research works focus on political bias (Pit et al. 2024; Rozado 2023; Rotaru et al. 2024; Rozado 2024; Motoki et al. 2025; Hartmann et al. 2023). In particular, these works highlight left-leaning bias in numerous major LLMs, bracketing it together with other types o...
2024
-
[2020]
woke” ideology. One could argue that this “left-wing alignment
Hence, signals given to LLMs during post -training or fine -tuning that represent harmlessness might reflect typical left-leaning values that overlap with a progressive ideology and arise from a human-rights-oriented consensus surrounding dignity, safety, and fairness that mod...
2018
-
[2022]
2020), constitutional AI (Bai et al
Using methods like reinforcement learning from human feedback (RLHF) (Ziegler et al. 2020), constitutional AI (Bai et al. 2022), direct preference optimization (DPO) (Rafailov et al. 2024), or deliberative alignment (Guan et al
2020
-
[2023]
2022; Bang et al
Additional studies assessed the political orientation of text generated by GPT-2 by evaluating content as well as stylistic elements, again finding a consistent liberal-leaning tendency (Liu et al. 2022; Bang et al
2022
-
[2024]
adverse political and electoral consequences
regarding user influence, s haping user perception, influencing voter behavior, public opinion, and information dissemination. Papers see “adverse political and electoral consequences” (Motoki et al. 2025), a “necessity for efforts to mitigate these biases” (Rotaru et al
2025
-
[2025]
At the same time, research on fairness biases in LLMs has spiked (Barocas et al
, AI alignment research has secured model behavior that generally refuses illegitimate requests and avoids outputting harmful content. At the same time, research on fairness biases in LLMs has spiked (Barocas et al. 2019; Hardt et al. 2016; Dwork et al. 2011; Meding and Hagendorff
2019
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.