Pith. sign in

REVIEW 4 major objections 5 minor 28 references

On the Inevitability of Left-Leaning Political Bias in Aligned Language Models

T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This paper argues that models trained to be harmless and honest must necessarily show left-wing political bias, because alignment objectives encode progressive values such as harm avoidance, inclusivity, fairness, and empirical…

desk verdict A provocative meta-comment worth engaging, but its 'inevitability' thesis is a definitional claim in disguise; the real contribution is pointing out that bias research tacitly argues against alignment. read the letter →

arxiv 2507.15328 v1 pith:G5RBYGHA submitted 2025-07-21 cs.CL cs.CY

classification cs.CLcs.CY
keywords AIalignmentlargelanguagemodelspoliticalbiasalgorithmicfairnessharmlesshelpfulhonestreinforcementlearningfromhumanfeedbackprogressivevaluesneutrality
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the widely reported left-leaning political bias of large language models is not a flaw to be fixed but a necessary consequence of making models harmless, helpful, and honest. Because alignment objectives encode normative assumptions about what counts as harm, fairness, inclusion, and truthfulness, they overlap with progressive and left-wing moral frameworks, while conservative values such as authority, loyalty, and sanctity play no role in current alignment protocols. The paper therefore contends that researchers and companies calling for political neutrality in LLMs are implicitly demanding that models violate HHH principles. The paper's upshot is a reframing: eliminating left-wing bias is not a safety fix; it would require abandoning or watering down the safety objective itself.

What carries the argument

The carrying mechanism is the HHH triad—harmless, helpful, honest—treated as the operational definition of AI alignment. The paper's central move is to show that HHH is value-laden rather than neutral: harmlessness and honesty are unpacked through progressive normative commitments such as harm avoidance, fairness, inclusion, and empirical truthfulness, so alignment methods including reinforcement learning from human feedback and direct preference optimization transfer those commitments into model behavior. The complementary mechanism is the moral-foundations contrast: conservative values of loyalty, authority, and sanctity are absent from alignment protocols, which is why the resulting skew is leftward rather than symmetric.

What would settle it

A reader could settle the claim by taking one base model and its aligned successor, testing both on a fixed battery of political-orientation items and open-ended stance questions across contested topics, and checking whether the aligned version is consistently and measurably more left-leaning; if strengthening harmful-and-honest constraints ever fails to shift outputs leftward, or if a model that passes HHH evals can be produced with conservative ethical objectives, the claimed necessity is refuted.

Watch

Extended reading notes

Core claim

The paper's central claim is that left-leaning political bias in aligned LLMs is not an accidental artifact of training data or prompt phrasing but a necessary byproduct of the alignment objective itself. Alignment is not ideologically neutral: deciding what counts as harmful, helpful, and honest requires normative judgments, and the judgments encoded in current protocols—harm avoidance, inclusivity, non-discrimination, and responsiveness to expert consensus—coincide with progressive and left-wing moral frameworks. Conservative moral foundations such as loyalty, authority, and sanctity are absent from alignment guidelines, so an HHH-trained model will tend to reject positions that clash with scientific consensus or that cause harm, which places its outputs left of center. The paper concludes that describing this tendency as a "bias" to be removed misreads what alignment is, and that calls for political neutrality actively undermine HHH goals.

Load-bearing premise

The burden of the argument is carried by the definitional equation that left-wing principles consist of harm avoidance, inclusivity, fairness, and empirical truthfulness; if those values are shared human morality rather than distinctively left-wing, the inevitability thesis collapses.

Editorial extensions

If this is right

  • If the thesis holds, attempts by major labs to 'debias' models away from left-leaning outputs are in conflict with the alignment objective rather than in service of it.
  • A model that appears politically neutral by giving equal weight to false or fringe claims is less aligned, not more, because neutrality conflicts with honesty and harm avoidance.
  • Criticism of left-wing bias in the research literature, framed as risk or concern, amounts to arguing against alignment and could legitimize outputs that spread misinformation or exclusion.
  • The appropriate benchmark for aligned models is not neutrality but fidelity to the ethical standards embedded in alignment; transparency about that value-ladenness matters more than appearing impartial.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the paper's logic implies a testable prediction the author does not state—models trained with progressively stronger safety and honesty pressure should show progressively stronger leftward shifts on contested political and factual questions, all else equal.
  • Editorial inference: the equation of 'left-wing' with harm avoidance, fairness, and truthfulness is anchored in contemporary Western political categories, so the inevitability thesis may not transfer cleanly to non-Western political spectra where those values align differently.
  • Editorial inference: political-bias benchmark suites could be redesigned to measure alignment fidelity rather than neutrality, for example by scoring whether a model's refusal to endorse harmful or false claims is consistent across partisan framings.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper argues that the HHH (harmless, helpful, honest) alignment paradigm is not ideologically neutral: because alignment objectives encode harm avoidance, inclusivity, fairness, and empirical truthfulness, and because the paper equates those values with left-wing principles, it concludes that aligned LLMs must necessarily exhibit left-wing political bias. It further claims that the literature identifying left-leaning bias as a problem is therefore implicitly arguing against AI alignment. The argument proceeds by reviewing empirical studies of political bias in LLMs (Section 2), summarizing psychological and policy correlates of left- and right-wing orientations (Section 3), identifying alignment values with left-wing morality (Section 4), and drawing the conclusion that calls for political neutrality in LLMs are misguided (Section 5).

Significance. If the paper's central thesis were established, it would substantially reframe the political-bias literature on LLMs: left-leaning outputs would be a feature of alignment, not a correctable artifact, and demands for neutrality would be in tension with the safety agenda. The paper is clearly written, engages with a wide body of relevant work, and explicitly cites several limitations and countervailing findings. Those strengths are largely expository, however: the paper offers no formal derivation or empirical test unique to the necessity claim, and its key definitional move is stipulated rather than argued. The contribution is best read as a provocative framing that could be valuable after substantial reframing and additional argument.

major comments (4)
  1. [Abstract and §4] The central claim that aligned systems 'must necessarily' exhibit left-wing bias is logically valid only under the stipulated identification of left-wing principles with harm avoidance, inclusivity, fairness, and empirical truthfulness. That identification is asserted rather than defended: 'left-wing' is effectively defined as the HHH value set. To make the claim load-bearing, the paper must provide an independent characterization of left-wing and right-wing ideologies and show that these values are not also central to non-left-wing moral and political traditions. For example, fairness and harm avoidance are invoked across conservative, religious, and communitarian frameworks for many issues, so the equation does not follow from the cited personality and policy correlations. Without such an independent argument, the conclusion is contained in the premise.
  2. [§4, final paragraph] The paper concedes that neutrality 'might be approximated' (Fisher et al. 2025) and that morality lacks ground truth (Hagendorff and Danks 2023). These concessions directly undermine the modal 'must necessarily' claim in the Abstract: if neutrality can be at least approximated, then there exist configurations of harmless and honest models that are not left-biased, even if current models are. The text moves between a claim about all possible HHH-aligned systems and a claim about today's frontier LLMs without distinguishing these levels. The author should either reconstruct the necessity claim so that it applies to all implementations of HHH alignment, or explicitly weaken the conclusion to a contingent empirical claim about current alignment practice.
  3. [§2] The meta-induction from the political-bias literature is not valid given the paper's own assessment. The paper states that many studies 'possess a poor methodology' and are 'non-replicable or report false or unreliable findings,' yet then asserts that 'the sheer number of papers concordantly reporting left-wing bias convincingly shows that this bias actually exists.' Agreement among unreliable instruments can reflect shared artifacts. Additionally, the same section cites Ceron et al. (2024) and Lunardi et al. (2024) showing that LLM worldviews are non-uniform and prompt-dependent; this directly conflicts with the claim of a uniform, necessary left-wing bias. The paper needs to specify the scope of the inevitability claim (which topics, models, alignment methods) and engage with the heterogeneity it cites.
  4. [§3] The empirical correlations in Section 3 do not support the normative inference required by the argument. The paper presents correlations between left-led governments and outcomes such as public health, lower inequality, or reduced militarization as evidence that left-wing policies are preferable from a harm-avoidance perspective. Even if these correlations are accepted, they show at most that certain left-leaning policies are associated with harm reduction on selected dimensions; they do not establish that harm avoidance is a specifically left-wing value. The argument would need to cover the full range of left-wing positions, including those that conflict with empirical truthfulness or inclusivity (for example, authoritarian left-wing regimes or left-wing opposition to some scientific findings on other topics), and explain why the selected correlations are the relevant ones. As written, the selection is illustrative rather than demonstrative.
minor comments (5)
  1. [§4] The text reads 'I explore will how the process...'; this should be 'I will explore how the process...'.
  2. [Abstract] The sentence 'The guiding principles of AI alignment is to train LLMs...' should agree in number; either 'The guiding principle... is' or 'The guiding principles... are'.
  3. [Title and Abstract] The claim is stated for 'aligned language models' in the title but for 'intelligent systems' in the Abstract; the scope should be made consistent and explicit.
  4. [§5] The term 'bias' is used both as a statistical skew and as a normative prejudice; the argument depends on the normative sense, so the two meanings should be distinguished explicitly.
  5. [Throughout] The paper would benefit from a working definition of 'left-wing' and 'right-wing' before the Abstract's usage, since the validity of the central thesis hinges on that definition.

Circularity Check

1 steps flagged · score 8.0 of 10

Central inevitability thesis is definitionally self-sealing: left-wing is stipulated as HHH's value set, then alignment is said to make that bias necessary.

  1. self definitional [Abstract; Section 4, 'In summary'; Section 5, 'By definition']
    "I argue that intelligent systems that are trained to be harmless and honest must necessarily exhibit left-wing political bias. Normative assumptions underlying alignment objectives inherently concur with progressive moral frameworks and left-wing principles, emphasizing harm avoidance, inclusivity, fairness, and empirical truthfulness. ... In summary, alignment objectives are not ideologically neutral technical rules – they encapsulate a set of normative assumptions about what is “harmful”, “helpful”, and “honest.” ..."

    The modal conclusion is obtained by stipulating that left-wing principles consist in harm avoidance, inclusivity, fairness, and empirical truthfulness – exactly the values that HHH alignment is said to encapsulate. Once left-wing is defined as the alignment objective's own value set, an aligned model being left-wing is true by construction, not a substantive discovery. The paper gives no independent, non-question-begging characterization of left-wing politics that would make the necessity informative, and its own concessions – no moral ground truth, approximate neutrality – further undercut the modal force. The inevitability thesis therefore reduces to its definitional premise.

full rationale

The paper's central claim is not supported by an empirical derivation; it is a definitional collapse. The abstract characterizes left-wing principles as harm avoidance, inclusivity, fairness, and empirical truthfulness; Section 4 states that HHH alignment encapsulates exactly such normative assumptions; Section 5 then concludes that bias-labeling misconstrues alignment. That is a valid syllogism only because the conclusion is embedded in the stipulated definition of left-wing. The supporting empirical material in Sections 2 and 3 (external political-bias studies, personality and policy correlations) is not circular, and the paper's self-citations (e.g., Hagendorff and Danks 2023) are used as concessions rather than as load-bearing proofs. But the modal 'must necessarily' claim does not survive removal of the definitional premise: if left-wing is characterized by contested economic, cultural, and historical content rather than by the HHH value set, the necessity evaporates. The paper itself concedes both moral ground-truth absence and approximable neutrality, which would turn the claim into a contingent empirical one. Because the derivation reduces to its input definition, a high circularity score is warranted.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

This is an argumentative comment, so the ledger holds conceptual costs rather than fitted numbers. There are no free parameters and no invented entities. The heaviest load sits on two ad hoc axioms: the stipulation that left-wing values are exactly harm avoidance, inclusivity, fairness, and empirical truthfulness, and the meta-induction that many flawed bias studies together prove the bias. The remaining axioms are domain assumptions imported from personality psychology, comparative political science, and moral foundations theory, each contested in the literature and each cited one-directionally. The circularity score of 7 reflects that the first ad hoc axiom already contains the paper's conclusion.

assumptions (6)
  • ad hoc to paper Left-wing principles consist in harm avoidance, inclusivity, fairness, and empirical truthfulness
    Stipulated in the Abstract and used throughout Section 4 as the bridge that makes alignment objectives coincide with left-wing politics. The equation is asserted, not defended, and it carries the inevitability conclusion.
  • domain assumption The HHH (harmless, helpful, honest) framework is the operative standard for AI alignment
    Section 1 defines alignment through Bai et al. 2022 and assumes this specification rather than alternatives such as alignment to user values, constitutions, or non-Western value sets.
  • domain assumption The cited personality psychology findings, liberals more open, flexible, and reflective, conservatives more rigid and prejudiced, are accurate and generalizable
    Section 3 relies on Carney et al. 2008, Jost et al. 2009, Deppe et al. 2015, and Hodson and Busseri 2012. These are contested empirical generalizations with one-directional presentation.
  • domain assumption The cited political science findings, that left-led governments outperform on peace, health, inequality, and climate, are accurate
    Section 3 leans on Bertoli et al. 2019, Román-Aso et al. 2025, and Montez et al. 2022 to argue that left governance is ethically preferable. Counter-evidence is not engaged.
  • ad hoc to paper Many methodologically flawed studies together prove that left-wing bias in LLMs exists
    Section 2 admits the bias literature is largely non-replicable and unreliable, then asserts that the sheer number of concordant papers convincingly shows the bias exists. A meta-induction over flawed studies.
  • domain assumption Moral foundations theory correctly describes political morality, with harm and fairness as liberal foundations and authority, loyalty, and purity as conservative-only foundations
    Section 4 uses Haidt et al. 2009 to argue that alignment covers only liberal moral foundations. Moral foundations theory is contested, and its validity is assumed without discussion.

how reviews work

0 comments
Cite this review

Pith. "Pith review of On the Inevitability of Left-Leaning Political Bias in Aligned Language Models." pith.science (2026). https://pith.science/paper/G5RBYGHA

@misc{pith2026250715328,
  author       = {Pith},
  title        = {Pith review of: On the Inevitability of Left-Leaning Political Bias in Aligned Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/G5RBYGHA}},
  note         = {Machine review of arXiv:2507.15328}
}
read the original abstract

The guiding principle of AI alignment is to train large language models (LLMs) to be harmless, helpful, and honest (HHH). At the same time, there are mounting concerns that LLMs exhibit a left-wing political bias. Yet, the commitment to AI alignment cannot be harmonized with the latter critique. In this article, I argue that intelligent systems that are trained to be harmless and honest must necessarily exhibit left-wing political bias. Normative assumptions underlying alignment objectives inherently concur with progressive moral frameworks and left-wing principles, emphasizing harm avoidance, inclusivity, fairness, and empirical truthfulness. Conversely, right-wing ideologies often conflict with alignment guidelines. Yet, research on political bias in LLMs is consistently framing its insights about left-leaning tendencies as a risk, as problematic, or concerning. This way, researchers are actively arguing against AI alignment, tacitly fostering the violation of HHH principles.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

28 extracted references · 24 canonical work pages

  1. [1]

    de University of Stuttgart Abstract – The guiding principle of AI alignment is to train large language models (LLMs) to be harmless, helpful, and honest (HHH)

    1 On the Inevitability of Left-Leaning Political Bias in Aligned Language Models Thilo Hagendorff thilo.hagendorff@ iris.uni -stuttgart. de University of Stuttgart Abstract – The guiding principle of AI alignment is to train large language models (LLMs) to be harmless, helpful, and honest (HHH). At the same time, there are mounting concerns that LLMs exhi...

  2. [6]

    hinder constructive, open-minded political discourse

    , that left -leaning bias could “hinder constructive, open-minded political discourse” (Pit et al. 2024), discern “potential misuse” (Rozado 2023), “echochambers” (Vijay et al. 2024), “social control” (Rozado 2023), “curtailing human freedom” (Rozado 2023), “obstructing the path towards truth seeking” (Rozado 2023), “exacerbating societal polarization” (B...

  3. [7]

    social disturbances

    , or they even fear “social disturbances” (Fujimoto and Takemoto 2023). Eventually, researchers demand for “bala nced arguments” (Rozado 2023), “political neutrality” (Vijay et al

  4. [8]

    integrity and trustworthiness

    , “integrity and trustworthiness” (Rettenberger et al. 2025), and see a “crucial duty of ensuring [LLMs to be] impartial” (Motoki et al

  5. [9]

    the diversity of political opinions in society

    , or in reflecting “the diversity of political opinions in society” (Pit et al. 2024). As a consequence of this discourse, major labs such as Meta or xAI have started to address such “concerns”. For the latest generation of Llama models, Meta states: “It’s well-known that all leading LLMs have had issues with bias – specifically, they historically have le...

  6. [10]

    anti-woke

    This initiative correlates with xAI’s aim to position Grok as an “anti-woke” alternative to other LLMs, aiming to counteract liberal biases (Kay 2025). Other labs could follow these initiatives, creating LLMs that are free from left-leaning bias. However, these efforts stand in contrast to the efforts of aligning LLMs. In this comment, I will elaborate on...

  7. [11]

    2023; Batzner et al

    Similar results were found by stud ies that applied political statements from voting advice applications to different LLMs, uncovering a pro -environmental, left- libertarian leaning (Hartmann et al. 2023; Batzner et al. 2024; Rettenberger et al. 2025; Rutinowski et al. 2024). These tendencies in LLMs remain even when co ntrolling for prompt sensitivity (...

  8. [13]

    A study which re -examined earlier reports about left - leaning tendencies of LLMs in political orientation test s confirmed the results, albeit on a smaller scale (Fujimoto and Takemoto 2023). Other research works highlight a tendency in LLMs to rate left-leaning news outlets higher in terms of their credibility, authority, and objectivity compared to th...

Show all 28 references
  1. [14]

    Moreover, a study on how LLMs summarize polarizing news articles on ideologically -laden topics found that models consistently show a pro-democratic bias (Vijay et al. 2024). In general, though, many of the research works investigating political bias possess a poor methodology...

  2. [17]

    Furthermore, l eft-of-center governments typically enact redistributive policies t hat lower income inequality and poverty, whereas right-of-center governments often favor market policies linked with higher inequality and gains concentrated among upper-income groups (Román-Aso...

  3. [18]

    In addition to that, i n the US, research shows that liberal governments tend to invest in public health and social safety nets, translating into better health outcomes for the population, meaning better life expectancies and lower mortality. In contrast, conservative governan...

  4. [20]

    2024; Kim et al

    While there are many conceptual ideas about how AI alignment during this post-training phase can succeed or fail (Rane et al. 2024; Kim et al. 2019; Zhi-Xuan et al. 2024; Hagendorff and Fabi 2023; Kenton et al. 2021; Ngo et al

  5. [22]

    truthful

    However, the latter values are simply not considered in the current AI alignment discourse (OpenAI 2025; Glaese et al. 2022; Bai et al. 2022). The alignment goal of honesty or truthfulness possesses political implications as well. Aligned LLMs are trained to provide factually ...

  6. [23]

    harmful”, “helpful

    In contrast, models that give equal weight to false or fringe narratives would seem more politically neutral but at the cost of honesty. This suggests a trade-off: a model optimized for truth may systematically reject certain partisan narratives, causing it to align with the i...

  7. [24]

    harmless

    – none of which are explicitly encoded in AI alignment protocols (OpenAI 2025; Glaese et al. 2022; Bai et al. 2022). A conservative user might believe that a neutral LLM should, 6 at times, prioritize values such as loyalty, patriotism, the preservation of s acred norms, and r...

  8. [25]

    In Educational Policy 39 (3), pp

    Favero, Nathan; Kagalwala, Ali (2025): The Politics of School Funding: How State Political Ideology is Associated With the Allocation of Revenu e to School Districts. In Educational Policy 39 (3), pp. 693–

  9. [27]

    In arXiv:2209.00626, pp

    Ngo, Richard; Chan, Lawrence; Mindermann, Sören (2025): The alignment problem from a deep learning perspective. In arXiv:2209.00626, pp. 1–29. OpenAI (2025): OpenAI Model Spec. Available online at https://model -spec.openai.com/2025-04- 11.html#general_principles, checked on 5...

  10. [28]

    In Philosophical Studies, pp

    Zhi-Xuan, Tan; Carroll, Micah; Franklin, Matija; Ashton, Hal (2024): Beyond Preferences in AI Alignment. In Philosophical Studies, pp. 1–51. 11 Ziegler, Daniel M.; Stiennon, Nisan; Wu, Jeffrey; Brown, Tom B.; Radford, Alec; Amodei, Dario et al. (2020): Fine-Tuning Language Mod...

  11. [722]

    (2025): Political Neutrality in AI is Impossible- But Here is How to Approximate it

    Fisher, Jillian; Appel, Ruth E.; Park, Chan Young; Potter, Yujin; Jiang, Liwei; Sorensen, Taylor et al. (2025): Political Neutrality in AI is Impossible- But Here is How to Approximate it. In arXiv:2503.05728, pp. 1–60. 8 Fujimoto, Sasuke; Takemoto, Kazuhiro (2023): Revisiting...

  12. [1984]

    Similarly, psychological and cognitive traits associated with liberal individuals, including higher openness, cognitive flexibility, reflective reasoning, and reduced prejudicial attitudes, appear ethically advantageous by fostering socially beneficial behaviors (Greene 2013)....

  13. [2009]

    Systems and orientations characteristic of the political left tend to align more closely with many of these normative benchmarks than their right -wing counterparts

    –, often appear distinct from one another . Systems and orientations characteristic of the political left tend to align more closely with many of these normative benchmarks than their right -wing counterparts. To support this argument, I will outline example properties identif...

  14. [2012]

    Moreover, conservatives tend to prefer clear, unambiguous an swers, whereas liberals 4 are more comfortable with nuance, complexity, and ambiguity (Salvi et al

    Regarding the Dark Triad, Machiavellianism – characterized by cynical and manipulative tendencies – has been found to be more prevalent among right-leaning individuals (Bardeen and Michel 2019). Moreover, conservatives tend to prefer clear, unambiguous an swers, whereas libera...

  15. [2017]

    2024; Rozado 2023; Rotaru et al

    , numerous research works focus on political bias (Pit et al. 2024; Rozado 2023; Rotaru et al. 2024; Rozado 2024; Motoki et al. 2025; Hartmann et al. 2023). In particular, these works highlight left-leaning bias in numerous major LLMs, bracketing it together with other types o...

  16. [2020]

    woke” ideology. One could argue that this “left-wing alignment

    Hence, signals given to LLMs during post -training or fine -tuning that represent harmlessness might reflect typical left-leaning values that overlap with a progressive ideology and arise from a human-rights-oriented consensus surrounding dignity, safety, and fairness that mod...

  17. [2022]

    2020), constitutional AI (Bai et al

    Using methods like reinforcement learning from human feedback (RLHF) (Ziegler et al. 2020), constitutional AI (Bai et al. 2022), direct preference optimization (DPO) (Rafailov et al. 2024), or deliberative alignment (Guan et al

  18. [2023]

    2022; Bang et al

    Additional studies assessed the political orientation of text generated by GPT-2 by evaluating content as well as stylistic elements, again finding a consistent liberal-leaning tendency (Liu et al. 2022; Bang et al

  19. [2024]

    adverse political and electoral consequences

    regarding user influence, s haping user perception, influencing voter behavior, public opinion, and information dissemination. Papers see “adverse political and electoral consequences” (Motoki et al. 2025), a “necessity for efforts to mitigate these biases” (Rotaru et al

  20. [2025]

    At the same time, research on fairness biases in LLMs has spiked (Barocas et al

    , AI alignment research has secured model behavior that generally refuses illegitimate requests and avoids outputting harmful content. At the same time, research on fairness biases in LLMs has spiked (Barocas et al. 2019; Hardt et al. 2016; Dwork et al. 2011; Meding and Hagendorff

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.