Pith. sign in

REVIEW 10 cited by

Linear Representations of Political Perspective Emerge in Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.02080 v2 pith:DJB5QLE2 submitted 2025-03-03 cs.CL cs.AIcs.CYcs.HCcs.LG

classification cs.CLcs.AIcs.CYcs.HCcs.LG
keywords linearllmsperspectivespoliticalheadsmodelstextattention
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) have demonstrated the ability to generate text that realistically reflects a range of different subjective human perspectives. This paper studies how LLMs are seemingly able to reflect more liberal versus more conservative viewpoints among other political perspectives in American politics. We show that LLMs possess linear representations of political perspectives within activation space, wherein more similar perspectives are represented closer together. To do so, we probe the attention heads across the layers of three open transformer-based LLMs (Llama-2-7b-chat, Mistral-7b-instruct, Vicuna-7b). We first prompt models to generate text from the perspectives of different U.S. lawmakers. We then identify sets of attention heads whose activations linearly predict those lawmakers' DW-NOMINATE scores, a widely-used and validated measure of political ideology. We find that highly predictive heads are primarily located in the middle layers, often speculated to encode high-level concepts and tasks. Using probes only trained to predict lawmakers' ideology, we then show that the same probes can predict measures of news outlets' slant from the activations of models prompted to simulate text from those news outlets. These linear probes allow us to visualize, interpret, and monitor ideological stances implicitly adopted by an LLM as it generates open-ended responses. Finally, we demonstrate that by applying linear interventions to these attention heads, we can steer the model outputs toward a more liberal or conservative stance. Overall, our research suggests that LLMs possess a high-level linear representation of American political ideology and that by leveraging recent advances in mechanistic interpretability, we can identify, monitor, and steer the subjective perspective underlying generated text.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Temporal Preference Concepts and their Functions in a Large Language Model

    cs.LG 2026-05 unverdicted novelty 6.5 of 10

    Temporal preference in Qwen3-4B-Instruct-2507 localizes to layers 17–35 (especially L24 attention), has curved residual-stream geometry, is behaviorally unstable, and can be bidirectionally steered.

  2. Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy

    cs.AI 2026-02 conditional novelty 6.0 of 10

    Linear probes on LLM residual streams classify Bloom's Taxonomy levels with high accuracy, but the result may reflect prompt lexico-semantic cues rather than a general cognitive-complexity representation.

  3. Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Bias in GPT-2 and Llama-2 is localized to a small set of edges, and ablation of those edges reduces bias while impairing unrelated NLP tasks.

  4. Fine-Grained Interpretation of Political Opinions in Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Four-dimensional political concept vectors learned from LLM internals can detect and partially steer political leanings better than a single left-right axis.

  5. Evaluating Intra-firm LLM Alignment Strategies in Business Contexts

    cs.CY 2025-05 conditional novelty 6.0 of 10

    Firms should intentionally align AI assistants' embedded perspectives using supportive, adversarial, or diverse strategies to protect workplace culture and moral norms.

  6. Linear representations of grammaticality in neural language models

    cs.CL 2026-07 conditional novelty 5.0 of 10

    Grammaticality is linearly decodable from language model sentence representations and generalizes across phenomena and languages in larger models.

  7. Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint

    cs.CL 2025-09 conditional novelty 5.0 of 10

    ProCon anchors each sample's hidden-state projection onto the LLM's initial refusal direction during instruction fine-tuning, reducing refusal-direction drift and safety risks with limited task-performance loss.

  8. When Models Refuse: Political Steerability and Feature Richness as Measures of Ideological Depth

    cs.CL 2025-08 reject novelty 5.0 of 10

    Comparing two open LLMs, the paper claims that greater internal political feature richness predicts steerability and that refusals on benign prompts reflect capability deficits, but the causal evidence is missing.

  9. PII Jailbreaking in LLMs via Activation Steering Reveals Personal Information Leakage

    cs.CR 2025-07 reject novelty 5.0 of 10

    Activation steering on probe-selected attention heads flips LLM privacy refusals into disclosures, with claims of high rates of factually correct personal information leakage.

  10. LegiGPT: Party Politics and Transport Policy with Large Language Model

    cs.CL 2025-06 reject novelty 4.0 of 10

    The party composition of a bill's sponsors, along with district area and population, predicts a Korean lawmaker's political affiliation in transportation bills, though the sponsor features are derived from the same af...

Pith tools