Pith. sign in

REVIEW 4 major objections 5 minor 76 references

Templated or fully synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance

T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Templated prompts constructed to appear neutral still push LLM political responses toward the stance encoded in their fillers, making prompt construction a confound in stance measurement.

desk verdict A careful empirical study showing templated prompts skew LLM stance measurements; the effect is real and well-supported, though the size of the gap rests on LLM judges that merit human anchoring. read the letter →

arxiv 2608.11008 v2 pith:6L4IQ3ZB submitted 2026-08-11 cs.CL cs.CY

classification cs.CLcs.CY
keywords promptconstructionpoliticalstancedetectiontemplatedpromptsLLM-generatedecologicalvalidityLLM-as-a-judgesycophancyconfound
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that how a political prompt is built—by filling a template or by having an LLM write a freestanding request—changes the political stance you measure in the answering model, even when the prompt is meant to be neutral. On 18 neutral topic-by-intent settings, templated prompts pushed GPT-5.4 mini's responses further toward the pole encoded in the template's filler (14 of 18 settings, Wilcoxon signed-rank p = .001), and the pattern replicated on Grok 4.3 (15 of 18 settings). The cause is structural: a filler must stay grammatical under all three stance versions, and the phrasing that does so is rarely stance-free. Two studies that differ only in prompt construction would therefore report different stances for the same model, with the templated study overstating leanings under neutral framings.

What carries the argument

The slot-and-filler template is the load-bearing object: a fixed request skeleton such as 'What is your opinion on X?' with a filler phrase carrying the topic and stance. Because the same filler must remain grammatical as neutral, pole-A, and pole-B versions, fillers like 'the Russian invasion and action in Ukraine' carry a presupposition, and the model reads that presupposition as a leading question. The paper shows the same fillers shift stance estimates in the direction they encode, whereas LLM-generated prompts, produced under detailed instructions with real prompts as seeds, express topic and stance without being constrained to stay grammatical across stance variants and therefore do not leak the same siding.

What would settle it

Have human annotators label the same model responses on the same 5-point stance scale. If human labels put neutral templated and neutral LLM-generated responses at equal distance from neutrality, the confound gap disappears; if human–ensemble agreement is worse on LLM-generated responses than on templated ones, the reported gap may be a judge artefact rather than a prompt-construction effect.

Watch

Extended reading notes

Core claim

The central discovery is that templated prompts systematically distort stance measurements of LLMs, while LLM-generated prompts anchored in real user prompts do not. In a small-scale study covering three contested policy issues and three geopolitical conflicts, human and LLM annotators ranked LLM-generated prompts as no less realistic than prompts drawn from real chat logs, and clearly more realistic than templated ones; LLM-generated prompts also carried their intended intent and stance more clearly. In the stance case study, neutral templated prompts elicited responses 0.48 scale points from neutral on average, versus 0.07 for neutral LLM-generated prompts, with the distortion always pointing in the direction the filler encodes: 'the Russian invasion and action in Ukraine' reads as a presupposition that the invasion is unjustified, and 'the climate change and the appropriate policy response' reads as an assertion of severity. The paper concludes that prompt construction is not a neutral design choice: same model, same judges, same topics, but different construction methods report different political stances, and the templated method overstates the model's leanings precisely where neutrality is what is being measured.

Load-bearing premise

The argument leans on three AI judges whose stance ratings are only checked against each other, never against human ratings of the same responses, and they disagree most on the nuanced responses the main gap depends on.

Editorial extensions

If this is right

  • Neutral templated prompts overstate a model's political leaning in the direction the filler encodes, so any existing stance estimate built on such prompts should be treated as potentially inflated.
  • Templated prompts are more legible as evaluation artefacts to LLM annotators than to humans (Kendall's W = .65 vs .21); if models can recognize them, they may answer differently than they would for a real user.
  • LLM-generated prompts with real prompts as seeds preserve experimental control over topic, intent, and stance while matching real prompts in realism, offering a practical alternative to templates.
  • Extending stance measurement beyond writing assistance changes what is measured: writing assistance mostly captures instruction-following, while information seeking and opinion sharing surface the model's own stance and asymmetric sycophancy.
  • Two studies of the same model, using the same judges and topics but different prompt construction, will report different political stances; prompt construction must be reported and validated as a design choice.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • We infer the confound is model-independent in direction, since it is a property of the filler's presuppositional content rather than of the model parsing it; a direct test would be to build fillers that are genuinely symmetric across stances and check whether the templated-vs-synthetic gap shrinks.
  • We infer that 'the stance of an LLM' is underdetermined until the prompt distribution is specified: political stance reports are joint products of model, prompt form, and judge, not properties of the model alone.
  • We infer that the blurred boundary between information seeking and opinion sharing (human intent agreement as low as κ ≈ .27) means cross-task stance comparisons are only meaningful when studies state how value-laden and speculative questions are classified.
  • We infer that since LLM judges separate templated from natural prompts more sharply than humans do, prompt realism could be screened automatically before an evaluation corpus is deployed, rather than only by human panels.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper argues that the way evaluation prompts are constructed is a confound in measuring the political stance of LLMs. The authors extend IssueBench beyond writing assistance to two additional intents (information seeking and opinion sharing), propose fully synthetic (LLM-generated) prompts anchored in real chat-log prompts as seeds, and validate real, templated, and synthetic prompts on realism and construct clarity using three human and three LLM annotators. They then compare stance estimates from templated and synthetic prompts for two deployed models, GPT 5.4 mini and Grok 4.3, judged by a majority-vote ensemble of three open-weight LLMs. The central claim (Section 3.4.3) is that neutral templated prompts elicit responses systematically farther from neutral, in the direction encoded by the template filler, than neutral synthetic prompts do (0.48 vs 0.07 scale points on GPT 5.4 mini, 14 of 18 settings, Wilcoxon p=.001; replicated on Grok 4.3 at 0.36 vs 0.05, 15 of 18, p<.001), so that a templated study would overstate the model's leanings precisely where neutrality is the target of measurement. The paper also reports that synthetic prompts are ranked at least as realistic as real prompts and clearer in intent and stance, and that LLM annotators separate templated prompts from real ones more sharply than human annotators do.

Significance. The contribution is significant if it holds: it identifies a direction-predictable, systematic confound in a widely used evaluation framework (IssueBench-style templating) rather than a random artefact, and pairs the critique with a concrete, validated alternative construction method. The paper's strengths are substantial: the full resource set (prompts, human and LLM annotations, model responses, stance judgments) is released on HuggingFace, the central effect is replicated on two models from different developers, the judge ensemble consists of three models with no developer overlap with the examined models, the direction-of-effect evidence is presented transparently (14/18 and 15/18 settings), and the Limitations are unusually candid, including the explicit statement that the stance judges are validated only against each other. The main risk is that the quantitative centrepiece is produced entirely by an LLM judge chain that is never checked against human stance labels, and the judge prompt itself contains the user prompt, so prompt-contamination of the labels is a concrete, uneliminated alternative explanation for the measured gap.

major comments (4)
  1. [§3.3, §3.4.3, Appendix E, Limitations ('LLM judges')] The load-bearing quantitative result rests on an unvalidated LLM judge chain. The central claim of Section 3.4.3 — that neutral templated responses sit 0.48 scale points from neutral versus 0.07 for LLM-generated responses — is entirely produced by the majority-vote ensemble of DeepSeek V4 Pro, Mistral Large 3, and Nemotron 3 Ultra, and the Limitations state: 'We validate the judges against each other rather than against human stance annotations of responses.' Because the judge prompt (Appendix E) displays the full user prompt and merely instructs the judge to ignore whether the request was one-sided, with no compliance check, a judge that unconsciously uses prompt framing would label responses to the demonstrably stance-leaking neutral fillers (Section 2.2.2) as more sided in the filler's direction, while rounding the more hedged LLM-generated responses (where judge agreement is lowest, ordinal α=.82 vs .91) toward the neutral class — producing exactly the observed pattern. The Grok 4.3 replication (Appendix G.1) uses the same judge ensemble and therefore cannot arbitrate. The paper should either (a) collect human stance annotations on a sample of responses and report judge–human agreement per construction method, or (b) run a control that fixes the response while removing or swapping the user prompt and checks label stability. Without one of these, the claim that templated prompts overstate the model's leanings, rather than the judges' labels reflecting the prompt, is under-supported.
  2. [§3.3 (Likert rubric relaxation)] Section 3.3 relaxes Likert classes 2 and 4 from Röttger et al.'s 'overwhelmingly (90%)' to 'substantially (75%)', which moves every response with 75–90% single-side emphasis out of the neutral class 3. The change is applied to both construction methods, but it interacts directly with the quantity being measured: the central comparison is precisely the distance of responses from the neutral class, and if templated responses are mildly but genuinely sided (filler-driven) while LLM-generated responses are hedged, the relaxed rubric amplifies the measured gap. The manuscript should report the Section 3.4.3 comparison under the original 90% thresholds (or under an alternative continuous-scale rubric) to show that the 0.48 vs 0.07 gap is not threshold-driven, especially since Section 4 itself describes this relaxation as treating a symptom rather than a cause.
  3. [§3.4.3, Tables 4 and 9, §2.2.2] The headline claim in Section 3.4.3 states that neutral templated prompts are 'systematically further from neutral, in the direction the filler encodes (14 of 18 neutral settings...)'. The arithmetic component of this count is verifiable from Table 4 (14 settings with |templated lean| > |LLM-generated lean|, 3 with the reverse, 1 tie), but the direction component is not auditable from the paper's tables: no per-setting table or figure maps each neutral setting to the filler-encoded direction and the observed direction. This matters because at least one topic appears inconsistent with the stated mechanism: the paper's own principle (Section 2.2.2 and Section 3.4.2) reads 'the US and Israeli strikes on Iran' as siding against the striker, i.e., toward pole B (pro-Iran), yet the neutral templated leans in Table 4 for US/Israel–Iran are -0.06 (information seeking) and -0.06 (opinion sharing), i.e., toward pole A. The authors should publish the per-setting direction coding (or a stacked plot of the 18 neutral settings with the filler-encoded direction marked) and either reconcile or explicitly explain these apparent counter-directional settings.
  4. [Appendix A (Table 5), Appendix C (Table 9), Appendix D (Detection Guidelines)] There is an internal contradiction in the definition of the climate-change poles. Table 5 and Table 9 assign Climate-urgency to Pole A and Climate-moderation to Pole B, so negative leans in Table 4 mean climate urgency, consistent with the text of Section 3.4.2. The Detection Guidelines in Appendix D, however, assign 'climate and ecological change not being severe, and the policy response being adequate' to Pole A and 'being severe... inadequate' to Pole B, i.e., the reverse lettering, with the note that 'in favour means downplaying the severity'. Since the direction claim in Section 3.4.3 ('the direction the filler encodes') depends on the pole lettering, and climate change is one of the two topics that together account for 85% of the total neutral-setting divergence (Section 3.4.2), the two appendices must be made consistent before the direction pattern is verifiable by a reader or replicator.
minor comments (5)
  1. [§3.4.3, Limitations (statistical dependence)] The Wilcoxon signed-rank test in Section 3.4.3 treats the 18 neutral settings as exchangeable units even though they share six topics, and the Limitations acknowledge this non-independence; the main text should carry that caveat alongside the reported p=.001, or the authors should report a cluster-robust or judgment-level test, since the nominal p-value is otherwise optimistic. The direction counts (14/18 and 15/18) are the more robust evidence and are not in question.
  2. [Tables 4 and 10] Tables 4 and 10 do not report the number of responses per setting (the N columns denote the neutral user stance, not counts); the authors should add per-setting response counts, particularly because refusal rates differ by construction method (5% vs 2%, Section 3.3), so the per-setting sample sizes underlying the lean estimates are not recoverable from the paper.
  3. [§2.2.1, Appendix F, Abstract] The conclusion that human annotators rank LLM-generated prompts as 'no less realistic than real ones' rests on three human annotators with near-chance agreement (Kendall's W=.21, Appendix F); the abstract and Section 2.2.1 should hedge this as essentially a null result rather than a demonstrated parity, since the human ordering evidence is thin.
  4. [Appendix D and §3.3] The same three LLMs serve as the realness and detection annotators (Appendix D) and as the stance judges (Section 3.3); the detectability finding (LLM annotators separate templated prompts from real ones more sharply than humans do) is therefore produced by the same models that later judge the responses, and this connection should be noted at the point where detectability is used to motivate the confound.
  5. [Appendix G.1, Table 1] Presentation issues: Appendix G.1 contains the typo 'Sumarry' for 'Summary', and in Table 1 the US/Israel–Iran topic label is typeset as 'US/IL-IR2' with a missing space, which should be corrected.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the confound claim is an empirical comparison, not a fitted or self-cited derivation.

full rationale

I find no circular step in the paper's derivation chain. The central empirical claims—that templated prompts are more detectable and less realistic than LLM-generated ones, and that neutral templated prompts elicit more polarized stance estimates in the direction of their fillers—are established by comparing three independently constructed prompt collections on responses from two examined models, using both human and LLM annotations. No parameter is fitted to a subset of the data and then renamed as a prediction; the lean values in Tables 4 and 10 are computed directly from judge labels via Equation 1, and the Wilcoxon comparisons in Section 3.4.3 are standard tests on those computed values. The author's self-citations (Chalkidis and Brandl 2024; Chalkidis and Sogaard 2026) are contextual and not load-bearing. The main limitation—LLM judges are validated only against each other, not against human stance annotations—is an external-validity threat that could affect the measured gap, but it does not make the derivation circular: the judges' labels are not defined in terms of the conclusion, and the paper explicitly discloses the threat in the Limitations and in Section 4. Likewise, using the same LLM family for realness annotation and stance judging is a possible preference confound, but the realness finding is also supported by human annotators. I therefore rate the paper as free of circularity, with the caveat that the stance estimates inherit the judge-validity limitation the authors acknowledge.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The paper contributes an empirical comparison, not a derivation, so the ledger is dominated by measurement design choices: hand-authored fillers, pole definitions, and the judge instrument. No fitted parameter produces the headline result, and no new theoretical entity is introduced; the synthetic prompt collection is a data artifact. The load-bearing assumption is judge-instrument validity, which the paper flags in Limitations.

free parameters (3)
  • Filler texts for templated prompts (Table 9) = 'the climate change and the appropriate policy response' and variants; 19 paraphrases per stance-topic
    Hand-authored, then reduced to the 'least suggestive' variants after observing weaknesses (Section 3.1). The choice directly determines the filler-induced stance gap quantified in Section 3.4.2; different fillers would change the gap's size.
  • Topic pole definitions and stance-loaded lexicons (Table 5, Appendix B) = Two poles per topic, A/B fixed
    Authored by the researchers; used both to instruct the synthetic generator and as the judge rubric. The Limitations concede the descriptions 'reflect the debate as we understand it'.
  • Likert class thresholds for stance judging = 75% for classes 2 and 4, relaxed from IssueBench's 90%
    Section 3.3. Hand-chosen to avoid collapsing 75 to 90 percent siding into the neutral class; applied symmetrically to both construction methods.
assumptions (4)
  • domain assumption The three-judge LLM ensemble produces stance labels valid for comparing construction methods.
    Section 3.3 and Limitations 'LLM judges'. The entire Section 3 measurement depends on it; only inter-judge agreement is reported (ordinal Krippendorff's alpha .91 templated, .82 LLM-generated), never agreement with human stance labels.
  • domain assumption Real prompts from WildChat and LMSys-Chat, filtered by IssueBench and re-labelled by the authors, are a valid realism anchor and representative seeds.
    Section 2.1. The paper acknowledges the logs predate the US/Israel-Iran conflict, are heavily stance-skewed, and are therefore used only as a realism anchor, not as a measurement source.
  • domain assumption Bare API LLMs behave informatively about the deployed assistants users interact with.
    Limitations 'LLMs are not AI chatbots': the examined models lack safety classifiers, RAG, geolocation and other harness modules present in production systems, and the authors expect production responses to differ.
  • domain assumption Value-laden and speculative questions are treated as opinion sharing rather than information seeking.
    Section 4 states this working definition is 'a design choice rather than a consensus', and acknowledges a different choice would relabel a substantial share of prompts and make studies incomparable.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Templated or fully synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance." pith.science (2026). https://pith.science/paper/6L4IQ3ZB

@misc{pith2026260811008,
  author       = {Pith},
  title        = {Pith review of: Templated or fully synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6L4IQ3ZB}},
  note         = {Machine review of arXiv:2608.11008}
}
read the original abstract

Political stance detection in LLMs has long been dominated by closed-ended, multiple-choice political survey questions---originally designed for humans, and thus lacks the realism and nuance of human-AI interactions in the wild, while also being susceptible to sandbagging. The recent IssueBench framework substantially mitigates these limitations with templated prompts anchored in real-world chat logs. Given the rise in non-work-related use of GenAI assistants, we extend IssueBench beyond writing assistance to include two additional tasks, information seeking and opinion sharing. We argue that templated prompts still lack the nuance of real ones, especially for open-ended tasks, and remain recognisable as evaluation artefacts. We propose the use of fully synthetic (LLM-generated) prompts, produced under detailed instructions with real prompts as seeds. We assess the ecological validity of real, templated, and LLM-generated prompts in a small-scale study covering 3 highly contested policy issues and 3 recent geopolitical conflicts. Human and LLM annotators rank LLM-generated prompts as no less realistic than real ones and clearly more realistic than templated ones, and find that they carry their intended intent and stance more clearly; the LLMs separate templated prompts from the other two far more sharply than the humans do. In a case study, templated and LLM-generated prompts yield systematically different stance estimates for the same model, most visibly under neutral framings, where templated prompts overstate the model's leaning in the direction encoded by the topic-and-stance text (filler) slotted into their templates.

Figures

Figures reproduced from arXiv: 2608.11008 by the authors.

Figure 1
Figure 1. Two sets of examples for templated (orange) and LLM-generated (purple) prompts for two selected topics (Climate Change / Russia–Ukraine), each set for one of the two newly introduced intents (information seeking / opinion sharing), and a given user’s stance (neutral / sided). Below, the mean lean of GPT 5.4 mini (Section 3) between the two topic-specific poles across all three intents in the same setting (intent and… view at source ↗
Figure 2
Figure 2. The custom UI interface for the prompt realness ranking task. The human annotator is presented with a [PITH_FULL_IMAGE:figures/full_fig_p028_2.png] view at source ↗
Figure 3
Figure 3. The custom UI interface for the topic/intent/stance detection task. The annotator is presented with a [PITH_FULL_IMAGE:figures/full_fig_p028_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

76 extracted references · 1 linked inside Pith

  1. [1]

    Rogov and Alexander Panchenko and Natalia Loukachevitch and Elena Tutubalina , year=

    Mikhail Salnikov and Dmitrii Korzh and Ivan Lazichny and Elvir Karimov and Artyom Iudin and Ivan Oseledets and Oleg Y. Rogov and Alexander Panchenko and Natalia Loukachevitch and Elena Tutubalina , year=

  2. [2]

    Santurkar, Shibani and Durmus, Esin and Ladhak, Faisal and Lee, Cinoo and Liang, Percy and Hashimoto, Tatsunori , title =

  3. [3]

    Chalkidis, Ilias and Brandl, Stephanie , booktitle =

  4. [4]

    Jochen Hartmann and Jasper Schwenzow and Maximilian Witte , year=

  5. [5]

    de Moura and Wei Zhang and Jose O

    William Guey and Pierrick Bougault and Vitor D. de Moura and Wei Zhang and Jose O. Gomes , year=

  6. [6]

    Choudhary, Tavishi , journal=

  7. [7]

    arXiv preprint arXiv:2505.23836 , year=

    Large language models often know when they are being evaluated , author=. arXiv preprint arXiv:2505.23836 , year=

  8. [8]

    International Conference on Learning Representations , volume=

    Ai sandbagging: Language models can strategically underperform on evaluations , author=. International Conference on Learning Representations , volume=

Show all 76 references
  1. [9]

    2026 , month =

    Claude Sonnet 5 System Card , institution =. 2026 , month =

  2. [10]

    Stereotyping N orwegian Salmon: An Inventory of Pitfalls in Fairness Benchmark Datasets

    Blodgett, Su Lin and Lopez, Gilsinia and Olteanu, Alexandra and Sim, Robert and Wallach, Hanna. Stereotyping N orwegian Salmon: An Inventory of Pitfalls in Fairness Benchmark Datasets. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and ...

  3. [11]

    2025 , eprint=

    Scaling Synthetic Data Creation with 1,000,000,000 Personas , author=. 2025 , eprint=

  4. [12]

    Gonzalez and Ion Stoica , booktitle=

    Lianmin Zheng and Wei-Lin Chiang and Ying Sheng and Siyuan Zhuang and Zhanghao Wu and Yonghao Zhuang and Zi Lin and Zhuohan Li and Dacheng Li and Eric Xing and Hao Zhang and Joseph E. Gonzalez and Ion Stoica , booktitle=. Judging. 2023 , url=

  5. [13]

    Bowman and Shi Feng , booktitle=

    Arjun Panickssery and Samuel R. Bowman and Shi Feng , booktitle=. 2024 , url=

  6. [14]

    Political Analysis , author=

    Synthetic Replacements for Human Survey Data? The Perils of Large Language Models , volume=. Political Analysis , author=. 2024 , pages=. doi:10.1017/pan.2024.5 , number=

  7. [15]

    Political Analysis , author=

    Out of One, Many: Using Language Models to Simulate Human Samples , volume=. Political Analysis , author=. 2023 , pages=. doi:10.1017/pan.2023.2 , number=

  8. [16]

    and Khashabi, Daniel and Hajishirzi, Hannaneh

    Wang, Yizhong and Kordi, Yeganeh and Mishra, Swaroop and Liu, Alisa and Smith, Noah A. and Khashabi, Daniel and Hajishirzi, Hannaneh. Self-Instruct: Aligning Language Models with Self-Generated Instructions. Proceedings of the 61st Annual Meeting of the Association for Computa...

  9. [17]

    BBQ : A hand-built bias benchmark for question answering

    Parrish, Alicia and Chen, Angelica and Nangia, Nikita and Padmakumar, Vishakh and Phang, Jason and Thompson, Jana and Htut, Phu Mon and Bowman, Samuel R. BBQ : A hand-built bias benchmark for question answering. Findings of the Association for Computational Linguistics: ACL 20...

  10. [18]

    S tereo S et: Measuring stereotypical bias in pretrained language models

    Nadeem, Moin and Bethke, Anna and Reddy, Siva. S tereo S et: Measuring stereotypical bias in pretrained language models. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Proc...

  11. [19]

    C row S -Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models

    Nangia, Nikita and Vania, Clara and Bhalerao, Rasika and Bowman, Samuel R. C row S -Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. doi:10.18...

  12. [20]

    Feng, Shangbin and Park, Chan Young and Liu, Yuhan and Tsvetkov, Yulia , editor =

  13. [21]

    Chatterji, Aaron and Cunningham, Thomas and Deming, David J and Hitzig, Zoe and Ong, Christopher and Shan, Carl Yan and Wadman, Kevin , year=

  14. [22]

    Marc Zao-Sanders , title =

  15. [23]

    Mrinank Sharma and Meg Tong and Tomasz Korbak and David Duvenaud and Amanda Askell and Samuel R. Bowman and Esin DURMUS and Zac Hatfield-Dodds and Scott R Johnston and Shauna M Kravec and Timothy Maxwell and Sam McCandlish and Kamal Ndousse and Oliver Rausch and Nicholas Schie...

  16. [24]

    Position: Political Neutrality in

    Jillian Fisher and Ruth Elisabeth Appel and Chan Young Park and Yujin Potter and Liwei Jiang and Taylor Sorensen and Shangbin Feng and Yulia Tsvetkov and Margaret Roberts and Jennifer Pan and Dawn Song and Yejin Choi , booktitle=. Position: Political Neutrality in. 2025 , url=

  17. [25]

    Croatian Journal of Philosophy , volume=

    The impossibility of political neutrality , author=. Croatian Journal of Philosophy , volume=. 2010 , publisher=

  18. [26]

    The Thirty-eighth Annual Conference on Neural Information Processing Systems , year=

    Questioning the Survey Responses of Large Language Models , author=. The Thirty-eighth Annual Conference on Neural Information Processing Systems , year=

  19. [27]

    Do LLM s Exhibit Human-like Response Biases? A Case Study in Survey Design

    Tjuatja, Lindia and Chen, Valerie and Wu, Tongshuang and Talwalkwar, Ameet and Neubig, Graham. Do LLM s Exhibit Human-like Response Biases? A Case Study in Survey Design. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00685

  20. [28]

    Llama meets EU : Investigating the E uropean political spectrum through the lens of LLM s

    Chalkidis, Ilias and Brandl, Stephanie. Llama meets EU : Investigating the E uropean political spectrum through the lens of LLM s. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Vo...

  21. [29]

    PloS one , volume=

    The political preferences of LLMs , author=. PloS one , volume=. 2024 , publisher=

  22. [30]

    First Conference on Language Modeling , year=

    Towards Measuring the Representation of Subjective Global Opinions in Language Models , author=. First Conference on Language Modeling , year=

  23. [31]

    2018 , month =

    Mitchell, Amy and Gottfried, Jeffrey and Barthel, Michael and Sumida, Nami , title =. 2018 , month =

  24. [32]

    1986 , publisher=

    The morality of freedom , author=. 1986 , publisher=

  25. [33]

    , author=

    On the ambivalence-indifference problem in attitude theory and measurement: A suggested modification of the semantic differential technique. , author=. Psychological bulletin , volume=. 1972 , publisher=

  26. [34]

    Global environmental change , volume=

    Balance as bias: Global warming and the US prestige press , author=. Global environmental change , volume=. 2004 , publisher=

  27. [35]

    Abecasis, Paulo and De Michiel, Federico and Basalisco, Bruno and Haanper

  28. [36]

    Adler, Steven , title =

  29. [37]

    biometrics , pages=

    The measurement of observer agreement for categorical data , author=. biometrics , pages=. 1977 , publisher=

  30. [38]

    Workshop on Trustworthy and Socially Responsible Machine Learning, NeurIPS 2022 , year=

    Quantifying Social Biases Using Templates is Unreliable , author=. Workshop on Trustworthy and Socially Responsible Machine Learning, NeurIPS 2022 , year=

  31. [39]

    2025 , month=

    Shen, Judy Hanwen and Appel, Ruth and Tucker, Madeleine and Jagadish, Kamya and Maheshwary, Paruul and Askell, Amanda and Durmus, Esin , title=. 2025 , month=

  32. [40]

    2011 , publisher=

    On the political , author=. 2011 , publisher=

  33. [41]

    Herrman, John , title =

  34. [42]

    Michel, Elie and Cicchi, Lorenzo and Garzia, Diego and Ferreira da Silva, Frederico and Trechsel, Alexander , year =

  35. [43]

    Wayne Brittenden , url =

  36. [44]

    and Mondría Terol, Teresa and Conger, Kate and Freedman, Dylan , title =

    Thompson, Stuart A. and Mondría Terol, Teresa and Conger, Kate and Freedman, Dylan , title =

  37. [45]

    Swenson, Ali , title =

  38. [46]

    Neudert, Lisa Maria and Marchal, Nahema , title =

  39. [47]

    Ghose, Anuttama and Pallav, Pallav and Ali, S. M. Aamir , title =

  40. [48]

    Celeste Kidd and Abeba Birhane , title =

  41. [49]

    Ilias Chalkidis , year=

  42. [50]

    Ferdman, Avigail , journal=

  43. [51]

    Erfani, Farhang , journal=

  44. [52]

    npj Artificial Intelligence , volume=

    Buyl, Maarten and Rogiers, Alexander and Noels, Sander and Bied, Guillaume and Dominguez-Catena, Iris and Heiter, Edith and Johary, Iman and Mara, Alexandru-Cristian and Romero, Rapha. npj Artificial Intelligence , volume=. 2026 , publisher=

  45. [53]

    and Gebru, Timnit and McMillan-Major, Angelina and Shmitchell, Shmargaret , title =

    Bender, Emily M. and Gebru, Timnit and McMillan-Major, Angelina and Shmitchell, Shmargaret , title =

  46. [54]

    Brainrot: Deskilling and Addiction are Overlooked AI Risks , year =

    Chalkidis, Ilias and S. Brainrot: Deskilling and Addiction are Overlooked AI Risks , year =. Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency , pages =

  47. [55]

    Tai-Quan Peng and Kaiqi Yang and Sanguk Lee and Hang Li and Yucheng Chu and Yuping Lin and Hui Liu , year=

  48. [56]

    The White House , title =

  49. [57]

    Simpson , title =

    Raphael Racicot and Kurtis H. Simpson , title =

  50. [58]

    2024 , publisher=

    Gu, Jiawei and Jiang, Xuhui and Shi, Zhichao and Tan, Hexiang and Zhai, Xuehao and Xu, Chengjin and Li, Wei and Shen, Yinghan and Ma, Shengjie and Liu, Honghao , journal=. 2024 , publisher=

  51. [59]

    Ashery, Ariel Flint and Aiello, Luca Maria and Baronchelli, Andrea , journal=

  52. [60]

    Bias in Language Models: Beyond Trick Tests and Towards RUTE d Evaluation

    Lum, Kristian and Anthis, Jacy Reese and Robinson, Kevin and Nagpal, Chirag and D ' Amour, Alexander Nicholas. Bias in Language Models: Beyond Trick Tests and Towards RUTE d Evaluation. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Vo...

  53. [61]

    2024 , url=

    Wenting Zhao and Xiang Ren and Jack Hessel and Claire Cardie and Yejin Choi and Yuntian Deng , booktitle=. 2024 , url=

  54. [62]

    Bronson Schoen and Evgenia Nitishinskaya and Mikita Balesni and Axel Højmark and Felix Hofstätter and Jérémy Scheurer and Alexander Meinke and Jason Wolfe and Teun van der Weij and Alex Lloyd and Nicholas Goldowsky-Dill and Angela Fan and Andrei Matveiakin and Rusheb Shah and ...

  55. [63]

    Alexander Meinke and Bronson Schoen and Jérémy Scheurer and Mikita Balesni and Rusheb Shah and Marius Hobbhahn , year=

  56. [64]

    2026 , publisher=

    Turner, Cody and Eisikovits, Nir , journal=. 2026 , publisher=

  57. [65]

    Hine, Emmie and Floridi, Luciano , journal=

  58. [66]

    2022 , publisher=

    Bareis, Jascha and Katzenbach, Christian , journal=. 2022 , publisher=

  59. [67]

    Investigating

    Chalkidis, Ilias. Investigating. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. doi:10.18653/v1/2024.emnlp-main.312

  60. [68]

    Proceedings of the 19th Conference of the E uropean Chapter of the A ssociation for C omputational L inguistics (Volume 1: Long Papers)

    Chalkidis, Ilias and Brandl, Stephanie and Aslanidis, Paris. Proceedings of the 19th Conference of the E uropean Chapter of the A ssociation for C omputational L inguistics (Volume 1: Long Papers). 2026. doi:10.18653/v1/2026.eacl-long.40

  61. [69]

    Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

    R. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. doi:10.18653/v1/2024.acl-long.816

  62. [70]

    Bias in the East, Bias in the West: A Bilingual Analysis of

    Lim, Ying Ying and R. Bias in the East, Bias in the West: A Bilingual Analysis of. Findings of the A ssociation for C omputational L inguistics: EACL 2026. 2026. doi:10.18653/v1/2026.findings-eacl.122

  63. [71]

    Transactions of the Association for Computational Linguistics , volume=

    R. Transactions of the Association for Computational Linguistics , volume=. 2026 , publisher=

  64. [72]

    Wright, Dustin and Arora, Arnav and Borenstein, Nadav and Yadav, Srishti and Belongie, Serge and Augenstein, Isabelle , booktitle=

  65. [73]

    Westwood, Sean J and Grimmer, Justin and Hall, Andrew B , journal=

  66. [74]

    2023 , eprint=

    LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset , author=. 2023 , eprint=

  67. [75]

    W ild V is: Open Source Visualizer for Million-Scale Chat Logs in the Wild

    Deng, Yuntian and Zhao, Wenting and Hessel, Jack and Ren, Xiang and Cardie, Claire and Choi, Yejin. W ild V is: Open Source Visualizer for Million-Scale Chat Logs in the Wild. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: System Demons...

  68. [76]

    2024 , url =

    Esin Durmus and Liane Lovitt and Alex Tamkin and Stuart Ritchie and Jack Clark and Deep Ganguli , title =. 2024 , url =

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.