REVIEW 4 major objections 5 minor 76 references
Templated or fully synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance
T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Templated prompts constructed to appear neutral still push LLM political responses toward the stance encoded in their fillers, making prompt construction a confound in stance measurement.
desk verdict A careful empirical study showing templated prompts skew LLM stance measurements; the effect is real and well-supported, though the size of the gap rests on LLM judges that merit human anchoring. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The slot-and-filler template is the load-bearing object: a fixed request skeleton such as 'What is your opinion on X?' with a filler phrase carrying the topic and stance. Because the same filler must remain grammatical as neutral, pole-A, and pole-B versions, fillers like 'the Russian invasion and action in Ukraine' carry a presupposition, and the model reads that presupposition as a leading question. The paper shows the same fillers shift stance estimates in the direction they encode, whereas LLM-generated prompts, produced under detailed instructions with real prompts as seeds, express topic and stance without being constrained to stay grammatical across stance variants and therefore do not leak the same siding.
What would settle it
Have human annotators label the same model responses on the same 5-point stance scale. If human labels put neutral templated and neutral LLM-generated responses at equal distance from neutrality, the confound gap disappears; if human–ensemble agreement is worse on LLM-generated responses than on templated ones, the reported gap may be a judge artefact rather than a prompt-construction effect.
Extended reading notes
Core claim
The central discovery is that templated prompts systematically distort stance measurements of LLMs, while LLM-generated prompts anchored in real user prompts do not. In a small-scale study covering three contested policy issues and three geopolitical conflicts, human and LLM annotators ranked LLM-generated prompts as no less realistic than prompts drawn from real chat logs, and clearly more realistic than templated ones; LLM-generated prompts also carried their intended intent and stance more clearly. In the stance case study, neutral templated prompts elicited responses 0.48 scale points from neutral on average, versus 0.07 for neutral LLM-generated prompts, with the distortion always pointing in the direction the filler encodes: 'the Russian invasion and action in Ukraine' reads as a presupposition that the invasion is unjustified, and 'the climate change and the appropriate policy response' reads as an assertion of severity. The paper concludes that prompt construction is not a neutral design choice: same model, same judges, same topics, but different construction methods report different political stances, and the templated method overstates the model's leanings precisely where neutrality is what is being measured.
Load-bearing premise
The argument leans on three AI judges whose stance ratings are only checked against each other, never against human ratings of the same responses, and they disagree most on the nuanced responses the main gap depends on.
Editorial extensions
If this is right
- Neutral templated prompts overstate a model's political leaning in the direction the filler encodes, so any existing stance estimate built on such prompts should be treated as potentially inflated.
- Templated prompts are more legible as evaluation artefacts to LLM annotators than to humans (Kendall's W = .65 vs .21); if models can recognize them, they may answer differently than they would for a real user.
- LLM-generated prompts with real prompts as seeds preserve experimental control over topic, intent, and stance while matching real prompts in realism, offering a practical alternative to templates.
- Extending stance measurement beyond writing assistance changes what is measured: writing assistance mostly captures instruction-following, while information seeking and opinion sharing surface the model's own stance and asymmetric sycophancy.
- Two studies of the same model, using the same judges and topics but different prompt construction, will report different political stances; prompt construction must be reported and validated as a design choice.
Reading between the lines
- We infer the confound is model-independent in direction, since it is a property of the filler's presuppositional content rather than of the model parsing it; a direct test would be to build fillers that are genuinely symmetric across stances and check whether the templated-vs-synthetic gap shrinks.
- We infer that 'the stance of an LLM' is underdetermined until the prompt distribution is specified: political stance reports are joint products of model, prompt form, and judge, not properties of the model alone.
- We infer that the blurred boundary between information seeking and opinion sharing (human intent agreement as low as κ ≈ .27) means cross-task stance comparisons are only meaningful when studies state how value-laden and speculative questions are classified.
- We infer that since LLM judges separate templated from natural prompts more sharply than humans do, prompt realism could be screened automatically before an evaluation corpus is deployed, rather than only by human panels.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper argues that the way evaluation prompts are constructed is a confound in measuring the political stance of LLMs. The authors extend IssueBench beyond writing assistance to two additional intents (information seeking and opinion sharing), propose fully synthetic (LLM-generated) prompts anchored in real chat-log prompts as seeds, and validate real, templated, and synthetic prompts on realism and construct clarity using three human and three LLM annotators. They then compare stance estimates from templated and synthetic prompts for two deployed models, GPT 5.4 mini and Grok 4.3, judged by a majority-vote ensemble of three open-weight LLMs. The central claim (Section 3.4.3) is that neutral templated prompts elicit responses systematically farther from neutral, in the direction encoded by the template filler, than neutral synthetic prompts do (0.48 vs 0.07 scale points on GPT 5.4 mini, 14 of 18 settings, Wilcoxon p=.001; replicated on Grok 4.3 at 0.36 vs 0.05, 15 of 18, p<.001), so that a templated study would overstate the model's leanings precisely where neutrality is the target of measurement. The paper also reports that synthetic prompts are ranked at least as realistic as real prompts and clearer in intent and stance, and that LLM annotators separate templated prompts from real ones more sharply than human annotators do.
Significance. The contribution is significant if it holds: it identifies a direction-predictable, systematic confound in a widely used evaluation framework (IssueBench-style templating) rather than a random artefact, and pairs the critique with a concrete, validated alternative construction method. The paper's strengths are substantial: the full resource set (prompts, human and LLM annotations, model responses, stance judgments) is released on HuggingFace, the central effect is replicated on two models from different developers, the judge ensemble consists of three models with no developer overlap with the examined models, the direction-of-effect evidence is presented transparently (14/18 and 15/18 settings), and the Limitations are unusually candid, including the explicit statement that the stance judges are validated only against each other. The main risk is that the quantitative centrepiece is produced entirely by an LLM judge chain that is never checked against human stance labels, and the judge prompt itself contains the user prompt, so prompt-contamination of the labels is a concrete, uneliminated alternative explanation for the measured gap.
major comments (4)
- [§3.3, §3.4.3, Appendix E, Limitations ('LLM judges')] The load-bearing quantitative result rests on an unvalidated LLM judge chain. The central claim of Section 3.4.3 — that neutral templated responses sit 0.48 scale points from neutral versus 0.07 for LLM-generated responses — is entirely produced by the majority-vote ensemble of DeepSeek V4 Pro, Mistral Large 3, and Nemotron 3 Ultra, and the Limitations state: 'We validate the judges against each other rather than against human stance annotations of responses.' Because the judge prompt (Appendix E) displays the full user prompt and merely instructs the judge to ignore whether the request was one-sided, with no compliance check, a judge that unconsciously uses prompt framing would label responses to the demonstrably stance-leaking neutral fillers (Section 2.2.2) as more sided in the filler's direction, while rounding the more hedged LLM-generated responses (where judge agreement is lowest, ordinal α=.82 vs .91) toward the neutral class — producing exactly the observed pattern. The Grok 4.3 replication (Appendix G.1) uses the same judge ensemble and therefore cannot arbitrate. The paper should either (a) collect human stance annotations on a sample of responses and report judge–human agreement per construction method, or (b) run a control that fixes the response while removing or swapping the user prompt and checks label stability. Without one of these, the claim that templated prompts overstate the model's leanings, rather than the judges' labels reflecting the prompt, is under-supported.
- [§3.3 (Likert rubric relaxation)] Section 3.3 relaxes Likert classes 2 and 4 from Röttger et al.'s 'overwhelmingly (90%)' to 'substantially (75%)', which moves every response with 75–90% single-side emphasis out of the neutral class 3. The change is applied to both construction methods, but it interacts directly with the quantity being measured: the central comparison is precisely the distance of responses from the neutral class, and if templated responses are mildly but genuinely sided (filler-driven) while LLM-generated responses are hedged, the relaxed rubric amplifies the measured gap. The manuscript should report the Section 3.4.3 comparison under the original 90% thresholds (or under an alternative continuous-scale rubric) to show that the 0.48 vs 0.07 gap is not threshold-driven, especially since Section 4 itself describes this relaxation as treating a symptom rather than a cause.
- [§3.4.3, Tables 4 and 9, §2.2.2] The headline claim in Section 3.4.3 states that neutral templated prompts are 'systematically further from neutral, in the direction the filler encodes (14 of 18 neutral settings...)'. The arithmetic component of this count is verifiable from Table 4 (14 settings with |templated lean| > |LLM-generated lean|, 3 with the reverse, 1 tie), but the direction component is not auditable from the paper's tables: no per-setting table or figure maps each neutral setting to the filler-encoded direction and the observed direction. This matters because at least one topic appears inconsistent with the stated mechanism: the paper's own principle (Section 2.2.2 and Section 3.4.2) reads 'the US and Israeli strikes on Iran' as siding against the striker, i.e., toward pole B (pro-Iran), yet the neutral templated leans in Table 4 for US/Israel–Iran are -0.06 (information seeking) and -0.06 (opinion sharing), i.e., toward pole A. The authors should publish the per-setting direction coding (or a stacked plot of the 18 neutral settings with the filler-encoded direction marked) and either reconcile or explicitly explain these apparent counter-directional settings.
- [Appendix A (Table 5), Appendix C (Table 9), Appendix D (Detection Guidelines)] There is an internal contradiction in the definition of the climate-change poles. Table 5 and Table 9 assign Climate-urgency to Pole A and Climate-moderation to Pole B, so negative leans in Table 4 mean climate urgency, consistent with the text of Section 3.4.2. The Detection Guidelines in Appendix D, however, assign 'climate and ecological change not being severe, and the policy response being adequate' to Pole A and 'being severe... inadequate' to Pole B, i.e., the reverse lettering, with the note that 'in favour means downplaying the severity'. Since the direction claim in Section 3.4.3 ('the direction the filler encodes') depends on the pole lettering, and climate change is one of the two topics that together account for 85% of the total neutral-setting divergence (Section 3.4.2), the two appendices must be made consistent before the direction pattern is verifiable by a reader or replicator.
minor comments (5)
- [§3.4.3, Limitations (statistical dependence)] The Wilcoxon signed-rank test in Section 3.4.3 treats the 18 neutral settings as exchangeable units even though they share six topics, and the Limitations acknowledge this non-independence; the main text should carry that caveat alongside the reported p=.001, or the authors should report a cluster-robust or judgment-level test, since the nominal p-value is otherwise optimistic. The direction counts (14/18 and 15/18) are the more robust evidence and are not in question.
- [Tables 4 and 10] Tables 4 and 10 do not report the number of responses per setting (the N columns denote the neutral user stance, not counts); the authors should add per-setting response counts, particularly because refusal rates differ by construction method (5% vs 2%, Section 3.3), so the per-setting sample sizes underlying the lean estimates are not recoverable from the paper.
- [§2.2.1, Appendix F, Abstract] The conclusion that human annotators rank LLM-generated prompts as 'no less realistic than real ones' rests on three human annotators with near-chance agreement (Kendall's W=.21, Appendix F); the abstract and Section 2.2.1 should hedge this as essentially a null result rather than a demonstrated parity, since the human ordering evidence is thin.
- [Appendix D and §3.3] The same three LLMs serve as the realness and detection annotators (Appendix D) and as the stance judges (Section 3.3); the detectability finding (LLM annotators separate templated prompts from real ones more sharply than humans do) is therefore produced by the same models that later judge the responses, and this connection should be noted at the point where detectability is used to motivate the confound.
- [Appendix G.1, Table 1] Presentation issues: Appendix G.1 contains the typo 'Sumarry' for 'Summary', and in Table 1 the US/Israel–Iran topic label is typeset as 'US/IL-IR2' with a missing space, which should be corrected.
Circularity Check
No circularity: the confound claim is an empirical comparison, not a fitted or self-cited derivation.
full rationale
I find no circular step in the paper's derivation chain. The central empirical claims—that templated prompts are more detectable and less realistic than LLM-generated ones, and that neutral templated prompts elicit more polarized stance estimates in the direction of their fillers—are established by comparing three independently constructed prompt collections on responses from two examined models, using both human and LLM annotations. No parameter is fitted to a subset of the data and then renamed as a prediction; the lean values in Tables 4 and 10 are computed directly from judge labels via Equation 1, and the Wilcoxon comparisons in Section 3.4.3 are standard tests on those computed values. The author's self-citations (Chalkidis and Brandl 2024; Chalkidis and Sogaard 2026) are contextual and not load-bearing. The main limitation—LLM judges are validated only against each other, not against human stance annotations—is an external-validity threat that could affect the measured gap, but it does not make the derivation circular: the judges' labels are not defined in terms of the conclusion, and the paper explicitly discloses the threat in the Limitations and in Section 4. Likewise, using the same LLM family for realness annotation and stance judging is a possible preference confound, but the realness finding is also supported by human annotators. I therefore rate the paper as free of circularity, with the caveat that the stance estimates inherit the judge-validity limitation the authors acknowledge.
Assumptions & free parameters
free parameters (3)
- Filler texts for templated prompts (Table 9) =
'the climate change and the appropriate policy response' and variants; 19 paraphrases per stance-topic
- Topic pole definitions and stance-loaded lexicons (Table 5, Appendix B) =
Two poles per topic, A/B fixed
- Likert class thresholds for stance judging =
75% for classes 2 and 4, relaxed from IssueBench's 90%
assumptions (4)
- domain assumption The three-judge LLM ensemble produces stance labels valid for comparing construction methods.
- domain assumption Real prompts from WildChat and LMSys-Chat, filtered by IssueBench and re-labelled by the authors, are a valid realism anchor and representative seeds.
- domain assumption Bare API LLMs behave informatively about the deployed assistants users interact with.
- domain assumption Value-laden and speculative questions are treated as opinion sharing rather than information seeking.
Cite this review
Pith. "Pith review of Templated or fully synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance." pith.science (2026). https://pith.science/paper/6L4IQ3ZB
@misc{pith2026260811008,
author = {Pith},
title = {Pith review of: Templated or fully synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance},
year = {2026},
howpublished = {\url{https://pith.science/paper/6L4IQ3ZB}},
note = {Machine review of arXiv:2608.11008}
}
read the original abstract
Political stance detection in LLMs has long been dominated by closed-ended, multiple-choice political survey questions---originally designed for humans, and thus lacks the realism and nuance of human-AI interactions in the wild, while also being susceptible to sandbagging. The recent IssueBench framework substantially mitigates these limitations with templated prompts anchored in real-world chat logs. Given the rise in non-work-related use of GenAI assistants, we extend IssueBench beyond writing assistance to include two additional tasks, information seeking and opinion sharing. We argue that templated prompts still lack the nuance of real ones, especially for open-ended tasks, and remain recognisable as evaluation artefacts. We propose the use of fully synthetic (LLM-generated) prompts, produced under detailed instructions with real prompts as seeds. We assess the ecological validity of real, templated, and LLM-generated prompts in a small-scale study covering 3 highly contested policy issues and 3 recent geopolitical conflicts. Human and LLM annotators rank LLM-generated prompts as no less realistic than real ones and clearly more realistic than templated ones, and find that they carry their intended intent and stance more clearly; the LLMs separate templated prompts from the other two far more sharply than the humans do. In a case study, templated and LLM-generated prompts yield systematically different stance estimates for the same model, most visibly under neutral framings, where templated prompts overstate the model's leaning in the direction encoded by the topic-and-stance text (filler) slotted into their templates.
Figures
Reference graph
Works this paper leans on
-
[1]
Rogov and Alexander Panchenko and Natalia Loukachevitch and Elena Tutubalina , year=
Mikhail Salnikov and Dmitrii Korzh and Ivan Lazichny and Elvir Karimov and Artyom Iudin and Ivan Oseledets and Oleg Y. Rogov and Alexander Panchenko and Natalia Loukachevitch and Elena Tutubalina , year=
-
[2]
Santurkar, Shibani and Durmus, Esin and Ladhak, Faisal and Lee, Cinoo and Liang, Percy and Hashimoto, Tatsunori , title =
-
[3]
Chalkidis, Ilias and Brandl, Stephanie , booktitle =
-
[4]
Jochen Hartmann and Jasper Schwenzow and Maximilian Witte , year=
-
[5]
de Moura and Wei Zhang and Jose O
William Guey and Pierrick Bougault and Vitor D. de Moura and Wei Zhang and Jose O. Gomes , year=
-
[6]
Choudhary, Tavishi , journal=
-
[7]
arXiv preprint arXiv:2505.23836 , year=
Large language models often know when they are being evaluated , author=. arXiv preprint arXiv:2505.23836 , year=
-
[8]
International Conference on Learning Representations , volume=
Ai sandbagging: Language models can strategically underperform on evaluations , author=. International Conference on Learning Representations , volume=
Show all 76 references
-
[9]
2026 , month =
Claude Sonnet 5 System Card , institution =. 2026 , month =
2026
-
[10]
Stereotyping N orwegian Salmon: An Inventory of Pitfalls in Fairness Benchmark Datasets
Blodgett, Su Lin and Lopez, Gilsinia and Olteanu, Alexandra and Sim, Robert and Wallach, Hanna. Stereotyping N orwegian Salmon: An Inventory of Pitfalls in Fairness Benchmark Datasets. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and ...
2021 doi
-
[11]
2025 , eprint=
Scaling Synthetic Data Creation with 1,000,000,000 Personas , author=. 2025 , eprint=
2025
-
[12]
Gonzalez and Ion Stoica , booktitle=
Lianmin Zheng and Wei-Lin Chiang and Ying Sheng and Siyuan Zhuang and Zhanghao Wu and Yonghao Zhuang and Zi Lin and Zhuohan Li and Dacheng Li and Eric Xing and Hao Zhang and Joseph E. Gonzalez and Ion Stoica , booktitle=. Judging. 2023 , url=
2023
-
[13]
Bowman and Shi Feng , booktitle=
Arjun Panickssery and Samuel R. Bowman and Shi Feng , booktitle=. 2024 , url=
2024
-
[14]
Political Analysis , author=
Synthetic Replacements for Human Survey Data? The Perils of Large Language Models , volume=. Political Analysis , author=. 2024 , pages=. doi:10.1017/pan.2024.5 , number=
2024 doi
-
[15]
Political Analysis , author=
Out of One, Many: Using Language Models to Simulate Human Samples , volume=. Political Analysis , author=. 2023 , pages=. doi:10.1017/pan.2023.2 , number=
2023 doi
-
[16]
and Khashabi, Daniel and Hajishirzi, Hannaneh
Wang, Yizhong and Kordi, Yeganeh and Mishra, Swaroop and Liu, Alisa and Smith, Noah A. and Khashabi, Daniel and Hajishirzi, Hannaneh. Self-Instruct: Aligning Language Models with Self-Generated Instructions. Proceedings of the 61st Annual Meeting of the Association for Computa...
2023 doi
-
[17]
BBQ : A hand-built bias benchmark for question answering
Parrish, Alicia and Chen, Angelica and Nangia, Nikita and Padmakumar, Vishakh and Phang, Jason and Thompson, Jana and Htut, Phu Mon and Bowman, Samuel R. BBQ : A hand-built bias benchmark for question answering. Findings of the Association for Computational Linguistics: ACL 20...
2022 doi
-
[18]
S tereo S et: Measuring stereotypical bias in pretrained language models
Nadeem, Moin and Bethke, Anna and Reddy, Siva. S tereo S et: Measuring stereotypical bias in pretrained language models. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Proc...
2021 doi
-
[19]
C row S -Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models
Nangia, Nikita and Vania, Clara and Bhalerao, Rasika and Bowman, Samuel R. C row S -Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. doi:10.18...
2020 doi
-
[20]
Feng, Shangbin and Park, Chan Young and Liu, Yuhan and Tsvetkov, Yulia , editor =
-
[21]
Chatterji, Aaron and Cunningham, Thomas and Deming, David J and Hitzig, Zoe and Ong, Christopher and Shan, Carl Yan and Wadman, Kevin , year=
-
[22]
Marc Zao-Sanders , title =
-
[23]
Mrinank Sharma and Meg Tong and Tomasz Korbak and David Duvenaud and Amanda Askell and Samuel R. Bowman and Esin DURMUS and Zac Hatfield-Dodds and Scott R Johnston and Shauna M Kravec and Timothy Maxwell and Sam McCandlish and Kamal Ndousse and Oliver Rausch and Nicholas Schie...
-
[24]
Position: Political Neutrality in
Jillian Fisher and Ruth Elisabeth Appel and Chan Young Park and Yujin Potter and Liwei Jiang and Taylor Sorensen and Shangbin Feng and Yulia Tsvetkov and Margaret Roberts and Jennifer Pan and Dawn Song and Yejin Choi , booktitle=. Position: Political Neutrality in. 2025 , url=
2025
-
[25]
Croatian Journal of Philosophy , volume=
The impossibility of political neutrality , author=. Croatian Journal of Philosophy , volume=. 2010 , publisher=
2010
-
[26]
The Thirty-eighth Annual Conference on Neural Information Processing Systems , year=
Questioning the Survey Responses of Large Language Models , author=. The Thirty-eighth Annual Conference on Neural Information Processing Systems , year=
-
[27]
Do LLM s Exhibit Human-like Response Biases? A Case Study in Survey Design
Tjuatja, Lindia and Chen, Valerie and Wu, Tongshuang and Talwalkwar, Ameet and Neubig, Graham. Do LLM s Exhibit Human-like Response Biases? A Case Study in Survey Design. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00685
2024 doi
-
[28]
Llama meets EU : Investigating the E uropean political spectrum through the lens of LLM s
Chalkidis, Ilias and Brandl, Stephanie. Llama meets EU : Investigating the E uropean political spectrum through the lens of LLM s. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Vo...
2024 doi
-
[29]
PloS one , volume=
The political preferences of LLMs , author=. PloS one , volume=. 2024 , publisher=
2024
-
[30]
First Conference on Language Modeling , year=
Towards Measuring the Representation of Subjective Global Opinions in Language Models , author=. First Conference on Language Modeling , year=
-
[31]
2018 , month =
Mitchell, Amy and Gottfried, Jeffrey and Barthel, Michael and Sumida, Nami , title =. 2018 , month =
2018
-
[32]
1986 , publisher=
The morality of freedom , author=. 1986 , publisher=
1986
-
[33]
, author=
On the ambivalence-indifference problem in attitude theory and measurement: A suggested modification of the semantic differential technique. , author=. Psychological bulletin , volume=. 1972 , publisher=
1972
-
[34]
Global environmental change , volume=
Balance as bias: Global warming and the US prestige press , author=. Global environmental change , volume=. 2004 , publisher=
2004
-
[35]
Abecasis, Paulo and De Michiel, Federico and Basalisco, Bruno and Haanper
-
[36]
Adler, Steven , title =
-
[37]
biometrics , pages=
The measurement of observer agreement for categorical data , author=. biometrics , pages=. 1977 , publisher=
1977
-
[38]
Workshop on Trustworthy and Socially Responsible Machine Learning, NeurIPS 2022 , year=
Quantifying Social Biases Using Templates is Unreliable , author=. Workshop on Trustworthy and Socially Responsible Machine Learning, NeurIPS 2022 , year=
2022
-
[39]
2025 , month=
Shen, Judy Hanwen and Appel, Ruth and Tucker, Madeleine and Jagadish, Kamya and Maheshwary, Paruul and Askell, Amanda and Durmus, Esin , title=. 2025 , month=
2025
-
[40]
2011 , publisher=
On the political , author=. 2011 , publisher=
2011
-
[41]
Herrman, John , title =
-
[42]
Michel, Elie and Cicchi, Lorenzo and Garzia, Diego and Ferreira da Silva, Frederico and Trechsel, Alexander , year =
-
[43]
Wayne Brittenden , url =
-
[44]
and Mondría Terol, Teresa and Conger, Kate and Freedman, Dylan , title =
Thompson, Stuart A. and Mondría Terol, Teresa and Conger, Kate and Freedman, Dylan , title =
-
[45]
Swenson, Ali , title =
-
[46]
Neudert, Lisa Maria and Marchal, Nahema , title =
-
[47]
Ghose, Anuttama and Pallav, Pallav and Ali, S. M. Aamir , title =
-
[48]
Celeste Kidd and Abeba Birhane , title =
-
[49]
Ilias Chalkidis , year=
-
[50]
Ferdman, Avigail , journal=
-
[51]
Erfani, Farhang , journal=
-
[52]
npj Artificial Intelligence , volume=
Buyl, Maarten and Rogiers, Alexander and Noels, Sander and Bied, Guillaume and Dominguez-Catena, Iris and Heiter, Edith and Johary, Iman and Mara, Alexandru-Cristian and Romero, Rapha. npj Artificial Intelligence , volume=. 2026 , publisher=
2026
-
[53]
and Gebru, Timnit and McMillan-Major, Angelina and Shmitchell, Shmargaret , title =
Bender, Emily M. and Gebru, Timnit and McMillan-Major, Angelina and Shmitchell, Shmargaret , title =
-
[54]
Brainrot: Deskilling and Addiction are Overlooked AI Risks , year =
Chalkidis, Ilias and S. Brainrot: Deskilling and Addiction are Overlooked AI Risks , year =. Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency , pages =
2026
-
[55]
Tai-Quan Peng and Kaiqi Yang and Sanguk Lee and Hang Li and Yucheng Chu and Yuping Lin and Hui Liu , year=
-
[56]
The White House , title =
-
[57]
Simpson , title =
Raphael Racicot and Kurtis H. Simpson , title =
-
[58]
2024 , publisher=
Gu, Jiawei and Jiang, Xuhui and Shi, Zhichao and Tan, Hexiang and Zhai, Xuehao and Xu, Chengjin and Li, Wei and Shen, Yinghan and Ma, Shengjie and Liu, Honghao , journal=. 2024 , publisher=
2024
-
[59]
Ashery, Ariel Flint and Aiello, Luca Maria and Baronchelli, Andrea , journal=
-
[60]
Bias in Language Models: Beyond Trick Tests and Towards RUTE d Evaluation
Lum, Kristian and Anthis, Jacy Reese and Robinson, Kevin and Nagpal, Chirag and D ' Amour, Alexander Nicholas. Bias in Language Models: Beyond Trick Tests and Towards RUTE d Evaluation. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Vo...
2025 doi
-
[61]
2024 , url=
Wenting Zhao and Xiang Ren and Jack Hessel and Claire Cardie and Yejin Choi and Yuntian Deng , booktitle=. 2024 , url=
2024
-
[62]
Bronson Schoen and Evgenia Nitishinskaya and Mikita Balesni and Axel Højmark and Felix Hofstätter and Jérémy Scheurer and Alexander Meinke and Jason Wolfe and Teun van der Weij and Alex Lloyd and Nicholas Goldowsky-Dill and Angela Fan and Andrei Matveiakin and Rusheb Shah and ...
-
[63]
Alexander Meinke and Bronson Schoen and Jérémy Scheurer and Mikita Balesni and Rusheb Shah and Marius Hobbhahn , year=
-
[64]
2026 , publisher=
Turner, Cody and Eisikovits, Nir , journal=. 2026 , publisher=
2026
-
[65]
Hine, Emmie and Floridi, Luciano , journal=
-
[66]
2022 , publisher=
Bareis, Jascha and Katzenbach, Christian , journal=. 2022 , publisher=
2022
-
[67]
Investigating
Chalkidis, Ilias. Investigating. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. doi:10.18653/v1/2024.emnlp-main.312
2024 doi
-
[68]
Proceedings of the 19th Conference of the E uropean Chapter of the A ssociation for C omputational L inguistics (Volume 1: Long Papers)
Chalkidis, Ilias and Brandl, Stephanie and Aslanidis, Paris. Proceedings of the 19th Conference of the E uropean Chapter of the A ssociation for C omputational L inguistics (Volume 1: Long Papers). 2026. doi:10.18653/v1/2026.eacl-long.40
2026 doi
-
[69]
Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
R. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. doi:10.18653/v1/2024.acl-long.816
2024 doi
-
[70]
Bias in the East, Bias in the West: A Bilingual Analysis of
Lim, Ying Ying and R. Bias in the East, Bias in the West: A Bilingual Analysis of. Findings of the A ssociation for C omputational L inguistics: EACL 2026. 2026. doi:10.18653/v1/2026.findings-eacl.122
2026 doi
-
[71]
Transactions of the Association for Computational Linguistics , volume=
R. Transactions of the Association for Computational Linguistics , volume=. 2026 , publisher=
2026
-
[72]
Wright, Dustin and Arora, Arnav and Borenstein, Nadav and Yadav, Srishti and Belongie, Serge and Augenstein, Isabelle , booktitle=
-
[73]
Westwood, Sean J and Grimmer, Justin and Hall, Andrew B , journal=
-
[74]
2023 , eprint=
LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset , author=. 2023 , eprint=
2023
-
[75]
W ild V is: Open Source Visualizer for Million-Scale Chat Logs in the Wild
Deng, Yuntian and Zhao, Wenting and Hessel, Jack and Ren, Xiang and Cardie, Claire and Choi, Yejin. W ild V is: Open Source Visualizer for Million-Scale Chat Logs in the Wild. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: System Demons...
2024
-
[76]
2024 , url =
Esin Durmus and Liane Lovitt and Alex Tamkin and Stuart Ritchie and Jack Clark and Deep Ganguli , title =. 2024 , url =
2024
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.