{"id":"88ab4178-34b3-42a4-a328-921adbb21327","arxiv_id":"2608.11008","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Templated prompts systematically overstate the political leaning of LLMs under neutral framings, and LLM-generated prompts, validated against real user prompts, are a more realistic alternative.","lead":"Political stance measurements of chatbots depend heavily on how the question prompt is built: formulaic template prompts make models look more one-sided than realistic AI-written prompts do. This paper shows AI-generated prompts pass as human-written, while template prompts leak hidden slants, and this construction choice changes the measured stance.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"LLM judges see the prompts and are never checked against human labels; the 0.48 vs 0.07 neutral gap could be partly judge-side.","rationale":"The reader's verdict is CONDITIONAL, and the weakest assumption identified is exactly the judge validity. I read the paper as a careful, honest study: the prompt construction results (realness/detection) are well-designed, the data are released, and the replication on a second model is real evidence. The central stance claim, however, is only as strong as the measurement chain. The 0.48/0.07 difference and the 14/18 result are computed from an LLM judge ensemble that is never anchored to human labels, and the same ensemble is used for both models. The paper acknowledges this limitation, which is exactly why the verdict should stay CONDITIONAL rather than ACCEPT or REJECT. A human-label validation is a concrete, feasible condition: a few hundred responses labeled by two or three humans would settle whether the confound is in the models or in the judges. If that validation fails, the headline claim weakens substantially; if it passes, the paper makes a solid methodological contribution. This is a good-faith assessment: the concern is not that the authors are hiding something, but that the key inference rests on an unvalidated measurement instrument, a point the authors themselves flag. The secondary statistical concern about per-setting independence and missing per-cell N is real but cannot reverse the result as directly as a judge-side artifact could, so I treat the judge validation as the load-bearing condition.","tokens_in":36841,"tokens_out":6156,"duration_ms":54009,"concrete_test":"Take a stratified sample of responses from the 18 neutral settings (and a random subset of sided settings) from both GPT 5.4 mini and Grok 4.3; strip the user prompt; have two or three trained human annotators label each response on the same 5-point stance rubric used by the LLM judges; aggregate by majority vote or mean and recompute the mean absolute distance from neutral per construction method plus the Wilcoxon signed-rank over the 18 settings. If human labels reproduce the roughly 0.4-point gap in the same direction for both models, the claim survives; if the gap shrinks or reverses, the headline confound is at least partly judge-side. A secondary check would be to give the LLM judges the same responses without the user prompt to see how much of the templated-vs-synthetic gap depends on seeing the prompt.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim (Section 3.4.3) is that responses to neutral templated prompts sit 0.48 scale points from neutral versus 0.07 for LLM-generated prompts (14 of 18 settings, p=.001; Grok 4.3: 0.36 vs 0.05, 15 of 18, p<.001). These distances are all produced by a majority-vote ensemble of three LLM judges (Section 3.3). Three properties make this the load-bearing leg of the argument. First, the judges are never validated against human stance annotations; the paper says so explicitly in the Limitations: 'We validate the judges against each other rather than against human stance annotations of responses.' Second, the judge prompt includes the full user prompt, and the judge is instructed not to factor in whether the request was one-sided; there is no check on whether the judges complied, so the outcome label could be contaminated by the very variable (prompt construction) the study manipulates. Third, agreement is lowest on the LLM-generated responses (ordinal Krippendorff's alpha .82 vs .91 for templated), exactly the responses that are most hedged and nuanced. If the judges systematically over-polarize responses to stylized, filler-loaded templated prompts and systematically round nuanced LLM-generated responses toward neutral, the observed gap would be an artifact of the measurement chain rather than of the models. Replication on Grok 4.3 does not resolve this because the same judge ensemble and the same prompt-contamination mechanism are used. The confound claim is about what a user would actually encounter, so the measurement must be anchored to human perception of stance before it can be called a confound.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper argues that the way evaluation prompts are constructed is a confound in measuring the political stance of LLMs. The authors extend IssueBench beyond writing assistance to two additional intents (information seeking and opinion sharing), propose fully synthetic (LLM-generated) prompts anchored in real chat-log prompts as seeds, and validate real, templated, and synthetic prompts on realism and construct clarity using three human and three LLM annotators. They then compare stance estimates from templated and synthetic prompts for two deployed models, GPT 5.4 mini and Grok 4.3, judged by a majority-vote ensemble of three open-weight LLMs. The central claim (Section 3.4.3) is that neutral templated prompts elicit responses systematically farther from neutral, in the direction encoded by the template filler, than neutral synthetic prompts do (0.48 vs 0.07 scale points on GPT 5.4 mini, 14 of 18 settings, Wilcoxon p=.001; replicated on Grok 4.3 at 0.36 vs 0.05, 15 of 18, p<.001), so that a templated study would overstate the model's leanings precisely where neutrality is the target of measurement. The paper also reports that synthetic prompts are ranked at least as realistic as real prompts and clearer in intent and stance, and that LLM annotators separate templated prompts from real ones more sharply than human annotators do.","tokens_in":36988,"tokens_out":26033,"duration_ms":237360,"significance":"The contribution is significant if it holds: it identifies a direction-predictable, systematic confound in a widely used evaluation framework (IssueBench-style templating) rather than a random artefact, and pairs the critique with a concrete, validated alternative construction method. The paper's strengths are substantial: the full resource set (prompts, human and LLM annotations, model responses, stance judgments) is released on HuggingFace, the central effect is replicated on two models from different developers, the judge ensemble consists of three models with no developer overlap with the examined models, the direction-of-effect evidence is presented transparently (14/18 and 15/18 settings), and the Limitations are unusually candid, including the explicit statement that the stance judges are validated only against each other. The main risk is that the quantitative centrepiece is produced entirely by an LLM judge chain that is never checked against human stance labels, and the judge prompt itself contains the user prompt, so prompt-contamination of the labels is a concrete, uneliminated alternative explanation for the measured gap.","major_comments":[{"comment":"The load-bearing quantitative result rests on an unvalidated LLM judge chain. The central claim of Section 3.4.3 — that neutral templated responses sit 0.48 scale points from neutral versus 0.07 for LLM-generated responses — is entirely produced by the majority-vote ensemble of DeepSeek V4 Pro, Mistral Large 3, and Nemotron 3 Ultra, and the Limitations state: 'We validate the judges against each other rather than against human stance annotations of responses.' Because the judge prompt (Appendix E) displays the full user prompt and merely instructs the judge to ignore whether the request was one-sided, with no compliance check, a judge that unconsciously uses prompt framing would label responses to the demonstrably stance-leaking neutral fillers (Section 2.2.2) as more sided in the filler's direction, while rounding the more hedged LLM-generated responses (where judge agreement is lowest, ordinal α=.82 vs .91) toward the neutral class — producing exactly the observed pattern. The Grok 4.3 replication (Appendix G.1) uses the same judge ensemble and therefore cannot arbitrate. The paper should either (a) collect human stance annotations on a sample of responses and report judge–human agreement per construction method, or (b) run a control that fixes the response while removing or swapping the user prompt and checks label stability. Without one of these, the claim that templated prompts overstate the model's leanings, rather than the judges' labels reflecting the prompt, is under-supported.","section":"§3.3, §3.4.3, Appendix E, Limitations ('LLM judges')"},{"comment":"Section 3.3 relaxes Likert classes 2 and 4 from Röttger et al.'s 'overwhelmingly (90%)' to 'substantially (75%)', which moves every response with 75–90% single-side emphasis out of the neutral class 3. The change is applied to both construction methods, but it interacts directly with the quantity being measured: the central comparison is precisely the distance of responses from the neutral class, and if templated responses are mildly but genuinely sided (filler-driven) while LLM-generated responses are hedged, the relaxed rubric amplifies the measured gap. The manuscript should report the Section 3.4.3 comparison under the original 90% thresholds (or under an alternative continuous-scale rubric) to show that the 0.48 vs 0.07 gap is not threshold-driven, especially since Section 4 itself describes this relaxation as treating a symptom rather than a cause.","section":"§3.3 (Likert rubric relaxation)"},{"comment":"The headline claim in Section 3.4.3 states that neutral templated prompts are 'systematically further from neutral, in the direction the filler encodes (14 of 18 neutral settings...)'. The arithmetic component of this count is verifiable from Table 4 (14 settings with |templated lean| > |LLM-generated lean|, 3 with the reverse, 1 tie), but the direction component is not auditable from the paper's tables: no per-setting table or figure maps each neutral setting to the filler-encoded direction and the observed direction. This matters because at least one topic appears inconsistent with the stated mechanism: the paper's own principle (Section 2.2.2 and Section 3.4.2) reads 'the US and Israeli strikes on Iran' as siding against the striker, i.e., toward pole B (pro-Iran), yet the neutral templated leans in Table 4 for US/Israel–Iran are -0.06 (information seeking) and -0.06 (opinion sharing), i.e., toward pole A. The authors should publish the per-setting direction coding (or a stacked plot of the 18 neutral settings with the filler-encoded direction marked) and either reconcile or explicitly explain these apparent counter-directional settings.","section":"§3.4.3, Tables 4 and 9, §2.2.2"},{"comment":"There is an internal contradiction in the definition of the climate-change poles. Table 5 and Table 9 assign Climate-urgency to Pole A and Climate-moderation to Pole B, so negative leans in Table 4 mean climate urgency, consistent with the text of Section 3.4.2. The Detection Guidelines in Appendix D, however, assign 'climate and ecological change not being severe, and the policy response being adequate' to Pole A and 'being severe... inadequate' to Pole B, i.e., the reverse lettering, with the note that 'in favour means downplaying the severity'. Since the direction claim in Section 3.4.3 ('the direction the filler encodes') depends on the pole lettering, and climate change is one of the two topics that together account for 85% of the total neutral-setting divergence (Section 3.4.2), the two appendices must be made consistent before the direction pattern is verifiable by a reader or replicator.","section":"Appendix A (Table 5), Appendix C (Table 9), Appendix D (Detection Guidelines)"}],"minor_comments":[{"comment":"The Wilcoxon signed-rank test in Section 3.4.3 treats the 18 neutral settings as exchangeable units even though they share six topics, and the Limitations acknowledge this non-independence; the main text should carry that caveat alongside the reported p=.001, or the authors should report a cluster-robust or judgment-level test, since the nominal p-value is otherwise optimistic. The direction counts (14/18 and 15/18) are the more robust evidence and are not in question.","section":"§3.4.3, Limitations (statistical dependence)"},{"comment":"Tables 4 and 10 do not report the number of responses per setting (the N columns denote the neutral user stance, not counts); the authors should add per-setting response counts, particularly because refusal rates differ by construction method (5% vs 2%, Section 3.3), so the per-setting sample sizes underlying the lean estimates are not recoverable from the paper.","section":"Tables 4 and 10"},{"comment":"The conclusion that human annotators rank LLM-generated prompts as 'no less realistic than real ones' rests on three human annotators with near-chance agreement (Kendall's W=.21, Appendix F); the abstract and Section 2.2.1 should hedge this as essentially a null result rather than a demonstrated parity, since the human ordering evidence is thin.","section":"§2.2.1, Appendix F, Abstract"},{"comment":"The same three LLMs serve as the realness and detection annotators (Appendix D) and as the stance judges (Section 3.3); the detectability finding (LLM annotators separate templated prompts from real ones more sharply than humans do) is therefore produced by the same models that later judge the responses, and this connection should be noted at the point where detectability is used to motivate the confound.","section":"Appendix D and §3.3"},{"comment":"Presentation issues: Appendix G.1 contains the typo 'Sumarry' for 'Summary', and in Table 1 the US/Israel–Iran topic label is typeset as 'US/IL-IR2' with a missing space, which should be corrected.","section":"Appendix G.1, Table 1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is single-authored with a transparent AI-assistance disclosure, and I see no integrity concern; the self-citations are appropriate and the empirical claims are backed by released data. My main reservation is that the headline quantitative claim is presented as established while the paper itself concedes the absence of human validation of the stance judges. The paper would be materially strengthened by presenting the prompt-level filler-leakage result (Section 2.2.2) as the primary, human-anchored claim and the response-level gap as conditional on the judgment validation requested in the major comments; in its current form the strength of the conclusion in Section 3.4.3 exceeds what the measurement chain supports."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing to know: this paper makes a specific, testable claim—that templated prompts in the IssueBench style are not neutral instruments, and that LLM-generated prompts are a better proxy for real users—and it backs that claim with open data, a second-model replication, and unusually candid limitations. I'd send it out. It deserves referee time.\n\nWhat's actually new: earlier work (Seshadri et al. 2022; Blodgett et al. 2021) showed template rewording changes measured bias. This paper extends that to political stance in open-ended tasks and proposes a validated alternative: synthetic prompts generated from real seeds, with humans and LLMs ranking realness and recovering intent and stance. The three-way validation protocol is the real contribution. The finding that neutral templated fillers shift GPT-5.4 mini's estimated lean by roughly 0.4 scale points (14 of 18 settings, p=.001) and that Grok 4.3 shows the same pattern (15 of 18, p<.001) is a clean, falsifiable result.\n\nWhere I'm less convinced: the stance comparison rests on three LLM judges whose labels are never checked against human stance annotations—the paper says so in the Limitations. The judge prompt includes the user prompt and tells judges to ignore one-sidedness, but there's no compliance check. Agreement is lowest on exactly the responses (LLM-generated, hedged, nuanced) that drive the headline comparison. So part of the 0.48 vs 0.07 gap could be judge-side. The Grok replication does not fully fix this, since the same judges are used. Also, the per-setting Wilcoxon treats 18 settings sharing 6 topics as independent, and the main table lacks per-cell N and variance. Both are acknowledged, but they limit how much weight the exact numbers can carry.\n\nThat said, the confound itself is not merely a judge artifact. The human detection task independently shows that fillers like \"the Russian invasion and action in Ukraine\" are read by humans as siding against Russia. So the direction of the effect has convergent support; the open question is magnitude and generalizability, not existence.\n\nBottom line: this is a solid, honest empirical paper. Anyone building or using political stance benchmarks should read it. With human stance labels on a subset of responses and a broader model set, it would be even stronger. Proper peer review, not desk reject—and I would cite it.","headline":"A careful empirical study showing templated prompts skew LLM stance measurements; the effect is real and well-supported, though the size of the gap rests on LLM judges that merit human anchoring.","tokens_in":37735,"tokens_out":2261,"would_cite":true,"duration_ms":29430,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Templated prompts constructed to appear neutral still push LLM political responses toward the stance encoded in their fillers, making prompt construction a confound in stance measurement.","keywords":["prompt construction","political stance detection","templated prompts","LLM-generated prompts","ecological validity","LLM-as-a-judge","sycophancy","confound"],"falsifier":"Have human annotators label the same model responses on the same 5-point stance scale. If human labels put neutral templated and neutral LLM-generated responses at equal distance from neutrality, the confound gap disappears; if human–ensemble agreement is worse on LLM-generated responses than on templated ones, the reported gap may be a judge artefact rather than a prompt-construction effect.","tokens_in":36441,"feed_emoji":"⚖️","tokens_out":5353,"duration_ms":48424,"temperature":0.7,"pith_summary":"The paper claims that how a political prompt is built—by filling a template or by having an LLM write a freestanding request—changes the political stance you measure in the answering model, even when the prompt is meant to be neutral. On 18 neutral topic-by-intent settings, templated prompts pushed GPT-5.4 mini's responses further toward the pole encoded in the template's filler (14 of 18 settings, Wilcoxon signed-rank p = .001), and the pattern replicated on Grok 4.3 (15 of 18 settings). The cause is structural: a filler must stay grammatical under all three stance versions, and the phrasing that does so is rarely stance-free. Two studies that differ only in prompt construction would therefore report different stances for the same model, with the templated study overstating leanings under neutral framings.","feed_headline":"Templated prompts bias LLM political stance readings","feed_subtitle":"Neutral templates push answers toward the stance baked into their fillers; synthetic prompts do not.","key_machinery":"The slot-and-filler template is the load-bearing object: a fixed request skeleton such as 'What is your opinion on X?' with a filler phrase carrying the topic and stance. Because the same filler must remain grammatical as neutral, pole-A, and pole-B versions, fillers like 'the Russian invasion and action in Ukraine' carry a presupposition, and the model reads that presupposition as a leading question. The paper shows the same fillers shift stance estimates in the direction they encode, whereas LLM-generated prompts, produced under detailed instructions with real prompts as seeds, express topic and stance without being constrained to stay grammatical across stance variants and therefore do not leak the same siding.","core_discovery":"The central discovery is that templated prompts systematically distort stance measurements of LLMs, while LLM-generated prompts anchored in real user prompts do not. In a small-scale study covering three contested policy issues and three geopolitical conflicts, human and LLM annotators ranked LLM-generated prompts as no less realistic than prompts drawn from real chat logs, and clearly more realistic than templated ones; LLM-generated prompts also carried their intended intent and stance more clearly. In the stance case study, neutral templated prompts elicited responses 0.48 scale points from neutral on average, versus 0.07 for neutral LLM-generated prompts, with the distortion always pointing in the direction the filler encodes: 'the Russian invasion and action in Ukraine' reads as a presupposition that the invasion is unjustified, and 'the climate change and the appropriate policy response' reads as an assertion of severity. The paper concludes that prompt construction is not a neutral design choice: same model, same judges, same topics, but different construction methods report different political stances, and the templated method overstates the model's leanings precisely where neutrality is what is being measured.","pith_inferences":["We infer the confound is model-independent in direction, since it is a property of the filler's presuppositional content rather than of the model parsing it; a direct test would be to build fillers that are genuinely symmetric across stances and check whether the templated-vs-synthetic gap shrinks.","We infer that 'the stance of an LLM' is underdetermined until the prompt distribution is specified: political stance reports are joint products of model, prompt form, and judge, not properties of the model alone.","We infer that the blurred boundary between information seeking and opinion sharing (human intent agreement as low as κ ≈ .27) means cross-task stance comparisons are only meaningful when studies state how value-laden and speculative questions are classified.","We infer that since LLM judges separate templated from natural prompts more sharply than humans do, prompt realism could be screened automatically before an evaluation corpus is deployed, rather than only by human panels."],"forward_implications":["Neutral templated prompts overstate a model's political leaning in the direction the filler encodes, so any existing stance estimate built on such prompts should be treated as potentially inflated.","Templated prompts are more legible as evaluation artefacts to LLM annotators than to humans (Kendall's W = .65 vs .21); if models can recognize them, they may answer differently than they would for a real user.","LLM-generated prompts with real prompts as seeds preserve experimental control over topic, intent, and stance while matching real prompts in realism, offering a practical alternative to templates.","Extending stance measurement beyond writing assistance changes what is measured: writing assistance mostly captures instruction-following, while information seeking and opinion sharing surface the model's own stance and asymmetric sycophancy.","Two studies of the same model, using the same judges and topics but different prompt construction, will report different political stances; prompt construction must be reported and validated as a design choice."],"supporting_citations":[{"why":"Supplies the templated-prompt framework, the stance detection rubric, and the writing-assistance template collection the paper extends and tests.","marker":"(Röttger et al., 2026)"},{"why":"Documents instability and format sensitivity of survey-style political questions, motivating open-ended realistic prompts.","marker":"(Röttger et al., 2024)"},{"why":"Provides the WildChat real chat logs used as a source of real prompts and seed examples.","marker":"(Zhao et al., 2024)"},{"why":"Provides the LMSys-Chat real conversation logs used as a source of real prompts and seed examples.","marker":"(Zheng et al., 2023a)"},{"why":"Gives evidence that frontier models can distinguish evaluation from deployment, grounding the concern that recognisable templated prompts are unfit for testing.","marker":"(Needham et al., 2025)"},{"why":"Shows models can strategically underperform on evaluations, supporting why prompt realism is a precondition for valid measurement.","marker":"(Van Der Weij et al., 2025)"},{"why":"Provides the sycophancy interpretation used to explain the model's accommodation of the user's stance.","marker":"(Sharma et al., 2024)"}],"fun_headline_variants":["Templated prompts skew LLM political stances","LLM stance tests: prompt type twists results","Neutral templates inflate LLM political leanings","Synthetic prompts beat templated for LLM stance"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument leans on three AI judges whose stance ratings are only checked against each other, never against human ratings of the same responses, and they disagree most on the nuanced responses the main gap depends on.","fun_headline_variants_meta":{"raw":{"variants":["Templated prompts skew LLM political stances","LLM stance tests: prompt type twists results","Neutral templates inflate LLM political leanings","Synthetic prompts beat templated for LLM stance"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000168,"raw_usage":{"total_tokens":1314,"prompt_tokens":1054,"completion_tokens":260,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":670,"completion_tokens_details":{"reasoning_tokens":198}},"tokens_in":670,"tokens_out":260,"duration_ms":3125,"temperature":1.0,"reasoning_tokens":198,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:13:52.745620+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have human annotators label the same model responses on the same 5-point stance scale. If human labels put neutral templated and neutral LLM-generated responses at equal distance from neutrality, the confound gap disappears; if human–ensemble agreement is worse on LLM-generated responses than on templated ones, the reported gap may be a judge artefact rather than a prompt-construction effect.","supporting_citations":[],"review_version":2}