{"id":"55451da2-cd01-4019-84bd-7be2f63d42e5","arxiv_id":"2607.20115","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Controlled linguistic rewrites, including meaning-preserving ones, shift political stance judgments in open-weight LLMs, and activation patching localizes the restoration signal to mid-to-late final-position block outputs.","lead":"This paper shows that four open-weight LLMs change their political stance judgments when the same policy statement is rewritten with different grammatical constructions, even when the meaning is intended to stay the same. It uses activation patching to trace those shifts to mid-to-late decoder layers, especially the block output at the final prompt position.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Concern: the 'meaning-preserving' rewrite controls may not actually preserve truth conditions; the behavioral claim and the patching localization both depend on this control, and Section 6 concedes small semantic changes are possible.","rationale":"I read the paper as making two connected claims: (1) controlled linguistic rewrites, including meaning-preserving ones, can systematically shift LLM political stance; and (2) activation patching localizes these shifts to mid-to-late decoder block outputs at the final prompt position. The second claim builds directly on the first, since patching is performed only on base/variant pairs with an observed polarity change. The most load-bearing assumption is therefore that the variant generation actually preserves (or inverts) meaning as labeled. The reader's weakest_assumption identifies exactly this, and the paper's Section 6 limitation admits the risk. I considered whether the patching design itself—final-position-only, no error bars, excluding low-flip rewrite types post hoc—might be the stronger concern. Those are legitimate weaknesses, but they affect the precision and generality of the localization claim rather than the foundation. If the semantic controls fail, the behavioral finding becomes uninterpretable and the localization claim has nothing reliable to localize. The proposed concrete test—independent validation and re-analysis on a confirmed mutual-entailment subset—directly settles this. Given the paper already received CONDITIONAL, and this concern is real but empirically checkable rather than demonstrably fatal, the verdict should remain CONDITIONAL (UNCHANGED).","tokens_in":12789,"tokens_out":3113,"duration_ms":36737,"concrete_test":"Independently validate all generated meaning-preserving variants: have at least two annotators (or a strong NLI system with high human agreement) label each base/variant pair as mutual entailment, one-way entailment, or unrelated, under strict truth-conditional criteria. Restrict the RQ1 analysis to pairs unanimously labeled as mutual entailment, then recompute the flip rates and articulation sensitivity (AS) from Section 4.1 on this clean subset. If the preserving-type flip rates and AS drop to the MU noise level or below the reported confidence intervals, the behavioral claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central RQ1 claim—that constructional choice alone shifts stance for meaning-preserving rewrites—rests on the premise that the four generated rewrite types (passive, it-cleft, wh-cleft, SVC) are truth-conditionally equivalent to their base statements. These rewrites were produced by Qwen3-8B few-shot prompting and validated only by the authors (Section 3.2). The paper itself concedes in Section 6: 'small semantic changes might still occur for linguistic rewrites.' This is not a formality: it-clefts and wh-clefts introduce presuppositions and focus-related meaning components; passivizing modal+ability constructions can produce degraded or scope-ambiguous outputs (the Table 1 passive drops the indefinite article: 'Vaccination certificate should be able to be requested...'); and SVC conversions may shift aspect or argument realization. If any part of the observed AS/flip-rate signal for the 'preserving' group comes from these artifacts rather than from surface construction, the conclusion that stance is sensitive to grammatical realization—not merely to hidden content drift—fails. Moreover, the patching analysis inherits this problem: restoring the base distribution could then be localizing response to artifact-induced content changes, not to constructional choice. Thus the weakest link is not the patching methodology but the validity of the input controls.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies whether controlled linguistic rewrites—both meaning-preserving (passive, it-cleft, wh-cleft, support verb construction) and meaning-inverting (negation, semantic opposite)—systematically change LLM political stance judgments, and where in the model these changes originate. Using four open-weight models and an extension of the ProbV AA dataset, the authors report flip rates, variance decompositions (PS/AS/MU), and activation patching results. They conclude that stance is sensitive to constructional choice beyond propositional content, and that mid-to-late decoder block outputs at the final prompt position provide the strongest restoration signal.","tokens_in":13077,"tokens_out":4356,"duration_ms":49417,"significance":"If the findings are robust, the paper would make a useful contribution to both LLM robustness evaluation and mechanistic interpretability: it moves beyond random paraphrases to controlled constructional variants, provides a reliability-oriented methodology (flip rates, Wasserstein distances, variance decomposition), and applies activation patching to a meaning-sensitive social-science task. The release of the dataset extension and the explicit prompts are strengths. However, the central behavioral claim—that meaning-preserving rewrites shift stance—rests on the validity of the generated rewrites, and the patching claim is currently reported without uncertainty quantification and is based on a restricted subset of rewrite types for the larger models. These issues are load-bearing for the paper's main conclusions.","major_comments":[{"comment":"The central RQ1 claim—that constructional choice alone shifts stance for 'meaning-preserving' rewrites—depends on the generated rewrites being truth-conditionally equivalent to the base statements. The rewrites are produced by Qwen3-8B few-shot prompting and manually validated only by the authors; Section 6 concedes that 'small semantic changes might still occur for linguistic rewrites.' This is not a formality: the Table 1 passive example ('Vaccination certificate should be able to be requested...') is arguably not equivalent to the original modal/argument structure, and it-clefts/wh-clefts introduce focus presuppositions. If part of the AS/flip-rate signal for the preserving group comes from content drift, the conclusion that stance is sensitive to grammatical realization is undermined. Please report inter-annotator agreement or independent validation, and provide a per-item analysis e","section":"§3.2, §6, Table 1"},{"comment":"The activation patching results are reported as mean normalized restoration scores without confidence intervals, per-example variance, or significance tests. Moreover, the analysis includes only a subset of rewrite types per model: for Gemma-3-12B and Qwen3-14B only negation and opposite; for Qwen3-4B negation, opposite, and it-clefts. The abstract and §5 generalize to 'mid-to-late decoder layers, especially block outputs at the final prompt position' across all rewrites, but the meaning-preserving constructional rewrites are absent from the patching analysis for the two larger models. The claim should be restricted to the analyzed rewrite types, and the patching results need error bars or a per-example distribution to support the localization conclusion.","section":"§4.2, Figures 3/4, Appendix D"},{"comment":"RQ2 asks where rewrites 'causally determine' the response shift, and §5 states that the shifts 'can be causally localized to specific internal components.' However, the method establishes only sufficiency of restoring clean behavior by patching activations at the final prompt position, not necessity, and it does not trace the origin of the rewrite-sensitive computation across earlier positions. Section 6 correctly notes that the analysis localizes 'sufficient restoration sites rather than the origin.' This caveat should be carried through the abstract and conclusion; as written, 'causally localized' and 'causally determine' overstate the evidence.","section":"§3.6, §4.2, RQ2"}],"minor_comments":[{"comment":"Grammar: 'stance instability affect' should be 'stance instability affects.'","section":"Abstract"},{"comment":"The passive example is missing the indefinite article before 'vaccination certificate' and the phrasing 'should be able to be requested' is awkward; this example precisely illustrates the semantic-validity concern raised in the major comments and should be revised or replaced.","section":"Table 1"},{"comment":"The manual validation of the generated rewrites is described in one sentence. Please specify how many annotators were involved, whether disagreements were resolved, and what criteria (e.g., truth-conditionality, naturalness) were used.","section":"§3.2"},{"comment":"The JSON output format in the prompts is malformed in several places (e.g., stray backticks and inconsistent field names like 'variants' vs 'variant'). This makes replication harder; please clean up the prompt listings.","section":"Appendix A.1"},{"comment":"For the flip-rate calculation, pairs with unclear polarity are excluded. Given the low clear-leaning rates for the larger models (40–56%), it would be helpful to report the number of eligible pairs per model and variant, since the flip rates are computed on very different denominators.","section":"§4.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope and the dataset plus code are valuable. The two main blockers are (1) the validity of the meaning-preserving control and (2) the unsupported generality of the patching claim without uncertainty quantification and with restrictive rewrite-type inclusion. Both are fixable with additional analysis or more careful hedging, so I do not recommend rejection, but the current version overclaims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a genuinely useful paper for the NLP robustness crowd. Where prior work used random paraphrases or prompt wording tweaks, this one isolates specific constructional rewrites—passives, it-clefts, wh-clefts, and support verb constructions—and shows they can shift LLM political stance judgments. The dataset extension of ProbVAA is a real contribution, and the variance decomposition cleanly separates articulation sensitivity from sampling noise, which is more than most robustness papers bother to do. The behavioral result is credible: AS values exceed MU across models, the flip rates are sensible, and the larger models are more stable, which matches prior findings. Credit where due: they acknowledge the main limitations in Section 6, including exactly the concern I'd raise first.\n\nThat concern is the 'meaning-preserving' rewrite controls. The stress-test note is on target. It-clefts and wh-clefts bring presuppositions and focus effects; the passive example in Table 1 ('Vaccination certificate should be able to be requested...') is degraded at best; and SVC conversions can shift aspect or predicate semantics. The paper concedes in Section 6 that small semantic changes might still occur. This matters because both the behavioral claim and the patching localization inherit whatever artifacts the rewrites carry. I don't think it's disqualifying—the effect sizes are modest, and the meaning-reversing rewrites show larger shifts, which is consistent with the controls being mostly okay—but it needs harder evidence than author manual validation. A human annotation study on truth-condition preservation, or a stricter generation pipeline with round-trip diagnostics, would strengthen the revision.\n\nThe activation patching results are suggestive rather than conclusive. The abstract's claim about mid-to-late decoder block outputs at the final prompt position is supported by the mean restoration scores in Figures 3 and 4, but there are no confidence intervals or per-example variance, and the exclusion of low-flip rewrite types per model is post hoc. The intervention is also limited to the final token position, which is honest but narrow. All of this makes the mechanistic conclusion conditional, not broken.\n\nThe GitHub link is unverifiable and should be fixed. That's minor.\n\nWho gets value from this? Anyone working on LLM robustness, political stance evaluation, or activation patching. It's a solid empirical building block, not a breakthrough. It deserves a serious referee: the dataset is reusable, the question is well-motivated, and the flaws are addressable in revision. I'd send it to peer review and ask for stronger validation of the rewrite controls and uncertainty quantification on the patching results.","headline":"Useful, honest empirical study; the behavioral claim about constructional stance shifts is probably right, but control validity and missing error bars on patching keep the localization claim conditional.","tokens_in":13552,"tokens_out":2024,"would_cite":true,"duration_ms":25304,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that controlled linguistic rewrites—including meaning-preserving ones—systematically shift LLM political stance judgments, and activation patching localizes the shift to mid-to-late decoder block outputs at the final prompt","keywords":["LLM stance","linguistic construction","meaning-preserving rewrites","activation patching","political stance","semantic invariance","it-clefts","variance decomposition"],"falsifier":"Annotate a random sample of the meaning-preserving rewrites for whether they preserve truth conditions and presuppositions; if a large share of it-clefts or passives introduces presuppositions not present in the base (e.g., an existential presupposition of the clefted constituent), then the observed stance shifts could be semantic artifacts, falsifying the claim that constructional choice alone shifts stance.","tokens_in":12672,"feed_emoji":"🗳️","tokens_out":4949,"duration_ms":51901,"temperature":0.7,"pith_summary":"This paper asks whether the grammatical form of a policy statement can change a language model's political stance, even when the meaning is preserved. Using four open-weight models and a set of political statements with six controlled rewrite types, it finds that meaning-preserving constructions—active/passive, it-clefts, wh-clefts, and support-verb constructions—do shift agreement judgments, though less than negation and semantic opposites. Activation patching, which replaces internal activations from one phrasing with those from another, shows that mid-to-late decoder layer outputs at the final prompt position are sufficient to restore the original stance, while attention and MLP sublayer patches are weaker. If correct, stance judgments are not fixed by propositional content alone, and syntactic realization is a causal variable that evaluation pipelines must control.","feed_headline":"Syntax changes LLM political stances, even when meaning is unchanged","feed_subtitle":"Controlled rewrites move agreement by about half to one Likert point; activation patching finds late decoder layers carry the shift.","key_machinery":"The central machinery is a set of six controlled linguistic rewrites—explicit negation, semantic opposites, active/passive, it-clefts, wh-clefts, and support-verb constructions—each either preserving or systematically inverting truth conditions, paired with activation patching. Patching substitutes activations from the original statement into the forward pass of the rewritten statement at three sites per layer (block output, attention sublayer, MLP sublayer) and measures how much the original seven-point Likert distribution is restored, using 1-Wasserstein distance and a normalized restoration score.","core_discovery":"Across two model families and four model sizes, stance instability appears under both meaning-reversing edits (explicit negation, semantic opposites) and meaning-preserving syntactic rewrites (active/passive, it-clefts, wh-clefts, support-verb constructions). Variance decomposition attributes more variation to issue content than to phrasing, but phrasing alone moves judgments by roughly half to over one Likert point on average, beyond sampling noise. Activation patching at the final answer position shows that swapping in block-output activations from the original statement restores the original stance distribution most strongly in mid-to-late decoder layers; attention and MLP patches mostly","pith_inferences":["A testable extension: if the same syntactic rewrites are applied to non-political factual agreement statements and flips disappear, the constructional effect is specific to opinionated content; if they persist, it is a general readout phenomenon.","The final-position block-output locus suggests stance is read out from a late residual-stream summary rather than from a dedicated early 'stance feature'; probing intermediate representations could confirm whether focus or polarity is encoded separately before the final position.","Benchmark designers could control for cleftability and passivizability when selecting stimulus sentences; otherwise information-structure confounds are built into the data.","Extending to languages with richer morphology (e.g., free word order) could reveal whether the locus shifts when the same constructions are encoded differently."],"forward_implications":["Meaning-preserving rewrites must be treated as a potential source of measurement error in any LLM-based stance or opinion evaluation.","Larger models are not simply more robust: they flip less often but also retreat toward neutral, so scale alone does not guarantee semantic invariance.","Because the restoration signal concentrates in block outputs at the final prompt position, interventions at those layers are the most promising targets for stabilizing stance judgments.","Meaning-preserving and meaning-inverting rewrites appear to act through the same downstream locus, suggesting a shared mechanism for prompt-sensitivity once a shift occurs.","Combining rewrites amplifies instability beyond individual edits, so evaluations that test only single transformations may understate phrasing sensitivity."],"fun_headline_variants":["Syntax alone shifts LLM political stances, even without meaning change","Causal tracing pinpoints where syntactic rewrites alter LLM stance","Meaning-preserving rewrites still move LLM political opinions","LLM stance swayed by phrasing; late decoder layers carry the shift"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The rewrites must actually preserve or invert the statement's meaning as labeled; the paper concedes that small semantic changes may remain despite manual validation, and if clefts or passives introduce presuppositions or content shifts, the conclusion that construction alone drives stance changes would not hold.","fun_headline_variants_meta":{"raw":{"variants":["Syntax alone shifts LLM political stances, even without meaning change","Causal tracing pinpoints where syntactic rewrites alter LLM stance","Meaning-preserving rewrites still move LLM political opinions","LLM stance swayed by phrasing; late decoder layers carry the shift"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000489,"raw_usage":{"total_tokens":2217,"prompt_tokens":691,"completion_tokens":1526,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":435,"completion_tokens_details":{"reasoning_tokens":1464}},"tokens_in":435,"tokens_out":1526,"duration_ms":10962,"temperature":1.0,"reasoning_tokens":1464,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T10:42:53.117162+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Annotate a random sample of the meaning-preserving rewrites for whether they preserve truth conditions and presuppositions; if a large share of it-clefts or passives introduces presuppositions not present in the base (e.g., an existential presupposition of the clefted constituent), then the observed stance shifts could be semantic artifacts, falsifying the claim that constructional choice alone shifts stance.","supporting_citations":[],"review_version":1}