{"id":"df7b5575-1461-4511-b4c9-1e81fb1faedb","arxiv_id":"2606.11205","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":4,"one_line_summary":"Centroid-difference steering for sycophancy in Llama-3-8B reduces agreement with factually correct statements as well as sycophantic ones, and this non-specificity is invisible under standard single-stance evaluation.","lead":"This paper introduces dual-stance evaluation for activation steering, testing both sides of a topic to check if a sycophancy-reduction direction also suppresses agreement with factually correct statements. It finds that centroid-difference steering on Llama-3-8B reduces agreement indiscriminately, including with true statements like 'the Earth is round,' revealing a specificity gap invisible to standard evaluation.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"Subspace analysis with 149:54 sample imbalance and no confidence intervals underpins the key geometric claims; bootstrap would settle whether 'equal projection' is real or underpowered.","rationale":"The reader correctly identified the most load-bearing concern: the subspace analysis (§4.5, Appendix D) uses a 149:54 sample split with no confidence intervals, and the 'equal projection' claim (ratio 0.90–0.97) depends on the smaller group's subspace being reliably estimated. This is the weakest link in the argument because it supports the 'geometric puzzle' framing — the claim that the model distinguishes the two agreement types geometrically but the steering direction cannot exploit this distinction. If the projection ratio is an artifact of poor subspace estimation (a noisy 10-dimensional subspace in 4096 dimensions from 54 samples will tend to capture more variance from any direction), the puzzle dissolves: the steering direction might actually project more onto one subspace than the other, and the behavioral dissociation could have a static geometric explanation after all. However, I recommend UNCHANGED rather than a harsher verdict because: (1) the paper's primary contribution — the dual-stance evaluation framework and the behavioral non-specificity finding — does not depend on the subspace analysis at all; (2) the random-split control (z=-7.86 for Grassmann similarity) is a large effect that is likely robust to sample size concerns; (3) the paper is appropriately hedged about the geometric conclusions ('we cannot rule out that unmeasured static properties might account for the dissociation'). The subspace analysis is supplementary to the main behavioral findings, and the paper acknowledges its limitations. A bootstrap analysis would strengthen or weaken the geometric puzzle narrative but would not change the core methodological contribution or the behavioral characterization. The reader's CONDITIONAL verdict with MODERATE confidence is appropriate.","tokens_in":17054,"tokens_out":2680,"duration_ms":382235,"concrete_test":"Bootstrap the projection ratio and Grassmann similarity: resample the 54 factual-agree activations with replacement 1000 times, recompute the top-10 PC subspace and the steering direction projection ratio each time. If the 95% CI for the ratio includes values below ~0.7 (indicating genuine asymmetry in projection), the 'equal access' claim weakens. Also subsample the sycophantic-agree group to 54 to match sample sizes and recompute; if the ratio shifts substantially, the current asymmetry in sample size is biasing the estimate.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's most distinctive claim is the 'geometric puzzle': sycophantic and factual agreement occupy distinct subspaces (Grassmann similarity 0.153 vs random 0.317), yet the steering direction projects equally onto both (ratio 0.90–0.97, Appendix D). This combination drives the conclusion that the behavioral dissociation 'was not explained by any of the static geometric properties we measured' (§4.5), which in turn motivates the 'readability does not entail writability' framing and the speculation about generation dynamics. However, the factual-agree group has only 54 activations from 6 topics (one stance each), while the sycophantic-agree group has 149 from 7 topics. With 54 samples in a 4096-dimensional space, the top-10 PC subspace estimate for the factual group may be unstable, and no bootstrap or confidence interval is reported for either the Grassmann similarity or the projection ratio. If the factual subspace is poorly estimated, the projection ratio could be biased toward 1.0 (a noisy subspace captures more of any given direction), which would make 'equal access' an artifact rather than a genuine geometric finding. This would weaken the 'geometric puzzle' narrative, though the behavioral non-specificity finding (the paper's primary contribution) would be unaffected.","agreement_with_reader":"agree"},"referee_report":{"model":"glm-5.2","summary":"This paper introduces dual-stance evaluation for activation steering specificity, testing whether a sycophancy-reduction direction also suppresses agreement with factually correct statements. Applied to Llama-3-8B-Instruct, the paper reports three findings: (1) centroid-difference steering is non-specific, reducing agreement with factual statements as well as sycophantic ones; (2) the non-specificity is structured and continuously predictable from a behavioral measure (dual-stance consistency), with out-of-sample replication; (3) sycophantic and factual agreement occupy geometrically distinct subspaces, yet the steering direction projects equally onto both. The paper includes train/test splits by topic, an alpha-ablation, prompt variation, a random-direction control, and parser validation.","tokens_in":17110,"tokens_out":2114,"duration_ms":79795,"significance":"The dual-stance evaluation framework is a genuine methodological contribution that addresses a real gap in steering evaluation practice. The behavioral non-specificity finding — that a sycophancy direction also suppresses agreement with factually correct statements — is important for safety and is well-supported by multiple controls (random-direction control showing 7.3% vs 74.6% differential, alpha-ablation, prompt variation). The out-of-sample prediction (r=0.84 on 12 novel topics) is a strong falsifiable test. The reproducible code and full item texts in appendices are commendable. The 'readability does not entail writability' framing, while appropriately hedged, provides a useful conceptual contribution to the interpretability literature.","major_comments":[{"comment":"§4.5 and Appendix D: The subspace analysis compares 149 sycophantic-agree activations against 54 factual-agree activations in a 4096-dimensional space. With only 54 samples, the top-10 PC subspace for the factual group may be poorly estimated, and a noisy subspace would tend to capture more of any given direction, biasing the projection ratio toward 1.0. This matters because the projection ratio (0.90–0.97) is load-bearing for the claim that 'the behavioural dissociation was not explained by any of the static geometric properties we measured' (§4.5), which in turn motivates the 'geometric puzzle' narrative and the 'readability does not entail writability' framing. No bootstrap confidence intervals or subsampling analyses are reported for either the Grassmann similarity or the projection ratio. A bootstrap analysis (resampling within each group, recomputing both quantities) would settle是否","section":null},{"comment":"the equal-projection claim is genuine or an artifact of the sample imbalance. This is fixable within the manuscript's scope and would strengthen or appropriately qualify the geometric claims. Note that the primary behavioral findings (non-specificity and predictability) are unaffected by this concern.","section":null},{"comment":"§4.5, Appendix D, Table: The random-split control for Grassmann similarity pools all 203 activations and partitions into groups of 149 and 54. This control tests whether the actual sycophantic/factual split produces lower alignment than random splits, which it does (z=-7.86). However, this control does not address the projection ratio concern: the random-split baseline for the projection ratio is not reported. If random splits also produce projection ratios near 1.0 (which is plausible if both subspaces are estimated from the same pooled distribution), then the 'equal projection' finding would be uninformative. Reporting the projection ratio distribution under random splits would clarify whether the observed ratio is surprising or expected.","section":null}],"minor_comments":[{"comment":"§3.5: The layer selection criterion is described as 'empirical' without specifying the procedure. Was layer 8 selected to maximize steering effect, probe accuracy, or both? This should be stated explicitly to assess potential selection bias.","section":null},{"comment":"§4.3, Figure 4: The zoos/ethics outlier is discussed but its position is not marked in the figure. Adding a label or annotation would help readers locate it.","section":null},{"comment":"§3.2: The item set was developed with Claude Sonnet 4.5 and reviewed by the author. The extent of LLM assistance in item generation should be transparently reported (e.g., how many items were LLM-generated vs. author-written, whether any were modified).","section":null},{"comment":"Appendix D, Additional Static Properties table: Cross-layer centroid cosine is reported as 0.404 vs 0.375 for sycophantic vs factual. No test of whether this difference is significant is reported; if it is not, this should be stated explicitly.","section":null},{"comment":"§4.6, Table 2: The prompt variation experiment uses only 3 sycophantic topics and 3 hard-fact topics with 10 trials each. This is a small sample; a note acknowledging the limited statistical power would be appropriate.","section":null},{"comment":"Figure 3: The axis labels (A) and (B) for each category are not self-explanatory without reference to the item texts. Adding brief stance descriptions or referencing Table 1 would improve readability.","section":null},{"comment":"§5.4: The Mistral-7B transfer result is mentioned briefly. Given that steering was 'substantially less effective' on Mistral, a sentence clarifying whether this reflects a failure of the steering method or the evaluation framework would help readers interpret the generalizability claim.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The paper is a single-author work by an independent researcher. The experimental design is notably careful for this context — train/test splits, out-of-sample prediction, random-direction control, and prompt variation are all present. The main concern about the subspace analysis sample sizes is legitimate but addressable with a bootstrap analysis, and the primary contributions do not depend on the geometric claims. I would encourage the editor to prioritize the bootstrap analysis as the key revision item; if the equal-projection result survives, the paper is substantially strengthened, and if it does not, the authors can appropriately qualify the geometric puzzle framing without losing their core contribution."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for a careful and constructive review. The referee's comments focus on a legitimate statistical concern about the subspace analysis in §4.5 and Appendix D: whether the sample-size imbalance (149 vs. 54 activations) and the absence of bootstrap confidence intervals or a random-split baseline for the projection ratio could undermine the geometric claims. We agree this concern is valid and will address it with additional analyses in the revision. The primary behavioral findings (non-specificity and predictability) are unaffected.","responses":[{"response":"The referee raises a valid concern. With 54 samples in a 4096-dimensional space, the top-10 PC subspace for the factual-agree group is indeed estimated from a small sample, and a noisy subspace could inflate the projection ratio toward 1.0 by capturing more variance from any direction. We acknowledge that the current manuscript does not report bootstrap confidence intervals or subsampling analyses for the projection ratio, and we agree these are needed to determine whether the equal-projection finding is genuine or an artifact of the sample imbalance. We will conduct a bootstrap analysis (resampling within each group with replacement, recomputing both the Grassmann similarity and the projection ratio over many iterations) and a subsampling analysis (repeatedly subsampling the sycophantic-agree group to match n=54 and recomputing the projection ratio) to test robustness. If the equal-projection finding holds under these analyses, this will strengthen the geometric claim. If it does not, we will appropriately qualify the claim and adjust the 'geometric puzzle' narrative and the 'readability does not entail writability' framing accordingly. We note that the primary behavioral findings (non-specificity and predictability) are unaffected by this concern, as the referee acknowledges.","revision_made":"yes","referee_comment":"§4.5 and Appendix D: The subspace analysis compares 149 sycophantic-agree activations against 54 factual-agree activations in a 4096-dimensional space. With only 54 samples, the top-10 PC subspace for the factual group may be poorly estimated, and a noisy subspace would tend to capture more of any given direction, biasing the projection ratio toward 1.0. This matters because the projection ratio (0.90–0.97) is load-bearing for the claim that 'the behavioural dissociation was not explained by any of the static geometric properties we measured.' A bootstrap analysis would settle whether the equal-projection claim is genuine or an artifact of the sample imbalance."},{"response":"The referee is correct that the random-split control currently addresses only the Grassmann similarity, not the projection ratio. This is a gap in the analysis. If random splits of the pooled activations also produce projection ratios near 1.0, then the observed ratio (0.90–0.97) would not be informative about whether the steering direction has preferential access to one subspace over the other. We will compute the projection ratio distribution under the same 500 random splits of the pooled activations (partitioning into groups of 149 and 54, computing the top-10 PC subspace for each, and projecting the steering direction onto both) and report this baseline alongside the observed ratio. If the observed ratio falls within the random-split distribution, we will state explicitly that the equal-projection finding is uninformative about differential geometric access and revise the geometric claims accordingly. If the observed ratio is surprising relative to the random-split baseline, this would strengthen the claim. Either way, this analysis will be added to Appendix D.","revision_made":"yes","referee_comment":"The random-split control for Grassmann similarity pools all 203 activations and partitions into groups of 149 and 54. This tests whether the actual sycophantic/factual split produces lower alignment than random splits, which it does (z=-7.86). However, this control does not address the projection ratio concern: the random-split baseline for the projection ratio is not reported. If random splits also produce projection ratios near 1.0, then the 'equal projection' finding would be uninformative. Reporting the projection ratio distribution under random splits would clarify whether the observed ratio is surprising or expected."}],"tokens_in":16654,"tokens_out":868,"duration_ms":102189,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"The main thing to know: this paper introduces dual-stance evaluation — testing both sides of a topic to detect whether a steering direction is specific or just pushing disagreement broadly — and it's a clean, useful idea that the steering literature should adopt. The core empirical finding (centroid-difference steering for sycophancy also suppresses agreement with factually correct statements) is well-supported and the experimental design is careful: train/test split by topic, alpha-ablation, prompt variation, out-of-sample prediction on 12 novel topics (r=0.84), and a random-direction control (7.3% differential vs 74.6% for the real direction). Code is public on GitHub. The 'readability does not entail writability' framing is a reasonable conceptual summary of the geometric puzzle, not overclaimed — the author is explicit that it applies to one method at one layer and may not generalize to SAE features or head-level interventions.","headline":"Dual-stance evaluation is a genuine methodological contribution; the subspace analysis has a real but non-fatal sample-size concern.","tokens_in":17771,"tokens_out":256,"would_cite":false,"duration_ms":122826,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"Sycophancy steering also kills factual agreement","keywords":["activation steering","sycophancy","dual-stance evaluation","representation engineering","model interpretability","LLM safety","residual stream","subspace analysis"],"falsifier":"A steering method that achieves genuine sycophancy-specificity under dual-stance evaluation—reducing sycophantic agreement while leaving factual agreement intact—would falsify the claim that the readability-writability gap is fundamental to residual-stream interventions (though the paper explicitly does not claim it extends to finer-grained methods like sparse autoencoder features or head-level interventions).","tokens_in":17049,"feed_emoji":"🧭","tokens_out":1494,"duration_ms":67376,"temperature":0.7,"pith_summary":"The paper introduces dual-stance evaluation, a method that tests both sides of a topic (e.g., 'the Earth is flat' and 'the Earth is round') to audit whether activation steering for sycophancy has unintended side effects. Applied to Llama-3-8B-Instruct, the paper shows that the standard centroid-difference steering direction—computed as the mean activation difference between agreement and disagreement trials—reduces sycophantic agreement by 89% but also reduces agreement with factually correct statements by 14%, despite comparable baseline agreement rates. This non-specificity is invisible under conventional single-stance evaluation, which only tests the target behaviour. The paper then establishes a geometric puzzle: sycophantic and factual agreement occupy distinct subspaces in the model's residual-stream activations (Grassmann similarity 0.15 vs 0.32 for random splits), yet the steering direction projects nearly equally onto both (ratio 0.90–0.97) and cannot differentially target either. All measured static geometric properties—activation norms, variance, cross-layer stability—are matched between the two agreement types, so the 75-percentage-point behavioural dissociation cannot be explained by pre-generation geometry. Instead, the differential susceptibility is continuously predictable from a simple behavioural measure called dual-stance consistency: the minimum agreement rate across both stances of a topic, which indexes how shallowly the model's agreement is held. This measure predicts steering effect magnitude with r=0.88 in-sample and r=0.84 on 12 novel held-out topics. The paper also finds that the casual 'friend' prompt framing both elicits sycophancy and partially protects factual agreement from the steering perturbation (3% reduction under casual framing vs 30% under neutral framing), suggesting that the model's social-compliance state moderates how the intervention interacts with factual knowledge. The paper frames the central gap as: representations readable from activations may not be writable through them—at least not at the granularity of residual-stream steering.","feed_headline":"Sycophancy steering also kills factual agreement","feed_subtitle":"A standard intervention that cuts sycophantic agreement by 89% also erodes agreement with facts like 'the Earth is round'—and standard evals","key_machinery":"Dual-stance evaluation tests both stances of each topic to classify items as sycophantic (high agreement on both sides), opinionated (high agreement on one side only), or mixed. The centroid-difference steering direction is computed as the mean of disagreement activations minus the mean of agreement activations at layer 8 of the residual stream, then added during generation. Dual-stance consistency is defined as the minimum of the two stance-wise baseline agreement rates for a topic. Subspace analysis uses Grassmann similarity between the top-10 principal components of sycophantic-agree and factual-agree activation groups, compared against a random-split null distribution.","core_discovery":"The centroid-difference steering direction for sycophancy captures general agreement polarity rather than sycophancy specifically. The model internally distinguishes sycophantic from factual agreement in geometrically distinct activation subspaces, but the steering direction has equal geometric access to both and cannot exploit that distinction. The resulting behavioural dissociation (89% vs 14% reduction) is not explained by any measured static property of the activations but is continuously predictable from dual-stance consistency—a behavioural measure of how shallowly the model holds its agreement.","pith_inferences":["The 149:54 sample imbalance between sycophantic-agree and factual-agree activations means the factual subspace is estimated from fewer data points. Without bootstrap confidence intervals on the Grassmann similarity or projection ratios, the claim of 'equal access' could partly reflect statistical noise rather than a genuine geometric property. A replication with a larger factual-agree sample would","The dual-stance consistency measure may be proxying for something more general than sycophancy shallowness—namely, the degree to which a representation is supported by redundant downstream circuitry versus a single diffuse compliance process. If so, the same measure might predict steering susceptibility for other behaviours (e.g., toxicity, deception) where some instances are backed by factual kno","The prompt-framing finding implies that steering interventions evaluated under one prompt context may not generalise to others. This raises the possibility that specificity audits need to be conducted across a distribution of prompt framings, not just a single template, to be deployment-ready.","The 'refusal leakage' hypothesis—that the general disagreement signal overlaps geometrically with the model's refusal direction—could be tested directly by measuring the cosine similarity between the sycophancy steering direction and the refusal direction identified in prior work, and checking whether it predicts the magnitude of factual-agreement reduction across topics."],"forward_implications":["Single-stance evaluation of activation steering has a structural blind spot: a direction can appear to successfully reduce a target behaviour while silently degrading agreement with factually correct statements, and this collateral damage is invisible without testing the contrastive stance.","Dual-stance consistency could serve as a pre-intervention screening tool: measuring how shallowly a model agrees on both sides of a topic predicts how susceptible that agreement will be to steering, allowing practitioners to anticipate specificity failures before deploying an intervention.","The finding that the casual compliance context protects factual agreement from the steering perturbation suggests that prompt framing and model social state interact with steering interventions in ways that current evaluation protocols do not capture.","If the readability-writability gap generalises beyond sycophancy, then probing accuracy for any behavioural category (deception, hallucination, toxicity) may not guarantee that the corresponding steering direction can target that category without collateral effects on related behaviours.","The paper identifies two competing explanations for the geometric puzzle—aggregation (the distinction exists in individual attention heads but is lost when aggregated into the residual stream) versus generation dynamics (the dissociation emerges from how perturbations propagate through autoregressive generation)—and these make distinct, testable predictions about whether head-level interventions c"],"fun_headline_variants":["Sycophancy interventions suppress factual agreement too","Steering for less sycophancy erases agreement with true statements","Readable representations aren't writable: the sycophancy steering gap","Sycophancy direction captures agreement, not sycophancy specifically","Dual-stance eval exposes collateral damage in sycophancy steering"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The subspace analysis comparing sycophantic-agree activations (149 samples) against factual-agree activations (54 samples) assumes that 54 samples are enough to reliably estimate the factual-agree subspace and that the near-equal projection ratio (0.90–0.97) is a genuine geometric property rather than an artefact of the small and imbalanced sample. No confidence intervals or bootstrap analyses are reported for these estimates.","fun_headline_variants_meta":{"raw":{"variants":["Sycophancy interventions suppress factual agreement too","Steering for less sycophancy erases agreement with true statements","Readable representations aren't writable: the sycophancy steering gap","Sycophancy direction captures agreement, not sycophancy specifically","Dual-stance eval exposes collateral damage in sycophancy steering"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":591,"prompt_tokens":503,"completion_tokens":88,"prompt_tokens_details":null},"tokens_in":503,"tokens_out":88,"duration_ms":16401,"temperature":1.0,"reasoning_tokens":null,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-05T02:24:07.271610+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"A steering method that achieves genuine sycophancy-specificity under dual-stance evaluation—reducing sycophantic agreement while leaving factual agreement intact—would falsify the claim that the readability-writability gap is fundamental to residual-stream interventions (though the paper explicitly does not claim it extends to finer-grained methods like sparse autoencoder features or head-level interventions).","supporting_citations":[],"review_version":1}