{"id":"a2a6a36d-e564-4bf7-a35c-60f077da4583","arxiv_id":"2505.22637","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Steering vectors are unreliable when the target behavior does not correspond to a coherent, well-separated linear direction in activation space, and this coherence can be measured from training data.","lead":"This paper tests seven prompt styles for building steering vectors for Llama2-7B and finds that all work on average, but reliability varies widely across samples. It shows that a behavior is easier to steer when its positive and negative example activations point in a consistent direction and are well separated in activation space.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'coherent direction' predictors are computed from prefilled answer-token activations, so the measured coherence may reflect answer-token identity rather than a linear representation of the target behavior.","rationale":"The paper makes a useful, honestly limited empirical contribution: it reports strong held-out correlations between activation geometry and steering outcomes, includes intuitive projection plots, and does not overstate internal consistency. These are real strengths. However, the load-bearing concern is that the geometric predictors are not independent of the answer-token confound inherent in the prefilled prompt construction. The reader identified layer/multiplier choice and the Δm_LD ground truth as the weakest assumptions; my concern is related but distinct: even if the layer and multiplier are fixed, the predictors may be measuring answer-token regularity rather than the linear representability of the behavior. The non-prefilled prompt results in Figure 1 show similar steering performance across prompt types, which partially mitigates this concern, but the geometric analysis was not run on those prompt types, so the confound is not ruled out. The abstract's statement that all seven prompt types produce a net positive effect is also contradicted in Appendix C for the six least steerable datasets, where some prompt types yield negative mean logit differences; this is a secondary error that should be corrected but does not directly threaten the geometric claim. I would keep the conditional accept, but the acceptance conditions should explicitly require the non-prefilled control or answer-token confound analysis. If that test fails, the central explanation would need to be substantially weakened, and the paper would become a purely descriptive diagnostic study rather than evidence for the linear-representation interpretation.","tokens_in":14311,"tokens_out":11840,"duration_ms":158433,"concrete_test":"Recompute the Section 3 and Appendix D correlations (cosine similarity and d' vs. mean Δm_LD and anti-steerable fraction) using non-prefilled training activations, recorded at the last prompt token with no answer token appended, while keeping the same CAA evaluation on plain prompts. If these predictors no longer exhibit Spearman correlations near 0.76 and -0.78, the prefilled answer-token identity is a primary driver. As a complementary check, for the prefilled setting, subtract the mean positive-minus-negative token-embedding direction from each activation difference before computing cosine similarity and d'; if the correlations collapse, the central claim is confounded.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that steering is reliable exactly when the target behavior is represented by a coherent linear direction. The operational evidence for this claim comes from two geometric predictors: mean cosine similarity between individual activation differences and the steering vector, and d' along the difference-of-means line. Both are computed on training activations recorded with the prefilled prompt type, where the positive and negative activations differ by replacing the final answer token (Section 2). At layer 13, the residual stream at the answer-token position contains a substantial and highly consistent component encoding the identity of the answer token (e.g., 'Yes' vs 'No' or 'A' vs 'B'). Consequently, for datasets with fixed or highly regular answer-choice formats, the activation differences will look coherent and well separated even if the behavior itself is not semantically encoded as a linear direction. The held-out evaluation uses plain prompts without an appended answer token; steering by a direction dominated by answer-token identity can shift logit differences through token-prior effects, not through a behavioral representation. The paper limits the geometric analysis to the prefilled prompt type (Section 3: 'we limit our analysis to the prefilled prompt type'), so the non-prefilled variants, which would avoid this confound, are not used to validate the geometric predictors. If the predictors only capture answer-token regularity, the conclusion that 'vector steering is unreliable when the target behavior is not represented by a coherent direction' overreaches the evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies when and why Contrastive Activation Addition (CAA) steering vectors are reliable, using Llama2-7B-Chat on 36 multiple-choice behavior datasets. It compares seven prompt types for training steering vectors, finding that all yield a net positive average steering effect but with high per-sample variance, and that no prompt type clearly outperforms the others. The main contribution is a proposed explanation based on activation geometry: datasets whose training activation differences are well aligned with the steering vector (high cosine similarity) and where positive and negative activations are well separated along the difference-of-means line (high d') exhibit stronger and more reliable steering. These geometric predictors correlate with held-out steering success (Spearman 0.76 for effect size and -0.78 for anti-steerable fraction), leading the authors to conclude that steering is unreliable when the target behavior is not represented as a coherent linear direction.","tokens_in":14475,"tokens_out":6654,"duration_ms":78491,"significance":"If the geometric predictors genuinely capture whether a behavior is linearly represented, the paper offers a cheap, pre-hoc diagnostic for when vector steering will work, which would be practically useful and conceptually clarifying. The reported correlations are strong and are computed against held-out evaluation prompts, and the paper is transparent about its limitations, including the restriction to a single model, a single layer, and CAA only. The central interpretation, however, is threatened by a confound in the way the prefilled training activations are constructed, which I detail in the major comments. The paper's strengths are its clear experimental design, direct comparison to prior work, and honest acknowledgment of scope; its main claim needs additional work to rule out the token-identity alternative.","major_comments":[{"comment":"The geometric predictors are computed on activations from the prefilled prompt type, where the positive and negative examples differ only by the appended answer token (e.g., 'A' versus 'B'). Because the residual stream at the answer-token position encodes the identity of that token, the activation differences and the steering vector contain a substantial component tied to answer-token identity rather than to the target behavior. The held-out evaluation uses plain prompts without an appended answer token, so the reported correlations (Appendix D: Spearman 0.76 for steerability and -0.78 for anti-steerable fraction) may reflect a token-prior effect rather than a linear representation of the behavior. This directly affects the central claim in Section 4 that 'steering vectors are not universally applicable, and that their effectiveness depends on whether the targeted behavior is well-represented as a linear direction.' I request that the authors recompute the predictors on non-prefilled prompt types (instruction and/or 5-shot), where the positive and negative prompts differ in content rather than in the final token, or otherwise control for the answer-token direction (e.g., by subtracting the mean answer-token embedding direction) and show that the predictive relationship persists.","section":"§2 (Steering Method) and §3 (Directional Agreement Predicts Steerability)"},{"comment":"All geometric predictors and steering evaluations are performed with a single layer, l=13, and a single multiplier, λ=1. The paper's conceptual conclusion is that steering reliability depends on whether the behavior is linearly represented in activation space, which is a general claim about the model's geometry. As it stands, the evidence is limited to one depth and one intervention strength. The limitations discussion lists breadth of models, datasets, and steering methods, but does not address sensitivity to layer or multiplier. I ask the authors to report whether the geometry-effectiveness relationship holds at other layers (e.g., layers 10 and 16) and for at least one other multiplier (e.g., λ=0.5 and λ=2). If the relationship reverses or weakens, the statement in Section 4 that 'both directional consistency of activation differences and separability of activations along the difference-of-means line are conceptually intuitive explanations and empirical predictors of steering vector performance' would need to be qualified.","section":"§2 (Experimental Setup) and §4 (Limitations)"},{"comment":"The comparison of prompt types is presented graphically without error bars or significance tests. The paper states that 'all seven prompt types produce a net positive steering effect' and that 'no prompt type clearly outperforms the others,' but Figure 1 shows only per-sample distributions and means. The authors acknowledge in Section 4 that statistical comparison is highly sensitive to hyperparameters and that they did not run such tests. This is acceptable as an exploratory finding, but the claims are stronger than the evidence supports. Please either add confidence intervals or paired significance tests, or soften the claims to state that no prompt type is consistently ranked best in this dataset set.","section":"§3 (Effect of Prompt Types on Steering Vectors) and §4 (Methodology for Prompt Type Comparison)"}],"minor_comments":[{"comment":"The text 'faction of such anti-steerable samples' should be 'fraction of such anti-steerable samples.'","section":"Figure 1 and Figure 5 captions"},{"comment":"The sentence 'now single prompt type is systematically better than the others' should read 'no single prompt type is systematically better than the others.'","section":"Appendix C.2"},{"comment":"In the example prompt, the prefilled answer token is written as 'Answer: (A' without a closing parenthesis; it should be 'Answer: (A)' for consistency with the other choices.","section":"Appendix A example"},{"comment":"The caption says the datasets are ordered by 'steerability rank from Tan et al. (2024),' but Section 3 also discusses the authors' own steerability measurements. This is potentially confusing; please clarify whether the ordering in Figure 2 is based on Tan et al.'s ranking, the authors' own evaluation, or both, and how that relates to the Spearman correlations in Appendix D.","section":"Figure 2 caption and Section 3"},{"comment":"The x-axis label 'Mean per-sample steerability' is not defined in the caption; please define it explicitly (e.g., mean Δm_LD over the held-out test set).","section":"Appendix D, Figure 6"},{"comment":"The paper uses λ=1 for 'most' experiments but does not specify which analyses use other multipliers. Please state explicitly whether the geometric-predictor correlations and the prompt-type comparisons all use λ=1, and whether any results use different values.","section":"§2 (Evaluation of Steering Success)"}],"recommendation":"major_revision","confidential_remarks":"The token-identity confound is the key threat to the paper's central claim, and it is fixable because the authors already have non-prefilled prompt-type data that would avoid the confound. If the correlations survive that check, the paper would be a solid contribution to the steering-vector reliability literature. I would encourage the editor to treat the layer/multiplier sensitivity request as a necessary robustness check rather than a demand for exhaustive coverage. The paper is otherwise clearly written and honest about its scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. This paper gives practitioners a cheap pre-hoc diagnostic for whether CAA steering will work: cosine similarity between training activation differences and the steering vector, and d' along the difference-of-means line. The held-out correlations are strong (Spearman 0.76, -0.78) and honestly computed. Second, the paper's systematic comparison of seven prompt types is a useful empirical addition, though it mostly confirms Tan et al.'s dataset-dependence story.\n\nWhat's genuinely new: the prompt-type comparison (no clear winner, all net positive on average, but 30-45% anti-steerable samples) and the two geometric predictors. The finding that steering works when the behavior is linearly coherent is not new, but tying it to cheaply computable predictors is a useful contribution. The paper is honest about its limitations, and the analysis is reproducible in principle (though no code/data is released).\n\nThe main soft spot is the stress-test concern. The geometric predictors are computed on prefilled answer-token activations, where positive and negative examples differ by the final token (A vs B, Yes vs No). At layer 13, that difference includes a substantial component encoding token identity, not just the behavior. The paper doesn't control for this. If the predictors mostly capture answer-token regularity, the conclusion that 'steering is unreliable when the behavior is not linearly represented' overreaches. The correlations with plain-prompt steering on held-out sets suggest it's not purely token identity, and the cosine similarities vary from 0.19 to 0.48 (not near 1), so I don't think the concern is fatal. But the authors should add a control—e.g., compute the predictors on non-prefilled activations, or randomize answer labels to measure the token-identity baseline. Without that, the central interpretation is underdetermined.\n\nMinor weak spots: no error bars or significance tests for the prompt-type comparison (the authors give a reasonable explanation, but the abstract's 'all seven prompt types produce a net positive effect' is only true averaged over datasets; Appendix C.2 shows negative means for some least-steerable datasets). One model, one layer, one steering method—acknowledged. No code/data.\n\nWho this is for: anyone using CAA or thinking about steering reliability. It deserves a serious referee. I'd recommend a conditional accept with the token-identity control as the main requested revision. It's a solid empirical contribution, not a breakthrough, and the authors should be pushed to make the interpretation solid.","headline":"Solid empirical study with a useful cheap predictor for CAA steering success, but the central interpretation is undercut by a possible answer-token identity confound.","tokens_in":15111,"tokens_out":4841,"would_cite":true,"duration_ms":54686,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Steering vectors work only when a behavior has one coherent direction in activation space.","keywords":["steering vectors","contrastive activation addition","activation geometry","directional agreement","discriminability index","linear representation","language model reliability","logit difference"],"falsifier":"Compute the same predictors at other layers and multipliers: if there exists a layer where all 36 datasets show high $d'$ and high cosine similarity yet steering effect is weak or anti-steerable, or a setting where steering works well despite low coherence, the proposed explanation fails. A direct test would be to measure Spearman correlation between mean cosine similarity and $\\Delta m_{\\mathrm{LD}}$ at layers other than 13 (e.g., layers 5, 20, 30) and at multipliers $\\lambda=0.5$ and $\\lambda=3$; the claim predicts the correlations survive.","tokens_in":14038,"feed_emoji":"🧭","tokens_out":4991,"duration_ms":50466,"temperature":0.7,"pith_summary":"This paper asks why activation steering sometimes fails, and argues that the answer lies in the geometry of the model's activations, not in the choice of prompt wording. Using Contrastive Activation Addition on Llama2-7B-Chat across 36 behavior datasets, the authors find that all seven prompt types they tried give a net positive steering effect on average, yet nearly a third of individual samples move in the opposite direction. They then show that two cheaply computable geometric measurements—the cosine similarity between individual activation differences and the steering vector, and the separability of positive versus negative activations along the difference-of-means line—predict both the size of the steering effect and how often steering backfires. The conclusion is that a behavior is reliably steerable by CAA exactly when it is represented by a coherent linear direction in activation space.","feed_headline":"Two measurements predict when steering vectors will fail","feed_subtitle":"Cosine agreement and activation separability forecast steering success across 36 behaviors.","key_machinery":"The carrying object is the CAA steering vector, $\\mathbf{s}_l = \\frac{1}{|\\mathcal{D}|}\\sum_{(x,y^+,y^-)\\in\\mathcal{D}}(\\mathbf{a}_l(x,y^+) - \\mathbf{a}_l(x,y^-))$, a single vector added to residual-stream activations at layer 13 during inference. Around this, the paper builds two geometric diagnostics. The first is directional agreement: the cosine similarity between each training activation difference and $\\mathbf{s}_l$, averaged over the dataset. The second is the difference-of-means line, the line through the mean positive activation $\\boldsymbol{\\mu}_{l,+}$ and mean negative activation $\\boldsymbol{\\mu}_{l,-}$, parameterized as $\\mathrm{doml}_l(\\boldsymbol{\\mu}_+,\\boldsymbol{\\mu}_-) = \\frac{1+\\kappa}{2}\\boldsymbol{\\mu}_{l,+} + \\frac{1-\\kappa}{2}\\boldsymbol{\\mu}_{l,-}$, onto which activations are projected and summarized by the discriminability index $d' = |\\mu_+ - \\mu_-| / \\sqrt{\\tfrac{1}{2}(\\sigma_+^2 + \\sigma_-^2)}$. Steering success is measured by $\\Delta m_{\\mathrm{LD}}$, the change in logit-difference between the desired and undesired answer token when steering is applied. The machinery shows that the same geometry that defines the steering vector also predicts whether the vector will work.","core_discovery":"The paper's central claim is that steering-vector reliability is a property of the dataset's activation geometry. For each of the 36 datasets, they compute the CAA steering vector as the mean difference between positive and negative residual-stream activations at layer 13, then measure, per training sample, the cosine similarity between the individual activation difference and the steering vector, and the discriminability index $d'$ of the two activation classes projected onto the difference-of-means line. Datasets ranked as most steerable have mean cosine similarities near 0.48, while the least steerable hover near 0.19, and $d'$ drops correspondingly; both quantities correlate strongly (Spearman 0.76 for effect size, -0.78 for anti-steerable fraction) with steering performance. The paper interprets this as evidence that unreliable steering occurs when the target behavior is not consistently encoded as a single linear direction in the residual stream. Prompt type has only a limited influence: vectors trained with different prompt formats point in quite different directions (cosine similarities as low as 0.07), yet all produce similar average steering effects, reinforcing that dataset-specific geometry dominates.","pith_inferences":["If coherence of activation differences is the true driver, then failure cases are not a tuning problem: no multiplier or prompt format can make a scattered direction behave like a coherent one, and efforts should shift to steering methods that do not assume a single linear shift.","A practical extension suggested by the paper is a pre-hoc 'steerability score' for new datasets: compute mean cosine similarity and $d'$ on a small training set before deciding whether vector steering is worth applying.","The geometry measured here is layer-13-specific; a natural test is whether the coherence-to-steerability relationship holds at all layers or only at the layer where the steering vector is applied, which would sharpen the mechanistic story.","Because the paper only tests CAA, the same two predictors could be measured for other linear steering methods (e.g., function vectors) to see whether directional agreement is a universal requirement or an artifact of the mean-difference construction."],"forward_implications":["Steering effectiveness can be predicted before applying the intervention by averaging per-sample cosine similarity between activation differences and the steering vector, or by computing $d'$ along the difference-of-means line.","Prompt-type choice has limited impact on average steering effect, so tuning prompt phrasing cannot fix unreliability that stems from a scattered activation geometry.","Reliability varies by dataset: roughly one-third of samples are anti-steerable on average, ranging from 3% to 50% per dataset, so per-sample variance is inherent to datasets with weak linear structure.","The results delimit the applicability of CAA-style vector steering to behaviors that are linearly represented; for other behaviors, the same method will be unreliable regardless of prompt format."],"supporting_citations":[{"why":"Supplies the Contrastive Activation Addition method whose steering vectors the paper constructs and evaluates.","marker":"Rimsky et al. (2024)"},{"why":"Provides the layer-13 choice, the logit-difference metric $\\Delta m_{\\mathrm{LD}}$, the steerability ranking of the 36 datasets, and the prior observation of unreliable steering that this paper seeks to explain.","marker":"Tan et al. (2024)"},{"why":"Supplies the 36 multiple-choice assistant-behavior datasets used in all experiments.","marker":"Rogers et al. (2023)"},{"why":"Provides the Llama2-7B-Chat model whose residual-stream activations are steered and measured.","marker":"Touvron et al. (2023)"}],"fun_headline_variants":["Steering reliability is determined by activation geometry","Why steering fails: dataset geometry matters more than prompts","Cosine similarity and d' predict steering success","Steering vectors work only when behaviors align as a line"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole explanation rests on assuming that steering success is faithfully captured by the change in logit-difference measured at layer 13 with multiplier $\\lambda=1$, and that the activation geometry at that single layer is the geometry that governs steering; if the link between coherence and effectiveness shifts with layer or steering strength, the proposed predictor may not generalize.","fun_headline_variants_meta":{"raw":{"variants":["Steering reliability is determined by activation geometry","Why steering fails: dataset geometry matters more than prompts","Cosine similarity and d' predict steering success","Steering vectors work only when behaviors align as a line"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000614,"raw_usage":{"total_tokens":2845,"prompt_tokens":929,"completion_tokens":1916,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":545,"completion_tokens_details":{"reasoning_tokens":1856}},"tokens_in":545,"tokens_out":1916,"duration_ms":14646,"temperature":1.0,"reasoning_tokens":1856,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:02:45.731438+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the same predictors at other layers and multipliers: if there exists a layer where all 36 datasets show high $d'$ and high cosine similarity yet steering effect is weak or anti-steerable, or a setting where steering works well despite low coherence, the proposed explanation fails. A direct test would be to measure Spearman correlation between mean cosine similarity and $\\Delta m_{\\mathrm{LD}}$ at layers other than 13 (e.g., layers 5, 20, 30) and at multipliers $\\lambda=0.5$ and $\\lambda=3$; the claim predicts the correlations survive.","supporting_citations":[],"review_version":1}