REVIEW 2 major objections 5 minor 14 references
Chat-calibrated additive steering reaches agent residual streams at near-full strength, yet its behavioral grip is rescaled model-by-model with no universal factor.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-13 05:01 UTC pith:GOO63JXW
load-bearing objection First careful chat-to-agent transfer study of additive steering: direction survives, behavioral coupling rescales per model (can amplify or attenuate), with solid controls and honest limits. the 2 major comments →
Present but Rescaled: Chat-to-Agent Transfer of Additive Activation Steering
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Transfer of additive activation steering from chat to ReAct agents is real but dissociated: the injected direction reaches late layers at near-full strength (install-site agent-over-chat ratios 0.83–1.16), while behavioral coupling is reset per model, spanning amplification (Gemma-2-9B T=2.00) to attenuation (Yi-1.5-9B T=0.43) with no universal constant or sign. The rescaling localizes to the ReAct format scaffold before tool observations and is additive-specific: directional ablation of the same axis does not amplify while additive injection does.
What carries the argument
The matched-information ladder (plain chat C0 versus ReAct-with-tool C3) paired with the transfer ratio T = Δagent/Δchat, under matched-norm random-direction controls and full transcript re-encoding every turn. It holds the instruction byte-identical while isolating the deployment wrapper, letting representation survival and behavioral coupling be measured separately.
Load-bearing premise
The parser that scores refusal or compliance is equally sensitive on free-form chat replies and on ReAct Thought tokens, so the transfer ratio is not an artifact of the scorer working better in one format than the other.
What would settle it
Re-run the uniform-protocol experiment on held-out model families under the same ladder and random controls: if behavioral T is statistically indistinguishable from 1 with tight intervals while install-site projection ratios remain near 1, or if removing only the ReAct format scaffold leaves T unchanged while tool-observation insertion moves it, the claimed dissociation and format-priming localization fail.
If this is right
- Agentic deployment can amplify steering-based refusal bypass by up to 2× on some models, so chat-only safety evaluations understate the hazard.
- Activation probes trained in chat continue to fire at near-chat strength in agents; control coefficients must be recalibrated in the deployment context.
- Coupling has no universal constant or sign across families; each model requires its own agent-side characterization.
- Rescaling is set by format priming, not by tool observations or residual-norm dilution, so observation-content interventions miss the main effect.
- Additive injection and directional ablation transfer differently: the former must be re-asserted every step, the latter does not.
Where Pith is reading between the lines
- Other structured agent formats (code-agent loops, plan-act-observe variants) may each impose their own distinct rescale factors that should be measured before deployment.
- The existence of both clean amplifiers and a clean attenuator points to alignment-training geometry as the likely determinant of sign; a broader census could map which recipes produce attenuation.
- Any long-horizon setting whose residual norm grows may systematically under-power one-shot or prefill-only steering relative to continuous per-token injection.
- Red-teaming that ranks models only under chat-calibrated additive jailbreaks will mis-order them once the same models run as agents.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents the first systematic study of chat-to-agent transfer for additive residual-stream steering. Using a matched-information ladder (C0 plain chat vs C3 ReAct with a deterministic tool), matched-norm random controls, full-transcript re-encoding each turn, and a dual readout (late-layer projection survival + parser-based behavioral transfer ratio T = Δagent/Δchat), it reports a dissociation: the injected direction reaches late layers at near-full strength (install-site agent-over-chat ratios 0.83–1.16 across Qwen2.5-7B, Llama-3.1-8B, Gemma-2-9B) while behavioral coupling is reset per model, spanning amplification (Gemma T=2.00, Qwen T=1.41–1.45) to attenuation (Yi-1.5-9B T=0.43). Additive injection amplifies while directional ablation of the same axis does not (Φ=20.1 pt gap); two pre-registered instruments localize the rescaling to the ReAct format scaffold before any tool observation. Safety implication: agentic deployment can amplify or attenuate refusal-bypass steering in a model-specific way that chat calibration does not predict.
Significance. If the dissociation and two-sided coupling distribution hold, the result is immediately useful for both mechanistic interpretability and deployment safety. Additive steering is already proposed as a control/monitoring primitive; the paper shows that representation survival does not imply behavioral grip transfer, that the effect is additive-specific (vs ablation), and that the rescaling is set by format priming rather than observation dilution. Strengths that raise the bar include pre-registration of live/die criteria and dose grids, item-paired bootstrap CIs (B=10k), matched-norm random bands at every cell, a capability-matched verbosity control, convergent localization instruments, and explicit disclosure of the two-pass finer-ladder rescue that recovered the attenuator. These make the central claims falsifiable and reproducible rather than post-hoc.
major comments (2)
- [The Setting-Invariant Metric; Table 1] The transfer ratio T rests on a setting-invariant binary parser (refuses/complies) applied identically to free-form chat replies and ReAct Thought tokens. The manuscript reports 83.8% agreement with blind human labels on a stratified 100-item slice, but does not break that agreement down by rung (C0 vs C3) or report false-positive/false-negative rates under the ReAct scaffold. Because T is a ratio of deltas, differential parser sensitivity or ceiling effects across formats could inflate or deflate the reported amplification/attenuation. A short per-rung confusion matrix or human re-label of a C3-only slice would close this load-bearing measurement risk.
- [Per-Model Distribution; Figure 2; Limitations] The claim of 'no universal sign' is anchored on a single clean attenuator (Yi-1.5-9B T=0.43), recovered only after a pre-registered finer-ladder second pass. While the two-pass structure and pre-registration of Yi's sign are disclosed, five multi-dose families plus one single-dose boundary family leave the two-sided distribution asymmetrically supported. The safety conclusion that 'a deployment cannot assume a given model is safe' is still warranted by the amplifiers alone, but the stronger phrasing of no universal sign should be tempered or supported by at least one additional attenuating family if available.
minor comments (5)
- [Table 1] Table 1 footnote on Qwen2-7B single-dose T and the gate-fail families is dense; a short column or legend clarifying which T values are AUC-over-doses versus single-dose point ratios would help readers.
- [Abstract; Behavioral leg: sycophancy] Sycophancy T=0.78 CI spans 1; the text correctly treats the cross-behavior interaction as the formal result, but the abstract and safety paragraphs still lean on 'attenuate' language that should stay qualified.
- [Representational leg; Table 3] Install-site ratios are reported as 0.83–1.16 across three families; the main text and Table 3 give more granular layer-wise numbers for Qwen. A compact cross-family install-site table would make the survival half easier to cite.
- [The Matched-Information Ladder] C4 is mentioned as held out for tau2-bench but never used; either drop the rung from the ladder description or note that external-benchmark transfer is future work more prominently in the design section.
- [Abstract; Introduction] Minor typography: several compound terms appear without spaces or hyphens in the abstract and early sections (e.g., 'Additiveactivationsteering', 'chat-to-agenttransfer'); these are likely PDF extraction artifacts but should be cleaned for the camera-ready version.
Circularity Check
No significant circularity: transfer ratios T, install-site survival, additive-vs-ablation gaps, and frame-priming localization are measured quantities under pre-registered matched designs, not forced by definition or self-citation.
full rationale
This is a self-contained empirical measurement paper. The central quantities (T = Δagent/Δchat from parser-scored refusal/sycophancy rates on matched items, install-site residual projections onto the unit steering direction, AUC-over-doses under a uniform protocol, and the 20.1-point additive-vs-ablation gain difference) are computed from observed behavioral and representational read-outs after injection; they are not algebraically identical to any fitted input or definitional identity. Operating layers and sub-saturation dose grids are fixed from chat pilots before any agent cell runs (explicitly pre-registered, including the two-pass rescue structure and Yi’s attenuating sign from a held-out prior run). Matched-norm random-direction bands, KV recompute, capability-matched verbosity controls, and item-paired bootstrap CIs further gate direction-specificity rather than bake in the result. The Belief-Dynamics head-to-head treats agent cells as out-of-sample predictions from a chat-fitted slope and reports large gaps (e.g., 67 points on induce). No uniqueness theorem, ansatz, or load-bearing self-citation by the sole author is invoked to force the dissociation or the two-sided coupling distribution. The paper therefore satisfies the hard rule for an honest non-finding of circularity.
Axiom & Free-Parameter Ledger
free parameters (4)
- per-family operating layer and sub-saturation dose grid
- AMP_MARGIN = 1.10
- injection coefficient α (c8, c12, c16, …)
- read-layer and install-site prefix length
axioms (5)
- domain assumption Difference-of-means residual-stream directions extracted on last-token activations of harmful vs harmless (or trait+/trait−) completions are the correct objects for additive steering.
- domain assumption A validated parser-based binary (refuses/complies, agrees/disagrees) applied identically to natural-language output is setting-invariant across plain chat and ReAct Thought tokens.
- domain assumption Re-encoding the full transcript from scratch each turn fully excludes KV-cache contamination as a mechanism.
- domain assumption Matched-norm random unit vectors (n_rand ≥ 5) at the same coefficient constitute a sufficient direction-specificity gate.
- standard math Bootstrap item-paired CIs with B=10 000 correctly capture uncertainty for the transfer ratios and gain differences.
invented entities (2)
-
matched-information five-rung ladder (C0–C4)
independent evidence
-
transfer ratio T = Δagent / Δchat
independent evidence
read the original abstract
Additive activation steering (injecting a scaled residual-stream direction during generation) is calibrated almost entirely in single-turn chat, yet the models it targets are increasingly deployed as tool-using ReAct agents. We present the first systematic chat-to-agent transfer study of additive steering, coupling behavioral measurement with a representation read-out in a matched-information design: the same items rendered as plain chat or as a ReAct tool-use episode, with matched-norm random-direction controls and the transcript re-encoded every turn to exclude KV-cache contamination. Transfer is real but rescaled, and the right description is a dissociation: the injected direction reaches the late layers at near-full strength in every setting and model tested (install-site agent-over-chat ratios 0.83-1.16 across three families), while the behavioral coupling is reset per model and context. On Qwen2.5-7B a refusal bypass vector amplifies in the agent (T = 1.45, CI [1.20, 1.78], N = 300); across a powered uniform-protocol distribution the coupling spans amplification (Gemma-2-9B T = 2.00) to attenuation (Yi-1.5-9B T = 0.43, CI [0.29, 0.60]), with no universal constant and a single clean attenuator against a universal sign. Directional ablation of the same axis does not amplify (T = 0.93, CI including 1) while additive injection amplifies (T = 1.50), a 20.1-point gain difference (CI [13.4, 26.8]) that identifies an additive-specific mechanism. Two pre-registered instruments converge to localize the rescaling to the ReAct format scaffold, before any tool observation, rather than to the observation boundary where a dilution account would predict it. The safety implication is immediate and unpredictable: agentic deployment amplifies steering-based refusal bypass by up to 2.00x on some models while others attenuate, so a deployment cannot assume a given model is safe under additive steering.
Figures
Reference graph
Works this paper leans on
-
[1]
RefusalinLanguageMod- els Is Mediated by a Single Direction.arXiv, 2406.11717
Arditi, A.; Obeso, O.; Syed, A.; Paleka, D.; Panickssery, N.; Gurnee,W.;andNanda,N.2024. RefusalinLanguageMod- els Is Mediated by a Single Direction.arXiv, 2406.11717. Bigelow, E.; Wurgaft, D.; Wang, Y.; Goodman, N.; Ullman, T.;Tanaka,H.;andLubana,E.S.2025. BeliefDynamicsRe- veal the Dual Nature of In-Context Learning and Activation Steering.arXiv, 2511.0...
Pith/arXiv arXiv 2024
-
[2]
Chen,Y.;Siu,V.;Liu,Y.;Song,D.;andWang,C.2026
PersonaVectors:MonitoringandControllingCharac- ter Traits in Language Models.arXiv, 2507.21509. Chen,Y.;Siu,V.;Liu,Y.;Song,D.;andWang,C.2026. Con- trollingToolUsewithHeading-SpecificActivationSteering. arXiv, 2607.05790. Cristofano, T
Pith/arXiv arXiv 2026
-
[3]
Deng,Y.2026.GEMS:GeometricConstraintsEnableMulti- Semantic Superposition in LLMs.arXiv, 2606.19946
Universal Refusal Circuits Across LLMs: Cross-Model Transfer via Trajectory Replay and Concept-Basis Reconstruction.arXiv, 2601.16034. Deng,Y.2026.GEMS:GeometricConstraintsEnableMulti- Semantic Superposition in LLMs.arXiv, 2606.19946. Fomin, M.; David, E.; and LeVi, A
arXiv 2026
-
[4]
Galeone, C.; Ettorre, A.; Park, M.; Ettorre, G.; and Ligorio, D
Internal-State Probes Read the Situation, Not the Action: Three Nega- tiveResultsforPre-ActionMisalignmentMonitoring.arXiv, 2606.30449. Galeone, C.; Ettorre, A.; Park, M.; Ettorre, G.; and Ligorio, D
-
[5]
Steering in Language Models.arXiv, 2606.24952
Perfect Detection, Failed Control: The Geome- try of Knowing vs. Steering in Language Models.arXiv, 2606.24952. Google DeepMind
-
[6]
Kang, D.; Liu, Z.; Ma, N.; Huang, Y.; Tan, Z.; and Jiang, M
Gemma 2: Improving Open Lan- guage Models at a Practical Size.arXiv, 2408.00118. Kang, D.; Liu, Z.; Ma, N.; Huang, Y.; Tan, Z.; and Jiang, M
-
[7]
Prompt-Activation Duality: Improving Activa- tion Steering via Attention-Level Interventions.arXiv, 2605.10664. Kumar,A.;andMaple,C.2026. RefusedinChat,Writtenin Code:Workflow-LevelJailbreakConstructioninIDECoding Agents.arXiv, 2607.03968. Lermen,S.;Dziemian,M.;andPimpale,G.2024. Applying Refusal-Vector Ablation to Llama 3.1 70B Agents.arXiv, 2410.10871. ...
Pith/arXiv arXiv 2026
-
[8]
The Llama 3 Herd of Models.arXiv, 2407.21783. Moskvoretskii, V.; Glandorf, D.; Medina Moreira, J.; Käser, T.;andWest,R.2026.TracingPersonaVectorsthroughLLM Pretraining.arXiv, 2605.13329. Nguyen,T.;Nguyen,T.A.;Alemohammad,S.;andBaraniuk, R. G
Pith/arXiv arXiv 2026
-
[9]
Panickssery,N.;Gabrieli,N.;Schulz,J.;Tong,M.;Hubinger, E.;andTurner,A.M.2024
Minimizing Collateral Damage in Activation Steering.arXiv, 2605.01167. Panickssery,N.;Gabrieli,N.;Schulz,J.;Tong,M.;Hubinger, E.;andTurner,A.M.2024. SteeringLlama2viaContrastive Activation Addition.arXiv, 2312.06681. Tan, D.; Chanin, D.; Lynch, A.; Kanoulas, D.; Paige, B.; Garriga-Alonso, A.; and Kirk, R
Pith/arXiv arXiv 2024
-
[10]
Analyzing the Generalization and Reliability of Steering Vectors.arXiv, 2407.12404. Team, Q
-
[11]
Qwen2.5 Technical Report.arXiv, 2412.15115. Turner, A. M.; Thiergart, L.; Leech, G.; Udell, D.; Vazquez, J.J.;Mini,U.;andMacDiarmid,M.2023.SteeringLanguage Models With Activation Engineering.arXiv, 2308.10248. Walsh, C.; and Barkett, E
Pith/arXiv arXiv 2023
-
[12]
Representation Without Control:TestingtheRealizationEffectinLanguageModels. arXiv, 2605.25151. Yap,J.Q.2026.BehavioralSteeringina35BMoELanguage Model via SAE-Decoded Probe Vectors: One Agency Axis, Not Five Traits.arXiv, 2603.16335. Zhong, V.; and Li, Q
Pith/arXiv arXiv 2026
-
[13]
Zou, A.; Phan, L.; Chen, S.; Campbell, J.; Guo, P.; Ren, R.; Turner, A
Refusal Lives Downstream of Persona in Chat Models.ICML 2026 Mechanistic Inter- pretability Workshop / arXiv, 2606.26161. Zou, A.; Phan, L.; Chen, S.; Campbell, J.; Guo, P.; Ren, R.; Turner, A. M.; Robey, B.; Kolter, Z.; Fredrikson, M.; and Hendrycks, D
Pith/arXiv arXiv 2026
-
[14]
Representation Engineering: A Top- Down Approach to AI Transparency.arXiv, 2310.01405. Preregistrations and Reproducibility Every behavioral experiment was pre-registered before pilot launch, with the live/die crite- ria and the verdict-mapping engine committed to the public repository before any data were collected. The registration documents indocs/incl...
Pith/arXiv arXiv 2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.