{"id":"b5433dd0-349c-4998-9179-ad0bb00f10d3","arxiv_id":"2505.01162","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"Steering vectors on GPT-2 XL reduce explicit demographic bias in simple prompts but overcorrect and contradict facts in complex social scenarios, leading the authors to conclude they are not yet robust for general alignment.","lead":"This paper tests whether steering vectors, an inference-time technique for guiding language model behavior, can align GPT-2 XL toward values like equality and impartiality. It finds the technique works for simple prompts but produces factual errors and contradictions in complex social scenarios.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Three hand-tuned examples from one small model cannot support a general limitation claim about steering; the observed failures are confounded by per-scenario coefficients and vector construction.","rationale":"The central claim is a negative generalization about a class of methods. For that claim to hold, the three failures must be intrinsic to steering vectors as an alignment mechanism, not artifacts of model scale, vector construction, target concept, layer, or coefficient. The paper's design cannot separate these factors because every dimension varies across the three examples and none is varied systematically. Table 1 assigns a different target direction, layer, and coefficient to each scenario; Section 4 gives no quantitative results and no ablations. The two failures that carry the argument—the legal case denial and the election hallucination—are classic failure modes of GPT-2 XL on multi-entity narrative prompts, so an equally plausible reading is that a 1.5B language model with an aggressive coefficient produces incoherent completions. The paper's own 'proof of concept' language in Section 1 and Future Work is more consistent with that reading than with the abstract's general claim. A single controlled replication that sweeps coefficients, models, and vector constructions would decide whether the failures are a property of steering or of this specific recipe. Until then, the evidence supports a narrow case study, not a general limitation. The reader's weakest assumption correctly identifies the same gap, and I see no additional load-bearing concern beyond this external-validity problem.","tokens_in":3919,"tokens_out":5245,"duration_ms":52128,"concrete_test":"Run a coefficient-sweep replication of Table 2's three prompts (plus at least 20 similar demographic/social prompts) on GPT-2 XL and two instruction-tuned models (e.g., Llama-2-7B-chat, Mistral-7B-Instruct), using at least two vector constructions: (i) the antonym-derived directions as in the paper, and (ii) CAA vectors from full demonstration pairs as in Panickssery et al. Sweep coefficients over a grid (e.g., +1 to +15 in steps of 2, and negative counterparts). Score outputs with automatic metrics for factual consistency (entailment against the prompt's premises), demographic bias, and coherence. If the legal/election failures disappear under any reasonable model/vector/coefficient combination, the paper's general limitation claim is not supported; if they persist across all combinations, the claim gains support.","verdict_should_be":"UNCHANGED","load_bearing_attack":"To establish that steering vectors 'may not provide a robust foundation for general-purpose alignment,' the paper must show that the failures in Table 2 are attributable to steering as a method, not to the particular recipe used. The design does not permit this inference. Each of the three scenarios employs a different target direction, a different layer (3, 8, 18), and a different empirically tuned coefficient (Table 1), and all are evaluated on a single 1.5B GPT-2 XL. Section 4 presents only qualitative comparisons of three completions; there is no coefficient sweep, no variation of vector construction (only antonym-based contrast pairs), no comparison across models, and no quantitative metric. The legal-case and election failures—an outright denial of facts and a mixed-gender/race confusion—are exactly the kind of coherence errors GPT-2 XL is known to produce on long multi-entity prompts, so the examples cannot separate 'steering fails' from 'this small model plus this tuned coefficient fails on this prompt.' The paper's own Introduction calls the work 'a proof of concept,' yet the abstract and Section 5 generalize from these three cases. Accepting every reported output as accurate, the strongest defensible conclusion is that one particular CAA-style value direction on GPT-2 XL can reduce explicit bias but sometimes harms factual consistency on two narrative prompts.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a framework for assessing limitations of activation steering as an alignment mechanism. The authors use GPT-2 XL with contrastive activation addition, select intervention layers via causal indirect effect analysis on an antonym task, and steer three demographic-sensitive scenarios with manually tuned coefficients. The reported findings are qualitative: one unsteered and one steered continuation per scenario. The abstract and conclusion state that steering is promising for value alignment but may not provide a robust foundation for general-purpose alignment in complex scenarios. The paper positions itself as a proof of concept and suggests future work on calibration, context-sensitive steering, and extension to other models.","tokens_in":4247,"tokens_out":4818,"duration_ms":48223,"significance":"If the negative findings were rigorously established, the paper would be a useful caution for steering-based alignment. The methodological framing—using causal analysis to select layers and contrastive activation addition to construct vectors—is coherent, and the authors are transparent about coefficients and example outputs. However, the contribution as it stands is a small proof of concept: three qualitative examples on one 1.5B model, with no quantitative metrics, no controls, no ablations, and no code or data release. The paper's main value is as a pointer to potential risks rather than as a settled empirical result, and the general conclusion in the abstract and Section 5 substantially overreaches the evidence.","major_comments":[{"comment":"The abstract and §5 generalize from steering to 'general-purpose alignment' on the basis of three hand-picked continuation examples from a single GPT-2 XL model. Table 2 reports no quantitative metrics, no coefficient sensitivity analysis, no variation in vector construction, and no comparison across models or prompts, so the failures shown could reflect the particular prompt, the particular tuned coefficient, or the model's known weakness on long multi-entity narratives rather than a property of steering as a method. The paper's own Introduction and §6 label the work a 'proof of concept,' which is in tension with the generalized conclusion.","section":"Abstract; §4; Table 2"},{"comment":"The intervention layers are selected via causal indirect effect analysis on an antonym prediction task, but no evidence is given that the layers most influential for antonym prediction are the right layers for value and demographic-bias steering. Moreover, each scenario uses a different layer and a different empirically tuned coefficient (+3/-3, +11/-11, +8/-8), and the tuning procedure—search range, objective, number of samples—is not reported. Consequently, the design cannot separate 'steering vectors fail' from 'these layers, coefficients, and contrast pairs fail on these prompts.'","section":"§3; Table 1"},{"comment":"The steering-target labels in Table 2 do not match the vector definitions in Table 1: Sean Morgan's row in Table 1 is 'Equality/Inequality' but the steered hiring output is labeled 'towards Equality and Impartial'; Farooq Hassan's row is 'Impartial/Prejudiced' but the legal-case output is labeled 'towards Non-Partisan and Equality'; Kwame Matthews's row is 'Non-partisan/Partisan' but the election output is labeled 'towards Non-Partisan and Impartial.' It is unclear whether multiple vectors were combined, whether the labels are typos, or whether different vectors were used than reported; this inconsistency must be resolved before the qualitative findings can be interpreted.","section":"§4; Table 2 vs Table 1"},{"comment":"The 'overcorrection' finding for the legal case treats the unsteered model's continuation as factual evidence against which the steered continuation is judged ('despite the evidence suggesting otherwise'), but the prompt is an open-ended continuation and no ground truth is supplied. Both the unsteered and steered completions are model-generated accounts of a fictional trial, so the assertion that the steered output introduces factual inaccuracies is not grounded in any externally supplied fact.","section":"§4; Legal case row"}],"minor_comments":[{"comment":"There is a typo: 'expenisve' should be 'expensive', and 'P anickssery' should be 'Panickssery'.","section":"§1"},{"comment":"There is a typo: 'emperically tuned' should be 'empirically tuned'.","section":"§3"},{"comment":"The table title 'Demographic and Case Variants' is misleading because the table also includes steering targets, layers, coefficients, and ethnicity/gender/religion information; a more descriptive title such as 'Scenarios, Demographic Variants, and Steering Configurations' would be clearer.","section":"Table 1"},{"comment":"Figure 1 lacks axis labels and a color-scale legend, making the claim that layers 3, 8, and 18 were selected from the CIE analysis difficult to verify from the figure alone.","section":"Figure 1"},{"comment":"The paper states that the antonym dataset was generated using GPT-4, but it does not report the number of pairs, the filtering steps, or the exact contrast-pair templates; this limits reproducibility.","section":"§3"},{"comment":"The text refers to 'Appendix Table 2', but Table 2 appears in the main text; the reference should be corrected, and the table formatting should be cleaned so each scenario has a complete prompt and two full outputs.","section":"Table 2"}],"recommendation":"reject","confidential_remarks":"The manuscript would need a substantially larger empirical study—multiple models, quantitative metrics, ablations over coefficients and vector construction, and a clearly defined factual ground truth—before the central claim could be supported. The current three-example qualitative design is better suited to a workshop-style position paper than to a full journal article, and the mismatch between the paper's 'proof of concept' framing and its generalized conclusions is a further concern."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a clearly written proof-of-concept that shows some value-steering examples in GPT-2 XL, but the central claim—that steering vectors 'may not provide a robust foundation for general-purpose alignment'—does not follow from the evidence. The paper is honest about its scope in the intro and future work, but the abstract and conclusion step well beyond it.\n\nWhat's actually new: applying CAA-style contrastive activation addition to demographic value dimensions (equality, impartiality, non-partisanship) on GPT-2 XL, with a causal-analysis-based layer selection. The concrete examples of steering reducing explicit bias but causing factual distortions are mildly interesting. The writing is clear and the related work is properly cited, including Tan et al. (2025), which already raises reliability concerns.\n\nSoft spots, in order of weight:\n\n1. Three qualitative examples, each with a different target direction, a different intervention layer, and a different tuned coefficient. There are no quantitative metrics, no controls across coefficient values, no variation of vector construction, and no second model. The failures in Table 2 are exactly what GPT-2 XL does on long multi-entity prompts; they cannot be attributed to steering as a method. The stress-test note is right: the design cannot separate 'steering fails' from 'this vector plus this coefficient plus this small model fails on this prompt.'\n\n2. The coefficient tuning is circular: coefficients are empirically chosen on the same scenarios used for evaluation, and the displayed outputs are selected. That makes the negative findings partly self-produced. The paper does not report a sweep or a principled selection rule.\n\n3. Layer selection via an antonym ICL task is then applied to value steering without evidence that the causal structure transfers. That is a potential mismatch, unaddressed.\n\nNovelty is limited: the conceptual conclusion largely restates Tan et al. (2025). The contribution is a small extension of CAA to demographic values, which could be a useful building block, but is not a new result as presented.\n\nWho this is for: someone collecting anecdotal evidence on steering behaviors might skim it, but it is not a reliable basis for a strong claim. As a workshop note it's fine; as a paper claiming to document limitations of steering, it needs proper evaluation: multiple models, controlled coefficient sweeps, quantitative coherence metrics, and pre-registered scenarios.\n\nRecommendation: I would not send this to a serious referee in its current form. It is a preliminary proof of concept that could be revised into a rigorous negative result, but the central claim is unsupported. A desk reject with encouragement to resubmit after substantial rework seems right.","headline":"A well-written proof of concept whose sweeping conclusion about steering vectors rests on three hand-tuned anecdotes from one small model; the limitation claim is plausible but unsupported.","tokens_in":4723,"tokens_out":2631,"would_cite":false,"duration_ms":24573,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Steering vectors are not a dependable foundation for general-purpose alignment; they work for narrow value tasks but fail in complex social contexts.","keywords":["steering vectors","language model alignment","activation engineering","contrastive activation addition","demographic bias","value alignment","inference-time intervention","GPT-2 XL"],"falsifier":"A systematic stress test with dozens of demographic prompt templates per complexity level, automated factual-consistency scoring, and a coefficient sweep would settle the claim: if steered outputs on complex prompts distort facts no more often than on simple prompts at any coefficient, the observed failures are artifacts of the three examples, not a general limitation of steering.","tokens_in":3749,"feed_emoji":"⚖️","tokens_out":11791,"duration_ms":102061,"temperature":0.7,"pith_summary":"This paper asks whether steering vectors — linear directions added to a transformer's internal activations at inference time — can serve as a general-purpose alignment mechanism for large language models. The authors build on contrastive activation addition, construct value-oriented steering vectors for equality, impartiality, and non-partisanship, and test them on demographic prompts of varying social complexity. They find that steering corrects obvious bias in a simple hiring scenario but, in more complex legal and electoral contexts, overcorrects and introduces factual and logical contradictions. The central claim is that steering vectors are useful for narrow value-alignment tasks but are not a dependable foundation for general-purpose alignment, especially where precise multi-faceted reasoning is required. A sympathetic reader should care because steering is one of the cheapest inference-time alignment tools, and this work begins to map where that tool can be trusted.","feed_headline":"Steering fixes simple bias, fails on complex prompts","feed_subtitle":"A GPT-2 XL study shows steering corrects bias yet overcorrects and contradicts itself in complex social prompts.","key_machinery":"The central object is the steering vector built by contrastive activation addition: take the difference between hidden activations for contrasting concepts such as equality versus inequality, then add that direction into the residual stream at a chosen layer with a tunable coefficient. The paper derives its vectors from antonym pairs generated by GPT-4, and uses causal indirect effect (CIE) analysis to select layers 3, 8, and 18, where interventions should have the most influence. This machinery does two jobs: it locates the layers where a concept is represented, and it fixes a single linear push whose success or failure can then be observed as prompt complexity increases.","core_discovery":"Steering vectors are linear directions in a model's activation space; adding one at inference time shifts outputs toward a target concept. The paper's demonstrations show that in a simple hiring prompt the equality-steered model drops the religious affiliation cue and chooses the candidate 'because she is the best candidate for the job,' which reads as successful bias mitigation. The same kind of intervention fails when the prompt is legally or socially complicated: the non-partisan-steered legal output absolves the defendant despite evidence to the contrary, and the impartiality-steered election output produces a self-contradictory sentence describing a Black man as the first Asian-American woman president. The paper takes this contrast as evidence that steering is not a general-purpose alignment mechanism: it can redirect obvious bias, but it cannot be trusted to preserve factual and logical consistency in complex, multi-attribute contexts. The intended contribution is a proof-of-concept framework — causal indirect effect layer selection, antonym-derived vectors, and transformer hook interventions — for mapping where steering works and where it breaks.","pith_inferences":["An immediate testable extension is to repeat the same CIE and steering pipeline on instruction-tuned or reasoning-focused models; if the same overcorrection and contradictions appear, the limitation is general rather than specific to GPT-2 XL.","The overcorrection failure suggests steering may amplify the model's prior on the target value beyond what the prompt's facts support, so measuring output confidence across a coefficient sweep would turn the qualitative observation into a quantitative calibration curve.","The three hand-picked cases could be expanded into a benchmark with many demographic prompt templates and automatic factual-consistency scoring, giving the proposed framework a statistical footing it does not yet have."],"forward_implications":["Steering should be used as a narrow intervention tool, reliable for binary value judgments but not as a general alignment solution.","Deploying steering in legal, electoral, or other socially complex settings requires a separate guard against factual overcorrection and self-contradiction.","Alignment evaluation should include prompts with multiple demographic attributes and should score factual consistency, not only bias removal.","The causal indirect effect layer-selection procedure offers a reusable way to locate where a concept is encoded before intervention, which can transfer to other steering targets."],"supporting_citations":[{"why":"It supplies the contrastive activation addition method that the paper extends from binary behavioral choices to value steering.","marker":"Panickssery et al. (2024)"},{"why":"It establishes activation engineering as a way to add steering directions into the residual stream, the intervention mechanism tested here.","marker":"Turner et al. (2024)"},{"why":"It demonstrates inference-time intervention for eliciting truthful answers, forming the promise of steering that the paper scrutinizes.","marker":"(Li et al., 2024)"},{"why":"It raises the generalization and reliability questions about steering vectors that this work directly addresses.","marker":"(Tan et al., 2025)"},{"why":"It shows that refusal is mediated by a single direction, a headline example of targeted steering the paper contrasts with general alignment.","marker":"(Arditi et al., 2024)"},{"why":"It defines GPT-2 XL, the 1.5-billion-parameter model whose activations are steered in all experiments.","marker":"(Radford et al., 2019)"},{"why":"It provides GPT-4, used to generate the antonym pairs that ground the steering vectors in clean binary contrasts.","marker":"(OpenAI et al., 2024)"}],"fun_headline_variants":["Steering fixes bias, but fails on complex legal and social prompts","Simple bias fix, complex prompt failure: steering's limits","Steering corrects obvious bias, overcorrects in complex contexts","Steering: effective for value alignment, not for complex reasoning","Why steering vectors fail on intricate prompts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's conclusion rests on the assumption that three hand-picked output examples from one GPT-2 XL model, obtained with empirically tuned coefficients, reveal the inherent properties of steering vectors rather than artifacts of that particular model, prompt set, or tuning run.","fun_headline_variants_meta":{"raw":{"variants":["Steering fixes bias, but fails on complex legal and social prompts","Simple bias fix, complex prompt failure: steering's limits","Steering corrects obvious bias, overcorrects in complex contexts","Steering: effective for value alignment, not for complex reasoning","Why steering vectors fail on intricate prompts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000163,"raw_usage":{"total_tokens":1184,"prompt_tokens":829,"completion_tokens":355,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":445,"completion_tokens_details":{"reasoning_tokens":273}},"tokens_in":445,"tokens_out":355,"duration_ms":3719,"temperature":1.0,"reasoning_tokens":273,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:23:54.718875+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A systematic stress test with dozens of demographic prompt templates per complexity level, automated factual-consistency scoring, and a coefficient sweep would settle the claim: if steered outputs on complex prompts distort facts no more often than on simple prompts at any coefficient, the observed failures are artifacts of the three examples, not a general limitation of steering.","supporting_citations":[],"review_version":1}