{"id":"472b3c0d-3f97-4e3b-875b-cd071914c9d3","arxiv_id":"2608.08159","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"In 17 LLMs from five families, the apparent scaling of concept steerability vanishes under normalized, held-out controls, while a decodable world map persists and other neuroscience parallels are measurement-dependent.","lead":"A cross-family audit of 17 large language models shows that an apparent 'emergent' ability to steer concepts by scale is an artifact of uncalibrated measurements, disappearing when normalization and held-out selection are applied. It also finds that a linear world map is consistently decodable, while number-neuron shape and language localization depend on analysis choices.","discovery_kind":"replication","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No-trend conclusion rests on the untested assumption that residual-norm strengths are functionally comparable doses; if the per-model dose-response peak shifts off the common grid with scale, the corrected pipeline can hide a real scaling trend.","rationale":"The paper's central claim has two parts: the raw-unit scaling is uninterpretable, and the corrected comparison shows no significant trend. The first part is solid: it follows from the scale-independent units argument (the raw injection fraction varies non-monotonically with size) and is reinforced by the 32B quantization control. The second part is the load-bearing positive contribution, and it is an underpowered null whose validity depends on the residual-norm dose being functionally comparable across models. The paper's own Limitations explicitly concede that residual-norm normalization 'does not guarantee equal functional dose across models.' If functional dose varies systematically with scale, the corrected comparison could mask a real trend or create a false flatness, which would reverse the paper's headline message. The held-out operating-point selection mitigates the concern only if the common strength grid is wide enough for every model to reach its own functional optimum; the paper does not report the selected strengths or check grid boundaries, so this remains an open condition. The proposed grid-extension test is a direct, inexpensive check of this failure mode: if the selected c saturates at the grid edge for large models, or if peak effects grow outside the original grid, the flat slope is not trustworthy. A secondary and smaller gap is that the factor decomposition holds the readout metric fixed while varying units and operating point, so the abstract's 'correcting any one of these removes it' is not directly tested for all three factors; that gap is worth fixing but is less consequential than the dose-comparability assumption. Because the concern does not invalidate the artifact claim and the authors have already flagged the dose limitation, the reader's CONDITIONAL verdict remains appropriate; no change is needed, but the paper should run the grid-extension check or add a behavioral calibration before claiming the corrected result is measurement-invariant.","tokens_in":19937,"tokens_out":12168,"duration_ms":119249,"concrete_test":"Extend the operating-point scan for the Qwen3 ladder (0.6, 1.7, 4, 8, 14, 32B) and Qwen2.5-72B to a wider strength grid, e.g., c in {0.125, 0.25, 0.5, 1, 2, 4, 8, 16}, keeping the layer grid and the four-fold concept split unchanged. Record the selected (layer, c) per model and the held-out per-concept specificity effect at that point, then recompute the log2-size slope with a bootstrap CI. If the selected c saturates at the previous upper bound (4, or 2 on the coarse large-model grid) for the larger checkpoints, or if the peak effect rises with scale beyond the original grid, the reported flat slope is an artifact of grid truncation; if the extended grid leaves selected c interior and the slope materially unchanged, the residual-norm dose is a valid comparable unit and the no-trend reading stands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The artifact claim that raw-unit steerability trends are uninterpretable is well supported: Section 4.1 shows that the raw injection fraction ||d||/||h|| swings between 0.12 and 0.27 non-monotonically across the Qwen3 ladder, so a fixed raw coefficient is not a comparable dose. The load-bearing step is the positive companion claim that, after correction, there is no significant scaling trend. That claim depends on the residual-norm parameterization h' = h + c||h|| d_hat (Section 3.3, rule R2) being functionally comparable across models, which the paper's own Limitations section concedes is not guaranteed. Held-out operating-point selection (Section 3.5, Appendix C) mitigates this only partially: it lets each model pick its best (layer, strength) from a common grid (c in {0.25, 0.5, 1, 2, 4}, and coarser {0.5, 1, 2} for models of 24B and larger), but it does not test whether the grid spans each model's functional range. If the dose-response peak drifts toward larger c as width or depth grows, or if a fixed c*||h|| has a different behavioral effect in larger models, then the per-model selected effect is truncated or compressed at the grid ceiling, and the flat slope +0.31 per doubling with 95% CI [-0.11, +0.73] over the 0.6-14B ladder could be an artifact of the corrected pipeline rather than evidence against scale dependence. The paper reports only that peaks occur at low-to-intermediate c for 0.6B, 4B, and 14B (Figure 2c) and does not report the selected c per model or check grid-edge saturation, so the flatness is conditional on an unvalidated dose-comparability assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper audits four neuroscience-inspired interpretability claims across 17 open-weight LLMs from five families, 0.6B to 72B: concept steering, number-magnitude tuning, language localization by lesioning, and the linear world map. The central experiment shows that an apparent scaling law for concept steerability across the Qwen3 ladder is a measurement artifact: with raw activation units, a fixed layer, and a fixed coefficient, steerability appears to grow with scale, but the raw injection displaces a non-monotonic fraction of the residual norm (0.12 to 0.27), and the trend depends jointly on units, readout metric, and operating point. After residual-norm normalization and held-out operating-point selection, steering remains significant at every scale but shows no significant trend across the Qwen3 series (slope +0.31 per doubling, 95% CI [-0.11, +0.73]), with the paper explicitly noting that the null is underpowered. The other results are mixed: number magnitude is strongly encoded but bell-versus-monotonic shape depends on neuron selection; language localization is attribution-dependent; and a linear geographic map is consistently decodable in all 17 models. The paper releases protocol, stimuli, and code.","tokens_in":20185,"tokens_out":5919,"duration_ms":62854,"significance":"The paper makes a valuable methodological contribution to AI neuroscience and interpretability. Its strongest result is the demonstration that a published-style raw-unit steering pipeline manufactures an apparent emergent scaling law, supported by the residual-norm analysis in Fig. 2a, the dose-response and layer-sensitivity curves, held-out operating-point selection, bootstrap confidence intervals, and a bf16-versus-8-bit quantization control. The explicit audit rules R1-R4, the factor decomposition, and the replication across families up to 72B are concrete strengths, as is the candid treatment of nulls and grid-sensitivity. If the artifact claim holds, the paper provides a useful template for comparable cross-model intervention. However, the positive companion claim that no scaling trend exists after correction depends on an untested functional-dose assumption that the paper itself concedes; this limits the force of the headline 'no trend' statement and needs additional work before the claims can be accepted at face value.","major_comments":[{"comment":"The 'no detectable trend' conclusion rests on treating the residual-norm-normalized injection h' = h + c||h||_l d_hat as a functionally comparable dose across models, and the paper itself concedes in the Limitations section that residual-norm normalization 'improves comparability but does not guarantee equal functional dose across models.' The held-out operating-point selection in §3.5 and Appendix C mitigates this only partially: it lets each model choose a (layer, strength) cell from a common grid (c in {0.25, 0.5, 1, 2, 4}, with a coarser {0.5, 1, 2} for models at 24B and above), but it does not test whether that grid spans each model's functional dose-response range. If the inverted-U peak in Fig. 2c drifts toward larger c with scale, or if a fixed c||h|| produces a different behavioral change in larger models, the flat slope (+0.31; 95% CI [-0.11, +0.73]) could be an artifact of the corrected pipeline rather than evidence against scale dependence. Please report the selected (layer, c) for each model, check whether selected strengths lie at grid boundaries, expand the c grid for the Qwen3 ladder, and provide a functional-dose sensitivity analysis, for example by calibrating strengths to a common target-token probability lift or by effect-matching across models.","section":"§3.3 (R2) and §4.1"},{"comment":"The factor decomposition does not isolate which of the three listed confounds (raw units, readout metric, operating point) drives the raw-unit trend, because the readout metric is held fixed at the corrected specificity contrast in all four cells. In particular, the raw+fixed cell in Fig. 6 already uses the specificity contrast and is flat, so the decomposition cannot support the statement that the apparent scaling is 'not attributable to any single factor.' A proper decomposition would vary one factor at a time from the exact naive pipeline (raw units, first-token readout, fixed layer and strength) and report the slope for each single correction. This is important because the abstract and Section 4.1 claim that correcting any one of the three choices removes the trend.","section":"§4.1 and Fig. 6"}],"minor_comments":[{"comment":"The 'activation addition' citation appears as a literal '(?)' placeholder in the Related Work section; please supply the intended reference.","section":"§2"},{"comment":"The sentence describing the memory-efficient lesion attribution is garbled: 'requires grad' should be 'requires_grad', and the phrase 'requires grad removes that buffer' should be reworded for clarity.","section":"§3.6"},{"comment":"The token-by-token examples in Figures 11 and 12 contain unrendered glyph or token-id sequences (e.g., '/uni00000011') rather than readable text; please fix the rendering so the qualitative examples are legible.","section":"Figures 11 and 12 (Appendix D)"},{"comment":"The pass-rate column mixes fine-grid and coarse-grid values, and the table notes that absolute pass-rates are not comparable across these groups; consider adding a visual separator or repeating the caveat directly in the column header so it is not missed by readers.","section":"Table 1"},{"comment":"The 'no significant trend' wording is used in several summary locations without always carrying the power caveat. Since the 95% CI does not exclude +0.73 per doubling and Appendix C notes that 3-4x more concepts or models would be needed to exclude a +0.3 trend, please consistently phrase the result as 'no detectable trend, with a confidence interval that admits a moderate positive slope' in the abstract, Table 2, and the conclusion.","section":"§4.1 and Abstract"},{"comment":"The world-map R2 column lists values such as '0.53/0.67' without stating in the caption that the two numbers are latitude and longitude; please make that explicit.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a good fit for the journal as an audit paper with a strong negative result and a positive control (the world map). The main risk is over-interpretation of the flat steering slope: the functional-dose comparability assumption is untested and is acknowledged only as a limitation. I would require the sensitivity analysis described in major comment 1, or a consistent softening of the no-trend claim, before acceptance. The artifact claim itself is well supported and should survive revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this one. It shows, convincingly, that the apparent \"concept steerability emerges with scale\" result is a measurement artifact. A fixed raw steering coefficient injects a wildly varying fraction of the residual across the Qwen3 ladder (0.12 to 0.27, non-monotonically), and a hand-picked operating point produces erratic pass-rates. After normalizing the intervention to the residual norm and selecting the operating point on held-out concepts, steering remains significant at every scale but shows no significant trend (slope +0.31 per doubling, 95% CI [-0.11, +0.73]). The authors are careful to say \"no detectable trend,\" not \"no trend.\"\n\nThe cross-family audit to 72B is genuinely new, and the factor decomposition shows the artifact is a compound of several choices. The number-neuron result is a nice catch: bell-shaped units exist but are invisible under linear correlation selection. The language-lesion asymmetry flipping under a different attribution method is a clean example of measurement-dependence. The world map replicating to 72B is a good positive control, and the 8-bit versus bf16 control at 32B is responsible.\n\nThe soft spot is the no-trend claim. It rests on residual-norm normalization being functionally comparable across models, which the authors concede is imperfect. The stress-test worry is legitimate: if the dose-response peak shifts off the common c-grid with scale, the corrected pipeline could hide a real trend. They report peaks at low-to-intermediate c for three models but do not report the selected c per model or check for grid-edge saturation. This does not undermine the artifact argument, but it does mean the flatness is conditional. A second gap: the abstract says correcting any one of raw units, readout metric, or operating point removes the trend, but the factor decomposition only varies units and operating point while holding the metric fixed. The readout-metric claim is plausible but not directly tested. Minor issues: a missing citation marked '(?)' in Related Work, and the reproducibility link is anonymous with no commit hash.\n\nWorth a serious referee. The core artifact claim is solid and important for interpretability; the no-trend companion needs stronger dose-comparability evidence and a more careful statement. I'd send it to review and ask for those, not desk-reject. Bring it to reading group; the protocol is reusable.","headline":"A careful audit that convincingly demonstrates the raw-unit steering scaling law is a measurement artifact; the companion no-trend claim is conditional on unvalidated dose comparability.","tokens_in":20828,"tokens_out":4521,"would_cite":true,"duration_ms":39427,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Apparent 'concept steerability emerges with scale' in LLMs is a measurement artifact: correcting raw units, the readout metric, or the operating point removes the trend, while genuine steering remains significant but trend-free across the…","keywords":["concept steering","activation steering","measurement artifact","emergent capabilities","linear probing","AI neuroscience","cross-model audit","residual-norm normalization"],"falsifier":"On a single model family trained with a controlled recipe across sizes, rerun the audited steering protocol with residual-norm normalization and held-out operating-point selection; a significant positive slope with a tight confidence interval would overturn the paper's corrected picture. Alternatively, if the raw-unit apparent emergence persists when only the first-token readout is replaced by the specificity contrast, the artifact would not be joint with the readout metric as claimed.","tokens_in":19617,"feed_emoji":"📏","tokens_out":8117,"duration_ms":77864,"temperature":0.7,"pith_summary":"The paper sets out to test whether neuroscience-style findings in large language models survive stronger measurement controls, and it argues that the main threat is not missing phenomena but incomparable measurements. In its central case study, an apparent scaling law—concept steerability grows with model size—turns out to depend jointly on raw activation units, the readout metric, and a fixed operating point, and correcting any one of these removes the trend. After normalizing steering strength to the residual-stream norm and selecting the layer and coefficient on held-out concepts, the paper finds concept steering remains significant at every scale tested but shows no significant trend across the dense Qwen3 series, with a slope of +0.31 per doubling and a 95% confidence interval from -0.11 to +0.73. The broader audit finds a linear world map consistently decodable in all 17 checkpoints, number magnitude strongly encoded but with neuron shape depending on selection, and language localization that flips direction under a different attribution method. A sympathetic reader would take away that claims like 'emergence' need calibrated, comparable measurements before they are treated as real.","feed_headline":"Concept steering's 'scale emergence' is a measurement artifact","feed_subtitle":"After normalizing the intervention, steering stays significant at every scale but no longer trends with model size.","key_machinery":"The load-bearing device is residual-norm-normalized activation steering: instead of injecting $\\alpha d$ in raw units, the paper unit-normalizes the direction and scales the coefficient by the residual-stream norm, so the injection is $h' = h + c\\|h\\|_\\ell \\hat{d}$ and its size is exactly a fraction $c$ of the residual norm regardless of model. Around this, the audit protocol adds held-out operating-point selection (choosing layer and strength on one split, scoring on a disjoint split), specificity and directional controls (target direction must raise target over control, the control direction must reverse it, random directions must not), bootstrap confidence intervals, and null comparisons. For the other phenomena, the carrying objects are ridge probes with cross-validated regularization for the world map, a shape-agnostic quadratic-fit criterion plus digit-versus-word cross-format correlation for number tuning, and gradient-by-activation versus activation-magnitude attribution for the lesion study. The mechanism that separates artifact from genuine effect is the comparison of measurements before and after each confound is corrected.","core_discovery":"The paper's central claim is that the reported finding 'concept steerability emerges with scale' is a measurement artifact rather than a property of the models. Under the conventional recipe—a raw mean-difference direction injected at a raw coefficient into a fixed layer, read out as first-token probability—steerability looks monotone in scale across the Qwen3 ladder, but the injection displaces a fraction of the residual stream that swings non-monotonically between 0.12 and 0.27. Once the intervention is expressed as a fixed fraction of the residual norm, $h' = h + c\\|h\\|_\\ell \\hat{d}$, and the (layer, strength) operating point is selected on held-out concepts, the effect is significantly positive at every scale but shows no significant trend across model size (slope +0.31 per doubling, 95% CI $[-0.11, +0.73]$). The paper extends the same audit to three other neuroscience parallels: a linear world map decodes from every checkpoint, number tuning is strong but its bell-versus-monotonic shape depends on the neuron-selection criterion, and language-selective lesional asymmetry reverses under a different attribution method. Its conclusion is that comparability and controls, not new phenomena, are the binding constraint on AI neuroscience.","pith_inferences":["Beyond the paper: if the residual-norm dose-comparability assumption holds, the same artifact class should affect other causal representational claims wherever intervention strength is not matched to representation scale.","Beyond the paper: a direct test is to run the audited steering protocol with leave-name-out prompts and non-lexical readouts, which the paper lists as construct-validity work, and check whether the flatness across scale survives.","Beyond the paper: the four-way taxonomy suggests ranking future 'LLMs also show X' claims by how much unit selection, intervention, and operating-point choice the method involves, since those are the places controls bite.","Beyond the paper: the paper's cross-family heterogeneity could partly reflect measurement incomparability rather than genuine family differences; comparing families on matched normalization and matched grids is a natural next step."],"forward_implications":["Future scaling claims about steerability need calibrated units, held-out operating-point selection, and confidence intervals before the trend is interpretable.","Concept steering remains a real, causally effective phenomenon at every tested scale, so the corrected result is not a null result.","Pure decoding results, such as the world map, are less vulnerable to audit than intervention-based results, because no selection or operating-point choice enters.","The corrected Qwen3 result is an absence of detectable trend, not evidence of scale-invariance, since the confidence interval still admits a moderate positive slope."],"supporting_citations":[{"why":"Supplies the conventional raw-unit contrastive activation addition recipe that the paper audits and shows is uncalibrated.","marker":"Rimsky et al., 2024"},{"why":"Demonstrates that emergent abilities can disappear under alternative metrics, providing the template for the artifact argument.","marker":"Schaeffer et al., 2023"},{"why":"Provides the linear world map result that the paper replicates across families and scales.","marker":"Gurnee and Tegmark, 2024"},{"why":"Supplies the superposition argument that decodable directions need not be the units of computation, operationalized as the decodable-does-not-equal-used gap.","marker":"Elhage et al., 2022"},{"why":"Establishes the control-task requirement for linear probes that underlies the specificity and null controls.","marker":"Hewitt and Liang, 2019"},{"why":"Provides the existing cross-scale steering study that finds effectiveness diminishing with size, the counterexample to emergence as an established claim.","marker":"Ali et al., 2025"},{"why":"Formulates the linear representation hypothesis that justifies treating directions as causal concept knobs, the target being audited.","marker":"Park et al., 2024"}],"fun_headline_variants":["Scale emergence in concept steering is a measurement artifact","Normalize the intervention, and steerability stops scaling","What limits AI neuroscience? Measurements, not phenomena","Cross-family audit: steering 'emergence' fails after calibration","Measurement, not model size, drives reported steering emergence"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that an intervention scaled to a fixed fraction of the residual-stream norm delivers a functionally comparable dose across models of different sizes and architectures; the paper itself concedes in its Limitations that this improves comparability but does not guarantee equal functional dose.","fun_headline_variants_meta":{"raw":{"variants":["Scale emergence in concept steering is a measurement artifact","Normalize the intervention, and steerability stops scaling","What limits AI neuroscience? Measurements, not phenomena","Cross-family audit: steering 'emergence' fails after calibration","Measurement, not model size, drives reported steering emergence"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001123,"raw_usage":{"total_tokens":4767,"prompt_tokens":1136,"completion_tokens":3631,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":752,"completion_tokens_details":{"reasoning_tokens":3555}},"tokens_in":752,"tokens_out":3631,"duration_ms":26139,"temperature":1.0,"reasoning_tokens":3555,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T00:19:29.661746+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On a single model family trained with a controlled recipe across sizes, rerun the audited steering protocol with residual-norm normalization and held-out operating-point selection; a significant positive slope with a tight confidence interval would overturn the paper's corrected picture. Alternatively, if the raw-unit apparent emergence persists when only the first-token readout is replaced by the specificity contrast, the artifact would not be joint with the readout metric as claimed.","supporting_citations":[],"review_version":1}