{"id":"b01ed8d2-7692-49de-9524-d27fbdfad7dc","arxiv_id":"2608.00023","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A role-steering screening workflow on 275 roles shows role-specific activation directions beat a non-scale-matched assistant-direction control (63.2 vs 41.1 judged alignment) and flags 38 roles as 'anti-controllable'.","lead":"This paper builds a screening workflow to check whether a language-model agent really plays a given role before it is put into a social simulation. Steering most roles harder makes them score better on a model-judged alignment test, but 38 of 275 roles get worse on every dimension when steering is increased.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Self-referential judge, not norm mismatch, is the load-bearing weakness: GPT-4.1-mini both defines and scores role fidelity, so the headline gap and anti-controllable flags may be judge artifacts.","rationale":"The reader's weakest-assumption identification is correct: the evaluation is self-referential, with the same model generating the reference, filtering extraction positives, and scoring outcomes. This threatens the validity of both the headline comparison and the anti-controllable classification, which are the paper's main empirical contributions. The norm-matching concern is real and is the paper's own stated caveat, but it affects only the assistant-axis comparison and is not the deepest issue. The paper is unusually honest about its scope, provides extensive reliability checks, and releases artifacts, so the appropriate verdict remains CONDITIONAL: the workflow is plausible and useful, but the central measure needs external validation before the role-fidelity claims can be accepted. A human-rated subset is the most direct test because it breaks the closed loop of judge-generated targets and judge-assigned scores.","tokens_in":27012,"tokens_out":5266,"duration_ms":57442,"concrete_test":"Have 3–5 human annotators score a stratified sample of roughly 200 responses with the same 0–100 role-alignment rubric (or pairwise preference), spanning the role-vector and assistant-axis conditions and including both controllable and anti-controllable roles, without exposing the GPT-4.1-mini scores. Compare human–judge Spearman correlation and the human-observed mean gap. If human-judge agreement is weak (r<0.5) or the 63.2-vs-41.1 gap does not reproduce, the central claim is unsupported as stated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central empirical quantity—judged role-profile alignment—is produced entirely by one model, GPT-4.1-mini, in three connected roles. Judge-filtered vector extraction (§3) keeps only responses that GPT-4.1-mini labels 'fully role-playing'; the prompted role reference (§A.3) is generated by GPT-4.1-mini; and all six alignment scores are assigned by GPT-4.1-mini (Judge 1 and Judge 2, §B.1). The role vectors are therefore selected to satisfy this judge's notion of role-playing and then evaluated against a target built by the same judge. The reported reliability metrics (e.g., RDM split-half r=0.97, §E.6) show only self-consistency, not validity. Consequently, the abstract's 63.2-vs-41.1 gap over the assistant-axis control and the classification of 38 'anti-controllable' roles (§4.2) could be properties of GPT-4.1-mini's priors rather than of the steering directions. The paper honestly names this boundary in §6 and the Reproducibility Statement, and the practical per-role screen may survive a weaker interpretation, but the load-bearing condition for the central claim is that the judge measures role fidelity, and that condition is currently untested. The norm-matching issue is real but secondary: even a perfectly norm-matched assistant-axis comparison would not establish that the dependent measure tracks anything beyond the judge's own stereotype.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an activation-steering screening workflow for role-conditioned LLM agents used in social simulations. For each of 275 roles, it constructs a structured profile, extracts a layer-16 contrastive direction from judge-filtered responses, sweeps four steering coefficients α ∈ {1.0, 1.5, 2.0, 2.5}, and scores the resulting generations with GPT-4.1-mini judges on six alignment dimensions. The two headline claims are (i) that role-specific directions outperform an assistant-axis directional control in mean judged role-profile alignment (63.2 vs. 41.1) while preserving lexical diversity, and (ii) that the per-role screen identifies a meaningful minority of 38 'anti-controllable' roles that decline on all six axes as α increases. The paper frames the contribution as a pre-deployment calibration method rather than as a new steering estimator, and it repeatedly and honestly scopes its claims to the specific model, layer, coefficient grid, and judge pipeline used.","tokens_in":27320,"tokens_out":3916,"duration_ms":41471,"significance":"If the central finding is valid, the workflow is practically valuable: simulation builders could screen role agents before deployment, choose per-role coefficients, and flag roles where stronger steering is counterproductive. The scale is substantial — 275 roles, 228 role-agnostic questions, four coefficients, ~500,000 judged responses — and the paper ships evaluation artifacts, detailed appendices, and explicit validity boundaries. The reproducibility statement and the repeated caveats about the non-scale-matched control and the model-generated reference are commendable. However, the core empirical quantity, 'role-profile alignment', is produced entirely by one model in three connected roles: GPT-4.1-mini filters the examples used to extract the role vectors, generates the prompted reference against which outputs are scored, and serves as the judge for all six alignment dimensions. Consequently, the headline gap and the anti-controllable classification could reflect the judge's own priors rather than properties of the steering directions. The norm mismatch between role vectors and the assistant-axis control further confounds the direction-vs-magnitude comparison. These issues are","major_comments":[{"comment":"The dependent measure is self-referential. GPT-4.1-mini generates the prompted role reference (§A.3), filters the contrastive pairs used to extract each role vector (only responses it labels 'fully role-playing' are retained, §3), and then scores all six alignment dimensions (Judge 1 and Judge 2, §B.1). Role vectors are therefore selected to satisfy GPT-4.1-mini's notion of role-playing and evaluated against a target built by the same model. The reported 63.2-vs-41.1 gap over the assistant-axis control and the identification of 38 'anti-controllable' roles could be properties of the judge's stereotype prior rather than of the steering directions. The paper names this boundary in §6 and the Reproducibility Statement, but it remains untested. I would need at least one independent evaluation anchor — e.g., a human-rated subset, a judge model not used in extraction or reference construction,","section":"§3, §4, §A.3, §B.1"},{"comment":"The role-vs-control comparison is not norm-matched. The assistant-axis vectors have mean ℓ2 norm 9.68 versus 3.79 for role vectors (a factor of 2.56×), so equal coefficients apply substantially larger perturbation magnitudes in the control condition. The paper states this caveat repeatedly and appropriately narrows the conclusion to 'under the current extraction and evaluation setup', but the abstract's headline comparison is still presented without it. The control's sharp decline at α=2.5 and its negative median per-role correlation (r=−0.89) may be driven by the larger perturbation magnitude rather than by the direction being non-role-specific. A norm-matched assistant-axis baseline, or an equivalent analysis with perturbation magnitude held constant, is required to support the directional interpretation that the paper's framing implies.","section":"§3, §4, Appendix C"},{"comment":"The 'anti-controllable over the tested range' classification relies on Pearson r computed from only four coefficient values per role per axis, without per-role uncertainty estimates or significance thresholds. With n=4, an r<0 classification is fragile, and the observed mean score drop is modest (7.0 points from α=1.0 to 2.5, Figure 12). Since this classification is the main practical output for simulation builders, the paper should report bootstrap or permutation-based confidence intervals for the per-role correlations, show the stability of the 38-role set under resampling of questions or judge runs, and justify that the 'all six axes r<0' criterion is not dominated by measurement noise. As written, the flag list may over-identify roles as anti-controllable.","section":"§4.2, Appendix C, Eq. (1)"}],"minor_comments":[{"comment":"The norm-matching caveat appears in §3 and §4 but not in the abstract. Since the abstract's 63.2-vs-41.1 comparison is the headline result, a one-sentence caveat should be added there.","section":"Abstract"},{"comment":"The extraction details are incomplete: judge-filtering thresholds, retained-pair counts, parse-failure counts, exact model revisions, decoding parameters, and activation pooling details are deferred to released artifacts. These should be summarized in the paper itself for reproducibility.","section":"§3 / Reproducibility Statement"},{"comment":"The text reports median r=−0.89 for the assistant-axis control, but the figure caption and the discussion of the two distributions could more clearly separate the role-vector and control histograms. Consider showing both distributions on the same axes with matched binning.","section":"§4.1 / Figure 4"},{"comment":"The example doctor few-shot turns include catchphrases that appear semantically disconnected from the stated role ('Trust me, I'm a doctor'; 'the needs of the many outweigh the needs of the few'). If these are intentional caricatures, they should be flagged as such; otherwise they undermine the example's pedagogical value.","section":"Appendix A.3, Table 5"},{"comment":"The effective-rank analysis is interesting but the connection to the main screen could be stated more directly: the low rank of the behavioral readout means the judge-based screen may not resolve all representational variation. This is acknowledged, but a one-sentence practical implication in the main text would help.","section":"§E.6 / Appendix J"}],"recommendation":"major_revision","confidential_remarks":"The paper is transparent and the workflow is potentially useful, but the central empirical comparison is currently supported only by a self-referential judge and a non-scale-matched control. I do not think this requires rejection — the authors have already scoped many of the limitations in §6 — but the central claim needs an independent evaluation anchor and a norm-matched comparison before publication. The anti-controllable classification also needs uncertainty quantification. These are fixable within the manuscript's scope, hence major_revision rather than reject."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The real contribution here is practical: a per-role activation-steering screen for social-simulation agents, with a clean operational taxonomy (controllable / partial / anti-controllable) and a large 275-role empirical sweep. Simulation builders could actually use this — pick a role, sweep four alpha values, see which coefficients work and which roles should be flagged. That is new and useful, and the paper deserves a serious referee.\n\nWhat it does well: the workflow is straightforward, the eval battery is large (228 questions, roughly 500k judged responses), the anti-controllable finding (38 roles declining on all six axes) is the kind of heterogeneity the field needs more of, and the paper is unusually honest. It names the norm mismatch (assistant-axis vectors 2.56x larger), the four-point correlation limit, the model-generated prompted references, and the shared-prior concern. The pairwise subset check at alpha=2.5 is a nice addition, even if it uses the same judge family.\n\nThe soft spots are real, and the stress-test note is right that the self-referential judge is the load-bearing one. GPT-4.1-mini generates the prompted role reference, filters the extraction positives, and scores all six alignment dimensions. So the 63.2-vs-41.1 gap and the anti-controllable flags partly measure that model's own stereotype prior. A norm-matched control would not fix this — even a perfectly matched assistant-axis comparison would still be scored by the same judge against a target built by the same judge. The paper scopes its outcome as \"judged agreement with the constructed role description and prompted role reference,\" so it is not hiding the limitation, but the abstract presents the gap without that caveat. The other items are secondary: parse-failure defaults to 0 or 50 with counts unreported, the 39-role subset selection is unexplained, and vector normalization / token span details are deferred to anonymous artifacts. None of these is fatal, because the practical screen is a relative within-judge tool: the per-role curves and the anti-controllable flag survive a weaker interpretation of the absolute scores.\n\nFor peer review: I would send it out, with a request for a norm-matched control, a small human-rated validation subset, parse-failure counts, and a sentence on how the 39-role subset was chosen. The workflow is the contribution, and it stands up to scrutiny. The headline numbers need softening.\n\nFor a reading group: solid discussion piece about what empirical evidence activation-steering papers need. I'd cite it if I worked on LLM-agent simulation, mostly for the anti-controllable category and the screening recipe.","headline":"A genuinely useful per-role screening workflow, honestly scoped, but the headline numbers rest on a self-referential judge; the workflow survives the weakness while the abstract oversells it.","tokens_in":27896,"tokens_out":1376,"would_cite":true,"duration_ms":18350,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Activation steering can role-condition an LLM for social simulation, but only if each role is screened coefficient-by-coefficient.","keywords":["activation steering","role-conditioned agents","social simulation","role-profile alignment","controllability","assistant axis","LLM evaluation","screening workflow"],"falsifier":"Run the same 275-role extraction and coefficient sweep but replace GPT-4.1-mini as judge with a held-out, human-validated rubric for role fidelity (or a different model family such as a stronger general evaluator). If the mean role-vector overall score stops exceeding the assistant-axis control by more than a few points, or if many of the 38 anti-controllable roles fail to decline, the paper's central claim is falsified rather than the screen being merely internally consistent.","tokens_in":26856,"feed_emoji":"","tokens_out":1248,"duration_ms":29822,"temperature":0.7,"pith_summary":"The paper tries to establish that activation steering — adding a role-specific direction to a language model's internal activations — can be turned into a repeatable, pre-deployment screening workflow for building role-conditioned agents. It claims that across 275 roles, role-specific directions receive higher judged role-profile alignment than a generic assistant-axis directional control (mean 63.2 vs 41.1), and that most roles improve as the steering coefficient increases. Crucially, it shows that 38 roles decline across all six measured dimensions with stronger steering, so a one-size-fits-all high-strength setting is unsafe. The practical point is that simulation builders can measure, compare, and flag candidate agent configurations before use, rather than trusting a prompt or a single steering vector blindly.","feed_headline":"Steering fails on 38 roles: per-role screening beats uniform settings","feed_subtitle":"Across 275 roles, role vectors beat an assistant-axis control 63.2 vs 41.1, but 14% of roles deteriorate. Calibrate each role before deploym","key_machinery":"The central object is the role-specific activation direction vr, computed as the layer-16 mean-of-differences over judge-filtered pairs of 'fully role-playing' positive responses and 'Assistant-style' negative responses — the optimal pointwise-MSE estimator under the formulation of Im and Li (2025). Steering applies additive intervention h→h+αvr at the same layer, with a swept scalar coefficient α. The workflow around this object — define role profile, extract direction, sweep four α values, evaluate role-profile alignment with LLM judges, and pass or flag each configuration — is what carries the argument, turning a mechanism into a testable, per-role calibration procedure.","core_discovery":"The central claim is that role-specific activation directions, extracted via judge-filtered contrastive mean-difference at layer 16 of OLMo-3-7B-Instruct, produce higher judged role-profile alignment than a non-scale-matched assistant-axis directional control across a four-point coefficient grid (α∈{1.0,1.5,2.0,2.5}). On a 228-question role-agnostic battery, mean overall alignment is 63.2 for role vectors versus 41.1 for the assistant-axis control; the control's unique-bigram ratio also collapses at large coefficients (0.94→0.27) while role vectors stay high (0.95→0.92). The paper's key practical discovery is heterogeneous response: 74% of roles improve monotonically with α (median per-role","pith_inferences":["A norm-matched assistant-axis experiment could separate direction from perturbation magnitude; if a norm-matched control still declines, it would substantiate that direction, not scale, drives the 63.2-vs-41.1 gap.","The anti-controllable category is likely populated by roles that saturate early; lowering the starting coefficient or using a sub-linear ramp could recover usable configurations that the current grid misses.","Extending the same screening procedure to multi-turn interactions and agent-agent settings would test whether the role-fidelity gains persist beyond single-turn responses, which the paper explicitly leaves for future work.","Because the judge and the prompted reference share model priors, the pipeline may systematically reward stereotype-consistent output; an independent human-judged validation set could recalibrate the screen's thresholds."],"forward_implications":["If role-specific steering reliably raises judged role-profile alignment, simulation builders gain a measurable, tunable knob for role conditioning that does not depend on prompt wording alone.","The per-role screen means agents can be individually vetted: a candidate configuration passes only if it expresses the profile without increasing lexical repetition or deteriorating at stronger coefficients.","The 38 anti-controllable roles imply that deploying a single global steering strength across a population would silently degrade a meaningful minority of simulated agents.","Because role profiles are explicit modeling assumptions, the workflow allows builders to document provenance and disclose which roles were flagged, before agents enter a simulation.","The comparison to the assistant-axis control shows that a generic assistant-like direction is not interchangeable with role-specific directions under this setup, though the comparison is not norm-matched."],"fun_headline_variants":["Role steering beats assistant-axis control 63.2 vs 41.1, but 38 roles decline","Per-role screening beats uniform steering: 38 roles worsen at high alpha","Role-specific vectors score 63.2 vs 41.1, yet 14% of roles degrade","Screening workflow: 38 roles fail, so calibrate per role, not one size"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that GPT-4.1-mini's judged agreement with a GPT-4.1-mini-generated prompted reference is a valid measure of role fidelity — if the judge rewards its own stereotype rather than role-appropriate behavior, the headline score gap and the anti-controllable classification reflect the judge, not the steering.","fun_headline_variants_meta":{"raw":{"variants":["Role steering beats assistant-axis control 63.2 vs 41.1, but 38 roles decline","Per-role screening beats uniform steering: 38 roles worsen at high alpha","Role-specific vectors score 63.2 vs 41.1, yet 14% of roles degrade","Screening workflow: 38 roles fail, so calibrate per role, not one size"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000334,"raw_usage":{"total_tokens":1716,"prompt_tokens":796,"completion_tokens":920,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":540,"completion_tokens_details":{"reasoning_tokens":822}},"tokens_in":540,"tokens_out":920,"duration_ms":10012,"temperature":1.0,"reasoning_tokens":822,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T01:40:50.645220+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 275-role extraction and coefficient sweep but replace GPT-4.1-mini as judge with a held-out, human-validated rubric for role fidelity (or a different model family such as a stronger general evaluator). If the mean role-vector overall score stops exceeding the assistant-axis control by more than a few points, or if many of the 38 anti-controllable roles fail to decline, the paper's central claim is falsified rather than the screen being merely internally consistent.","supporting_citations":[],"review_version":1}