Pith. sign in

REVIEW 3 major objections 5 minor 34 references

Activation steering can role-condition an LLM for social simulation, but only if each role is screened coefficient-by-coefficient.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-04 01:40 UTC pith:MGK6NPTW

load-bearing objection A genuinely useful per-role screening workflow, honestly scoped, but the headline numbers rest on a self-referential judge; the workflow survives the weakness while the abstract oversells it. the 3 major comments →

arxiv 2608.00023 v1 pith:MGK6NPTW submitted 2026-07-09 cs.CL cs.AI

Role Steering of Language Models for Social Simulations

classification cs.CL cs.AI
keywords activation steeringrole-conditioned agentssocial simulationrole-profile alignmentcontrollabilityassistant axisLLM evaluationscreening workflow
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tries to establish that activation steering — adding a role-specific direction to a language model's internal activations — can be turned into a repeatable, pre-deployment screening workflow for building role-conditioned agents. It claims that across 275 roles, role-specific directions receive higher judged role-profile alignment than a generic assistant-axis directional control (mean 63.2 vs 41.1), and that most roles improve as the steering coefficient increases. Crucially, it shows that 38 roles decline across all six measured dimensions with stronger steering, so a one-size-fits-all high-strength setting is unsafe. The practical point is that simulation builders can measure, compare, and flag candidate agent configurations before use, rather than trusting a prompt or a single steering vector blindly.

Core claim

The central claim is that role-specific activation directions, extracted via judge-filtered contrastive mean-difference at layer 16 of OLMo-3-7B-Instruct, produce higher judged role-profile alignment than a non-scale-matched assistant-axis directional control across a four-point coefficient grid (α∈{1.0,1.5,2.0,2.5}). On a 228-question role-agnostic battery, mean overall alignment is 63.2 for role vectors versus 41.1 for the assistant-axis control; the control's unique-bigram ratio also collapses at large coefficients (0.94→0.27) while role vectors stay high (0.95→0.92). The paper's key practical discovery is heterogeneous response: 74% of roles improve monotonically with α (median per-role

What carries the argument

The central object is the role-specific activation direction vr, computed as the layer-16 mean-of-differences over judge-filtered pairs of 'fully role-playing' positive responses and 'Assistant-style' negative responses — the optimal pointwise-MSE estimator under the formulation of Im and Li (2025). Steering applies additive intervention h→h+αvr at the same layer, with a swept scalar coefficient α. The workflow around this object — define role profile, extract direction, sweep four α values, evaluate role-profile alignment with LLM judges, and pass or flag each configuration — is what carries the argument, turning a mechanism into a testable, per-role calibration procedure.

Load-bearing premise

The load-bearing premise is that GPT-4.1-mini's judged agreement with a GPT-4.1-mini-generated prompted reference is a valid measure of role fidelity — if the judge rewards its own stereotype rather than role-appropriate behavior, the headline score gap and the anti-controllable classification reflect the judge, not the steering.

What would settle it

Run the same 275-role extraction and coefficient sweep but replace GPT-4.1-mini as judge with a held-out, human-validated rubric for role fidelity (or a different model family such as a stronger general evaluator). If the mean role-vector overall score stops exceeding the assistant-axis control by more than a few points, or if many of the 38 anti-controllable roles fail to decline, the paper's central claim is falsified rather than the screen being merely internally consistent.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If role-specific steering reliably raises judged role-profile alignment, simulation builders gain a measurable, tunable knob for role conditioning that does not depend on prompt wording alone.
  • The per-role screen means agents can be individually vetted: a candidate configuration passes only if it expresses the profile without increasing lexical repetition or deteriorating at stronger coefficients.
  • The 38 anti-controllable roles imply that deploying a single global steering strength across a population would silently degrade a meaningful minority of simulated agents.
  • Because role profiles are explicit modeling assumptions, the workflow allows builders to document provenance and disclose which roles were flagged, before agents enter a simulation.
  • The comparison to the assistant-axis control shows that a generic assistant-like direction is not interchangeable with role-specific directions under this setup, though the comparison is not norm-matched.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A norm-matched assistant-axis experiment could separate direction from perturbation magnitude; if a norm-matched control still declines, it would substantiate that direction, not scale, drives the 63.2-vs-41.1 gap.
  • The anti-controllable category is likely populated by roles that saturate early; lowering the starting coefficient or using a sub-linear ramp could recover usable configurations that the current grid misses.
  • Extending the same screening procedure to multi-turn interactions and agent-agent settings would test whether the role-fidelity gains persist beyond single-turn responses, which the paper explicitly leaves for future work.
  • Because the judge and the prompted reference share model priors, the pipeline may systematically reward stereotype-consistent output; an independent human-judged validation set could recalibrate the screen's thresholds.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes an activation-steering screening workflow for role-conditioned LLM agents used in social simulations. For each of 275 roles, it constructs a structured profile, extracts a layer-16 contrastive direction from judge-filtered responses, sweeps four steering coefficients α ∈ {1.0, 1.5, 2.0, 2.5}, and scores the resulting generations with GPT-4.1-mini judges on six alignment dimensions. The two headline claims are (i) that role-specific directions outperform an assistant-axis directional control in mean judged role-profile alignment (63.2 vs. 41.1) while preserving lexical diversity, and (ii) that the per-role screen identifies a meaningful minority of 38 'anti-controllable' roles that decline on all six axes as α increases. The paper frames the contribution as a pre-deployment calibration method rather than as a new steering estimator, and it repeatedly and honestly scopes its claims to the specific model, layer, coefficient grid, and judge pipeline used.

Significance. If the central finding is valid, the workflow is practically valuable: simulation builders could screen role agents before deployment, choose per-role coefficients, and flag roles where stronger steering is counterproductive. The scale is substantial — 275 roles, 228 role-agnostic questions, four coefficients, ~500,000 judged responses — and the paper ships evaluation artifacts, detailed appendices, and explicit validity boundaries. The reproducibility statement and the repeated caveats about the non-scale-matched control and the model-generated reference are commendable. However, the core empirical quantity, 'role-profile alignment', is produced entirely by one model in three connected roles: GPT-4.1-mini filters the examples used to extract the role vectors, generates the prompted reference against which outputs are scored, and serves as the judge for all six alignment dimensions. Consequently, the headline gap and the anti-controllable classification could reflect the judge's own priors rather than properties of the steering directions. The norm mismatch between role vectors and the assistant-axis control further confounds the direction-vs-magnitude comparison. These issues are

major comments (3)
  1. [§3, §4, §A.3, §B.1] The dependent measure is self-referential. GPT-4.1-mini generates the prompted role reference (§A.3), filters the contrastive pairs used to extract each role vector (only responses it labels 'fully role-playing' are retained, §3), and then scores all six alignment dimensions (Judge 1 and Judge 2, §B.1). Role vectors are therefore selected to satisfy GPT-4.1-mini's notion of role-playing and evaluated against a target built by the same model. The reported 63.2-vs-41.1 gap over the assistant-axis control and the identification of 38 'anti-controllable' roles could be properties of the judge's stereotype prior rather than of the steering directions. The paper names this boundary in §6 and the Reproducibility Statement, but it remains untested. I would need at least one independent evaluation anchor — e.g., a human-rated subset, a judge model not used in extraction or reference construction,
  2. [§3, §4, Appendix C] The role-vs-control comparison is not norm-matched. The assistant-axis vectors have mean ℓ2 norm 9.68 versus 3.79 for role vectors (a factor of 2.56×), so equal coefficients apply substantially larger perturbation magnitudes in the control condition. The paper states this caveat repeatedly and appropriately narrows the conclusion to 'under the current extraction and evaluation setup', but the abstract's headline comparison is still presented without it. The control's sharp decline at α=2.5 and its negative median per-role correlation (r=−0.89) may be driven by the larger perturbation magnitude rather than by the direction being non-role-specific. A norm-matched assistant-axis baseline, or an equivalent analysis with perturbation magnitude held constant, is required to support the directional interpretation that the paper's framing implies.
  3. [§4.2, Appendix C, Eq. (1)] The 'anti-controllable over the tested range' classification relies on Pearson r computed from only four coefficient values per role per axis, without per-role uncertainty estimates or significance thresholds. With n=4, an r<0 classification is fragile, and the observed mean score drop is modest (7.0 points from α=1.0 to 2.5, Figure 12). Since this classification is the main practical output for simulation builders, the paper should report bootstrap or permutation-based confidence intervals for the per-role correlations, show the stability of the 38-role set under resampling of questions or judge runs, and justify that the 'all six axes r<0' criterion is not dominated by measurement noise. As written, the flag list may over-identify roles as anti-controllable.
minor comments (5)
  1. [Abstract] The norm-matching caveat appears in §3 and §4 but not in the abstract. Since the abstract's 63.2-vs-41.1 comparison is the headline result, a one-sentence caveat should be added there.
  2. [§3 / Reproducibility Statement] The extraction details are incomplete: judge-filtering thresholds, retained-pair counts, parse-failure counts, exact model revisions, decoding parameters, and activation pooling details are deferred to released artifacts. These should be summarized in the paper itself for reproducibility.
  3. [§4.1 / Figure 4] The text reports median r=−0.89 for the assistant-axis control, but the figure caption and the discussion of the two distributions could more clearly separate the role-vector and control histograms. Consider showing both distributions on the same axes with matched binning.
  4. [Appendix A.3, Table 5] The example doctor few-shot turns include catchphrases that appear semantically disconnected from the stated role ('Trust me, I'm a doctor'; 'the needs of the many outweigh the needs of the few'). If these are intentional caricatures, they should be flagged as such; otherwise they undermine the example's pedagogical value.
  5. [§E.6 / Appendix J] The effective-rank analysis is interesting but the connection to the main screen could be stated more directly: the low rank of the behavioral readout means the judge-based screen may not resolve all representational variation. This is acknowledged, but a one-sentence practical implication in the main text would help.

Circularity Check

2 steps flagged

Judge-filtered extraction and same-judge scoring make the headline alignment gap partly self-fulfilling

specific steps
  1. fitted input called prediction [§3 Judge-filtered vector extraction; §4 Evaluation Across 275 Roles]
    "We retain only fully role-playing positives and Assistant-style negatives, then form the role vector vr as the layer-16 mean-of-differences over the filtered pairs ... GPT-4.1-mini generates the prompted references and serves as all judges, so shared model priors may influence both the target and the score."

    The role vectors are constructed from response pairs that GPT-4.1-mini labels 'fully role-playing' versus 'Assistant-style,' and the reported role-profile alignment is then scored by GPT-4.1-mini against a GPT-4.1-mini-generated reference. The extraction is therefore fitted to that judge's notion of role-playing, and the evaluation measures the same judge's agreement. The headline 63.2-vs-41.1 gap is partly forced because the role-specific directions were selected to satisfy the judge while the assistant-axis control was not; the comparison is not an independent test of role fidelity.

  2. self definitional [§1 Scope of the measured outcome; §A.3 Prompted Role Reference]
    "role-profile alignment denotes judged agreement with the constructed role description and prompted role reference used by our evaluation pipeline ... It is assembled from three components, each generated by GPT-4.1-mini."

    The outcome construct is defined as agreement with a target that is itself produced by the same model family that assigns the alignment scores. Thus the target and the score share model priors by construction, as the paper concedes: 'shared model priors may influence both the target and the score.' This does not make the role-versus-assistant comparison meaningless, since both conditions go through the same judge, but it means absolute alignment values and the anti-controllable classification are properties of the judge/reference loop rather than of an external role-fidelity standard.

full rationale

The paper is methodologically self-contained in most respects: it does not rely on load-bearing self-citations, does not import a uniqueness theorem, and the assistant-axis control comes from non-overlapping prior work (Lu et al. 2026). The norm mismatch between role vectors and the assistant axis is a real confound but is not circularity. The central issue is the self-referential judge loop. GPT-4.1-mini performs three connected functions: it filters the responses used to extract each role vector, it generates the prompted role reference used as the evaluation target, and it scores all six alignment dimensions. The role vector is therefore fitted to GPT-4.1-mini's classification of 'fully role-playing,' and the dependent measure is the same model's agreement with its own generated target. The headline claim that role-specific directions outperform the assistant-axis control (63.2 vs 41.1) is consequently not an independent measurement of role fidelity; part of the gap is attributable to the fact that the role vectors were selected using the judge's priors while the control was not. The paper is admirably explicit about this boundary in §3 and §6, but disclosure does not break the circular dependency. The practical per-role screen may still be useful as a within-pipeline calibration tool, but the load-bearing validity claim—that the judge measures role fidelity—is untested. This is partial circularity of the 'fitted input called prediction' kind, warranting a score of 6 rather than a higher one because the role-versus-assistant comparison does apply the same judge to both conditions and the extracted vectors are not directly optimized to maximize the final alignment score.

Axiom & Free-Parameter Ledger

5 free parameters · 6 axioms · 0 invented entities

No fundamentally new physical or mathematical entities are introduced. The central claim rests on a chain of operational assumptions: activation steering at layer 16; LLM-generated role profiles; GPT-4.1-mini both generating the evaluation target and scoring alignment; and a non-scale-matched control used as the comparison.

free parameters (5)
  • Layer index 16 = 16 (fixed)
    The central claim depends on steering at layer 16, chosen by convention from Lu et al. (2026); different layers could change the steering outcome. Not fitted to these data.
  • Steering coefficient grid α = 1.0, 1.5, 2.0, 2.5
    All coefficient-response and anti-controllable claims are limited to this grid; α=1.0 is the lowest tested, so saturation before α=1.0 cannot be seen.
  • Judge filter threshold for 'fully role-playing' positives = unspecified (score threshold)
    The extracted role vector depends on which responses the GPT-4.1-mini judge labels fully role-playing; the exact cutoff is not reported.
  • Pairwise win margin = 10 points
    Ties are declared below a 10-point debiased advantage; changing the margin changes the reported 92.9% win rate.
  • Parse-failure default score = 0 (Judges 1-2), 50 (Judge 3)
    Responses that fail JSON/plain-integer parsing are scored 0, which biases aggregate scores downward; parse-failure counts are not reported.
axioms (6)
  • domain assumption Additive steering h → h + αv_r at layer 16 changes behavior along the intended role dimension.
    Invoked in §3; based on prior persona-vector/steering work (Lu et al. 2026; Chen et al. 2025; Turner et al. 2024) rather than proven here.
  • domain assumption GPT-4.1-mini judge scores (0-100) are a valid monotone measure of role-profile fidelity.
    The entire screen uses these scores; the paper states the evaluation is 'an operational screen, not a population-validity study' (§6).
  • domain assumption The prompted role reference is an adequate evaluation target for role expression.
    Generated by GPT-4.1-mini from description, catchphrases, and few-shot turns (§A.3); the paper notes it is not empirical ground truth.
  • standard math Mean-difference contrastive extraction is the optimal pointwise-MSE estimator under Im & Li (2025).
    Taken as an established result; used in §3 to justify the role-vector formula.
  • domain assumption O*NET-seeded kimi-generated role profiles operationalize the 275 role labels.
    Appendix A.1; the paper labels these 'explicit modeling assumptions rather than empirical descriptions.'
  • domain assumption Layer-16 residual directions transfer from the 50 elicitation questions to the 228-question battery.
    The vector extracted on role-specific prompts is evaluated on role-agnostic questions; this transfer is assumed.

pith-pipeline@v1.3.0-alltime-deepseek · 26621 in / 15404 out tokens · 175814 ms · 2026-08-04T01:40:50.645220+00:00 · methodology

0 comments
read the original abstract

Social simulations built from language-model agents need role-conditioned behavior that can be checked before agents are placed into a simulated population. We introduce an activation-steering screening workflow for role-conditioned agents: define a role profile, extract a role-specific direction, sweep four steering coefficients, evaluate role-profile alignment, and pass or flag each candidate configuration. On OLMo-3-7B-Instruct, we apply the workflow to a mixed 275-role inventory with 228 role-agnostic questions, GPT-4.1-mini prompted role references, and GPT-4.1-mini judges. Role-specific directions receive higher judged role-profile alignment than an assistant-axis directional control from prior persona-vector work, with mean overall scores of 63.2 versus 41.1 across the tested grid. They also preserve high lexical diversity, while the control drops sharply at larger coefficients. The role-level screen is the main practical output: most roles improve as steering increases, but 38 roles decline across all six measured dimensions, showing why simulation builders should choose coefficients per role rather than deploy a uniform high-strength setting. We make our code and evaluation artifacts available at https://anonymous.4open.science/r/anonymous-research-code-5F03/.

Figures

Figures reproduced from arXiv: 2608.00023 by Akhil Theerthala, Arjun Chatterjee, Emile Anand, Glenn Matlin, Isaac Song, Maria Kostylew, Mark Riedl, Mohammed Rehan Parwani, Sebastien Krier, Yonadav G. Shavit.

Figure 1
Figure 1. Figure 1: Pre-deployment calibration workflow for role-conditioned agents. A structured role profile defines a candidate behavior target; judge-filtered contrastive activations pro￾duce a candidate layer-16 direction; the direction is evaluated at four steering coefficients; and a behavioral screen either retains a documented configuration or flags the role for lower-strength use, profile revision, prompt conditioni… view at source ↗
Figure 2
Figure 2. Figure 2: Aggregate evaluation across n=275 roles. Left: mean judged role-profile alignment for role-specific vectors and the assistant-axis directional control at α ∈ {1.0, 1.5, 2.0, 2.5}; the dashed line marks the prompted role-reference level. Right: unique-bigram ratio, used as a lexical repetition proxy. Role-specific generations maintain high lexical diversity through α=2.5 (0.95 → 0.92), whereas the assistant… view at source ↗
Figure 3
Figure 3. Figure 3: Pairwise preference check for the 39-role subset at [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Per-role Pearson r between steering coefficient α and overall role-profile alignment score across four tested coefficients. Role-specific directions have median r=+0.98, and 74% of roles improve at every consecutive step; the assistant-axis directional control from Lu et al. (2026) has median r=−0.89. Because each correlation is based on four values and the two vector families are not norm matched, the plo… view at source ↗
Figure 5
Figure 5. Figure 5: Response of roles classified as anti-controllable over the tested range. [PITH_FULL_IMAGE:figures/full_fig_p007_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Debiased score distributions for role-specific responses (mean 78.7) and the [PITH_FULL_IMAGE:figures/full_fig_p017_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Per-role pairwise win rate. Steered responses win the majority of comparisons in [PITH_FULL_IMAGE:figures/full_fig_p018_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Per-role Pearson r (Eq. 1) across all six behavioral axes for all 275 roles, sorted by overall-score r. Each cell is the correlation between α and mean judge score for one role on one axis. Anti-controllable roles (negative r on all axes) cluster visibly at the bottom of the figure. Overall Emotional Register Vocabulary Choice Social Dynamic Motivation Worldview Align. 0.0 0.2 0.4 0.6 0.8 1.0 Fraction of R… view at source ↗
Figure 9
Figure 9. Figure 9: Fraction of roles exceeding Pearson r thresholds of 0.6, 0.8, and 0.95 per behavioral axis, and the fraction with strictly positive slope (i.e. r > 0) and strict monotonicity. Vocab choice falls below every other axis on all five metrics, confirming it as the weakest dimension of steerability. 22 [PITH_FULL_IMAGE:figures/full_fig_p022_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Mean score vs. α for all six behavioral axes across the controllable majority (the 237 roles outside the anti-controllable set). Shaded bands are 95% intervals of the group mean (1.96×SEM). All axes rise monotonically; vocab choice shows the shallowest slope, consistent with its lower median Pearson r and higher role-to-role variability. D Anti-Controllability Takeaway. Roles classified as anti-controllab… view at source ↗
Figure 11
Figure 11. Figure 11: Individual score-vs-α trajectories for the 38 anti-controllable roles (negative r on all six axes). Each thin line is one role; the black line is the group mean. Most trajectories start high at α = 1.0 and the characteristic pattern — a narrow low-α window of marginal gain followed by decline — varies substantially in rate across roles (OLS slopes range from near zero to −15.2 score units per α-unit). 23 … view at source ↗
Figure 12
Figure 12. Figure 12: Distribution of absolute score drop, defined as the score at [PITH_FULL_IMAGE:figures/full_fig_p024_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Per-role, per-axis score drop for the 38 anti-controllable roles (rows sorted by [PITH_FULL_IMAGE:figures/full_fig_p024_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Prompted-reference score vs. steered score at [PITH_FULL_IMAGE:figures/full_fig_p025_14.png] view at source ↗
Figure 15
Figure 15. Figure 15: Exploratory geometry summaries separated by statistic family. The previous [PITH_FULL_IMAGE:figures/full_fig_p027_15.png] view at source ↗
Figure 16
Figure 16. Figure 16: PC × metric Pearson r across the four steering coefficients (one panel per α; rows: overall score and the five sub-dimensions; columns: PC1–PC20 of the centered role cloud). PC1 has the strongest single-PC associations at α=1.0–1.5; PC5 has the strongest single-PC associations at α=2.0–2.5. These are exploratory summaries, not certified regimes. joint-OLS table, and a trait-PC pathway consistency check r=… view at source ↗
Figure 17
Figure 17. Figure 17: Per-role distance-to-assistant vs. per-role steering behavior. [PITH_FULL_IMAGE:figures/full_fig_p030_17.png] view at source ↗
Figure 18
Figure 18. Figure 18: Vector magnitude vs. α-response. Roles are binned into four norm quartiles; mean steered_score against α is averaged within each quartile (shading: 95% interval of the group mean). Low-norm vectors (Q1) start high and often peak early; high-norm vectors (Q4) start lowest and rise through the tested grid, with the quartiles roughly converging by α=2.5. The pattern is correlational. Full four-cell distance … view at source ↗
Figure 19
Figure 19. Figure 19: Pairwise representational vs. behavioral geometry. [PITH_FULL_IMAGE:figures/full_fig_p031_19.png] view at source ↗
Figure 20
Figure 20. Figure 20: Trait-axis projection × metric Pearson r at α=1.0. Emotional axes (empathetic– stoic, nurturing–hostile, serene–evil) load most strongly on emotional register (r up to +0.30); methodical–chaotic loads most on vocab choice (+0.25); diplomatic–dramatic on social dynamic (+0.17). Axis names are interpretive labels rather than psychological scales. Effective ranks are mismatched. The centered role-vector clou… view at source ↗
Figure 21
Figure 21. Figure 21: Why the geometry–behavior effect sizes are moderate. [PITH_FULL_IMAGE:figures/full_fig_p035_21.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

34 extracted references · 5 canonical work pages · 4 internal anchors

  1. [1]

    doi:10.48550/ARXIV.2602.04863 , urldate =

    Subliminal. doi:10.48550/ARXIV.2602.04863 , urldate =. arXiv , copyright =:2602.04863 , primaryclass =

  2. [2]

    Continuous Latent Contexts Enable Efficient Online Learning in Transformers

    Anand, Emile and Ateyeh, Abdullah and Cao, Xinyuan and Dabagia, Max , year = 2026, month = may, eprint =. Continuous. doi:10.48550/arXiv.2605.09867 , abstract =

  3. [3]

    2026 , eprint=

    Structure from Strategic Interaction & Uncertainty: Risk Sensitive Games for Robust Preference Learning , author=. 2026 , eprint=

  4. [4]

    2026 , eprint=

    Learning Approximate Nash Equilibria in Cooperative Multi-Agent Reinforcement Learning via Mean-Field Subsampling , author=. 2026 , eprint=

  5. [5]

    2026 , eprint=

    Training Generalizable Collaborative Agents via Strategic Risk Aversion , author=. 2026 , eprint=

  6. [6]

    Anand, Emile and Karmarkar, Ishani and Qu, Guannan , year = 2025, month = oct, number =. Mean-. doi:10.48550/arXiv.2412.00661 , urldate =. arXiv , keywords =:2412.00661 , primaryclass =

  7. [7]

    Bas, Tetiana and Novak, Krystian , year = 2026, month = jan, number =. What. doi:10.48550/arXiv.2511.18284 , urldate =. arXiv , keywords =:2511.18284 , primaryclass =

  8. [8]

    Peer-to-

    Chaudhari, Shreyas and Pranav, Srinivasa and Anand, Emile and Moura, Jos. Peer-to-. doi:10.1109/ICASSP49660.2025.10890126 , urldate =. arXiv , keywords =:2409.15267 , primaryclass =

  9. [9]

    Chen, Runjin and Arditi, Andy and Sleight, Henry and Evans, Owain and Lindsey, Jack , year = 2025, month = sep, number =. Persona. doi:10.48550/arXiv.2507.21509 , urldate =. arXiv , keywords =:2507.21509 , primaryclass =

  10. [10]

    Im, Shawn and Li, Sharon , year = 2025, month = feb, eprint =. A. doi:10.48550/arXiv.2502.02716 , abstract =

  11. [11]

    Prometheus:

    Kim, Seungone and Shin, Jamin and Cho, Yejin and Jang, Joel and Longpre, Shayne and Lee, Hwaran and Yun, Sangdoo and Shin, Seongjin and Kim, Sungdong and Thorne, James and Seo, Minjoon , year = 2024, eprint =. Prometheus:. The. doi:10.48550/arXiv.2310.08491 , abstract =

  12. [12]

    Measuring and

    Li, Kenneth and Liu, Tianle and Bashkansky, Naomi and Bau, David and Vi. Measuring and. doi:10.48550/arXiv.2402.10962 , urldate =. arXiv , keywords =:2402.10962 , primaryclass =

  13. [13]

    Online Adaptive Policy Selection in Time-Varying Systems: No-Regret via Contractive Perturbations

    Lin, Yiheng and Preiss, James A. and Anand, Emile and Li, Yingying and Yue, Yisong and Wierman, Adam , year = 2023, month = jun, number =. Online. doi:10.48550/arXiv.2210.12320 , urldate =. arXiv , keywords =:2210.12320 , primaryclass =

  14. [14]

    Online Policy Optimization in Unknown Nonlinear Systems

    Lin, Yiheng and Preiss, James A. and Xie, Fengze and Anand, Emile and Chung, Soon-Jo and Yue, Yisong and Wierman, Adam , year = 2024, month = apr, number =. Online. doi:10.48550/arXiv.2404.13009 , urldate =. arXiv , keywords =:2404.13009 , primaryclass =

  15. [15]

    Lu, Christina and Gallagher, Jack and Michala, Jonathan and Fish, Kyle and Lindsey, Jack , year = 2026, month = jan, number =. The. doi:10.48550/arXiv.2601.10387 , urldate =. arXiv , keywords =:2601.10387 , primaryclass =

  16. [16]

    Lutz, Marlene and Sen, Indira and Ahnert, Georg and Rogers, Elisa and Strohmaier, Markus , year = 2025, month = oct, number =. The. doi:10.48550/arXiv.2507.16076 , urldate =. arXiv , keywords =:2507.16076 , primaryclass =

  17. [17]

    Maiya, Sharan and Bartsch, Henning and Lambert, Nathan and Hubinger, Evan , year = 2025, month = nov, eprint =. Open. doi:10.48550/arXiv.2511.01689 , abstract =

  18. [18]

    Marks, Sean and Lindsey, Jack and Olah, Christopher , year = 2026, month = feb, urldate =. The

  19. [19]

    Efficient

    Mikolov, Tomas and Chen, Kai and Corrado, Greg and Dean, Jeffrey , year = 2013, month = sep, number =. Efficient. doi:10.48550/arXiv.1301.3781 , urldate =. arXiv , keywords =:1301.3781 , primaryclass =

  20. [20]

    Steering

    Panickssery, Nina and Gabrieli, Nick and Schulz, Julian and Tong, Meg and Hubinger, Evan and Turner, Alexander Matt , year = 2024, month = jul, number =. Steering. doi:10.48550/arXiv.2312.06681 , urldate =. arXiv , keywords =:2312.06681 , primaryclass =

  21. [21]

    Park, Kiho and Choe, Yo Joong and Veitch, Victor , year = 2024, month = jul, number =. The. doi:10.48550/arXiv.2311.03658 , urldate =. arXiv , keywords =:2311.03658 , primaryclass =

  22. [22]

    Potert. Can. Findings of the. doi:10.18653/v1/2025.findings-emnlp.963 , urldate =

  23. [23]

    Scalable and

    Shah, Rusheb and. Scalable and. doi:10.48550/arXiv.2311.03348 , urldate =. arXiv , keywords =:2311.03348 , primaryclass =

  24. [24]

    Shanahan, Murray and McDonell, Kyle and Reynolds, Laria , year = 2023, month = may, number =. Role-. doi:10.48550/arXiv.2305.16367 , urldate =. arXiv , keywords =:2305.16367 , primaryclass =

  25. [25]

    Taimeskhanov, Magamed and Vaiter, Samuel and Garreau, Damien , year = 2026, month = feb, eprint =. Towards. doi:10.48550/arXiv.2602.02712 , abstract =

  26. [26]

    Analysing the Generalisation and Reliability of Steering Vectors , booktitle =

    Tan, Daniel and Chanin, David and Lynch, Aengus and Paige, Brooks and Kanoulas, Dimitrios and. Analysing the Generalisation and Reliability of Steering Vectors , booktitle =. doi:10.52202/079017-4417 , urldate =

  27. [27]

    doi:10.48550/arXiv.2512.13961 , urldate =

    Olmo 3 , author =. doi:10.48550/arXiv.2512.13961 , urldate =. arXiv , keywords =:2512.13961 , primaryclass =

  28. [28]

    Persistent

    Tosato, Tommaso and Helbling, Saskia and. Persistent. doi:10.48550/arXiv.2508.04826 , urldate =. arXiv , langid =:2508.04826 , primaryclass =

  29. [29]

    and Mini, Ulisse and MacDiarmid, Monte , year = 2024, month = oct, number =

    Turner, Alexander Matt and Thiergart, Lisa and Leech, Gavin and Udell, David and Vazquez, Juan J. and Mini, Ulisse and MacDiarmid, Monte , year = 2024, month = oct, number =. Steering. doi:10.48550/arXiv.2308.10248 , urldate =. arXiv , keywords =:2308.10248 , primaryclass =

  30. [30]

    Wang, Miles and la Tour, Tom Dupr. Persona. doi:10.48550/arXiv.2506.19823 , urldate =. arXiv , keywords =:2506.19823 , primaryclass =

  31. [31]

    Zhang, Jifan and Sleight, Henry and Peng, Andi and Schulman, John and Durmus, Esin , year = 2025, month = oct, eprint =. Stress-. doi:10.48550/arXiv.2510.07686 , abstract =

  32. [32]

    and Zhang, Hao and Gonzalez, Joseph E

    Zheng, Lianmin and Chiang, Wei-Lin and Sheng, Ying and Zhuang, Siyuan and Wu, Zhanghao and Zhuang, Yonghao and Lin, Zi and Li, Zhuohan and Li, Dacheng and Xing, Eric P. and Zhang, Hao and Gonzalez, Joseph E. and Stoica, Ion , year = 2023, eprint =. Judging. Advances in. doi:10.48550/arXiv.2306.05685 , abstract =

  33. [33]

    doi:10.48550/arXiv.2508.10014 , urldate =

    Zhou, Lingfeng and Zhang, Jialing and Gao, Jin and Jiang, Mohan and Wang, Dequan , year = 2025, month = aug, number =. doi:10.48550/arXiv.2508.10014 , urldate =. arXiv , keywords =:2508.10014 , primaryclass =

  34. [34]

    and Wang, Zifan and Mallen, Alex and Basart, Steven and Koyejo, Sanmi and Song, Dawn and Fredrikson, Matt and Kolter, J

    Zou, Andy and Phan, Long and Chen, Sarah and Campbell, James and Guo, Phillip and Ren, Richard and Pan, Alexander and Yin, Xuwang and Mazeika, Mantas and Dombrowski, Ann-Kathrin and Goel, Shashwat and Li, Nathaniel and Byun, Michael J. and Wang, Zifan and Mallen, Alex and Basart, Steven and Koyejo, Sanmi and Song, Dawn and Fredrikson, Matt and Kolter, J. ...