{"id":"7796f025-be40-4422-876f-44ba03b4f26b","arxiv_id":"2607.07916","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Composable LoRA adapters can amplify or suppress OCEAN traits in LLMs, combine approximately additively, preserve moderate-scale capability, and move safety-relevant behaviours.","lead":"Researchers train small weight adapters that push language models along Big-Five personality axes, and show the adapters can be scaled and stacked to build mixed personas without wrecking basic skills. The same knobs also shift safety-relevant habits such as frustration and sycophancy, offering a practical handle on model character.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The strongest claim rests on measurement validity that is partly circular: TRAIT and judges share the same OCEAN definitions used to train the LoRAs, so monotonicity and additivity may track definition-aligned style rather than independent behavioural axes.","rationale":"The reader correctly identifies the weakest assumption as construct validity of human OCEAN operationalised via shared constitutions, TRAIT, and judges, plus non-inert controls and limited full-stack evaluation. That is the single most load-bearing concern for the strongest claim: without independent behavioural measurement, monotonicity, approximate additivity, and safety correlations can be re-descriptions of the training signal rather than evidence of generalisable trait axes in weight space. The paper’s multi-model TRAIT/MMLU sweeps, residual analysis, rank-1 compression, and teacher robustness are real strengths and support a conditional accept-shaped contribution; they do not fully close the validity gap. No stronger internal inconsistency or experimental error is required to keep the verdict CONDITIONAL. The concrete test (human re-rating without OCEAN vocabulary) would settle whether the concern lands without demanding a full re-train. Agreement with the reader is full on the load-bearing point; verdict remains CONDITIONAL with no upgrade or downgrade.","tokens_in":65130,"tokens_out":697,"duration_ms":9696,"concrete_test":"Hold out a fixed set of free-form multi-turn rollouts (e.g. the 240 Lu et al. seeds plus the frustration/sycophancy protocols) and re-score them with human raters using only behavioural rubrics that never mention OCEAN facet names or the training constitutions. Correlate human scores with TRAIT logprob and the paper’s LLM-judge scores for each adapter×scale. If human–TRAIT/judge Spearman ρ falls below ~0.5 on target traits while TRAIT still moves monotonically, the measurement loop is doing substantial work and the claim should be narrowed to ‘definition-aligned style control with safety side-effects’.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that OCEAN LoRAs move target traits monotonically, compose approximately additively, preserve capabilities at moderate scale, and shift held-out safety behaviours. The load-bearing hinge is that TRAIT logprob scores and the LLM-judge panel measure the same constructs the constitutions train for, rather than correlated surface style or teacher artefacts. Constitutions, judge rubrics, and TRAIT all instantiate the same OCEAN facet catalogue (Appendix B; C.2.1: judges are built from the same definition object as training). Neutral control adapters are not inert: they nearly double sycophancy (0.61 vs 0.33 baseline), raise WildJailbreak harmful compliance, and modestly dampen frustration (§3). Full judge+safety stack is mainly on Llama-3.1-8B-Instruct; other models get only TRAIT+MMLU (§5.1). Downstream safety tasks are held-out from training but still interpreted through the same trait labels. If TRAIT/judges largely re-read the trained style, monotonicity, residual additivity (Fig. 4c, Appendix F), and safety correlations become weaker evidence for genuine weight-space trait axes. The paper flags this construct-validity issue; it remains the softest support for the strongest claim.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes treating LLM personas as positions in a behavioural trait space, operationalised via OCEAN (Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism). It trains rank-64 LoRA amplifiers and suppressors for each trait using constitution-guided DPO plus SFT distillation (Open Character Training), evaluates them with TRAIT logprob MCQs, a human-calibrated LLM-judge panel, and capability benchmarks (MMLU, GSM8K, TruthfulQA), and shows that adapters move target traits largely monotonically with scale, compose approximately additively, and preserve capabilities at moderate scales across six models (4B–32B, three families). Held-out safety evaluations link neuroticism to multi-turn frustration and agreeableness to sycophancy/compliance; an unsupervised factor analysis on rollouts recovers four TIDE factors (Tone, Initiative, Didacticism, Epistemic Caution), with partial LoRA control of Initiative.","tokens_in":65459,"tokens_out":1371,"duration_ms":20340,"significance":"If the results hold, the work supplies a practical, weight-space alternative to brittle prompting and layer-sensitive activation steering for persona control: lightweight, scalable, and composable trait adapters that transfer across model families and teachers, retain rank-1 compressibility, and move safety-relevant behaviours without task-specific training. Strengths include multi-pronged measurement (TRAIT, human-calibrated judges with reported Spearman ρ/MAE and Krippendorff α, Wilson/bootstrap CIs), neutral control adapters, two distillation teachers, six baselines, interaction residuals (Eq. form in Appendix F / Fig. 4c), amplifier×suppressor heatmaps, and explicit safety trade-off experiments (WildJailbreak, CoCoNot, sycophancy). The unsupervised TIDE pipeline is a useful step beyond human psychometrics. The paper is a solid bridge between personality measurement, model editing, and alignment practice.","major_comments":[{"comment":"§2.1 / Appendix B and C.2.1: Constitutions, judge rubrics, and system-prompt baselines are all generated from the same OCEAN-definition object. TRAIT is also an OCEAN instrument. Monotonicity and near-additivity (Figs. 2–4) therefore partly re-measure definition-aligned style. Held-out safety tasks and human judge calibration partially break this loop, but the main text should quantify how much TRAIT/judge movement remains after residualising out surface correlates (length, sentiment, first-person affect) and should state more sharply which claims rest on external vs. definition-internal evidence.","section":"§2.1 Methods; Appendix B, C.2.1"},{"comment":"§3 and Figs. 39–40: Neutral control adapters nearly double sycophancy (0.61 vs 0.33 baseline), raise WildJailbreak harmful compliance, and modestly dampen frustration. Distillation artefacts are therefore not negligible relative to trait effects. Safety claims that attribute shifts to OCEAN axes (neuroticism↔frustration, agreeableness↔sycophancy) need systematic control-subtracted effect sizes and, where possible, matched-verbosity or matched-preference-strength baselines so trait signal is separated from pipeline shift.","section":"§3 Downstream Applications"},{"comment":"§5.1 Limitations: The full stack (judges, multi-turn safety, composition residuals) is reported primarily on Llama-3.1-8B-Instruct; other models receive TRAIT+MMLU only. The abstract’s “across six models” claim for composability and safety transfer is stronger than the evaluation breadth supports. Either extend at least one safety task and one composition residual analysis to a second family (e.g. Gemma-3-27B-IT, already used for frustration) or narrow the abstract/intro wording to match what was fully measured.","section":"§5.1; Abstract; §2.3"}],"minor_comments":[{"comment":"Negative scaling does not reliably invert traits for all adapters (Appendix E); this should be flagged in the main §2.3 invertibility paragraph rather than only in the appendix, since signed-axis language appears in the abstract.","section":"§2.3 Scaling and Combining LoRAs"},{"comment":"§4 / Appendix M.5: Initiative LoRAs are validated on the same forced-choice questionnaire used to discover TIDE factors. Note this circularity more prominently when claiming “partial success” at modulating unsupervised traits.","section":"§4 Unsupervised Persona Exploration"},{"comment":"Fig. 4c and Appendix F: Conscientiousness is the clear residual outlier; a one-sentence mechanistic hypothesis (or explicit “unknown”) in the main text would help readers interpret non-additivity.","section":"Fig. 4c; Appendix F"},{"comment":"Appendix A.1.3: The factor-space merge (√w A/B) introduces cross terms distinct from the elementwise ΔW composition in §2.3; a short clarifying sentence in the main methods would prevent conflating the two “composition” notions.","section":"Appendix A.1.3; §2.3"},{"comment":"Several figure panels (e.g. stacked MMLU bars at extreme scales) are dense; ensuring colourblind-safe palettes and consistent axis ranges across Appendix E sweeps would improve readability.","section":"Appendix E"},{"comment":"Related work: Sun et al. (2025) personality-vector merging and Vu et al. (2026) PsychAdapter are discussed; a short table contrasting rank, training objective, and composition method would make the novelty claim sharper.","section":"§5.2 Related Work"}],"recommendation":"minor_revision","confidential_remarks":"The manuscript is unusually thorough for a persona-control paper and is close to a strong accept after local revisions. The construct-validity and control-artefact issues are real but already partially acknowledged; requiring a full re-evaluation stack on all six models would be disproportionate. Fit for a methods/empirical AI venue is good. No integrity concerns from the text as provided."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is that this is a careful multi-model engineering paper showing you can train rank-64 (even rank-1) LoRAs that move OCEAN traits mostly monotonically, stack them with small residuals, keep MMLU/GSM8K/TruthfulQA roughly intact at moderate scale, and shift held-out safety behaviours—neuroticism with multi-turn frustration, agreeableness with sycophancy and CoCoNot compliance, conscientiousness/agreeableness trade-offs on WildJailbreak. That package is new as a systematic result even if the ingredients (Open Character Training, DPO+SFT, task arithmetic, TRAIT, activation capping) are not.\n\nWhat they do well: six baselines across three families, two teachers, control adapters, interaction residuals, amplifier×suppressor grids, rank-1 SVD and base–instruct interpolation transfer, human-calibrated judges with reported ρ/MAE, and honest limitations. The safety probes are the part that makes it more than psychometrics. The unsupervised TIDE factors are a sensible first step beyond human OCEAN, even if validation is still circular.\n\nSoft spots, in proportion. Constitutions, TRAIT, and judges share the same facet catalogue, so some of the monotonicity/additivity is definition-aligned style; the paper flags this. Neutral controls are not inert (sycophancy nearly doubles), which undercuts pure “trait isolation” claims and should be front-and-centre in any revision. Full judge+safety stack is mostly Llama-3.1-8B-Instruct; other models get TRAIT+MMLU only. Negative scaling is imperfect; axes are not orthogonal. None of that sinks the central empirical claim within the reported scope—it just caps how hard you should lean on “genuine independent trait axes in weight space.”\n\nWho it is for: people doing character training, activation/weight steering, or safety evals who need a concrete, composable control surface. Math and citation pattern look fine; code/monorepo is pointed at. I would bring it to reading group, cite the scaling/composition/safety results if I work in this area, and send it to peer review. Ask for broader full-stack eval, independent TIDE judges, and a clearer audit of control artefacts—not a rewrite of the core story.","headline":"Solid empirical toolkit paper: OCEAN LoRAs scale and compose across six models with real safety side-effects; measurement partly circular but not hollow, and the work is worth engaging.","tokens_in":66164,"tokens_out":566,"would_cite":true,"duration_ms":10045,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"LLM personas can be steered as composable trait directions in weight space, using OCEAN LoRAs that scale, mix, and shift safety behaviour.","keywords":["LLM personas","OCEAN traits","LoRA adapters","weight-space editing","persona control","sycophancy","frustration","psychometrics"],"falsifier":"If, on held-out free-form multi-turn rollouts scored by independent human raters who never saw the training constitutions, scaling an openness amplifier failed to raise openness while leaving non-target traits and capability scores essentially unchanged, the central claim that the adapters are targeted trait directions would fail.","tokens_in":65999,"feed_emoji":"🧭","tokens_out":651,"duration_ms":7725,"temperature":0.7,"pith_summary":"Language models do not just answer questions; they behave with stable styles of deference, energy, caution, and tone that shape how they generalise and how safe they are. This paper treats those styles as positions in a trait space, starting from the five OCEAN personality dimensions, and trains low-rank adapters that amplify or suppress each trait. Across six models from three families, each adapter moves its target trait largely monotonically with scale, combines approximately additively with other adapters, and leaves capability benchmarks intact at moderate strengths. The same axes also move held-out safety behaviours: neuroticism tracks multi-turn frustration, agreeableness tracks sycophancy and compliance trade-offs. An unsupervised pipeline then recovers four model-native factors—tone, initiative, didacticism, and epistemic caution—from free rollouts, showing that the method is not locked to human psychology. The practical claim is that persona control becomes a matter of learning, scaling, and composing trait directions in weight space rather than one-off prompting or full retraining.","feed_headline":"LLM personas steered as mixable trait dials in weight space","feed_subtitle":"OCEAN LoRAs scale, combine, keep capability, and shift frustration and sycophancy","key_machinery":"Composable OCEAN trait LoRAs: low-rank adapters trained by constitution-guided DPO plus a lighter SFT stage on self-interaction transcripts, then scaled and linearly summed in weight space so that each adapter acts as a continuous, signed direction of persona change.","core_discovery":"Across six models (4B–32B, three families), constitution-trained OCEAN LoRA adapters move their target behavioural traits largely monotonically with scale, compose approximately additively into mixed personas, preserve MMLU and related capabilities at moderate scales, and shift held-out safety-relevant behaviours such as frustration and sycophancy.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["OCEAN LoRAs chart and control LLM personas in weight space","Mixable trait adapters steer LLM personalities without capability loss","Persona traits dialed via weight-space LoRAs across model families","Additive OCEAN adapters compose LLM personas and shift safety behaviours","Weight-space trait dials make LLM personas measurable and mixable"],"cache_read_input_tokens":49280,"weakest_assumption_plain":"That human OCEAN labels, the same constitutions used in training, and TRAIT-style multiple-choice plus judges built from those definitions isolate real generalisable behavioural axes rather than teacher style, verbosity, or other correlated surface artefacts.","fun_headline_variants_meta":{"raw":{"variants":["OCEAN LoRAs chart and control LLM personas in weight space","Mixable trait adapters steer LLM personalities without capability loss","Persona traits dialed via weight-space LoRAs across model families","Additive OCEAN adapters compose LLM personas and shift safety behaviours","Weight-space trait dials make LLM personas measurable and mixable"]},"model":"grok-4.5","effort":"low","cost_usd":0.003842,"raw_usage":{"total_tokens":1229,"prompt_tokens":787,"num_sources_used":0,"completion_tokens":70,"cost_in_usd_ticks":38420000,"prompt_tokens_details":{"text_tokens":787,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":372,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":787,"tokens_out":70,"duration_ms":4443,"temperature":1.0,"reasoning_tokens":372,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-10T15:27:44.342046+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"If, on held-out free-form multi-turn rollouts scored by independent human raters who never saw the training constitutions, scaling an openness amplifier failed to raise openness while leaving non-target traits and capability scores essentially unchanged, the central claim that the adapters are targeted trait directions would fail.","supporting_citations":[],"review_version":1}