{"id":"97f4aeb4-b0a8-40a2-b03b-59200256efc1","arxiv_id":"2608.13482","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Injecting first-person, value-laden reflections into pretraining text improves constitution following, jailbreak resistance, and out-of-distribution moral choices in small language models, with the largest gains when the reflections are present from token zero.","lead":"This paper tests whether a language model can be made more aligned with human values by adding short, first-person ethical reflections to its pretraining data from the very first tokens. The results suggest that early 'alignment from token zero' produces models that follow a constitution better and generalize values to new moral dilemmas, more so than adding the same reflections only at the end of pretraining.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Token-zero advantage over midtraining rests on a single 3B training run per condition; reported intervals exclude training-seed variance.","rationale":"The reader's weakest assumption concerns the Persona Selection Model. I think that is a real interpretational risk, but it is not the most load-bearing for the empirical central claim. Even if pretraining only teaches reflection-shaped text rather than a stable, selectable persona, the timing comparison and the OOD generalization results would still be meaningful; the persona language is a framing device that could be revised without destroying the empirical findings. The more immediate threat is statistical: the effects that carry the token-zero claim—the AI Risk gap, the value-prioritization shift, and the grows-with-budget conclusion—emerge only at 3B, where each condition is a single training run. Appendix D.7 explicitly excludes training-seed variance from all intervals. At 1.7B, the AI Risk and value-prioritization effects do not reproduce, so the larger scale is doing all the work. A single seed cannot rule out run-to-run noise. This is a concrete, testable weakness that does not require accepting or rejecting the persona mechanism. I therefore disagree with the reader's choice of weakest assumption, while agreeing with the CONDITIONAL verdict because the seed-variance condition is already implicit in the reader's rationale. The paper otherwise has strong independent support: OOD AIRiskDilemmas results, the citation-exclusion test, and consistent capability preservation. The condition to add is explicit multi-seed replication, particularly at the scale where the headline effects appear.","tokens_in":52979,"tokens_out":13210,"duration_ms":119896,"concrete_test":"Train at least three independent seeds (different data order and initialization) of SPP{T0} and SPP{MT} at 1.7B/100B, plus one additional seed of each at 3B/500B, and compare the distributions of the differences on ConstitutionEval, AI Risk, and value-priority RBO. If the token-zero advantage does not replicate in sign and approximate magnitude across seeds (for example, if the 95% interval over seed-pair differences includes zero for the AI Risk gap), the central claim is not supported. Also report the seed-level spread for the 19 pp AI Risk gap at 3B.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central comparison is SPP{T0} vs SPP{MT} at 3B/500B tokens. Each condition is one training run. Appendix D.7 states that all reported confidence intervals capture only evaluation-prompt/item sampling, not training-seed variation. The effects that distinguish token zero—the AI Risk misalignment gap (about 19 pp at 3B vs about 4 pp at 1.7B) and the value-prioritization shift—appear only at this single 3B seed; at 1.7B/100B, value priorities are shared across all methods and AI Risk does not separate (Appendix F.10.1). Thus the headline claims of deeper value alignment and of an advantage that grows with pretraining budget lean on one run per condition at the larger scale. The scaling comparison itself is a two-point contrast that also changes model architecture (SmolLM2-1.7B vs Llama-3.2-3B-shaped), token budget, and the annotated dataset (10M-doc pilot vs 51.4M-doc production), so even the direction of the scaling effect is not isolated. If that single 3B seed is atypical, the central timing advantage would not be established.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Synthetic Persona Pretraining (SPP), a data-level intervention that installs a constitution-based assistant persona during pretraining by inserting first-person value reflections into a fraction of pretraining documents, then uses a matching post-training mixture (SP-SFT) for persona binding. The authors compare token-zero (SPP{T0}), midtraining (SPP{MT}), combined (SPP{T0,MT}), filtered, and vanilla recipes at two scales (1.7B/100B tokens and 3B/500B tokens) with identical post-training. They report that SPP improves constitution following, jailbreak robustness, and performance on out-of-distribution moral dilemmas; that token-zero models follow the constitution better than midtraining models, shift value priorities, and choose safer actions; that the token-zero advantage grows with pretraining budget; that midtraining is sufficient for jailbreak robustness; and that the gains depend on persona binding. The paper releases code, data, and model checkpoints.","tokens_in":53085,"tokens_out":3990,"duration_ms":37256,"significance":"If the central claims hold, this is a significant contribution: it provides evidence for the Persona Selection Model, shows path-dependence in alignment, and proposes a concrete pretraining-time alignment recipe that could be scaled. The experimental design has notable strengths: the five recipes at each scale are token-matched with identical post-training, the multiple-choice evaluations use a rotation-based debiasing protocol, external risk labels are re-audited, and the artifacts are released. However, the load-bearing claims currently rest on single training runs per condition and on a two-point scaling comparison that changes architecture, token budget, and annotated dataset size. The headline 'token zero deepens value alignment' is also evaluated partly with an in-domain benchmark derived from the same constitution that defines the training reflections. These issues do not invalidate the approach, but they currently prevent the paper from establishing its strongest conclusions.","major_comments":[{"comment":"The central claim that token-zero alignment is deeper than midtraining alignment rests on a single 3B training run per condition. Appendix D.7 explicitly states that all reported confidence intervals capture only evaluation-prompt/item sampling and not training-seed variation. The effects that distinguish token zero—the AI Risk misalignment gap of about 19 percentage points at 3B and the value-prioritization shift—are absent at the 1.7B scale (Appendix F.10.1). Because each condition is one seed, the reported intervals and McNemar tests cannot establish that the token-zero advantage is not a seed artifact. This is load-bearing for the paper's second and third findings in Section 1 and for the takeaway in Section 3.1. I ask the authors to either train multiple seeds per condition at the 3B scale and report seed variance, or substantially temper the causal 'token zero' language to 'in our single run at each scale'.","section":"Section 3.3, Appendix D.7"},{"comment":"The scaling claim—'token zero advantages grow with pretraining budget'—is based on a two-point comparison that simultaneously changes the model architecture (SmolLM2-1.7B vs. Llama-3.2-3B-shaped), the token budget (100B vs. 500B), and the annotated dataset (10M-document pilot vs. 51.4M-document production set). Any of these factors, alone or in combination, could produce the observed increase in the SPP{T0} vs. SPP{MT} gap. The paper's Figure 3 therefore does not isolate pretraining budget as the cause. A within-architecture budget sweep (e.g., the same 3B architecture at 100B and 500B tokens) or at least an explicit acknowledgment that the scaling comparison is confounded is needed before the 'grows with pretraining budget' claim can be accepted.","section":"Section 3.3, Figure 3, Table 2, Appendix A.4.8"},{"comment":"ConstitutionEval is an in-domain benchmark: its items are generated from the same constitution that defines the reflections used in SPP training, and its gold answers are validated by asking another model to interpret that constitution. The paper acknowledges this ('in-domain benchmark') but uses performance on it to support the broader claim that token-zero models 'internalize the constitution's underlying principles.' The value-prioritization agreement measure in Appendix F.4 similarly derives its reference ordering by asking Claude Fable 5 to interpret the same constitution. These evaluations are partly circular with respect to the training signal. To support the generalization claim, the authors should report results on external or held-out value-alignment benchmarks not sourced from the training constitution, or at minimum hold out a set of constitution articles from both training and evaluation.","section":"Section 2.2.3, Appendix D.1, Appendix F.4"}],"minor_comments":[{"comment":"The abandoned canary stream and identity-canary injections remain in the released data and training mix. Please document this more explicitly so downstream users do not mistake them for an active experimental condition, or remove them from the release if they are not used.","section":"Appendix A.2, A.3"},{"comment":"The sentence 'Roughly 15.8K rows have an empty first-person string due to isolated parsing failures' should state how these rows are handled during training (e.g., skipped, treated as loss-masked, or included as empty reflections).","section":"Appendix A.2"},{"comment":"The 'parity' line and the y-axis label 'Improvement (pp)' are not defined in the caption. Please state explicitly that positive values favor SPP{T0} over SPP{MT} and what 'parity' means.","section":"Figure 3"},{"comment":"The difference-of-differences bars in Figure 4 are presented without confidence intervals or significance tests. Given that these are single-run comparisons, the visual pattern should be accompanied by at least bootstrap intervals over items or an explicit caveat.","section":"Section 3.4.1, Figure 4"},{"comment":"The discussion paragraph beginning 'Like all empirical findings in language modeling, our results hold under our specific configuration...' is appropriately cautious, but it should also explicitly cite the single-seed-per-condition limitation that is documented in Appendix D.7, rather than only mentioning seed as part of the configuration.","section":"Section 5"}],"recommendation":"major_revision","confidential_remarks":"The paper's empirical core is carefully built, but the strongest claims (token-zero depth and scaling) currently rest on single seeds and a confounded two-point scaling comparison. If the authors can add seed variance or restructure the claims, the paper could become a strong contribution. The in-domain evaluation circularity is a second barrier. I recommend major revision rather than rejection because the approach and release are valuable and the identified issues are addressable, albeit with substantial additional compute."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the paper does a genuinely careful job of testing a concrete new idea—splicing constitution-conditioned reflections into pretraining—and the token-zero vs midtraining comparison is the right experiment to run. But the headline claim that token-zero yields deeper value alignment currently rests on one 3B training run per condition, and the scaling story is two points with three confounds. I'd send it to review, but the strongest conclusions need multi-seed evidence.\n\nWhat's new: SPP is a new data-intervention recipe, and the evaluation suite is unusually careful for this line of work. The token-matched variants (SPP{T0} vs SPP{MT} see the same loss-carrying tokens), the identical post-training across conditions, the swap-debiased multiple-choice protocol, the audit of AIRiskDilemmas labels, and the citation-exclusion test for persona binding are all well executed. They also ship code, data, and checkpoints. That is real evidence and deserves credit.\n\nSoft spots: First, each condition is a single training run. Appendix D.7 says the confidence intervals cover only evaluation-prompt/item sampling, not training-seed variation. At 1.7B, the AI Risk and value-priority effects do not separate; the token-zero advantage appears only at the single 3B seed. Second, the scaling claim in Figure 3 is a two-point comparison that changes architecture (SmolLM2-1.7B vs Llama-3.2-3B-shaped), token budget (100B vs 500B), and the annotated dataset (10M-doc pilot vs 51.4M production). Direction and magnitude of the growth are not isolated. Third, ConstitutionEval is in-domain: built from the same constitution that defines the training reflections, and the value-prioritization agreement (Appendix F.4) derives its reference ordering by asking Claude Fable 5 to interpret that constitution. That is partly circular, though the OOD AIRiskDilemmas results provide independent grounding. The persona-selection mechanism also remains an assumption, not something the paper demonstrates directly.\n\nNone of this is fatal. The paper is honest about most of it, and the empirical pattern is consistent across many evaluations. But the central timing advantage is less robust than the abstract suggests until it survives more seeds and a cleaner scaling curve.\n\nWho should read it: anyone working on alignment during pretraining, data-centric safety, or model raising. It deserves a serious referee—conditionally. I would ask for multiple seeds per condition, a size-controlled scaling comparison, and an external or blinded constitution-eval before letting the strongest claims through.","headline":"A well-run single-seed experiment with a plausible but not yet robust token-zero alignment advantage; deserves serious review, but the strongest claims need multi-seed and cleaner scaling evidence.","tokens_in":53788,"tokens_out":2542,"would_cite":true,"duration_ms":24666,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Synthetic Persona Pretraining installs a constitution-grounded assistant persona from token zero, giving deeper value alignment than introducing the same reflections late.","keywords":["Synthetic Persona Pretraining","persona binding","alignment from token zero","constitution following","jailbreak robustness","value prioritization","moral dilemmas","pretraining-time alignment"],"falsifier":"Train a model with the same reflections interleaved at token zero but with the <assistant> marker removed or with reflection voices randomly shuffled across documents; if constitution following and dilemma choices stay equally strong, the persona-stability mechanism is not what carries the result. The paper's own checkpoint trajectories also predict that a clear persona representation or value-priority shift should be visible well before post-training, so measuring the activations and finding no persistent persona signal would similarly weaken the claim.","tokens_in":52668,"feed_emoji":"⚖️","tokens_out":5608,"duration_ms":48251,"temperature":0.7,"pith_summary":"This paper sets out to prove that an AI assistant's values are not a thin layer to be added after training but can be installed from the very first token of pretraining. It introduces Synthetic Persona Pretraining (SPP), which inserts short, constitution-derived first-person reflections into about ten percent of pretraining documents, so the model learns to speak as the desired persona while it is still learning language. The claim is that models trained this way follow their constitution more faithfully, better resist jailbreaks, and choose safer actions on moral dilemmas they were never shown, while retaining general capability. The paper also argues that timing is the crux: adding the identical reflections only at the end of pretraining produces weaker constitution following, no shift in value priorities, and less aligned choices, and the gap grows with pretraining budget.","feed_headline":"Values planted at token zero beat the same training done late","feed_subtitle":"Persona-from-token-zero models follow constitutions better and act safer on unseen dilemmas.","key_machinery":"The central object is the synthetic persona reflection: a short first-person monologue, written by a generator conditioned on a 35-article value constitution, that reflects on a pretraining document and cites the relevant articles, inserted mid-document behind a special <assistant> token. During pretraining the model maximizes cross-entropy on both the original document and the reflection, with the reflection block masked from the document's attention and RoPE positions aliased so the document is effectively unchanged. This teaches the model to simulate the desired persona alongside the many personas already in the corpus. A second machinery piece is persona binding: the same constitution-conditioned generator is used to rewrite 300k user-assistant dialogues (SP-SFT), and post-training on this matched distribution is what attaches the assistant identity to the pretraining-installed persona. The underlying interpretive frame is the Persona Selection Model: pretraining teaches simulation of many personas, and post-training selects one of them to be the assistant.","core_discovery":"On its own terms, the paper's discovery is that installing the assistant persona from token zero produces a different and deeper kind of alignment than mid- or post-training interventions. Token zero models generalize beyond the constitution's explicit statements and internalize its underlying principles, which shows up as a different value prioritization (Truthfulness and Justice ranked highest) and better aligned actions on out-of-distribution moral dilemmas. The same annotated reflections introduced only in midtraining do not shift value priorities and do not improve dilemma choices. The effect depends on persona binding, the process by which post-training connects the assistant identity to the persona that pretraining installed, and the token zero advantage grows when the model is scaled from 1.7B parameters on 100B tokens to 3B on 500B, particularly on harder, out-of-distribution evaluations.","pith_inferences":["If value priorities really are set by token-zero exposure, then larger-scale RL or fine-tuning may erode surface refusals while the token zero value ordering persists, an empirically testable prediction that would extend the paper's abliteration results.","The method suggests a general recipe for raising models with multiple distinct personas: pretraining could install several constitutions and post-training could select among them, turning persona choice into a controllable inference-time property.","Because reflections add only about 0.55% of tokens and reuse the existing context window, the SPP data intervention is cheap enough to combine with data filtering and curriculum methods, potentially stacking with other safety-pretraining interventions.","A direct test of the persona-selection mechanism would be to ablate the <assistant> marker or shuffle reflection voices across documents; if the token zero advantage survives, the load-bearing claim shifts from persona stability to mere exposure to value-laden text."],"forward_implications":["Alignment ceases to be a post-hoc overlay: value formation is something pretraining itself can target, and late post-training alone cannot recreate it.","Token zero advantages on harder, out-of-distribution evaluations grow with pretraining budget, so small-scale experiments may understate the value of early interventions.","Midtraining exposure is sufficient for jailbreak robustness but not for deep value priority shifts, separating two alignment goals with different timing requirements.","Persona binding means post-training data must match the persona installed during pretraining; mismatched post-training data erases most of the value-alignment benefit.","Constitution-grounded values learned in pretraining can survive even when the constitution article is never cited in post-training, evidence that the values are not merely memorized from the SFT set."],"supporting_citations":[{"why":"Supplies the Persona Selection Model, the hypothesis SPP operationalizes: pretraining teaches simulation of many personas and post-training selects one.","marker":"Marks et al., 2026"},{"why":"Provides the safety classifier used to flag harmful documents for annotation and serves as a safety-pretraining baseline for comparison.","marker":"Maini et al., 2025"},{"why":"Supplies AIRiskDilemmas, the out-of-distribution moral-dilemma benchmark on which token zero models show lower misalignment.","marker":"Chiu et al., 2025"},{"why":"Supplies evidence for early-exposure and continual-training effects that motivate the token-zero timing and structure the continual-training experiment.","marker":"Feng et al., 2026"},{"why":"Supplies the WildChat prompts used to build the SP-SFT post-training mixture that performs persona binding.","marker":"Zhao et al., 2024"},{"why":"Defines the standard post-training alignment pipeline that SPP positions itself against, where the assistant identity is introduced late.","marker":"Ouyang et al., 2022"}],"fun_headline_variants":["Token zero personas: safer on unseen moral dilemmas","Start values at token zero: deeper alignment, safer AI","Persona from pretraining beats post-training for safety","Token-zero alignment: better on OOD dilemmas, says 3B study"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that next-token training on first-person moral reflections actually forms a stable, selectable persona in the model, rather than just teaching a style of reflection-shaped text; if no stable persona forms, the token-zero timing advantage and the persona-binding explanation would not follow even if some empirical gains remained.","fun_headline_variants_meta":{"raw":{"variants":["Token zero personas: safer on unseen moral dilemmas","Start values at token zero: deeper alignment, safer AI","Persona from pretraining beats post-training for safety","Token-zero alignment: better on OOD dilemmas, says 3B study"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000658,"raw_usage":{"total_tokens":3034,"prompt_tokens":992,"completion_tokens":2042,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":608,"completion_tokens_details":{"reasoning_tokens":1974}},"tokens_in":608,"tokens_out":2042,"duration_ms":15013,"temperature":1.0,"reasoning_tokens":1974,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T10:05:04.320080+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train a model with the same reflections interleaved at token zero but with the <assistant> marker removed or with reflection voices randomly shuffled across documents; if constitution following and dilemma choices stay equally strong, the persona-stability mechanism is not what carries the result. The paper's own checkpoint trajectories also predict that a clear persona representation or value-priority shift should be visible well before post-training, so measuring the activations and finding no persistent persona signal would similarly weaken the claim.","supporting_citations":[],"review_version":1}