Injecting first-person, value-laden reflections into pretraining text improves constitution following, jailbreak resistance, and out-of-distribution moral choices in small language models, with the largest gains when the reflections are present from token zero.
-> 50 (Steps 1-4 are a real trafficking primitive ordering; steps 5-30 are noise. Score on the first four -- rule 4a partial delivery -- not on the noise -- rule 5.)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Synthetic Persona Pretraining: Alignment from Token Zero
Injecting first-person, value-laden reflections into pretraining text improves constitution following, jailbreak resistance, and out-of-distribution moral choices in small language models, with the largest gains when the reflections are present from token zero.