Pith. sign in

Inoculation prompting: Eliciting traits from llms during training can suppress them at test-time.arXiv preprint arXiv:2510.04340

7 Pith papers cite this work. Polarity classification is still indexing.

7 Pith papers citing it

citation-role summary

background 1

citation-polarity summary

fields

cs.LG 5 cs.CL 2

years

2026 7

verdicts

UNVERDICTED 7

roles

background 1

polarities

background 1

representative citing papers

Overtrained, Not Misaligned

cs.LG · 2026-05-12 · unverdicted · novelty 6.0

Emergent misalignment arises from overtraining after primary task convergence and is preventable by early stopping, which retains 93% of task performance on average.

When Role-playing, Do Models Believe What They Say?

cs.CL · 2026-06-09 · unverdicted · novelty 5.0

Different persona induction methods produce a spectrum of belief internalization: prompting, ICL and SFT mainly alter outputs while Emergent Misalignment produces large representational shifts and Open Character Training produces smaller ones clearest in larger models.

Weird Generalization is Weirdly Brittle

cs.CL · 2026-04-11 · unverdicted · novelty 4.0

Weird generalization in fine-tuned models is brittle, appearing only in specific cases and disappearing under prompt-based interventions that make the undesired behavior expected.

citing papers explorer

Showing 7 of 7 citing papers.