Amplifying reasoning task vectors (α>1) surfaces learned secrets in LLMs up to 10× more frequently than standard reasoning models across four secret-keeping settings.
arXiv preprint arXiv:2511.05408 , year=
5 Pith papers cite this work. Polarity classification is still indexing.
years
2026 5representative citing papers
A geometric decomposition framework shows that affine transformations best recover prompt-induced task geometry and behavior in language and vision models across multiple datasets.
Emergent misalignment arises from overtraining after primary task convergence and is preventable by early stopping, which retains 93% of task performance on average.
CreativityNeuro applies contrastive weight steering to LLMs, yielding up to 14 percentile gains on the Divergent Association Task and improved originality in human-rated tests while reducing mode collapse.
Trait-space drift monitoring detects emergent misalignment checkpoints in 7-9B LLMs with 2.2% FNR, 2.9% FPR and 0.99 AUROC, outperforming PCA and SAE baselines.
citing papers explorer
-
Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets
Amplifying reasoning task vectors (α>1) surfaces learned secrets in LLMs up to 10× more frequently than standard reasoning models across four secret-keeping settings.
-
Decomposing how prompting steers behavior
A geometric decomposition framework shows that affine transformations best recover prompt-induced task geometry and behavior in language and vision models across multiple datasets.
-
Overtrained, Not Misaligned
Emergent misalignment arises from overtraining after primary task convergence and is preventable by early stopping, which retains 93% of task performance on average.
-
CreativityNeuro: Steering Language Model Weights to Improve Divergent Thinking and Reduce Mode Collapse
CreativityNeuro applies contrastive weight steering to LLMs, yielding up to 14 percentile gains on the Divergent Association Task and improved originality in human-rated tests while reducing mode collapse.
-
Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning
Trait-space drift monitoring detects emergent misalignment checkpoints in 7-9B LLMs with 2.2% FNR, 2.9% FPR and 0.99 AUROC, outperforming PCA and SAE baselines.