Pith. sign in

I will not harm you unless you harm me first

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.CL 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Normative Conflicts and Shallow AI Alignment

cs.CL · 2025-06-05 · conditional · novelty 6.0

A philosophical argument that preference fine-tuning produces shallow alignment, plus a new 'thought injection' attack that exploits reasoning traces in LLMs.

citing papers explorer

Showing 1 of 1 citing paper.

  • Normative Conflicts and Shallow AI Alignment cs.CL · 2025-06-05 · conditional · none · ref 11

    A philosophical argument that preference fine-tuning produces shallow alignment, plus a new 'thought injection' attack that exploits reasoning traces in LLMs.