Hidden-state probes detect hint reliance in explicit and latent chain-of-thought models about equally well, with task properties and internal access mattering more than the reasoning format.
Mechanistic Interpretability Workshop at NeurIPS 2025 , year=
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Does Out-of-Sight Equal Out-of-Mind in CoT Monitorability?
Hidden-state probes detect hint reliance in explicit and latent chain-of-thought models about equally well, with task properties and internal access mattering more than the reasoning format.