Behavior Cue Reasoning trains LLMs to emit special tokens before behaviors, enabling monitors to cut up to 50% wasted reasoning tokens and recover safe actions from 80% of unsafe traces, more than doubling success rates with no performance cost.
Backtracking for safety
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.AI 2years
2026 2representative citing papers
CROP applies conformal prediction to certify the longest contiguous prefix of an LLM reasoning trace that is guaranteed to contain no annotated errors under an exchangeability assumption.
citing papers explorer
-
Behavior Cue Reasoning: Monitorable Reasoning Improves Efficiency and Safety through Oversight
Behavior Cue Reasoning trains LLMs to emit special tokens before behaviors, enabling monitors to cut up to 50% wasted reasoning tokens and recover safe actions from 80% of unsafe traces, more than doubling success rates with no performance cost.
-
Conformal Certification of Reasoning Trace Prefixes
CROP applies conformal prediction to certify the longest contiguous prefix of an LLM reasoning trace that is guaranteed to contain no annotated errors under an exchangeability assumption.