GuardedRepair uses guarded best-of-N repair with symbolic checks, semantic diagnostics, and conservative policies to selectively replace LLM reasoning traces, raising GSM8K accuracy from 95.60% to 96.89% and ASDiv from 78.40% to 87.60% without breaking correct cases.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Guarded Repair for Harm-Aware Post-hoc Replacement of LLM Mathematical Reasoning
GuardedRepair uses guarded best-of-N repair with symbolic checks, semantic diagnostics, and conservative policies to selectively replace LLM reasoning traces, raising GSM8K accuracy from 95.60% to 96.89% and ASDiv from 78.40% to 87.60% without breaking correct cases.