A human-in-the-loop system that defers uncertain queries to humans based on reasoning-trace length cuts error rates of Qwen3 and DeepSeek R1 on hard MATH problems by about 2 percent, while fronting them with a faster model cuts latency and cost by roughly 40 to 50 percent.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Fail Fast, or Ask: Mitigating the Deficiencies of Reasoning LLMs with Human-in-the-Loop Systems Engineering
A human-in-the-loop system that defers uncertain queries to humans based on reasoning-trace length cuts error rates of Qwen3 and DeepSeek R1 on hard MATH problems by about 2 percent, while fronting them with a faster model cuts latency and cost by roughly 40 to 50 percent.