A new benchmark shows LLMs rarely detect and correct errors in user prompts unless explicitly instructed, and fine-tuning on error-handling examples greatly improves this ability.
How long does it take to drive from the university to the beach?
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Mis-prompt: Benchmarking Large Language Models for Proactive Error Handling
A new benchmark shows LLMs rarely detect and correct errors in user prompts unless explicitly instructed, and fine-tuning on error-handling examples greatly improves this ability.