A 13-way text-scrambling benchmark makes open-weight LLMs drop up to 54% average accuracy, and a multi-problem prompt makes their accuracy on the last question decay.
Enhancing robustness in large language models: Prompting for mitigating the impact of irrelevant information
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Robust Reasoning Benchmark
A 13-way text-scrambling benchmark makes open-weight LLMs drop up to 54% average accuracy, and a multi-problem prompt makes their accuracy on the last question decay.