A judge-extract-code-conclude workflow beats plain prompting on numeric long-context tasks and cuts API cost, but underperforms CoT on one of two benchmarks.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
An Effective Framework to Help Large Language Models Handle Numeric-involved Long-context Tasks
A judge-extract-code-conclude workflow beats plain prompting on numeric long-context tasks and cuts API cost, but underperforms CoT on one of two benchmarks.