A new 304-task benchmark shows that LLM agents rarely adjust their tool spending to match either the evidence or the budget, so economical action selection is a distinct, currently missing capability.
Transactions on Machine Learning Research , url=
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents
A new 304-task benchmark shows that LLM agents rarely adjust their tool spending to match either the evidence or the budget, so economical action selection is a distinct, currently missing capability.