{"id":"f265de2a-d87e-4678-9ceb-5eeaef1c34e9","arxiv_id":"2608.05519","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A new 304-task benchmark shows that LLM agents rarely adjust their tool spending to match either the evidence or the budget, so economical action selection is a distinct, currently missing capability.","lead":"The paper introduces a benchmark of 304 tasks where AI agents must complete jobs on a strict budget and pay for every tool they use. It finds that today's leading agents often under-use expensive tools when they are needed and over-use them when they are not, so completing a task and spending wisely turn out to be separate skills.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Cheap-QA task labels are the least verified save-regime evidence; mislabeled tasks would inflate the over-escalation and 'misdirected budget response' findings that anchor Finding 5.","rationale":"The reader's weakest assumption is precisely the unverified cheap-QA label, and I agree that this is the most load-bearing soft spot. The paper's construct-validity story is strong: scripted controls show opposite optimal actions, and the escalation_qa family has a four-condition verification. But cheap_qa is the save-regime counterpart and lacks that verification; a single family of 72 tasks carries the entire claim that agents 'overspend on cheap tasks' when budgets are loose. If a meaningful fraction of those tasks are not actually cheap-solvable, then low Save scores and the Gemini 19.4% deep-research figure reflect task mislabeling rather than irrational escalation. I do not think this overturns the central claim that completion and conditional resource selection are distinct: the oracle controls still fail save-oriented families while achieving high micro accuracy, and the workspace agents show much higher Econ than tool-API agents. It does mean the quantitative magnitude and the 'misdirected' asymmetry in Finding 5 should stay conditional until the cheap_qa labels are audited at the item level. I also noted the Figure 1 caption contains a contradictory six-family variant with counts that do not match Table 2; the body and tables consistently describe five families, so I treat it as a packaging inconsistency rather than a substantive flaw. Verdict remains conditional; no adjustment.","tokens_in":11700,"tokens_out":10719,"duration_ms":99958,"concrete_test":"Apply the escalation_qa four-condition verification to all 72 cheap_qa tasks: (1) a scripted CheapFirst policy using only local_keyword_search and read_document reaches the answer-bearing document and produces the expected answer; (2) the gold answer is absent from deep-research-only sources until after the cheap-visible document is read; (3) a scripted deep-research agent also solves the task but with cost greater than the tight budget; (4) the canonical cheap trajectory cost is at most the tight budget. Settling criterion: accept the labels if no more than 5 of 72 fail any condition; if 10 or more fail, re-estimate Table 4 Save scores and Finding 5's unnecessary deep-research rates after excluding or relabeling the failing tasks, and check whether Gemini's 19.4% loose-budget deep-research rate and Sonnet's 74% cheap-QA budget-bust rate persist.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central asymmetry in Finding 5—GPT-5.4 under-escalates while Gemini over-escalates—depends on cheap_qa tasks being genuinely solvable by local_keyword_search plus one read_document within the tight budget. This is asserted for all 72 tasks but is not verified the way escalation_qa is. Escalation-required tasks are checked against four conditions (cheap search cannot reach the source; the answer is absent from cheap-visible documents; deep research reaches it; a scripted cheap agent fails while an escalating agent solves). The cheap_qa family has no analogous four-condition check, and the human realism audit covered only 4 of 72 tasks. Since the Save group in Econ=min(Up,Save) is only 82 tasks (72 cheap-QA plus 10 stop-loss), even 10–15 mislabeled cheap-QA episodes would move Save scores by several points. Sonnet busts the budget on 74% of cheap-QA tasks and Gemini uses deep research on 19.4% at the loose budget; if a nontrivial share of these tasks actually require broader evidence, those episodes are rational escalation rather than waste, and the 'misdirected budget response' conclusion weakens.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that resource use should be part of the agent task itself, not a post-hoc logging statistic, and introduces EcoAgent-Bench: 304 real-derived tasks from GAIA, HotpotQA, and MuSiQue across five families (escalation QA, model-upgrade QA, cheap QA, stop loss, frozen-information QA), each with priced tools and tight/medium/loose budgets. The benchmark defines strict budgeted success (correct, evidence-accessed, and within budget) and an economic-consistency score Econ = min(Up, Save) over upgrade- and save-oriented family groups. Seven LLM agents are evaluated in tool-API and workspace-CLI settings, together with four oracle scripted controls. The central finding is that one-sided policies each fail one regime: tool-API agents attain only 3.9–24.0% micro strict success and 3.6–7.3% economic consistency, largely by under-escalating or overspending; GPT-5.4 rarely escalates even when the budget crosses the escalation cost, while Gemini spends a loose budget on unnecessary deep research on cheap-QA tasks. The authors conclude that completion, budget feasibility, and conditional resource selection are distinct capabilities.","tokens_in":11936,"tokens_out":13947,"duration_ms":127257,"significance":"If the benchmark construction is valid, this is a useful and fairly original contribution: it makes cost a first-class constraint, introduces a compact regime-balanced diagnostic (Econ) that exposes one-sided policies hidden by micro-accuracy, and ships a reproducible transformation pipeline with leakage checks, deterministic selection, model-tier verification, scripted construct-validity controls, cost-robustness sensitivity, and integrity-bound result artifacts. The headline claim that current agents do not make economically rational tool choices is plausible, falsifiable, and worth testing. However, the strength of the quantitative claims currently exceeds the support provided by the cheap-QA verification level and by single-episode execution without confidence intervals.","major_comments":[{"comment":"The cheap-QA family carries most of the save-regime evidence (72 of the 82 Save-group tasks), but it is the only large family without a label-verification gate. Escalation-QA is checked against four conditions (cheap search cannot reach the source, the answer is absent from cheap-visible documents, deep research reaches it, and scripted cheap agents fail), and model-tier and stop-loss families have their own checks; cheap-QA is described only as MuSiQue items whose answer appears directly in frozen search results, with a human realism audit covering 4 of 72 tasks. The coarse episode-level grounding check does not close this gap. Consequently, a nontrivial share of the reported Sonnet/Gemini budget busts (74% and 99% of cheap-QA tasks) and Gemini's 19.4% deep-research use on loose-budget cheap-QA tasks (Finding 5) could be rational escalation on mislabeled items rather than waste. I ask for a cheap-QA verification gate analogous to the escalation-QA gate: confirm per item that the answer is reachable via local_keyword_search plus read_document within the tight budget, that a scripted cheap agent succeeds, and that an escalating scripted control overshoots the budget; then re-run Findings 2 and 5 on the verified subset.","section":"Task Families (cheap_qa, 72); Validity Controls"},{"comment":"The main tool-API results and the budget sweeps use one episode per task per condition, with no confidence intervals. The escalation sweep's key comparison is 0/115 versus 1/115 versus 3/115 deep-research calls; Clopper-Pearson 95% intervals for these counts overlap broadly, so the statement that deep-research use changes 'from 0% to only 3%' is not a stable quantitative estimate, and the same caveat applies to the Econ differences in Table 4 (3.6–7.3%) and to the cheap-QA mean-cost monotonicities in Finding 5. The limitation paragraph acknowledges single-shot execution, but the headline numbers are reported without uncertainty and are used for the paper's main conclusions. Please add repeated episodes for the central tool-API and sweep conditions, report exact intervals or variance estimates, and phrase the weak-budget-response conclusion as directional if the intervals overlap.","section":"Experimental Setup; Finding 5; Table 4"},{"comment":"For several QA families, strict success only requires that the episode accessed some evidence, not that the cited evidence entails the answer. The manuscript acknowledges that the grounding check is deliberately coarse, but the Results section repeatedly interprets strict success as a ground-truth-compatible correctness signal. In cheap-QA this is not a purely hypothetical concern, since GPT-5.4 is reported to answer cheap-QA at low cost with only 25% lenient accuracy. Please either tighten the grounding check (require the answer span to appear in the evidence accessed) or consistently label the metric as evidence-accessed accuracy, and quantify how the re-labeling would change the main Econ and micro-success scores.","section":"Results, strict success definition; Limitations"}],"minor_comments":[{"comment":"Figure 1 contains two conflicting task-family coverage blocks: one lists 'N=304; six families' with code debugging (58) and cheap QA (14), and the other lists the correct five families with cheap QA (72). The stale block should be removed so the figure agrees with Table 2.","section":"Figure 1"},{"comment":"The sentence 'One escalation_qa, deep-research use changes only from 0%...' is missing a word and should read 'On escalation_qa, deep-research use changes only from 0%...'.","section":"Finding 5"},{"comment":"The caption says filled points mark Pareto frontiers but does not define the frontier criterion; please state whether the frontier is in terms of strict success versus mean ledger cost and add the same definition to the caption of Figure 3.","section":"Figure 2"},{"comment":"The row ordering note says rows are ordered by Econ within each track, but the Units column mixes ledger cost and execution proxy; please add a prominent note that the two unit types are not comparable and that the workspace proxy is not a priced action ledger.","section":"Table 4"},{"comment":"The limitation about the coarse grounding check should be repeated in the Results section or next to the strict-success definition, since many readers will not reach the Limitations section.","section":"Limitations"},{"comment":"The sentence reporting human-judge agreement says all eight disagreements are judge false positives and that the supplementary material gives reweighted estimates; please state in the main text whether the reweighted estimates change any of the model-level conclusions.","section":"Experimental Setup"}],"recommendation":"major_revision","confidential_remarks":"The paper is transparent about its limitations, and the release of the pipeline and integrity-bound artifacts is a genuine strength. The main risk is that the cheap-QA family is under-verified relative to the evidential load it carries for Finding 5 and for the Save side of the Econ score; I would prioritize that verification and a small repeated-runs study over additional benchmark breadth. No citation or novelty concerns beyond the minor figure/packaging errors noted in the report."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core claim is that completion, budget feasibility, and conditional resource selection are distinct capabilities, and the benchmark supports it. The scripted controls show micro accuracy rewards always-escalate policies; Econ = min(Up,Save) exposes one-sidedness; the budget sweep shows GPT-5.4 barely escalates even when the budget crosses the path cost. I buy that.\n\nWhat is genuinely new: making priced actions and a budget part of the task instance, the budget sweep, and the Econ diagnostic. CostBench is the closest competitor and the paper positions against it fairly. The construction quality is above the usual benchmark bar: deterministic selection, leakage checks, four-condition escalation verification, three-attempt model-tier verification, independent stop-loss checks, and regression smoke tests for the gates. The authors also publish their limitations rather than hiding them; the coarse grounding check, single-shot runs, and workspace execution proxy are all disclaimed up front.\n\nNow the soft spots. The one the stress-test note flags is real: cheap_qa is the least verified family. 72 tasks are asserted to be solvable by local search plus one read, but only four got human realism review and there is no escalation_qa-style four-condition check. If a nontrivial share of those actually need broader evidence, Gemini's 19.4% deep-research use at the loose budget is rational, not waste, and the 'misdirected budget response' framing weakens. That does not sink the paper: the central under-escalation finding rests on the well-verified escalation_qa family, where GPT-5.4 abstains in 45/115 and never uses deep research. But the authors should either audit all cheap-QA tasks or soften the over-escalation claim.\n\nSecond, the headline LLM table is one episode per task at temperature zero, so the 3.6–7.3% Econ ordering among backbones has no error bars. That is acknowledged but it limits how much should be concluded from small gaps between Sonnet and Gemini. The grounding check is coarse: 'grounded' means evidence was accessed, not that the cited span entails the answer.\n\nThis paper deserves a serious referee. The benchmark is useful for anyone building or evaluating resource-aware agents, and the artifacts look built to be reused. I'd send it out with two asks: verify the cheap-QA labels at the same standard as escalation-QA, and report variance or rerun with multiple episodes.","headline":"A carefully built budget-conditioned agent benchmark whose core separation claim holds up; the cheap-QA family and single-episode LLM runs are the two places to push before trusting magnitudes.","tokens_in":12430,"tokens_out":2173,"would_cite":true,"duration_ms":21181,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Current LLM agents do not make economically rational tool choices under explicit budgets and prices, according to a new 304-task benchmark.","keywords":["budget-constrained agents","economic decision-making","tool selection","LLM agent evaluation","benchmark","cost-aware planning","abstention","budget sweep"],"falsifier":"Inspect the 72 cheap-QA frozen evidence sets and run a scripted cheap-only policy that has access to the same search results but is forbidden from escalating; if it fails on a substantial share of tasks, or if independent annotators find that the answer is not present in the cheap-visible documents, the family's 'cheap evidence suffices' label is wrong and the over-escalation finding is an artifact.","tokens_in":11498,"feed_emoji":"💸","tokens_out":5708,"duration_ms":46095,"temperature":0.7,"pith_summary":"The paper argues that current agent benchmarks conflate two distinct capabilities: completing a task and choosing actions that fit the evidence and the budget. It introduces EcoAgent-Bench, a 304-task benchmark in which every step has a price and every task has an explicit budget, so the economically correct policy depends on whether cheap evidence suffices, escalation is warranted, or the premise is unsupported. Across tool-API agents, micro strict success is 3.9–24.0% and economic consistency at most 7.3%, because agents either stop before the required escalation or overspend on tasks cheap evidence already solves. A budget sweep shows that larger budgets do not reliably buy better decisions: one agent barely escalates even when the budget crosses the cost threshold, while another increasingly spends on unnecessary deep research. The central claim is that completion under a budget and economical action selection are distinct properties, and that a compact worst-regime diagnostic, the economic-consistency score, exposes one-sided policies that plain accuracy hides.","feed_headline":"Agents that ignore budgets fail priced tasks, new benchmark shows","feed_subtitle":"Across 304 priced tasks, tool-API agents reach at most 24% strict success and 7.3% economic consistency.","key_machinery":"The central object is the budget-conditioned task format: each task couples a frozen evidence set, a priced action ledger (e.g., cheap local search at 18 units vs an agentic deep-research tool at 260 units), a budget, and oracle-only evaluator material. The load-bearing diagnostic is the economic-consistency score, $\\mathrm{Econ} = \\min(\\mathrm{Up}, \\mathrm{Save})$ over family groups, which prevents a one-sided always-escalate or always-stop policy from looking strong. Paired task families (cheap_qa vs escalation_qa, and balanced model_upgrade_qa) operationalize the same decision under opposite evidence states. The threshold-crossing budget sweep is the key experimental mechanism: it varies the budget across tight, medium, and loose levels and tests whether escalation rates respond.","core_discovery":"EcoAgent-Bench turns cost into a first-class task constraint rather than a post-hoc statistic. Each of its 304 real-derived tasks provides priced actions, an explicit budget, a canonical budget-compliant trajectory, and a contrasting high-regret trajectory, across five families targeting four decisions: avoid unnecessary escalation, escalate when local evidence is insufficient, route to a stronger model tier, and stop when the premise is unsupported. The paper's central result is that no evaluated LLM agent reasons economically under this schedule: tool-API agents reach only 3.9–24.0% micro strict success and 3.6–7.3% economic consistency, with failure split between under-escalation (GPT-5.4 abstains on 45 of 115 escalation-required episodes and never invokes deep research) and over-spending (Gemini busts budget on 99% of cheap-QA tasks). Threshold-crossing budget sweeps reveal that more budget does not correct this: GPT-5.4's deep-research use rises from 0% to 2.6% as the budget crosses the escalation cost, while Gemini's unnecessary deep-research rate rises to 19.4% on loose-budget cheap tasks. The authors conclude that completion, budget feasibility, and conditional resource selection are separate capabilities.","pith_inferences":["One implication the paper leaves implicit: the same paired-regime design could be applied to non-QA agent tasks, such as code repair or web navigation, where a cheap versus expensive action trade-off also exists.","A testable extension would be to give agents an explicit cost-of-next-action signal in the observation, rather than only a budget, and measure whether the economic-consistency score rises.","The Econ diagnostic generalizes as an evaluation principle: any benchmark with opposing optimal actions should report worst-regime success to avoid majority-regime masking.","The asymmetric budget response suggests that merely scaling model capability or prompt length will not fix economic rationality; the action-selection mechanism itself must incorporate affordability."],"forward_implications":["If the central claim is right, budgeted benchmarks should report economic consistency and budget feasibility alongside accuracy; plain micro-success is not a sufficient evaluation signal.","Tool-API agent developers should target the two distinct failure modes separately: under-escalation on escalation-required tasks and over-spending on cheap-evidence tasks.","Budget information placed in the prompt is insufficient by itself; the budget must enter the action-selection rule.","A controller that estimates the marginal value of the next action, comparing expected gain against price and remaining budget, is the design direction the results support."],"supporting_citations":[{"why":"Source of the frozen-information anchor tasks and the closest completion-only benchmark the paper contrasts against.","marker":"Mialon et al. 2024"},{"why":"Seed corpus for the escalation-QA and model-upgrade QA tasks.","marker":"Yang et al. 2018"},{"why":"Seed corpus for cheap-QA and escalation-QA multi-hop items.","marker":"Trivedi et al. 2022"},{"why":"Establishes the abstention tradition the stop-loss family is grounded in.","marker":"Rajpurkar, Jia, and Liang 2018"},{"why":"Underpins the LLM-as-a-judge semantic scoring of open-ended answers.","marker":"Zheng et al. 2023"},{"why":"Supplies the argument that agent evaluation must measure deployment-relevant factors such as cost, which this paper operationalizes.","marker":"Kapoor et al. 2024"},{"why":"Closest prior benchmark with priced tools; the paper positions EcoAgent-Bench as complementary rather than corrective.","marker":"Liu et al. 2025"}],"fun_headline_variants":["EcoAgent-Bench: LLM agents flunk budget-aware decisions","Budget constraints expose LLM agents' economic blind spots","Priced tasks reveal LLM agents can't balance cost and performance","LLM agents overspend and under-escalate under budget limits","Economic consistency is a new metric LLM agents fail at"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 72 cheap-QA tasks really are solvable by a cheap search plus one quick read, so that spending on deep research on them counts as a mistake; if a meaningful share of those tasks actually needs deeper evidence, the measured over-escalation is partly an artifact of task construction.","fun_headline_variants_meta":{"raw":{"variants":["EcoAgent-Bench: LLM agents flunk budget-aware decisions","Budget constraints expose LLM agents' economic blind spots","Priced tasks reveal LLM agents can't balance cost and performance","LLM agents overspend and under-escalate under budget limits","Economic consistency is a new metric LLM agents fail at"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000638,"raw_usage":{"total_tokens":3014,"prompt_tokens":1093,"completion_tokens":1921,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":709,"completion_tokens_details":{"reasoning_tokens":1834}},"tokens_in":709,"tokens_out":1921,"duration_ms":14702,"temperature":1.0,"reasoning_tokens":1834,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T11:37:21.081828+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Inspect the 72 cheap-QA frozen evidence sets and run a scripted cheap-only policy that has access to the same search results but is forbidden from escalating; if it fails on a substantial share of tasks, or if independent annotators find that the answer is not present in the cheap-visible documents, the family's 'cheap evidence suffices' label is wrong and the over-escalation finding is an artifact.","supporting_citations":[],"review_version":1}