{"id":"5ae97c12-18f2-4b64-9c5e-b7373abe3d85","arxiv_id":"2608.09282","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A new benchmark for budget-constrained, coupon-optimized basket shopping agents, on which the best tested agent succeeds on only 61.2% of tasks.","lead":"This paper introduces ComboShoppingBench, a benchmark for evaluating AI shopping agents on basket-construction tasks with budgets and coupons. It combines deterministic rule checks with LLM judges to score semantic fit, response quality, and honest reporting, finding that even the strongest agent fails 39% of tasks.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"LLM-judge reliability is validated on only 30 Qwen No-think outputs; extrapolating 97% agreement to all 22 configurations is the weakest link for the 61.2% headline.","rationale":"The central claim has two parts: the benchmark reliably measures combo-shopping capability, and current agents fail it, with the best at 61.2% Overall success. The reader's weakest assumption targets exactly the reliability of the LLM judges, and I agree this is the most load-bearing concern. The human agreement study uses only 30 outputs from a single weak agent, while the headline is an intersection of four dimensions on 6,402 outputs from 22 configurations. Because the reported pass rates are between 60% and 90%, even a modest judge error rate can change the intersection by several points, potentially altering the qualitative conclusion. The Response Quality dimension is especially critical: human agreement is lower (94.67%), and the judge produces strikingly low pass rates for specific configurations (e.g., 19.6% for Claude-Opus-4.6 No-think) that were never validated against human labels. Other candidate concerns, such as the budget lower bound in approximate-target mode, are explicitly defended by a calibration study (Appendix D) and therefore are not as decisive. The lack of released code and data is a reproducibility issue but secondary to the numeric claim. A targeted human evaluation on diverse outputs would settle whether the 97% agreement generalizes; until then, the reader's conditional verdict is appropriate, and no verdict change is needed.","tokens_in":30920,"tokens_out":8984,"duration_ms":93928,"concrete_test":"Run a stratified human-annotation study on 100 outputs: 10 each from the 10 most diverse agent configurations, including GPT-5.5 (Think), Claude-Opus-4.6 (No-think), and DeepSeek-V4-Pro (No-think), covering all budget modes and both success and failure patterns. Have two annotators and an expert adjudicator label every rubric instance, then measure Gemini-3.1-Pro-Preview agreement against this reference. If per-dimension agreement, especially Response Quality, falls below 95%, or if overall agreement drops more than 2 points from Table 3's 97.57%, recompute Table 2's Overall success with a bias-corrected judge. If the corrected rates change the rank order or the 'only 61.2%' conclusion, the headline claim is not robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Table 3 reports human agreement for just 700 rubric decisions drawn from 30 outputs of Qwen3.6-27B (No-think). The full evaluation produces roughly 6,402 outputs x about 19 rubric decisions, approximately 120k judgments, and the headline Overall success is the strict intersection S AND V AND Q AND F. Even a small degradation in judge accuracy on outputs from other models could move the intersection materially: for GPT-5.5 (Think), S=83.8%, and a judge with 90% instead of 97% agreement can shift the true rate by several points. The Response Quality dimension is particularly fragile: human agreement is only 94.67% (Table 3, N=150), and the judge gives very low pass rates for some configurations, such as 19.6% for Claude-Opus-4.6 (No-think) and 14.1% for DeepSeek-V4-Pro (No-think), which are never checked against human labels. The claim that 'the strongest agent achieves only 61.2%' is therefore not strongly grounded until judge accuracy is demonstrated on the actual distribution of outputs across all 22 configurations.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces ComboShoppingBench, a benchmark of 291 tasks in a simulated e-commerce and takeout environment. Each task is constructed solution-first: an exploration agent finds a feasible hidden witness basket, and from that witness the pipeline synthesizes a coupon pack, a budget expression, a user query, and semantic rubrics. Evaluation combines three LLM judges (semantic satisfaction, response quality, claim faithfulness) with deterministic validation of SKU validity, coupon legality and optimality, and budget compliance. The authors evaluate 22 configurations of 11 agents and report that the strongest configuration, GPT-5.5 with thinking, achieves 61.2% overall success, defined as the strict conjunction of the four dimensions. Reliability of the LLM judges is supported by a human agreement study on 700 rubric decisions from 30 Qwen3.6-27B No-think outputs, and a witness-independence audit shows that successful outputs rarely reproduce the hidden witness.","tokens_in":31163,"tokens_out":5668,"duration_ms":59770,"significance":"If the results hold, ComboShoppingBench is a meaningful contribution to agent evaluation: it addresses open-ended basket construction while retaining objective verification of transactional constraints, avoiding both exact-match rigidity and purely semantic scoring. The strongest parts of the paper are the deterministic validator design, the basket-specific coupon optimality check, and the witness-independence audit, which together provide credible evidence against answer-key leakage. The human-agreement study is a genuine attempt to validate LLM judging, and the reported 97%+ raw agreement on the sampled subset is encouraging. The main risk is that the headline success rates depend on LLM judges whose reliability is demonstrated on only one agent's outputs; if judge accuracy differs on other configurations, the comparative and headline claims could shift. The benchmark itself and the deterministic pipeline are valuable regardless, and the paper is transparent about several of its design choices and limitations.","major_comments":[{"comment":"The reliability study is the load-bearing support for the LLM-judged dimensions, but it is confined to 30 outputs of a single agent (Qwen3.6-27B No-think) and 700 rubric decisions. The full evaluation applies the same three judges to approximately 6,402 outputs across 22 configurations. Response Quality is the weakest point: human agreement for Gemini-3.1-Pro is 94.67% (N=150), and the judge reports strikingly low Q pass rates for configurations never included in the human sample (e.g., 19.6% for Claude-Opus-4.6 No-think and 14.1% for DeepSeek-V4-Pro No-think in Table 2). Since Overall Success is the strict intersection S ∧ V ∧ Q ∧ F (Eq. (1)), even a small drop in judge accuracy on those outputs can move the headline rate by several points. I therefore ask for either human validation stratified across all or a representative subset of configurations, or a sensitivity analysis that recomputes the main table under conservative assumptions about judge error on low-Q configurations.","section":"§4, 'Are the LLM-Based Evaluators Reliable?' and Table 3"},{"comment":"Appendix D operationalizes the 'around N' budget as a two-sided target band: L ≤ P ≤ U, with L derived from the witness payable p* and a symmetric 15% tolerance (δ(p*) = max(0.15p*, 5)). The questionnaire of 35 participants gives median acceptable deviations of 17% below and 13% above, but the paper does not report the distribution or justify why a symmetric 15% band, rather than the observed median interval or a one-sided interpretation, is the correct formalization. This is consequential because Appendix H.2 reports that 36% of approximate-target outputs fall below L and treats this as underspending, and budget compliance is a deterministic V check. A reader who reads 'around N' as a soft target would not count those as failures. Please report the full questionnaire distribution and provide a robustness check showing how budget compliance and Overall success change under a one-sided or asymmetric band.","section":"Appendix D and Table D.1"}],"minor_comments":[{"comment":"The text contains the Unicode ligature typo 'diﬀiculty' in several places; please replace it with 'difficulty'.","section":"Appendix A.2 and Table A.1"},{"comment":"The labels such as '412' and '1712' are not explained in the caption; they appear to be task counts for recovered and lost tasks, and the caption should state this explicitly.","section":"Figure 6B"},{"comment":"The main text reports only raw agreement percentages; since the label distribution is heavily skewed toward Pass, consider also reporting balanced accuracy or Cohen's kappa in the main text, as Appendix G already does.","section":"§4 and Table G.2"},{"comment":"The sentence 'Our analysis identify cross-item constraint accumulation...' contains a subject-verb agreement error; it should be 'Our analysis identifies...'.","section":"§5"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the scope of an AI-agent benchmark venue, and the deterministic validator plus witness-independence audit are genuine strengths. The central concern is evaluator generalization: the 97% agreement figure is established on a narrow sample, and the low Q pass rates for several configurations make the extrapolation non-negligible. I would not reject on the current evidence, but the requested additional validation or sensitivity analysis is necessary before the headline comparison can be considered fully grounded. I would also encourage the authors to state clearly in the main text that the agreement study is limited to one agent's outputs."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"ComboShoppingBench is a genuinely useful addition to the LLM-agent evaluation stack. What's new is the combination: basket-level tasks with executable orders, settlement-based budgets, coupon legality and optimality, plus response-quality and claim-faithfulness checks. No prior benchmark covers all of those, and Table 1 makes the gap visible. The witness-based construction is the smartest part: an exploration agent builds a feasible basket first, then the query, coupons, budget, and rubrics are derived from it, so tasks are solvable by construction without making the witness the answer key. The independence audit supports this: only 7.2% of successful outputs exactly match the witness and 33.1% are entirely disjoint.\n\nDeterministic validation for SKUs, coupon legality, coupon optimality, and budget compliance is well specified, and coupon optimality is evaluated for the agent's own basket rather than against the witness. That is the right design. The budget calibration with a 35-person questionnaire is a reasonable attempt to ground “around N yuan.”\n\nMain soft spot: LLM-judge reliability evidence is thin where it matters most. The human agreement study uses 30 outputs from one No-think model, Qwen3.6-27B. Semantic and claim-faithfulness agreement look solid, but Response Quality agreement is 94.67% on 150 decisions, and the judge gives very low pass rates to some configurations (Claude Opus-4.6 No-think 19.6%, DeepSeek-V4-Pro No-think 14.1%) with no human labels on those outputs. The headline 61.2% is the strict intersection of four dimensions, so even modest judge degradation on other configs could move it by several points. This is not fatal — the paper is a benchmark, not a single-number claim about model capability — but the headline should be treated as provisional until judge accuracy is sampled across configurations. Second soft spot: no code or data released. That limits reproducibility and reuse, and for a benchmark that is a real cost.\n\nCitation pattern looks fair; the related-work comparison is informative. The think/no-think setup is approximate across providers, and the paper says so. Minor.\n\nWho it's for: people building shopping agents, and anyone designing verifiable open-ended agent benchmarks. It deserves a serious referee. I would ask for judge validation across configurations and artifact release before acceptance.","headline":"A solid, genuinely useful shopping-agent benchmark whose witness-based construction is the standout idea; the headline failure rate is provisional until LLM-judge validity is shown across configurations.","tokens_in":31678,"tokens_out":1835,"would_cite":true,"duration_ms":18991,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that current LLM shopping agents cannot reliably assemble multi-item baskets that satisfy compatibility, coupon, and budget constraints, with the best agent passing all checks on only 61.2% of tasks.","keywords":["ComboShoppingBench","LLM agents","shopping agents","benchmark","coupon optimization","budget constraints","basket construction","semantic evaluation"],"falsifier":"Take a random sample of 100 outputs from GPT-5.5 with thinking across the 291 tasks, have two expert annotators judge semantic satisfaction, response quality, and claim faithfulness against the same rubrics, and compare with the Gemini-3.1-Pro judge. If agreement on that sample falls below the 97% level measured on Qwen3.6-27B outputs, then the reported 61.2% overall success rate and the per-dimension pass rates are not reliable.","tokens_in":1617,"feed_emoji":"🛒","tokens_out":3437,"duration_ms":81883,"temperature":0.7,"pith_summary":"The paper introduces ComboShoppingBench, a benchmark that tests whether AI shopping agents can assemble a basket of multiple compatible items under budget, coupon, and store constraints. The central claim is that current LLM agents are far from reliable at this task: the strongest of 22 configurations, GPT-5.5 with thinking, passes all four evaluation dimensions on only 61.2% of 291 tasks. The authors argue that the benchmark's design, building each task from a verified hidden basket and then evaluating any valid alternative basket, makes the difficulty measurable rather than subjective. A sympathetic reader would care because successful agents would directly help with real multi-item purchases like building a PC or ordering a group meal.","feed_headline":"Best shopping AI fails 39% of combo baskets","feed_subtitle":"New benchmark with coupons and budgets shows top models miss constraints on nearly 40% of 291 tasks.","key_machinery":"The central mechanism is solution-first task construction paired with hybrid evaluation. An exploration agent first finds and validates a purchasable basket in a simulated commerce-and-takeout environment; this hidden witness guarantees the task is solvable but is never used as a reference answer. The witness's costs then drive the synthesis of coupons, an optimal-payable budget interval, and a user query with semantic rubrics. At evaluation, deterministic code checks SKU validity, order feasibility, coupon legality and optimality, and budget compliance, while LLM judges assess semantic satisfaction, response quality, and claim faithfulness. Overall success is the intersection of all four dimensions.","core_discovery":"The paper's central discovery is that combo shopping — selecting multiple products that jointly satisfy compatibility, availability, store-level rules, coupon legality and optimality, and budget — remains largely unsolved by current LLM agents, even though each individual dimension looks passable. The best agent reaches 83.8% semantic satisfaction and 83.8% rule-based validation, yet only 61.2% of its tasks pass every dimension simultaneously; the gap shows that failures accumulate across the pipeline. The benchmark also shows that thinking configurations help most agents but not uniformly, and that cross-item compatibility requirements and choosing among mutually exclusive coupons are the hardest parts. The paper further establishes that its evaluation is not tied to a single hidden answer: among 1,729 accepted outputs, only 7.2% reproduce the witness basket exactly, and 33.1% use a completely different set of products.","pith_inferences":["The hybrid design could be adapted to other constraint-satisfaction agent tasks where multiple valid outputs exist, such as travel itineraries or multi-step procurement, by using a witness only to certify feasibility.","The paper's 97% judge agreement may not transfer to outputs from stronger or different models; a stress-test extension would re-run the human agreement study on GPT-5.5 Think failures before relying on precise rankings.","The budget behavior result suggests a cheap intervention: rewriting approximate budgets as explicit ranges could close much of the observed gap without changing the underlying model.","If judge accuracy degrades on rare but important failure modes, such as subtle coupon miscalculations, the reported 61.2% is an upper bound on true capability; a deterministic-only audit on a subset would quantify that."],"forward_implications":["If the benchmark measures what it claims, then any agent that approaches 100% overall success would have to jointly handle compatibility reasoning, arithmetic, and faithful reporting, not just retrieve relevant items.","The failure analysis implies that improving individual capabilities, such as semantic matching or coupon ID validity, is insufficient because agents need to avoid constraint accumulation across ten or more simultaneous criteria.","The observation that approximate budgets are under-spent (only 49% pass rate versus 85–87% for caps or explicit ranges) suggests agents need explicit prompting or training to treat 'around N' as a two-sided spending target.","The finding that thinking helps most agents but hurts some (for example, Qwen uses fewer calculator calls and loses tasks) implies that enabling a reasoning mode is not a reliable fix by itself.","Because the benchmark accepts many alternative baskets, it can support future agent development without overfitting to a single reference answer."],"supporting_citations":[{"why":"Provides the WebShop environment and metric tradition that ComboShoppingBench builds beyond, establishing the baseline for shopping-agent evaluation.","marker":"[1]"},{"why":"Represents the individual-product shopping benchmark whose basket-level gap ComboShoppingBench fills.","marker":"[3]"},{"why":"Supplies a multi-store benchmark with partial basket support against which ComboShoppingBench positions its full basket-level coverage.","marker":"[4]"},{"why":"Is the closest prior benchmark with partial basket, budget, and coupon evaluation, serving as the main comparison point.","marker":"[5]"},{"why":"Adds long-horizon preference tasks but lacks executable orders and claim faithfulness, motivating the new four-dimension design.","marker":"[6]"}],"fun_headline_variants":["AI agents miss 39% of combo shopping baskets","Top shopping AI only passes 61% of basket tasks","Coupon combo shopping stumps even best AI agents","New benchmark shows AI struggles with multi-item shopping"],"cache_read_input_tokens":33920,"weakest_assumption_plain":"The benchmark's headline numbers assume that the three LLM judges, validated against only 700 human decisions from 30 outputs of a single weaker agent, stay at least 97% accurate on all 291 tasks and on outputs from stronger agents with different error patterns.","fun_headline_variants_meta":{"raw":{"variants":["AI agents miss 39% of combo shopping baskets","Top shopping AI only passes 61% of basket tasks","Coupon combo shopping stumps even best AI agents","New benchmark shows AI struggles with multi-item shopping"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000961,"raw_usage":{"total_tokens":4082,"prompt_tokens":924,"completion_tokens":3158,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":540,"completion_tokens_details":{"reasoning_tokens":3095}},"tokens_in":540,"tokens_out":3158,"duration_ms":21726,"temperature":1.0,"reasoning_tokens":3095,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T20:13:01.940318+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of 100 outputs from GPT-5.5 with thinking across the 291 tasks, have two expert annotators judge semantic satisfaction, response quality, and claim faithfulness against the same rubrics, and compare with the Gemini-3.1-Pro judge. If agreement on that sample falls below the 97% level measured on Qwen3.6-27B outputs, then the reported 61.2% overall success rate and the per-dimension pass rates are not reliable.","supporting_citations":[],"review_version":1}