{"id":"92d2d6a0-915f-4855-85b8-dd7bd41fe3d1","arxiv_id":"2608.08621","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Business Arena, a data-grounded B2B marketplace simulator, ranks 15 LLM agents from $20,856 to $188,488 mean final net worth, well below expert-designed strategies that reach $436,195.","lead":"This paper introduces Business Arena, a simulated marketplace where an AI agent runs a cross-border shop for 30 days and is scored on final net worth. On it, the best LLM reaches about $188,000 versus $436,000 for a human-designed strategy, showing large headroom in end-to-end business operation.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Arena realism rests on an engineered Google Trends signal: Appendix C/G makes the public signal correlate with hidden demand at pooled ρ=0.972, so the measured evidence-guided advantage may not transfer to real, noisier public data.","rationale":"The reader's weakest assumption correctly identifies the engineered Google Trends signal as the load-bearing point for the realism claim: the paper explicitly post-processes the public signal so that it correlates with oracle demand, and the reported ρ=0.972 makes the signal nearly as informative as hidden demand. This is not an internal inconsistency, but it is a correctness risk for the central claim that high scores reflect transferable business intelligence. The within-arena evidence remains strong: mechanism ablations, split-half stability (ρ=0.898), and ICC=0.944 support reliable and purposeful measurement inside the simulator. However, those results cannot validate the realism of the signal-construction choice. Since the reader already conditioned acceptance on clarifying the signal construction and releasing artifacts, no verdict adjustment is needed; the conditional stance is the right one. I agree with the reader's identification of the weakest assumption, and the proposed test would settle whether the concern actually lands.","tokens_in":23658,"tokens_out":2851,"duration_ms":32888,"concrete_test":"Rerun the demand-inference ablation and the top-model leaderboard with raw, unprocessed Google Trends data, or with controlled noise added to the post-processed signal so that pooled correlation drops from 0.972 to a realistic 0.5–0.7. Keep all other arena mechanics and seeds fixed; compare evidence-guided, blind bulk, and cheapest-SKU policies, plus Gemini 3.1 Pro and GPT-5.6 Sol. If the +$63.6k evidence advantage collapses or the model-vs-expert gap narrows materially, the realism claim is unsupported. Also report the raw Google-Trends-to-oracle correlation and the exact post-processing transformation used in Appendix C.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 6.5's central claim is that higher arena scores reflect 'genuine business intelligence rather than simulator-specific shortcuts.' This requires that agent-visible evidence be representative of real public signals, not manufactured to match hidden state. Appendix C states: 'We conduct minor post-processing of google trends data so that the results is correlated with oracle demand...' Appendix G reports pooled ρ=0.972 for public evidence and 93.3% country-category identification. The demand-inference ablation shows evidence-guided sourcing beats blind bulk buying by $63.6k. But this advantage is measured with a signal engineered to be nearly as informative as oracle demand. The post-processing is unexplained: no raw-correlation baseline, no noise model, no temporal or category structure. In real markets, Google Trends correlations with demand are substantially lower, time-lagged, and category-dependent. If the signal were cleaned to a realistic noise level, the capability gap and the claimed 'substantial headroom' (best model $188k versus expert $436k) could shrink or reorder. The paper therefore establishes internal validity but not the realism component of its central claim; the trustworthiness claim for real-world transfer remains conditional on unvalidated signal construction.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"Business Arena is a simulated cross-border B2B marketplace in which an LLM agent operates a shop over 30 simulated days, with supplier data grounded in Alibaba.com and demand, tariffs, and events calibrated from public sources. The paper evaluates 15 frontier models across 10 runs each, reporting a 9.0x spread in mean final net worth, with 51% of runs losing money, and compares model performance to a library of expert-designed strategies. It contributes diagnostic skill-level metrics, action-level attribution, stateful checkpointing, and nine mechanism ablations. The central claims are that the arena preserves the challenges of real business and that higher arena scores reflect genuine business intelligence rather than shortcuts, leaving large headroom between frontier LLM agents and expert strategies.","tokens_in":23929,"tokens_out":4844,"duration_ms":50556,"significance":"If the realism claim were fully supported, this would be a valuable benchmark for long-horizon, open-ended agent evaluation: it combines noisy evidence, delayed feedback, non-stationarity, and persistent obligations in one testbed, and the evaluation methodology is careful. The reliability analysis (ICC = 0.944 for ten-run means; split-half rho = 0.898) and the nine mechanism ablations with negative controls are strengths, as are the stateful save-fork-load pipeline and action-level attribution. However, the external validity of the demand signal and the reference-strategy headroom estimate are not yet established; the within-benchmark evidence is strong, but the realistic and trustworthy claim remains conditional.","major_comments":[{"comment":"The demand signal is post-processed to correlate with oracle demand: Appendix C states 'We conduct minor post-processing of google trends data so that the results is correlated with oracle demand,' and Appendix G reports a pooled correlation of 0.972 with 93.3% country-category identification. This engineered correlation means the evidence-guided sourcing advantage of +$63.6k in Table 1 is measured under a public signal nearly as informative as the hidden demand state. The manuscript does not report the raw Google Trends correlation before post-processing, nor does it provide a noise model, temporal lag structure, or category-specific degradation. Because real public demand signals are substantially noisier and time-lagged, the measured capability gap and the claimed headroom could shrink or reorder if the signal were cleaned to a realistic noise level. The authors should add a raw-correlation baseline, a sensitivity analysis across noise levels, or validation against out-of-sample real demand series before claiming that the arena preserves the noisy-evidence challenge. This issue is load-bearing for the Section 6.5 conclusion that higher scores reflect genuine business intelligence.","section":"Appendix C and Appendix G, Table 1"},{"comment":"The expert-designed strategies, including the Bayesian export compounder reaching $436,195, are authored by the same team with full knowledge of the arena's mechanisms and NPC behavior. Although they operate only on agent-visible information, this reference standard may encode privileged knowledge of which mechanisms matter, making the substantial-headroom conclusion partly internal to the authors' design. The paper should test the robustness of the headroom estimate, for example by using independently authored strategies or by varying NPC population and world seeds and showing that the expert advantage is stable. Without such evidence, the conclusion that frontier models fall substantially behind human-designed strategies is not independently validated.","section":"Section 6.5 and Appendix J"},{"comment":"The mechanism ablations compare hand-coded intended policies against hand-coded neglect and misuse variants; they demonstrate that the arena rewards the intended mechanisms, but they do not directly establish that the observed LLM leaderboard differences are caused by those mechanisms. The paper states that 'these ablations support interpreting higher Business Arena scores as evidence of stronger business intelligence,' yet the policies that earn higher scores in the ablations are not the LLM agents. To bridge this gap, the authors should explicitly acknowledge this limitation and, where possible, correlate model-level scores with the corresponding skill-level metrics (e.g., demand-inference accuracy or pricing precision) to show that the same mechanisms drive both the ablation ladder and the model ranking.","section":"Section 6.5 and Appendix G"}],"minor_comments":[{"comment":"The sentence 'We conduct minor post-processing of google trends data so that the results is correlated with oracle demand' contains a grammatical error; 'results is' should be 'results are'.","section":"Appendix C"},{"comment":"The word 'capitol' appears in 'capitol commitments' (Section 4) and 'redeploy capitol' (Appendix J); both should be 'capital'.","section":"Section 4 and Appendix J"},{"comment":"The text mentions '10 baselines (one per archetype, used for alpha computation)' but never defines what alpha is; please specify this quantity or remove the reference.","section":"Appendix D"},{"comment":"The sentence 'The contrast shows the what we expect from models' is ungrammatical and should be rephrased, for example to 'The contrast shows what we expect from models.'","section":"Appendix M"},{"comment":"Reference [4] is listed as 'arXiv preprint, 2026' without author names; please complete the citation so readers can locate the work.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a strong fit for an agent-evaluation venue, and the within-benchmark methodology is careful. The primary risk is overclaiming realism given the engineered demand signal; I would condition acceptance on the additional analyses described in major comments 1 and 2, and on an explicit treatment of the scope of the mechanism-ablation evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know before you read it. First, Business Arena is a genuinely new contribution: a full B2B cycle with noisy evidence, delayed feedback, compliance obligations, and a serious diagnostic toolkit (skill-level metrics, action-level attribution, stateful fork/rollback). No existing benchmark puts all of that together. Second, the paper's realism claim is softer than the abstract suggests because the Google Trends signal is post-processed to correlate with hidden demand (Appendix C, pooled rho=0.972 in Appendix G). The internal measurements are sound, but the external-transfer part is not established.\n\nWhat the paper does well is substantial. The environment is detailed and grounded in real sourcing and trade data. The evaluation is unusually careful: 150 runs, ICC=0.944 for ten-run means, split-half rank stability at 0.898, and a set of mechanism ablations where the intended policy beats neglect and misuse across nine mechanisms. That is real evidence that the score rewards the intended capabilities rather than simulator exploits. The expert-designed strategy reserve is a good idea for estimating opportunity, and the skill profiles do reveal distinct operating styles. The action-level attribution seems technically sound and useful for future training-data construction.\n\nSoft spots, in order of importance. The demand-signal post-processing is the biggest one. The paper says \"minor post-processing\" but provides no raw baseline, noise model, or temporal structure, so the agent-visible signal is nearly as informative as the oracle. This makes the demand-inference mechanism cleaner than real markets, and it means the \"realistic and trustworthy testbed\" phrasing overclaims. It does not kill the internal validity: the ablations still show that evidence-guided behavior beats neglect, and the headroom result might even widen with noisier real signals, but we simply don't know. Second, there is no external validation that arena success transfers to live business; the authors admit this in the Limitations section, so it's less a hidden flaw than an unaddressed question. Third, no code or data appear to be released, which is a real practical problem for a benchmark paper. Fourth, the expert-designed strategies are written by the same team with full knowledge of the arena, so the headroom estimate is an internal reference, not an absolute ceiling.\n\nWho is this for? Researchers working on LLM agents for business or long-horizon economic decision-making, and anyone developing evaluation methodology for open-ended environments. It deserves a serious referee. I'd send it to peer review, but with a request for artifact release and a proper sensitivity analysis on the demand signal construction.","headline":"A serious benchmark for end-to-end business agents with solid internal reliability, but the realism claim is undercut by an engineered demand signal and the lack of released artifacts; worth engaging and sending to peer review.","tokens_in":24453,"tokens_out":2772,"would_cite":true,"duration_ms":32438,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Business Arena claims to measure business intelligence in a realistic simulated marketplace, and finds the best LLM agents still fall far behind expert strategies.","keywords":["LLM agents","business intelligence","agent benchmark","marketplace simulation","cross-border e-commerce","action attribution","mechanism ablation","long-horizon evaluation"],"falsifier":"Run the same benchmark with the raw, unprocessed Google Trends series substituted for the post-processed signal: if evidence-guided policies no longer beat the blind baseline by tens of thousands of dollars in final net worth, the claim that scores reflect genuine business intelligence rather than an engineered signal would be undercut.","tokens_in":23459,"feed_emoji":"🏪","tokens_out":9387,"duration_ms":84363,"temperature":0.7,"pith_summary":"Business Arena is a simulated cross-border B2B marketplace where an AI agent runs a shop for 30 simulated days, buying from suppliers, setting prices, serving buyers, and satisfying compliance obligations before trading. The paper's central claim is that this environment preserves the four challenges that make real business difficult—noisy evidence, delayed and coupled feedback, a changing market, and persistent obligations—and that strong performance in the arena reflects genuine business intelligence rather than simulator-specific shortcuts. The claim matters because existing agent benchmarks measure fixed, verifiable tasks, while running a business is open-ended, capital-at-risk, and has no single correct answer. If the claim is right, Business Arena offers a controlled way to measure and improve end-to-end business agents before they are trusted with real money.","feed_headline":"Best AI business agent trails an expert strategy by 2.3x","feed_subtitle":"In a 30-day simulated shop, the top model earned $188,488 while an expert-designed playbook earned $436,195.","key_machinery":"The load-bearing object is the arena itself: a 30-day, data-grounded marketplace simulator with hidden state, scripted competitor archetypes, a financial ledger, compliance gates, and a save-fork-load stateful evaluation harness. Around this core, the diagnostic stack—an expert-designed strategy reserve that estimates available opportunity, skill-level metrics that expose operating styles, action-level attribution that traces gains and losses to decisions, and mechanism ablations that test intended behavior against neglect and misuse—is what turns a single profit number into evidence about business capability. The four design mechanisms of noisy evidence, delayed consequences, a changing market, and persistent obligations are what make the scores interpretable as business intelligence rather than as routine workflow completion.","core_discovery":"Business Arena places one LLM-operated shop inside a simulated cross-border B2B marketplace with 965 supplier offers across 135 SKUs, more than 60 tools, a persistent workspace, and scripted competitor sellers, over a 30-day episode. The central discovery is that the environment separates models sharply: mean final net worth across 15 frontier models ranges from $188,488 to $20,856, a $9.0\\times$ gap, and 51% of runs finish below the $80,000 starting capital. The strongest expert-designed strategy reaches $436,195 in the same world, more than twice the best model mean, showing available opportunity the agents are not yet capturing. Mechanism ablations over portfolio selection, market events, pricing, tariffs, and customer service show that evidence-guided behavior beats both neglect and misuse, which the paper takes as evidence that higher arena scores reflect genuine business intelligence rather than simulator-specific shortcuts.","pith_inferences":["Beyond the paper, a natural stress test is to replace the post-processed Google Trends signal with raw, unprocessed public data; because Appendix C reveals the signal is engineered to correlate with hidden demand, the measured capability gap may shrink under real-world noise.","Beyond the paper, the action-level attribution chains could be used directly as dense credit assignment for reinforcement learning, converting each sourcing, pricing, and recovery action into a reward signal.","Beyond the paper, a normalized opportunity-capture score—model final net worth divided by the best expert-strategy final net worth—would make the benchmark comparable across different world seeds and future arena versions.","Beyond the paper, the arena currently tests one business setting (cross-border B2B sourcing), so the generality of the business-intelligence claim to retail, services, or manufacturing remains an open question."],"forward_implications":["Frontier LLM agents are not yet reliable enough for autonomous business operation: 51% of runs lose money, and only four of fifteen models preserve starting capital in every trial.","The best model mean ($188,488) trails the best expert-designed strategy ($436,195) by more than a factor of two, indicating substantial headroom in end-to-end business operation.","Mechanism ablations show that evidence-guided behavior outperforms both neglect and misuse across five tested mechanisms, supporting the interpretation that higher scores reflect genuine business intelligence.","Skill-level metrics reveal stable operating styles—premium sellers, volume wholesalers, customer-service specialists—that a single final score hides, making targeted diagnosis and training-data construction possible.","Stateful save-fork-load evaluation enables test-time compute scaling: five-day trace search beat daily search and independent runs under the same rollout budget, evidence that the arena's feedback is genuinely delayed."],"supporting_citations":[{"why":"supplies the standard short-horizon software-agent benchmark that Business Arena contrasts with open-ended long-horizon operation.","marker":"[7]"},{"why":"supplies a realistic web-agent benchmark used as a comparison point for what agent evaluation has covered.","marker":"[17]"},{"why":"represents a narrower business workflow benchmark that lacks the full sourcing-to-compliance loop.","marker":"[2]"},{"why":"represents another partial business benchmark whose coverage gaps motivate the arena's end-to-end design.","marker":"[14]"},{"why":"supplies a long-term planning benchmark that Business Arena positions against for coordinated business execution.","marker":"[6]"},{"why":"covers a simulated company management scenario but omits the physical-commerce cycle that Business Arena adds.","marker":"[3]"},{"why":"supplies the intraclass-correlation method the paper uses to argue that ten-run leaderboard means are reliable.","marker":"[9]"},{"why":"supplies the agent-based computational economics perspective that justifies heterogeneous market dynamics in the arena.","marker":"[13]"},{"why":"grounds the arena's tariff rates in authoritative trade data, a load-bearing component of the realism claim.","marker":"[27]"},{"why":"provides the Google Trends demand signal whose post-processing underlies the demand-inference ablations.","marker":"[22]"}],"fun_headline_variants":["Best AI business agent trails expert playbook by 2.3x in simulation","Business Arena: top LLM earns $188K, expert strategy $436K","AI agents in business: 9x spread, best still lags human strategy","Simulated shop: AI best profit $188K, expert $436K, 2.3x gap","LLM agents fall short of human-designed business strategies in arena"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that higher arena scores reflect genuine business intelligence depends on hand-set simulator parameters and on a post-processed Google Trends signal that is deliberately correlated with hidden demand; if that signal were not so clean, evidence-guided agents might not outperform blind ones, and the measured capability gap could shrink.","fun_headline_variants_meta":{"raw":{"variants":["Best AI business agent trails expert playbook by 2.3x in simulation","Business Arena: top LLM earns $188K, expert strategy $436K","AI agents in business: 9x spread, best still lags human strategy","Simulated shop: AI best profit $188K, expert $436K, 2.3x gap","LLM agents fall short of human-designed business strategies in arena"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000288,"raw_usage":{"total_tokens":1725,"prompt_tokens":1019,"completion_tokens":706,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":635,"completion_tokens_details":{"reasoning_tokens":598}},"tokens_in":635,"tokens_out":706,"duration_ms":7323,"temperature":1.0,"reasoning_tokens":598,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:29:05.329613+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same benchmark with the raw, unprocessed Google Trends series substituted for the post-processed signal: if evidence-guided policies no longer beat the blind baseline by tens of thousands of dollars in final net worth, the claim that scores reflect genuine business intelligence rather than an engineered signal would be undercut.","supporting_citations":[{"cited_title":"SWE-bench: can language models resolve real-world GitHub issues? InICLR","cited_arxiv_id":null,"evidence_quote":"supplies the standard short-horizon software-agent benchmark that Business Arena contrasts with open-ended long-horizon operation."},{"cited_title":"WebArena: a realistic web environment for building autonomous agents","cited_arxiv_id":null,"evidence_quote":"supplies a realistic web-agent benchmark used as a comparison point for what agent evaluation has covered."},{"cited_title":"ShoppingBench: a real-world intent-grounded shopping benchmark for LLM-based agents.AAAI, 2024","cited_arxiv_id":null,"evidence_quote":"represents another partial business benchmark whose coverage gaps motivate the arena's end-to-end design."},{"cited_title":"Agent-based computational economics: a constructive approach to economic theory","cited_arxiv_id":null,"evidence_quote":"supplies the agent-based computational economics perspective that justifies heterogeneous market dynamics in the arena."},{"cited_title":"World integrated trade solution: bilateral tariff technical note","cited_arxiv_id":null,"evidence_quote":"grounds the arena's tariff rates in authoritative trade data, a load-bearing component of the realism claim."},{"cited_title":"Google trends","cited_arxiv_id":null,"evidence_quote":"provides the Google Trends demand signal whose post-processing underlies the demand-inference ablations."}],"review_version":1}