{"id":"729dc98c-027c-4d49-8b69-ec0ba1359774","arxiv_id":"2507.09083","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"LLM bidders replicate several classic auction experimental regularities at roughly three orders of magnitude lower cost than human experiments, supporting but not yet validating their use as proxies.","lead":"Researchers ran more than 1,000 simulated auctions in which GPT-4 agents bid against each other, and found the agents reproduced several classic human auction behaviors: overbidding in first-price auctions, better play in clock auctions, and the winner's curse in common-value auctions. The results suggest cheap AI agents could serve as test subjects for auction design, though they also diverge from humans in important ways.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The proxy-validity claim depends on LLM behavior being emergent strategic reasoning; the paper never tests whether its agreement with human auction experiments is trained-in memorization or semantic priming, so the central claim is not yet supported.","rationale":"Reading the paper in good faith, the authors are transparent about their protocol, robustness checks, and mismatches (e.g., SPSB underbidding versus Kagel-Levin overbidding; the AC versus AC-B null result). The framework and cost advantages are real, and the paper is best understood as a proof-of-concept. However, the proof-of-concept status depends entirely on whether the observed agreement with the literature reflects the model's in-context response to the economic environment or retrieval of benchmark conclusions from pretraining. The paper does not provide a test that separates these. The reader identified the same weakest assumption, and the proposed decontextualization check would settle the semantic-priming branch of that concern with the existing infrastructure. Because the reader already made this the basis for a CONDITIONAL verdict, my read does not change the verdict; conditional acceptance remains appropriate pending this or an equivalent decontamination test.","tokens_in":35832,"tokens_out":8159,"duration_ms":109877,"concrete_test":"Decontextualization test: re-run the three flagship comparisons (FPSB vs SPSB in IPV, sealed-bid vs clock in APV, and common-value SPSB for n=2-6) with the same plan-bid-reflect loop and payoff structure, but replace all auction vocabulary with neutral game language (e.g., 'bid' becomes 'chosen number,' 'value' becomes 'points,' 'second-highest bid' becomes 'the lower of the two other submitted numbers,' 'clock' becomes 'exit-threshold format'). Use the same GPT-4 configuration, temperature, and sample sizes. If the qualitative patterns persist under neutral framing, emergent strategic reasoning is supported. If the patterns weaken, invert, or disappear, the headline agreement is an artifact of auction-specific semantic priming from training data. Report the neutral-language prompts and pooled bid data alongside the main results.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that GPT-4/COT bidding is a valid low-cost proxy for human auction behavior because it reproduces established regularities: FPSB bids above the risk-neutral benchmark, clock formats induce closer-to-truthful play, and common-value SPSB auctions show the winner's curse (Sections 3.1.5, 3.2.4, 3.3.4). The load-bearing premise is that these regularities emerge from context-sensitive strategic reasoning rather than from the model having memorized the very experimental literature used as benchmarks (Kagel and Levin 1986/1993, Li 2017, Roth and Ockenfels 2002). The paper never tests this premise. Its protocol (Section 2.1) avoids persona anchoring, but it retains the full auction vocabulary and the exact benchmark settings, which can activate memorized content. The indirect evidence is ambiguous: Appendix D's out-of-the-box ablation and Section 5's intervention results show that behavior is highly sensitive to framing, which is consistent with semantic priming rather than with a stable, emergent bidding disposition. Because the intended use of the proxy is to learn about new auctions and designs, reproducing memorized results would not validate that use. This is therefore a correctness risk for the central claim, not a minor caveat.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes using large language model (LLM) agents as low-cost proxies for human bidders in auction experiments. The authors run a large number of simulated auctions with GPT-4 agents under chain-of-thought reasoning, covering first-, second-, third-price, all-pay, ascending clock, common-value, and eBay-style auctions. They report that LLM bidders overbid relative to the risk-neutral benchmark in first-price auctions, bid closer to truth in clock formats than in sealed-bid formats, suffer the winner's curse in common-value settings, and exhibit strategic behaviors such as sniping in eBay-style auctions. They also study how different prompt interventions affect bidding accuracy. The central claim is that LLM agents reproduce key empirical regularities from the human experimental auctions literature, thereby serving as a valid and cheap synthetic lab.","tokens_in":36096,"tokens_out":7253,"duration_ms":90498,"significance":"If the proxy-validity claim were established, this would be a useful methodological contribution: it would let researchers explore auction designs at a fraction of the cost of human experiments, and the open-source framework would enable broader use. The paper's strengths are its breadth of auction formats, the explicit planning-and-reflection protocol, the cost transparency, and the inclusion of robustness checks on prompts, languages, and currencies. However, the abstract and introduction overstate the agreement with the human literature, and the paper does not address the most important threat to its central claim: that GPT-4's training data includes the very experimental results used as benchmarks, so the observed agreement could reflect memorization rather than emergent strategic reasoning. This unresolved issue, together with an internal contradiction in the reported language-robustness results, means the paper currently supports a proof-of-concept but not yet the stronger proxy-validity claim.","major_comments":[{"comment":"The abstract and introduction state that LLM bidders 'agree with the experimental literature in auctions' and specifically 'produce results consistent with risk-averse human bidders.' However, Section 3.1.5 and Table 2 show that in the SPSB auction, LLMs underbid their value 65.78% of the time, while the paper itself notes that Kagel and Levin (1993) report humans overbid 67% of the time. The direction of the dominant deviation is therefore opposite to the human data, not merely different in magnitude. This is a direct contradiction of the unqualified 'agree' claim, and it is load-bearing because the paper uses this agreement as evidence that LLMs can serve as proxies. The authors should either substantially qualify the central claim in the abstract and introduction or reframe the SPSB result as a systematic difference between LLM and human behavior that itself requires explanation before proxy validity can be assumed.","section":"Abstract; Section 3.1.5, Table 2"},{"comment":"The central proxy-validity claim rests on the premise that GPT-4's bidding behavior reflects emergent, context-driven strategic reasoning rather than memorized reproduction of the auction-theory and experimental-literature content in its training data. The paper never tests this premise. The protocol (Section 2.1) intentionally avoids persona anchoring, but it still uses the exact benchmark settings and vocabulary (the same value ranges, the same APV and common-value structures, and descriptions that echo the experimental papers). Appendix D show that removing the plan-bid-reflect loop produces qualitatively different behavior (e.g., no overbidding in SPSB), which is consistent with the alternative explanation that the observed regularities are sensitive to priming and surface cues rather than reflecting a stable reasoning disposition. Because the intended use of the proxy is to learn about new auction designs, reproducing memorized results would not validate that use. The paper should include a control that distinguishes these hypotheses, for example: (i) auctions with novel formats or payment rules that are unlikely to appear verbatim in training data, (ii) paraphrase or distractor tests that remove auction-theory vocabulary while preserving the incentive structure, or (iii) an analysis of whether the LLM's stated reasoning in chain-of-thought actually derives the equilibrium logic rather than citing known results. Without such a control, the agreement with human experiments remains ambiguous evidence for proxy validity.","section":"Section 2.1; Section 3.1.5; Appendix D"},{"comment":"The abstract claims that LLMs are 'not very sensitive to naive changes in prompts (e.g., language, currency),' but the robustness check in Appendix C.3 reports that the Russian-language FPSB prompt produces a stark outlier: agents bid close to their true value rather than shading bids as in the English version. This is a qualitative change in the key treatment, not a small quantitative difference. The paper should either remove the overgeneralized claim, or report this outlier prominently in the main text and explain why it does not undermine the conclusion that behavior is robust to prompt variation. The current presentation gives the reader an inaccurate impression of the evidence.","section":"Section C.3, Figure 11; Abstract"}],"minor_comments":[{"comment":"The sentence 'In particular, 90bids exceeded their value in this auction' appears to be missing a percentage sign or a word (likely '90% of bids'); please correct this typo.","section":"Section 3.1.4"},{"comment":"The claim that LLMs replicate sniping 'with no prompting' is somewhat overstated, because the prompt explicitly tells bidders that they may place bids on the final day and that no bidder knows if they are the last; this is a design cue that likely encourages late bidding. Please phrase this as 'with a standard rules description that mentions final-day bidding' rather than 'with no prompting.'","section":"Section 4.1; Section 4.3"},{"comment":"The text says the risk-neutrality intervention 'has about the same performance as the proxy intervention,' but Table 7 reports R^2 = 0.4004 for the risk intervention, which is lower than the no-intervention R^2 = 0.4845 and much lower than the proxy intervention's 0.2966 (or the Nash intervention's 0.8365). The qualitative description should be tied more carefully to the reported fit statistics, or the table should be annotated to clarify that the comparison is limited to the overbidding component.","section":"Section 5.1, Table 7"},{"comment":"Several tables report pairwise t-tests or chi-square tests without any correction for multiple comparisons or clustering of bids within the 5 experiments per condition. Because the LLM agents share a common history across rounds, the effective sample size is smaller than the number of bids. Please either report cluster-robust standard errors or explicitly acknowledge this limitation in the statistical analysis.","section":"Section 3.1.5 and Section 5 use multiple statistical tests"}],"recommendation":"major_revision","confidential_remarks":"The paper is a good proof-of-concept with careful implementation and transparent cost reporting, but the central proxy-validity claim is currently supported only if one ignores the memorization alternative and the paper's own SPSB underbidding result. The memorization concern is the most important issue and should be addressed with a direct control experiment before publication. The authors' decision to include an explicit discussion of the SPSB underbidding difference is commendable, but the abstract does not reflect that nuance. I recommend major revision rather than rejection because the needed additions (a novelty control, a qualification of the overclaims, and a closer treatment of the Russian-language outlier) are within the scope of the manuscript."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Thanks for sending this. Quick take: it's a useful proof-of-concept with transparent reporting, but the headline claim overreaches. The paper's own SPSB data shows LLMs underbid 66% of the time, while Kagel-Levin humans overbid 67% - that's the opposite error. The abstract says they 'agree with the experimental literature,' which is only true in direction for FPSB and clock formats, not for the SPSB failure mode. To the paper's credit, they acknowledge this mismatch in the text, but the abstract still overstates.\n\nWhat's new: the systematic benchmark across FPSB/SPSB/TPSB/all-pay, the APV clock comparisons, the common-value winner's curse tests, the eBay-style simulations showing soft-close reduces sniping, and the six prompt interventions. The eBay results are genuinely new and interesting. The prompt intervention results are also useful for thinking about how description affects play.\n\nWhat the paper does well: the protocols are described in enough detail to replicate, the robustness checks (language, currency, one-shot) are thoughtful, and they openly report where LLMs diverge from human data (SPSB underbid, AC-B not differing from AC). That honesty earns credit.\n\nSoft spots, in proportion. The big one is the memorization concern. The load-bearing claim is that LLM bids reflect emergent strategic reasoning, not recitation of the benchmark literature. The paper never tests this. The out-of-the-box ablation and the prompt-sensitivity results suggest behavior is heavily framing-dependent, which cuts both ways, but the authors don't engage with the contamination question. An easy control would be to debias or to test on newer less-seen settings. Until that's done, the 'proxy validity' claim is conditional.\n\nSecond, the sample sizes are small - five experiments per format, 15 rounds, three bidders. The winner's curse results rely on boxplots from what looks like a handful of auctions per n. That's enough for a proof-of-concept, but not for the confidence in the abstract.\n\nThird, the code and data are promised but not yet available. For a paper whose main deliverable is a framework, that's a real issue.\n\nOverall: I think this deserves peer review. The empirical work is careful, the eBay findings are new, and the framework is useful. But the central claim needs to be pulled back to 'LLM behavior is consistent with some human regularities' rather than 'LLMs are valid proxies.' Suggested major revision, with a contamination control and code release.\n\nWould I cite? Not until the code and the control test exist. Bring to reading group though - good discussion of what LLM experiments can and can't tell us.","headline":"Useful proof-of-concept with honest reporting, but the proxy-validity claim is overstated and the memorization concern is untested.","tokens_in":36618,"tokens_out":2095,"would_cite":false,"duration_ms":24777,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91B26","91A90"],"pacs":[],"model":"deepseek-v4-flash","headline":"LLM bidders reproduce classic human auction behaviors — risk aversion, clock-auction clarity, and the winner's curse — in a synthetic lab.","keywords":["large language models","auction experiments","synthetic labs","chain-of-thought reasoning","winner's curse","obviously strategy-proof","mechanism design","GPT-4"],"falsifier":"Run the same protocol on a novel auction format with no published human benchmark, and record whether the model's chain-of-thought references specific known studies when it produces human-typical bids; absence of such references on novel formats, together with preserved behavioral patterns, would support emergent reasoning, while patterns that only appear for formats covered in the training literature would point to recitation.","tokens_in":35654,"feed_emoji":"🔨","tokens_out":5787,"duration_ms":68255,"temperature":0.7,"pith_summary":"This paper tries to establish that large language models, when given chain-of-thought reasoning and a plan-bid-reflect loop, can act as plausible stand-ins for human bidders in auction experiments. It finds that GPT-4 agents bid above the risk-neutral prediction in first-price auctions, matching risk-averse human behavior, bid closer to true value in ascending clock auctions than in strategically equivalent sealed-bid auctions, and fall victim to the winner's curse in common-value settings. If these results hold, auction researchers could test designs and interventions at a cost roughly three orders of magnitude lower than running human experiments. The paper runs over a thousand GPT-4 auctions for less than $400 and provides a codebase to extend the approach to any auction format an experimenter can describe.","feed_headline":"LLM agents replicate key human auction behaviors for under $400","feed_subtitle":"Chain-of-thought GPT-4 bidders match decades of experimental auction findings in a testbed costing 1,000 times less.","key_machinery":"The central machinery is the synthetic-lab protocol: a plan-bid-reflect loop in which each GPT-4 bidder independently writes a bidding plan, places a bid given its private value, observes the outcome and all bids, then reflects and revises its plan before the next round. This loop, powered by chain-of-thought prompting (asking the model to reason step-by-step before answering), is what generates the behavior under study. Its low per-round cost lets the authors run more than a thousand auctions and compare aggregate bid patterns against theoretical equilibria and published human data.","core_discovery":"The central claim is that LLM agents, when endowed with chain-of-thought planning and reflection, reproduce the main behavioral regularities of human auction experiments across multiple classic formats. In independent private-value auctions, their first-price bids sit above the risk-neutral Bayes-Nash equilibrium, consistent with the risk-averse overbidding documented in human subjects, while their second-price bids skew below value, an asymmetry not found in the human data. In affiliated private-value settings, LLM bidders drop out near their value in ascending clock auctions far more often than they bid their value in the strategically equivalent second-price sealed-bid auction, mirroring the obvious strategy-proofness effect. In common-value auctions, they bid as if their private signal were an unbiased estimate of the common value and suffer a winner's curse that deepens as the number of bidders grows. The paper also shows that naive prompt changes, such as language and currency, matter little, while prompting with the language of Nash deviations substantially improves play, and that an extended closing rule eliminates last-second sniping in an eBay-style environment.","pith_inferences":["A direct test of the emergent-reasoning interpretation would be to run the same protocol on a novel auction format with no published human benchmark; if the model still displays human-typical biases there, the proxy claim is strengthened.","The framework's flexibility suggests the same plan-bid-reflect loop could be applied to other mechanism-design settings, such as matching markets or contests, to generate behavioral data before expensive human validation.","The contrast between LLM underbidding and human overbidding in second-price sealed-bid auctions hints at a model-specific bias that prompt engineering might adjust, potentially letting LLM proxies match not just the magnitude but the direction of human error.","Because the paper runs multi-round auctions, the observed overbidding may be partly an artifact of the agent's explicit instruction to explore and gather data, rather than a stable preference; single-round ablation results in the appendix suggest this is the case."],"forward_implications":["If LLM proxies are valid, auction researchers can screen designs and interventions for about $10 per experiment instead of $15,000, making large parameter sweeps feasible.","The similarity with human risk aversion and the winner's curse means LLM agents can generate synthetic data on behavioral deviations, not just rational benchmarks.","The finding that clock formats improve truth-telling suggests that obvious strategy-proofness design principles carry over to LLM agents, and that prompting interventions can be tested in silico before human trials.","The eBay closing-rule result offers a low-cost path to revisit classic marketplace design debates, such as the competing ending rules that distinguished Amazon and eBay auctions.","Naive prompt robustness, across languages and currencies, increases confidence that the observed patterns are not artifacts of a specific wording choice."],"supporting_citations":[{"why":"Supplies the obvious-strategy-proofness framework and the APV experimental results that the clock-auction comparisons are designed to replicate.","marker":"Li (2017)"},{"why":"Experimental reference for the winner's curse and its dependence on number of bidders, which the common-value simulations corroborate.","marker":"Kagel and Levin (1986)"},{"why":"Provides the SPSB empirical benchmark for truth-telling and overbidding rates that the LLM second-price results are compared against.","marker":"Kagel and Levin (1993)"},{"why":"Documents risk-averse overbidding in first-price auctions across a large set of experiments, matched by the LLM findings.","marker":"Cox et al. (1988)"},{"why":"Provides the eBay versus Amazon closing-rule comparison that motivates the sniping and auction-extension treatments.","marker":"Roth and Ockenfels (2002)"},{"why":"Supplies the clock-description result and the AC-versus-AC-B comparison used in the obviously strategy-proofness section.","marker":"Breitmoser and Schweighofer-Kodritsch (2022)"},{"why":"Supplies the menu-description intervention design that the intervention section adapts and tests.","marker":"Gonczarowski et al. (2023)"},{"why":"Contributes the profit-maximization prompt style and prior LLM auction setup that the agent instructions build on.","marker":"Chen et al. (2023)"},{"why":"Provides the stay-or-drop prompting pattern used in clock-auction rounds and the framing of LLMs as auction participants.","marker":"Fish et al. (2024)"}],"fun_headline_variants":["LLM bidders mimic human risk aversion in auctions","Chain-of-thought turns LLMs into human-like auction bidders","Auction experiments on LLMs: cheap, fast, human-like","GPT-4 bidders reproduce auction quirks for under $400","LLMs feel the winner's curse when prompted right"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that GPT-4's bids come from reasoning about the described auction context rather than from memorizing the results of the auction-experiment literature during training.","fun_headline_variants_meta":{"raw":{"variants":["LLM bidders mimic human risk aversion in auctions","Chain-of-thought turns LLMs into human-like auction bidders","Auction experiments on LLMs: cheap, fast, human-like","GPT-4 bidders reproduce auction quirks for under $400","LLMs feel the winner's curse when prompted right"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000187,"raw_usage":{"total_tokens":1345,"prompt_tokens":979,"completion_tokens":366,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":595,"completion_tokens_details":{"reasoning_tokens":278}},"tokens_in":595,"tokens_out":366,"duration_ms":4723,"temperature":1.0,"reasoning_tokens":278,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:05:26.962413+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same protocol on a novel auction format with no published human benchmark, and record whether the model's chain-of-thought references specific known studies when it produces human-typical bids; absence of such references on novel formats, together with preserved behavioral patterns, would support emergent reasoning, while patterns that only appear for formats covered in the training literature would point to recitation.","supporting_citations":[],"review_version":1}