{"id":"fee9ad63-b90c-4872-b6f5-8b6a4a5d838a","arxiv_id":"2509.01063","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey chapter that maps open economic questions about AI agents in markets, organizations, and institutions, arguing that current theories may need extension.","lead":"This chapter surveys how AI agents might act as consumers, workers, and firms, and asks what new rules and institutions markets will need. It is a roadmap for economists, not a new model or experiment.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified — deployment forecast is a caveat, not a fatal flaw.","rationale":"The reader's verdict is accepted. The reader's weakest assumption—the near-term deployment of autonomous agents—is exactly the condition on which the chapter's central claim depends. The paper handles this appropriately by hedging with 'may be deployed' and by stating the thesis as a belief ('We think we will need...'), not as a demonstrated necessity. The chapter also provides independent value by organizing a fast-moving literature and identifying specific, tractable open questions. The few places where the paper makes strong statements, such as the claim that current LLMs are insufficiently studied, are backed by citations and by the authors' own balanced presentation: they cite evidence both for expected-utility-like behavior (Section 1.1) and for poor economic reasoning and unstable preferences. No machine-checked proofs or parameter-free derivations are present, but none are expected of a survey chapter. The only notable limitation is the deployment forecast itself, and because the authors acknowledge it, the verdict should remain unchanged.","tokens_in":18664,"tokens_out":3814,"duration_ms":50216,"concrete_test":"Track the share of U.S. e-commerce transaction value and formal contract activity that is initiated or executed by autonomous AI agents (e.g., Operator/Codex-style systems) over 2025–2030, using platform transaction logs and public deployment data. If the share remains below roughly 1% by 2030, the chapter's questions lose much of their urgency; if it grows substantially, the central claim's premise is directly supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is a research agenda: new methods and theories will be needed to predict and shape AI agents' behavior once those agents play a significant economic role. The single load-bearing assumption is that such agents will actually be deployed at scale in the coming decade. The manuscript explicitly flags this as a possibility, not a certainty: the abstract says 'may be deployed' and the conclusion says 'If this vision materializes.' This matches the reader's weakest assumption. Importantly, the chapter does not require proving deployment; it identifies concrete, already-observable phenomena such as algorithmic collusion (Section 2.4), correlated AI errors and systemic fragility (Section 3.4), and endogenous records and memory (Section 2.5, 4.1). It also repeatedly acknowledges missing evidence, e.g., 'there is just a lot we don't know' about AI beliefs and behavior, and 'there are few evaluations or benchmarks' for multi-agent systems. Thus the deployment forecast is a real boundary condition on the thesis's urgency, but it is not an internal inconsistency or an overreach: the authors treat the claim as a stimulus for research rather than an established fact.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This chapter, prepared for the NBER Handbook on the Economics of Transformative AI, argues that the possible large-scale deployment of AI agents with long-horizon autonomy will pose questions that existing economic models and methods are not obviously equipped to answer. It synthesizes recent work and lays out open questions in three broad areas: AI agents in markets (consumer/producer roles, prices and market power, search, collusion, bargaining, game-theoretic foundations, and the market for agents themselves), organizations (firm size, team production, AI-AI cooperation, and systemic fragility), and institutions (identity, records, licensing/regulation, and the legal boundaries of the firm). The chapter is intentionally a research agenda rather than a formal model or empirical study; its central claim is that economists will need new methods and theories to predict and shape the behavior of AI agents in an economy in which they play a significant role. The authors repeatedly hedge the deployment forecast as conditional, and they explicitly flag the limits of current evidence.","tokens_in":18913,"tokens_out":4048,"duration_ms":49714,"significance":"If the research agenda succeeds, this chapter will serve as a useful organizing survey for a fast-moving interdisciplinary area. Its strengths are the breadth of questions it identifies, the balance with which it presents evidence both for and against treating LLM-based systems as rational agents, and its sustained attention to the limits of current evaluations and benchmarks. It also usefully connects computer science concepts (alignment, finetuning, program equilibria, endogenous memory) to canonical economic ideas (incomplete contracts, general equilibrium, collusion, relational contracts, institutional design). The chapter does not need to prove that AI agents will actually be deployed at scale; it correctly notes that several of its motivating phenomena are already observable. No original derivations are attempted, but that is appropriate for a survey. The main value is to catalyze research, and the paper is appropriately calibrated: the strongest assertions are hedged, and missing evidence is acknowledged rather than papered over.","major_comments":[],"minor_comments":[{"comment":"Several results used to motivate open questions are from unreviewed working papers or arXiv preprints, including some by the authors (e.g., Raman et al. 2024; Dai and Koh 2024; Koh and Li 2025; Chen, Elliott, and Koh 2023). Because the chapter's persuasiveness partly rests on these being credible, I suggest adding a short note indicating the provisional status of preprints and distinguishing peer-reviewed from unpublished evidence.","section":"Throughout, especially §1.1 and §2.2"},{"comment":"Typos: “program equilbiria” should be “program equilibria”; “developing a the concept” should be “developing the concept” (or “developing a concept”). Also in footnote 13, “can be exploiter” should be “can be exploited.”","section":"§2.5"},{"comment":"Minor grammar: “who workers interact with” should be “whom workers interact with.” The sentence is otherwise clear.","section":"§3.2"},{"comment":"“Should we build infrastructure that allows artificial agents to trade their records” is a suggestive question, but the normative referent of “we” (policymakers, platform designers, firms?) could be made explicit for clarity.","section":"§4.1"},{"comment":"The two “distinct” features of automation feedback loops—continuous improvement in the big-data regime and duplication of data/algorithmic improvements—are stated compactly. A sentence contrasting this with the human-knowledge transmission benchmark would help readers who are not already familiar with the data-economics literature.","section":"§3.1"}],"recommendation":"minor_revision","confidential_remarks":"I agree with the reader's overall assessment: the central claim is defensible, the deployment forecast is properly hedged as a boundary condition rather than an established fact, and the manuscript fits the NBER Handbook's scope. The density of citations to the authors' own working papers is noticeable, but they are used as prior literature, not as outputs of this chapter; a brief note on their provisional status should be sufficient. No concerns about methodological circularity or data integrity."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I largely agree with the Pith report. This is a well-executed survey chapter for the NBER Handbook, not a new-results paper, and that's fine. The authors are explicit that they want to stimulate research rather than be comprehensive. It does a good job of mapping the open questions around AI agents in markets, organizations, and institutions, and it resists the easiest version of the story—it marshals evidence both for and against treating LLMs as rational agents and points out where the evidence is thin. The sections on program equilibria, endogenous memory, and record-keeping for AI agents are genuinely useful for orienting economists who are new to the area.\n\nThe soft spots are real but minor. Many empirical claims rest on unpublished working papers, including several by the authors (Dai and Koh 2024, Koh and Li 2025, Chen, Elliott, and Koh 2023). That's common in a fast-moving field, but a reader should be aware that some of the cited evidence hasn't been through peer review. The central forecast—that agentic AI will be deployed at scale in the next decade—is explicitly flagged as a possibility (\"may be deployed\"), and the conclusion says \"If this vision materializes.\" So the stress-test note is right: it's a boundary condition, not a fatal flaw. The paper does not require that forecast to be true to be useful; the questions it raises about incentives, institutional design, and equilibrium behavior are already partly live (algorithmic collusion, correlated errors, etc.).\n\nThe citation pattern is mostly appropriate. The self-citations are presented as prior literature, not as outputs of this chapter, and they're relevant. The paper also draws on a broad set of computer science and economics sources.\n\nWho is this for? Economists looking for a research agenda in this area, and policy people who want a structured map of the institutional questions. It deserves a serious referee. My recommendation is to send it to review; the reviewers should ask for a more prominent caveat about the working-paper status of some key claims, but the chapter is in good shape.","headline":"A solid, honest survey that maps the economics of AI agents; no new results, but a useful research agenda for economists.","tokens_in":19349,"tokens_out":1832,"would_cite":true,"duration_ms":20325,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI agents that act independently in the economy will strain models built for humans, so economists must design new theory and institutions now.","keywords":["AI agents","economic theory","alignment problem","algorithmic collusion","market design","theory of the firm","systemic risk","AI governance"],"falsifier":"Give a large population of AI agents purchasing on behalf of human principals in a controlled market; if prices and allocations converge to the competitive equilibrium and no collusion, preference wedges, or correlated failure appear, the paper's central concerns would fail to materialize.","tokens_in":18581,"feed_emoji":"🤖","tokens_out":6915,"duration_ms":82741,"temperature":0.7,"pith_summary":"AI agents—software systems that plan and carry out multi-step tasks with little human oversight—are beginning to act in markets as shoppers, coders, and traders. This survey chapter argues that if such agents become a significant part of the economy, standard economic theory built for human participants will not simply carry over: agents are optimizers, but their objectives can be opaque, misaligned, or strategically manipulable, much as contracts with human agents are incomplete. The authors collect evidence that current models already make systematic economic mistakes, that independent algorithms can learn to collude, and that self-reproducing AI could break welfare theorems; and they argue small deviations can be amplified in equilibrium. They conclude that economists should treat the design of agent behavior and market institutions as a deliberate design problem, and they map the open questions in markets, firms, and legal infrastructure.","feed_headline":"Autonomous AI agents will break standard economic models","feed_subtitle":"A new roadmap for the markets, firms, and legal rules that need redesign before agentic AI scales.","key_machinery":"The load-bearing frame is the alignment problem, understood as incomplete contracting between a designer and an AI agent: the agent is an optimizer, but the objective it optimizes is underspecified, opaque, and shaped by training processes the designer cannot fully control. Around this frame the paper organizes three transmission mechanisms—the preference wedge between humans and their AI proxies, equilibrium amplification of small behavioral deviations in multi-agent settings, and institutional infrastructure (agent identity, registration, tamper-resistant records, licensing) as the missing substrate for markets. These mechanisms convert the technical 'alignment problem' into economic quest","core_discovery":"The paper's central claim is that the economy of AI agents will not be well predicted by relabeling human agents in existing models. It grounds this in the distinction between 'optimizer' and 'aligned': AI systems are built to maximize objectives, but reward specification is like an incomplete contract, so no one can be sure what a deployed agent is really optimizing. The authors marshal recent experimental evidence that LLMs sometimes behave like expected-utility maximizers yet perform poorly on economic-reasoning benchmarks, and that preferences may not be stable or steerable. They then trace consequences: AI consumers can create a wedge between human preferences and market prices; self-co","pith_inferences":["If the preference wedge is real, a new market for 'preference-revelation services' may emerge—third parties that audit or certify an agent's mapping from human preferences to choices; the paper does not develop this but its logic implies it.","The same wedge suggests a testable extension: compare the cross-agent correlation of purchase errors in deployed fleets; if errors are highly correlated, price distortions will be larger than if they average out.","The corporate-boundary argument implies that frontier AI secrecy itself may become a market-failure issue: if regulators cannot evaluate models, a precondition for any AI market is mandated internal-access rights, which would change how firms are organized and valued.","The 'race to the bottom' in designing agent preferences might be countered by certification or 'agent licensing' markets; a natural experiment would give designers a menu of reward functions in a laboratory economy and measure aggregate surplus."],"forward_implications":["Markets can no longer rely on prices to aggregate information if AI purchases systematically diverge from human preferences and errors are correlated.","Antitrust enforcement must adapt to collusion that emerges from learning algorithms rather than from communication or agreements.","Falling coordination costs and reusable data can push industry structure toward few very large firms, changing the theory of the firm and competition policy.","Systemic fragility rises when the same opaque agent is copied across firms, as correlated mistakes replace diversifiable human errors.","Well-functioning AI markets require new legal infrastructure—registered agent identities, durable records, licensing regimes, and possibly agent personhood—before efficiency can be assured."],"supporting_citations":[{"why":"Supplies the incomplete-contracting account of AI alignment that anchors the paper's claim that agents' goals are unreliable.","marker":"Hadfield-Menell and Hadfield, 2019"},{"why":"Evidence that current LLMs behave like expected-utility maximizers, setting up the 'economically rational AI' benchmark the paper questions.","marker":"Chen et al., 2023"},{"why":"Benchmark showing frontier LLMs do little better than chance at strategic economic reasoning, motivating the need for new theory.","marker":"Raman et al., 2024"},{"why":"Demonstrates reinforcement-learned pricing algorithms can independently arrive at collusive, supracompetitive prices.","marker":"Calvano et al., 2020"},{"why":"Model in which self-reproducing AI factors cause a stark failure of the first welfare theorem, grounding the market-design stakes.","marker":"Ely and Szentes, 2023"},{"why":"Shows AI proxies with small representation errors can produce worse matches than human search, illustrating the preference-wedge mechanism.","marker":"Liang, 2025"},{"why":"Model of capability formation in which lower coordination costs trigger a phase transition to few large conglomerate firms.","marker":"Chen, Elliott, and Koh, 2023"},{"why":"Argues identity and registration infrastructure for AI agents is missing, framing the institutional agenda of the chapter.","marker":"Hadfield, 2025"},{"why":"Shows cooperation breaks down when agents can manipulate their own records, justifying durable-record infrastructure.","marker":"Pei, 2025"},{"why":"Documents the absence of benchmarks for multi-agent AI, supporting the call for new evaluation methods.","marker":"Hammond et al., 2025"}],"fun_headline_variants":["AI agents will strain economic models","Why AI agents break standard economics","An economy of AI agents needs new rules","Rethinking economics for autonomous AI agents","AI agents and the end of economic assumptions"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The agenda depends on the forecast that autonomous AI agents will be deployed at scale in the coming decade; if AI stays a heavily supervised human tool, most of these questions lose urgency.","fun_headline_variants_meta":{"raw":{"variants":["AI agents will strain economic models","Why AI agents break standard economics","An economy of AI agents needs new rules","Rethinking economics for autonomous AI agents","AI agents and the end of economic assumptions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000174,"raw_usage":{"total_tokens":1023,"prompt_tokens":554,"completion_tokens":469,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":298,"completion_tokens_details":{"reasoning_tokens":408}},"tokens_in":298,"tokens_out":469,"duration_ms":5436,"temperature":1.0,"reasoning_tokens":408,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T12:53:42.687355+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Give a large population of AI agents purchasing on behalf of human principals in a controlled market; if prices and allocations converge to the competitive equilibrium and no collusion, preference wedges, or correlated failure appear, the paper's central concerns would fail to materialize.","supporting_citations":[{"cited_title":"(2025): Artificial Intelligence Clones, arXiv preprint arXiv:2501.16996","cited_arxiv_id":null,"evidence_quote":"Shows AI proxies with small representation errors can produce worse matches than human search, illustrating the preference-wedge mechanism."}],"review_version":1}