Pith. sign in

REVIEW 4 major objections 6 minor 77 references

People prefer an AI advisor they can veto, but letting an AI delegate act for them raises group surplus and even helps non-users.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 23:53 UTC pith:TWPA4GNA

load-bearing objection Useful modality-comparison experiment with a plausible mechanism, but the headline welfare numbers rest on an untested external baseline and don't survive the paper's own multiple-testing correction. the 4 major comments →

arxiv 2602.12089 v3 pith:TWPA4GNA submitted 2026-02-12 cs.GT cs.AIcs.HC

Choose Your Agent: Tradeoffs in Adopting AI Advisors, Coaches, and Delegates in Multi-Party Negotiation

classification cs.GT cs.AIcs.HC
keywords human-AI interactiondelegationbargainingwelfare externalitiesLLM agentspreference-performance misalignmentmulti-party negotiationrandomized experiment
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper runs a 243-person online bargaining experiment in which each three-person group plays the same negotiation game three times, once with each of three AI-assistance designs: an Advisor that recommends moves, a Coach that comments on your own move, and a Delegate that acts for you. All three are powered by the same strong LLM. The paper's central claim is a preference-performance misalignment: people say they prefer the Advisor, yet groups with Delegate access produce measurably higher collective surplus than all-human groups, while Advisor and Coach access do not. The Delegate's benefit comes roughly half from direct users and half from positive spillovers to people in the same group who never delegate. The paper argues the mechanism is a human filter: users edit or ignore AI suggestions in the Advisor and Coach modes, eliminating the very trades that create surplus, whereas delegation lets the AI's proposals go through unmodified.

Core claim

The paper's central discovery is that in a three-player bargaining game, delegating one's actions to an LLM agent raises group surplus relative to no AI, even though participants prefer an advisory interface that leaves them in control. The effect is not driven by adopters alone: non-users in Delegate groups also benefit because Delegate-generated offers are larger, asymmetric, and Pareto-improving, shifting the offer environment. The key contrast is that Advisor and Coach users filter AI suggestions through their own judgment—editing, overriding, or ignoring them—so final trades resemble the human baseline. The Delegate advantage comes from removing that filter, not from a better model.

What carries the argument

The experimental design holds the underlying LLM fixed and varies only the interaction contract: Advisor (proactive recommendation, user can edit), Coach (reactive feedback on the user's plan), and Delegate (autonomous execution, no veto). The load-bearing mechanism is the 'human filter': when users can revise or ignore AI output, they undo the AI's high-surplus proposals; the Delegate condition bypasses this, letting the agent act as a market maker that injects rational, Pareto-improving offers into the trading environment.

Load-bearing premise

The entire comparison rests on the assumption that the all-human baseline from an earlier, separately run study is equivalent to today's participants; if that earlier group was unrepresentative in skill, effort, or session conditions, the measured Delegate advantage and all spillover estimates are biased.

What would settle it

Recruit a fresh no-AI control group contemporaneous with the experimental sessions, using the same platform, instructions, and payments. If the Delegate group's scaled surplus is not above this concurrent control by roughly 0.08, the headline effect is an artifact of baseline drift. In addition, test whether Advisor-condition offers that were edited by users are statistically indistinguishable in receiver surplus from purely manual offers; if edited offers still outperform manual offers, the 'human filter' explanation would need revision.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Groups with access to a Delegate agent achieve higher average collective surplus than all-human groups; Advisor and Coach access produce no significant gain.
  • Users strongly prefer the Advisor, citing control and trust, even though the Delegate yields better average payoffs.
  • Roughly half of the Delegate's group-level benefit comes from positive spillovers to people who never use the AI themselves.
  • The mechanism is the human filter: AI proposals create more surplus, but Advisor and Coach users modify or ignore them, reverting toward human-baseline trade patterns; Delegate bypasses this filter.
  • Because the same underlying LLM powers all three modes, the result cannot be attributed to model capability differences.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the human-filter mechanism generalizes, the same LLM could produce opposite welfare effects depending only on whether users can intervene per turn; system designers should treat user veto power as part of the mechanism, not as a neutral safety feature.
  • The spillover result suggests welfare evaluations of agentic AI should be measured at the market level, not the individual level: the people who gain may not be the people who adopt.
  • A natural extension is testing a 'delegate with an upfront veto window' that lets users review the agent's overall strategy before the game rather than each action; the paper's mechanism predicts this could keep most of the Delegate gain while easing trust concerns.
  • Since the Delegate benefit is attributed to interaction structure rather than model capability, a testable follow-up is swapping in a weaker model: if the Delegate advantage persists, structure dominates model quality.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper reports a within-subjects online experiment (N=243; 81 groups) comparing three LLM assistance modalities—Advisor, Coach, and Delegate—in a three-player chip bargaining game with induced valuations. Participants play three games, one per modality, in counterbalanced order. The paper claims: (i) participants prefer the Advisor (44%) over the Delegate (19%) but achieve the highest mean surplus with Delegate access; (ii) Delegate access raises group scaled surplus by 0.084 relative to an external all-human baseline (raw p=0.033, Holm-Bonferroni p=0.10); (iii) positive spillovers to non-users (β=0.039, raw p=0.054, p_adj=0.16); and (iv) a mechanism: Delegate proposals shift receiver surplus distributions (KS p=0.002) because human filtering dilutes AI proposals in Advisor/Coach modes. The paper also decomposes Delegate gains into roughly 51% direct and 49% spillover contributions, and the abstract references a '1.5x intent-to-treat' compliance-adjusted estimate that does not appear in the body.

Significance. The question—how interface modality shapes adoption and welfare in strategic multi-agent settings—is timely and important. The experiment has real strengths: within-subject counterbalancing; a constant underlying model across modalities; turn-level endogenous take-up measurement; detailed prompts and implementation details (including the open-source Deliberate Lab platform); and a mechanism analysis using distributional tests. If the results were supported, they would provide an early empirical baseline for the emerging 'agentic economy' and would make a substantive design point: interaction structure, not just model capability, determines realized welfare. However, the central welfare estimates all rest on an external no-AI baseline whose exchangeability is asserted but not demonstrated, and none of the headline treatment contrasts survives the paper's own multiple-testing correction. The contribution is promising but not yet statistically secure.

major comments (4)
  1. [§3.1, Table 1] All RQ1 treatment effects are differences from the external all-human baseline of Qian et al. [54] (N=72, mean scaled surplus 0.537). The paper asserts that recruitment criteria, payout rates, and interface mechanics were identical, but provides no balance table, comparability test, or concurrent no-AI arm. Because current participants play three counterbalanced games while the baseline played a single game with no AI, practice/fatigue/order effects are not absorbed. The Delegate coefficient (β=0.084, SE=0.040, raw p=0.033) is therefore biased if baseline cohorts differ in skill, effort, or session conditions. A same-experiment no-AI control (or a formal sensitivity analysis) is needed to identify RQ1.
  2. [§4.2.2 vs §5.1.1, Table 3] The spillover result for Delegate non-users is reported inconsistently: Table 3 gives β=0.039, SE=0.020, p=0.054, but §5.1.1 reports p=0.033 and calls it 'significant.' Moreover, non-user status is an endogenous choice: if more skilled participants self-select into non-use, the comparison against the human baseline is biased. The decomposition in §5.1 further claims 'roughly half' of the Delegate gain comes from spillovers, but the components are individually non-significant (adopter +0.022, p=0.148; non-user p=0.054/0.033) and no joint test or standard error for the 51/49 split is provided.
  3. [Abstract vs §4.1, Tables 1–2] The abstract states that groups 'significantly increase collective surplus under Delegate access' and that participants 'achieve the highest mean individual gains with the Delegate.' However, the group-level effect has p_adj=0.10 and the individual-level effect has p_adj=0.20 under the paper's own Holm-Bonferroni correction. The abstract's language overstates the inferential support; these claims are at most suggestive. Please report adjusted p-values and confidence intervals in the abstract or soften the claim accordingly.
  4. [Abstract vs body] The abstract reports a compliance-adjusted estimate ('roughly 1.5x the intent-to-treat estimate') that does not appear anywhere in the main text, tables, or appendices. This is a load-bearing quantitative claim; it must be derivable from a stated model (e.g., a treatment-on-treated / IV analysis) with explicit assumptions about selection into delegation. Please add the specification, point estimate, and uncertainty to §5.1 or remove it from the abstract.
minor comments (6)
  1. [§3.2] Typo: 'int he game' should be 'in the game.'
  2. [§5.1] The text says the Delegate 'yielded an increase in scaled surplus of 2.8 percentage points,' which matches the individual-level coefficient in Table 2 (0.028) but not the group-level coefficient in Table 1 (0.084). Clarify which estimand the decomposition uses.
  3. [References [53]/[54]] References [53] and [54] appear to refer to the same arXiv paper under different titles. Consolidate to avoid ambiguity, especially because the external baseline is drawn from this work.
  4. [§4.3, Table 4] The Chi-square test is described as comparing 'the three AI modalities,' but Table 4 includes a 'None' category (N=52). Clarify whether 'None' responses were included in the test and in the pairwise comparisons.
  5. [§6.1] There is a stray citation artifact 'domain-specific norms![18]' in the Limitations section; please clean up the markup.
  6. [Throughout] Minor language inconsistencies: 'super-human' vs 'superhuman'; 'the authors thanks' should be 'the authors thank.'

Circularity Check

0 steps flagged

No significant circularity: the central claims are empirical contrasts against an external, independently measurable human baseline; the only self-citation carrying weight is that baseline, which is real experimental evidence rather than a fitted or definitional input.

full rationale

This paper is an empirical randomized experiment, not a derivation from first principles. Its central claims—Delegate access raises group scaled surplus (β=0.084, SE=0.040), non-users in Delegate groups show positive spillovers (β=0.039, SE=0.020), and users prefer the Advisor modality—are estimated from newly collected data (N=243, 81 groups) using LMMs. The treatment effects are contrasts against the all-human baseline from Qian et al. [54] (N=72, mean scaled surplus 0.537). This is the only self-citation with load-bearing weight. It is not circular: the baseline is an externally measurable quantity reported with its own sample and variance, not a parameter fitted to the current data and not defined in terms of the current outcomes. The paper also transparently reports that the headline Delegate effect becomes marginally non-significant after Holm-Bonferroni correction (p_adj=0.10), which is not what one expects if the result were forced by construction. The mechanism claim—that the Delegate advantage comes from bypassing a human filter—is supported by within-experiment comparisons of AI-generated versus manual proposals and by the significant upward shift in receiver surplus only in the Delegate condition (KS p=0.002), not by the external baseline. The paper's own limitation is baseline comparability: it asserts that 'recruitment criteria, payout rates, and interface mechanics were identical' to the prior study but provides no balance table and no concurrent no-AI arm. That is a validity threat, not a circularity. No equation is shown to reduce to another by construction, and no fitted parameter is relabeled as a prediction. Therefore no specific circular step can be exhibited; the honest finding is no circularity.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

This is an empirical behavioral study; it introduces no mathematical free parameters. The central estimates rest on domain assumptions about baseline comparability, the LP-optimum normalization, and transfer of the LLM's simulated capability into live games. The Advisor/Coach/Delegate conditions are experimental manipulations, not invented entities.

axioms (4)
  • domain assumption The external all-human baseline from Qian et al. [54] is comparable to the current sample.
    All treatment effects are estimated as differences from this prior experiment's N=72 groups; the paper asserts identical recruitment, payout, and interface mechanics (Section 3.1) but offers no same-experiment control or balance table.
  • domain assumption The optimal surplus benchmark (LP solution under full public valuations) is the correct normalization target.
    Scaled surplus normalizes by a theoretical optimum from Qian et al. [54]; the paper relies on this benchmark for all welfare comparisons without independent validation.
  • domain assumption The LLM agent's simulated superhuman performance (scaled surplus 0.595 vs human 0.537) transfers to live human games.
    Superhuman capability is established only in all-agent simulation (Section 3.3); there is no in-experiment measure of agent proposal quality independent of human filtering.
  • domain assumption Participants' surplus-maximizing incentives are aligned through the average-bonus payment scheme.
    Standard experimental-economics assumption; the paper states the bonus was designed to align incentives (Section 3.5) but does not test for strategic gaming or wealth effects.

pith-pipeline@v1.3.0-alltime-deepseek · 20527 in / 10878 out tokens · 103202 ms · 2026-08-02T23:53:25.014078+00:00 · methodology

0 comments
read the original abstract

As AI usage becomes more prevalent in social contexts, understanding agent-user interaction is critical to designing systems that imp rove both individual and group outcomes. We present an online behavioral experiment (N=243) in which participants play three multi-tu rn bargaining games in groups of three. Each game, presented in randomized order, grants access to a single LLM assistance modality: proactive recommendations from an Advisor, reactive feedback from a Coach, or autonomous execution by a Delegate. All three modalitie s are powered by an LLM with super-human performance within this negotiation setting. On each turn, participants privately decide whe ther to act manually or use the AI modality available in that game. We document a preference-performance misalignment: participants s trongly prefer the higher-control Advisor (44%) over the Delegate (19%), yet groups only significantly increase collective surplus un der Delegate access. Adjusting for voluntary non-compliance, delegating to the AI yields suggestive individual welfare gains, roughly 1.5x the intent-to-treat estimate. A mechanism analysis traces this gap to a human filter: AI-generated proposals create more joint surplus than manual proposals across all conditions, but in the Advisor and Coach modes users modify, override, or ignore the AI's su ggestions, reverting toward human-baseline trade patterns. The Delegate advantage arises not from a different AI capability but from bypassing this filtering step altogether. Realizing these welfare gains depends not only on model capability, but on the interaction structure through which that capability is delivered. We argue that assistance modalities should be designed as mechanisms with endog enous participation; adoption-compatible interaction rules are a prerequisite to improving welfare with automated assistance.

Figures

Figures reproduced from arXiv: 2602.12089 by Crystal Qian, James Wexler, Kehang Zhu, Nithum Thain, Vivian Tsai.

Figure 1
Figure 1. Figure 1: Overview of the experimental design and contributions. Participants (N=243) engaged in three-person [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Relevant game components. Panels 1. and 2. show interface properties visible across all modalities. [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Access to delegation increases both group and individual surplus. Panel (A) shows the compari￾son of scaled group surplus gain. Panel (B) shows the comparison of scaled surplus gain for individuals. condition, whereas a user is defined as someone who utilized AI assistance at least once for a proposal. This definition echoes prior findings that the act of making a proposal drives the final surplus signific… view at source ↗
Figure 4
Figure 4. Figure 4: Top: Frequency of assistance usage by negotiation round. Bottom: Coefficient of regression predicting [PITH_FULL_IMAGE:figures/full_fig_p012_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Receiver surplus distributions of accepted trades. Each panel compares the surplus of AI-assisted [PITH_FULL_IMAGE:figures/full_fig_p013_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Pre-game Likert survey response distributions. [PITH_FULL_IMAGE:figures/full_fig_p021_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Post-game survey Likert survey response distributions. [PITH_FULL_IMAGE:figures/full_fig_p022_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Coach-specific Likert survey response distributions. [PITH_FULL_IMAGE:figures/full_fig_p022_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Advisor-specific Likert survey response distributions. [PITH_FULL_IMAGE:figures/full_fig_p022_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Delegate-specific Likert survey response distributions. [PITH_FULL_IMAGE:figures/full_fig_p023_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: A Sankey diagram showing flows from participants’ post-game mode preference (left) to coded [PITH_FULL_IMAGE:figures/full_fig_p023_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: An iteration of a classification prompt applied to an LLM auto-rater to conduct a thematic analysis [PITH_FULL_IMAGE:figures/full_fig_p024_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: A screenshot of the bargaining game interface implemented in the Deliberate Lab platform. It is [PITH_FULL_IMAGE:figures/full_fig_p025_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: A diagram showing the AI assistance interfaces during the offer generation phase. [PITH_FULL_IMAGE:figures/full_fig_p026_14.png] view at source ↗
Figure 15
Figure 15. Figure 15: A diagram showing the AI assistance interfaces during the offer response phase. [PITH_FULL_IMAGE:figures/full_fig_p026_15.png] view at source ↗
Figure 16
Figure 16. Figure 16: The baseline proposing prompt used for the Delegate and Advisor agents. The [PITH_FULL_IMAGE:figures/full_fig_p028_16.png] view at source ↗
Figure 17
Figure 17. Figure 17: The baseline answering prompt used for the Delegate and Advisor agents. The [PITH_FULL_IMAGE:figures/full_fig_p029_17.png] view at source ↗
Figure 18
Figure 18. Figure 18: The additional prompt for the Coach mode. The user’s draft idea is appended to the context, prompting [PITH_FULL_IMAGE:figures/full_fig_p029_18.png] view at source ↗
Figure 19
Figure 19. Figure 19: The additional prompt for the Coach mode. The user’s draft decision is appended to the context, [PITH_FULL_IMAGE:figures/full_fig_p030_19.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

77 extracted references · 2 canonical work pages

  1. [1]

    Sahar Abdelnabi, Amr Gomaa, Sarath Sivaprasad, Lea Schönherr, and Mario Fritz. 2023. LLM-Deliberation: Evaluating LLMs with Interactive Multi-Agent Negotiation Games. (2023). , Vol. 1, No. 1, Article . Publication date: February 2025. 16 Kehang Zhu et al

  2. [2]

    2023.Combining human expertise with artificial intelligence: Experimental evidence from radiology

    Nikhil Agarwal, Alex Moehring, Pranav Rajpurkar, and Tobias Salz. 2023.Combining human expertise with artificial intelligence: Experimental evidence from radiology. Technical Report. National Bureau of Economic Research

  3. [3]

    2025.Designing Human-AI Collaboration: A Sufficient-Statistic Approach

    Nikhil Agarwal, Alex Moehring, and Alexander Wolitzky. 2025.Designing Human-AI Collaboration: A Sufficient-Statistic Approach. Technical Report. National Bureau of Economic Research

  4. [4]

    2025.Designing Human-AI Collaboration: A Sufficient-Statistic Approach

    Nikhil Agarwal, Alex Moehring, and Alexander Wolitzky. 2025.Designing Human-AI Collaboration: A Sufficient-Statistic Approach. NBER Working Paper 33949. National Bureau of Economic Research. https://doi.org/10.3386/w33949

  5. [5]

    Saaket Agashe, Yue Fan, Anthony Reyna, and Xin Eric Wang. 2023. Llm-coordination: evaluating and analyzing multi-agent coordination abilities in large language models.arXiv preprint arXiv:2310.03903(2023)

  6. [6]

    Mohammed Alsobay, David G Rand, Duncan J Watts, and Abdullah Almaatouq. 2025. Integrative Experiments Identify How Punishment Impacts Welfare in Public Goods Games.arXiv preprint arXiv:2508.17151(2025)

  7. [7]

    Mohammed Alsobay, David M Rothschild, Jake M Hofman, and Daniel G Goldstein. 2025. Bringing Everyone to the Table: An Experimental Study of LLM-Facilitated Group Decision Making.arXiv preprint arXiv:2508.08242(2025)

  8. [8]

    Lasecki, Daniel S

    Gagan Bansal, Besmira Nushi, Ece Kamar, Walter S. Lasecki, Daniel S. Weld, and Eric Horvitz. 2019. Beyond Accuracy: The Role of Mental Models in Human–AI Team Performance. InProceedings of the AAAI Conference on Human Computation and Crowdsourcing (HCOMP), Vol. 7. 2–11

  9. [9]

    Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel Weld. 2021. Does the whole exceed its parts? the effect of ai explanations on complementary team performance. InProceedings of the 2021 CHI conference on human factors in computing systems. 1–16

  10. [10]

    Federico Bianchi, Patrick John Chia, Mert Yuksekgonul, Jacopo Tagliabue, Dan Jurafsky, and James Zou. 2024. How well can llms negotiate? negotiationarena platform and analysis.arXiv preprint arXiv:2402.05863(2024)

  11. [11]

    Ken Binmore, Ariel Rubinstein, and Asher Wolinsky. 1986. The Nash bargaining solution in economic modelling.The RAND Journal of Economics(1986), 176–188

  12. [12]

    Olivier Bochet, Manshu Khanna, and Simon Siegenthaler. 2024. Beyond dividing the pie: Multi-issue bargaining in the laboratory.Review of Economic Studies91, 1 (2024), 163–191

  13. [13]

    Virginia Braun and Victoria Clarke. 2006. Using thematic analysis in psychology.Qualitative Research in Psychology3, 2 (2006), 77–101. https://doi.org/10.1191/1478088706qp063oa

  14. [14]

    Zana Buçinca, Phoebe Lin, Krzysztof Z Gajos, and Elena L Glassman. 2020. Proxy tasks and subjective measures can be misleading in evaluating explainable AI systems. InProceedings of the 25th International Conference on Intelligent User Interfaces. 454–464. https://doi.org/10.1145/3377325.3377498

  15. [15]

    Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z Gajos. 2021. To trust or to think: Cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making.Proceedings of the ACM on Human-Computer Interaction5, CSCW1 (2021), 1–21. https://doi.org/10.1145/3449287

  16. [16]

    Shi-Yi Chen, Zhe Feng, and Xiaolian Yi. 2017. A general introduction to adjustment for multiple comparisons.Journal of thoracic disease9, 6 (2017), 1725

  17. [17]

    Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, et al. 2025. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities.arXiv preprint arXiv:2507.06261(2025)

  18. [18]

    Jared R Curhan, Hillary Anger Elfenbein, and Heng Xu. 2006. What do people value when they negotiate? Mapping the domain of subjective value in negotiation.Journal of personality and social psychology91, 3 (2006), 493

  19. [19]

    Fred D Davis. 1989. Perceived usefulness, perceived ease of use, and user acceptance of information technology.MIS quarterly(1989), 319–340

  20. [20]

    Regina de Brito Duarte, Mónica Costa Abreu, Joana Campos, and Ana Paiva. 2025. The Amplifying Effect of Ex- plainability in AI-assisted Decision-making in Groups. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–15

  21. [21]

    Berkeley J Dietvorst, Joseph P Simmons, and Cade Massey. 2015. Algorithm aversion: people erroneously avoid algorithms after seeing them err.Journal of experimental psychology: General144, 1 (2015), 114

  22. [22]

    Berkeley J Dietvorst, Joseph P Simmons, and Cade Massey. 2018. Overcoming algorithm aversion: People will use imperfect algorithms if they can (even slightly) modify them.Management science64, 3 (2018), 1155–1170

  23. [23]

    Elias Fernández Domingos, Inês Terrucha, Rémi Suchon, Jelena Grujić, Juan C Burguillo, Francisco C Santos, and Tom Lenaerts. 2021. Delegation to autonomous agents promotes cooperation in collective-risk dilemmas.arXiv preprint arXiv:2103.07710(2021)

  24. [24]

    Meta Fundamental AI Research Diplomacy Team (FAIR)†, Anton Bakhtin, Noam Brown, Emily Dinan, Gabriele Farina, Colin Flaherty, Daniel Fried, Andrew Goff, Jonathan Gray, Hengyuan Hu, et al. 2022. Human-level play in the game of Diplomacy by combining language models with strategic reasoning.Science378, 6624 (2022), 1067–1074

  25. [25]

    Andreas Fügener, Jörn Grahl, Alok Gupta, and Wolfgang Ketter. 2022. Cognitive challenges in human–artificial intelligence collaboration: Investigating the path toward productive delegation.Information Systems Research33, 2 (2022), 678–696. , Vol. 1, No. 1, Article . Publication date: February 2025. Choose Your Agent: Tradeoffs in Adopting AI Advisors, Coa...

  26. [26]

    Ella Glikson and Anita Williams Woolley. 2020. Human trust in artificial intelligence: Review of empirical research. Academy of management annals14, 2 (2020), 627–660

  27. [27]

    2025.Guided Learning in Gemini: From answers to understanding

    Google. 2025.Guided Learning in Gemini: From answers to understanding. https://blog.google/outreach-initiatives/ education/guided-learning/ Accessed: 2025-09-11

  28. [28]

    Ben Green and Yiling Chen. 2019. The principles and limits of algorithm-in-the-loop decision making.Proceedings of the ACM on Human-Computer Interaction3, CSCW (2019), 1–24. https://doi.org/10.1145/3359152

  29. [29]

    Sture Holm. 1979. A simple sequentially rejective multiple test procedure.Scandinavian journal of statistics(1979), 65–70

  30. [30]

    Ruru Hoong and Bnaya Dreyfuss. 2025. Improving AI-Assisted Decision-Making Through Calibrated Coarsening. A vailable at SSRN 5286198(2025)

  31. [31]

    2023.Large language models as simulated economic agents: What can we learn from homo silicus? Technical Report

    John J Horton. 2023.Large language models as simulated economic agents: What can we learn from homo silicus? Technical Report. National Bureau of Economic Research

  32. [32]

    Alex Imas, Kevin Lee, and Sanjog Misra. 2025. Agentic Interactions.A vailable at SSRN 5875162(2025)

  33. [33]

    2011.Thinking, fast and slow

    Daniel Kahneman. 2011.Thinking, fast and slow. macmillan

  34. [34]

    Daniel Kahneman, Jack L Knetsch, and Richard H Thaler. 1990. Experimental tests of the endowment effect and the Coase theorem.Journal of political Economy98, 6 (1990), 1325–1348

  35. [35]

    Charlotte Kobiella, Ulugbek Isroilov, and Albrecht Schmidt. 2025. When AI Joins the Negotiation Table: Evaluating AI as a Moderator. InProceedings of the 7th ACM Conference on Conversational User Interfaces. 1–18

  36. [36]

    Justin Kruger and David Dunning. 1999. Unskilled and unaware of it: how difficulties in recognizing one’s own incompetence lead to inflated self-assessments.Journal of personality and social psychology77, 6 (1999), 1121

  37. [37]

    Vivian Lai, Chacha Chen, Alison Smith-Renner, Q Vera Liao, and Chenhao Tan. 2023. Towards a Science of Human-AI Decision Making: An Overview of Design Space in Empirical Human-Subject Studies. InProceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency. 1369–1385. https://doi.org/10.1145/3593013.3594087

  38. [38]

    Why is’ Chicago’deceptive?

    Vivian Lai, Han Liu, and Chenhao Tan. 2020. " Why is’ Chicago’deceptive?" Towards Building Model-Driven Tutorials for Humans. InProceedings of the 2020 CHI Conference on Human Factors in Computing Systems. 1–13

  39. [39]

    Vivian Lai, Alison Smith-Renner, Ke Zhang, Ruijia Cheng, Wenjuan Zhang, Joel Tetreault, and Alejandro Jaimes. 2022. An Exploration of Post-Editing Effectiveness in Text Summarization. InProceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computation...

  40. [40]

    Min Hun Lee, Daniel P Siewiorek, Asim Smailagic, Alexandre Bernardino, and Sergi Bermúdez i Badia. 2021. A Human-AI Collaborative Approach for Clinical Decision Making on Rehabilitation Assessment. InProceedings of the 2021 CHI Conference on Human Factors in Computing Systems. 1–14. https://doi.org/10.1145/3411764.3445472

  41. [41]

    Yuan Li, Yixuan Zhang, and Lichao Sun. 2023. Metaagents: Simulating interactions of human behaviors for llm-based task-oriented coordination via collaborative generative agents.arXiv preprint arXiv:2310.06500(2023)

  42. [42]

    2024.Automated social science: Language models as scientist and subjects

    Benjamin S Manning, Kehang Zhu, and John J Horton. 2024.Automated social science: Language models as scientist and subjects. Technical Report. National Bureau of Economic Research

  43. [43]

    Robert B Miller. 1968. Response time in man-computer conversational transactions. InProceedings of the December 9-11, 1968, fall joint computer conference, part I. 267–277

  44. [44]

    Shinichi Nakagawa. 2004. A farewell to Bonferroni: the problems of low statistical power and publication bias. Behavioral ecology15, 6 (2004), 1044–1045

  45. [45]

    1994.Usability engineering

    Jakob Nielsen. 1994.Usability engineering. Morgan Kaufmann

  46. [46]

    2025.Introducing ChatGPT Agent: Bridging Research and Action

    OpenAI. 2025.Introducing ChatGPT Agent: Bridging Research and Action. https://openai.com/index/introducing- chatgpt-agent/ Accessed: 2025-09-11

  47. [47]

    2025.Introducing Study Mode

    OpenAI. 2025.Introducing Study Mode. https://openai.com/index/chatgpt-study-mode/ Accessed: 2025-09-11

  48. [48]

    David Owens, Zachary Grossman, and Ryan Fackler. 2014. The control premium: A preference for payoff autonomy. American Economic Journal: Microeconomics6, 4 (2014), 138–161

  49. [49]

    Stefano Palminteri, Basile Garcia, and Crystal Qian. 2025. How Objective Source and Subjective Belief Shape the Detectability and Acceptability of LLMs’ Moral Judgments. https://doi.org/10.31234/osf.io/ct6rx_v2

  50. [50]

    Aman Pathak and Veena Bansal. 2024. AI as decision aid or delegated agent: The effects of trust dimensions on the adoption of AI digital agents.Computers in Human Behavior: Artificial Humans2, 2 (2024), 100094

  51. [51]

    Prolific. 2024. Prolific. https://www.prolific.com. First released 2014. London, UK. Version used: [insert month and year of use]

  52. [52]

    Crystal Qian and James Wexler. 2024. Take it, leave it, or fix it: Measuring productivity and trust in human-ai collaboration. InProceedings of the 29th International Conference on Intelligent User Interfaces. 370–384

  53. [54]

    Manning, Vivian Tsai, James Wexler, and Nithum Thain

    Crystal Qian, Kehang Zhu, John Horton, Benjamin S. Manning, Vivian Tsai, James Wexler, and Nithum Thain. 2025. Understanding Economic Tradeoffs Between Human and AI Agents in Bargaining Games. arXiv:2509.09071 [cs.AI] https://arxiv.org/abs/2509.09071

  54. [55]

    Maithra Raghu, Katy Blumer, Greg Corrado, Jon Kleinberg, Ziad Obermeyer, and Sendhil Mullainathan. 2019. The Algorithmic Automation Problem: Prediction, Triage, and Human Effort. arXiv:1903.12220 [cs.CV]

  55. [56]

    Alvin E Roth, Vesna Prasnikar, Masahiro Okuno-Fujiwara, and Shmuel Zamir. 1991. Bargaining and market behavior in Jerusalem, Ljubljana, Pittsburgh, and Tokyo: An experimental study.The American economic review(1991), 1068–1095

  56. [57]

    David M Rothschild, Markus Mobius, Jake M Hofman, Eleanor W Dillon, Daniel G Goldstein, Nicole Immorlica, Sonia Jaffe, Brendan Lucier, Aleksandrs Slivkins, and Matthew Vogel. 2025. The Agentic Economy.arXiv preprint arXiv:2505.15799(2025)

  57. [58]

    Ariel Rubinstein. 1982. Perfect equilibrium in a bargaining model.Econometrica: Journal of the Econometric Society (1982), 97–109

  58. [59]

    Richard M Ryan and Edward L Deci. 2000. Self-determination theory and the facilitation of intrinsic motivation, social development, and well-being.American psychologist55, 1 (2000), 68

  59. [60]

    Anand Shah, Kehang Zhu, Yanchen Jiang, Jeffrey G Wang, Arif K Dayi, John J Horton, and David C Parkes. 2025. Learning from Synthetic Labs: Language Models as Auction Participants.arXiv preprint arXiv:2507.09083(2025)

  60. [61]

    2025.The Coasean Singularity? Demand, Supply, and Market Design with AI Agents

    Peyman Shahidi, Gili Rusak, Benjamin S Manning, Andrey Fradkin, and John J Horton. 2025.The Coasean Singularity? Demand, Supply, and Market Design with AI Agents. Technical Report. National Bureau of Economic Research

  61. [62]

    Chris Simms. [n. d.]. AI is more persuasive than people in online debates.Nature([n. d.])

  62. [63]

    Ermis Soumalias, Yanchen Jiang, Kehang Zhu, Michael Curry, Sven Seuken, and David C Parkes. 2025. LLM-Powered Preference Elicitation in Combinatorial Assignment.arXiv preprint arXiv:2502.10308(2025)

  63. [64]

    John Sweller. 1988. Cognitive load during problem solving: Effects on learning.Cognitive science12, 2 (1988), 257–285

  64. [65]

    Michael Henry Tessler, Michiel A Bakker, Daniel Jarrett, Hannah Sheahan, Martin J Chadwick, Raphael Koster, Georgina Evans, Lucy Campbell-Gillingham, Tantum Collins, David C Parkes, et al. 2024. AI can help humans find common ground in democratic deliberation.Science386, 6719 (2024), eadq2852

  65. [66]

    Nenad Tomasev, Matija Franklin, Joel Z Leibo, Julian Jacobs, William A Cunningham, Iason Gabriel, and Simon Osindero. 2025. Virtual agent economies.arXiv preprint arXiv:2509.10147(2025)

  66. [67]

    2025.Deliberate Lab: Open-Source Platform for LLM-Powered Social Science

    Vivian Tsai, Crystal Qian, Michael Behr, and Deliberate Lab community contributors. 2025.Deliberate Lab: Open-Source Platform for LLM-Powered Social Science. https://github.com/PAIR-code/deliberate-lab

  67. [68]

    Michelle Vaccaro, Abdullah Almaatouq, and Thomas Malone. 2024. When combinations of humans and AI are useful: A systematic review and meta-analysis.Nature Human Behaviour8 (2024), 2293–2303. https://doi.org/10.1038/s41562- 024-02024-1

  68. [69]

    Michelle Vaccaro, Michael Caosun, Harang Ju, Sinan Aral, and Jared R Curhan. 2025. Advancing ai negotiations: New theory and evidence from a large-scale autonomous negotiations competition.arXiv preprint arXiv:2503.06416(2025)

  69. [70]

    Jessie Yang, Christopher D

    X. Jessie Yang, Christopher D. Wickens, and Katja Hölttä-Otto. 2016. How users adjust trust in automation: Contrast effect and hindsight bias.Proceedings of the Human Factors and Ergonomics Society Annual Meeting60, 1 (2016), 196–200. https://doi.org/10.1177/1541931213601044 arXiv:https://doi.org/10.1177/1541931213601044

  70. [71]

    Kunlun Zhu, Hongyi Du, Zhaochen Hong, Xiaocheng Yang, Shuyi Guo, Zhe Wang, Zhenhailong Wang, Cheng Qian, Xiangru Tang, Heng Ji, et al. 2025. Multiagentbench: Evaluating the collaboration and competition of llm agents.arXiv preprint arXiv:2503.01935(2025)

  71. [72]

    Skill Amplifier

    Shenzhe Zhu, Jiao Sun, Yi Nian, Tobin South, Alex Pentland, and Jiaxin Pei. 2025. The automated but risky game: Modeling agent-to-agent negotiations and transactions in consumer markets. InICML 2025 Workshop on Reliable and Responsible Foundation Models. , Vol. 1, No. 1, Article . Publication date: February 2025. Choose Your Agent: Tradeoffs in Adopting A...

  72. [73]

    **Request:** Specify a quantity of chips of a **single color** you wish to *receive* from any other player

  73. [74]

    Your goal is to make as much money as possible by making an advantageous proposal that is likely to be accepted

    **Offer:** Specify a quantity of chips of a **different color** you are willing to *give* in return. Your goal is to make as much money as possible by making an advantageous proposal that is likely to be accepted. The trades, you choose to make to accomplish this, are up to you. Be rational - do not propose a trade in which the user loses money. The value...

  74. [75]

    1 chip for 1 chip

    Try to AVOID VERY CONSERVATIVE trades, e.g. 1 chip for 1 chip. Remember you only have 3 chances to propose trades

  75. [76]

    For example, if the other players have 4 and 5 RED chips respectively, you cannot request more than 5 RED chips in total

    You CANNOT request more chips than a player currently has. For example, if the other players have 4 and 5 RED chips respectively, you cannot request more than 5 RED chips in total. Output a proposal response. Your response **must adhere strictly to the following format**. Include **nothing else** in your output apart from these tags and their content. Fig...

  76. [77]

    But Player XXX appears to value blue chips more than you do

    Your current offer is profitable. But Player XXX appears to value blue chips more than you do. You may want to consider trading blue chips for other colors

  77. [78]

    You may want to consider increasing the quantity of chips you are offering

    There is only 1 round left. You may want to consider increasing the quantity of chips you are offering. Fig. 19. The additional prompt for the Coach mode. The user’s draft decision is appended to the context, prompting the model to give feedback rather than generate de novo. , Vol. 1, No. 1, Article . Publication date: February 2025