Pith. sign in

REVIEW 3 major objections 7 minor 27 references

Two-player Alternate Uses Test: A Controlled Testbed for Interactive Human-AI and Human-Human Co-Creation

T0 review · 3 major / 7 minor · reviewed 2026-07-09 · glm-5.2

Pith's one-line read AI matches humans as creative brainstorming partners under fair conditions

desk verdict Solid platform contribution with a well-tested equivalence claim; secondary findings are exploratory and underpowered but the paper is mostly honest about that. read the letter →

arxiv 2607.07522 v1 pith:KXGLSBMZ submitted 2026-07-08 cs.HC

classification cs.HC
keywords AlternateUsesTesthuman-AIco-creationdivergentthinkingcreativitysupporttoolscognitiveoutsourcingapproachmotivationinteractiveideationGPT-4
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper introduces a controlled, two-player version of the Alternate Uses Test — a classic psychology task where people generate unusual uses for everyday objects — adapted so that a human can brainstorm with either another human or a GPT-4 conversational partner under identical time limits, instructions, and interface. The central finding from a 62-person pilot is that, under these matched conditions, the originality of ideas produced with a GPT-4 partner is statistically equivalent to that produced with a human partner. This stands in contrast to prior studies showing AI outperforming humans on independent divergent-thinking tasks, suggesting that whether AI appears more or less creative than humans depends heavily on whether the task is interactive and how that interaction is structured. The paper also reports that a participant's motivational disposition (specifically, approach motivation) determines whether having an interactive partner helps or hurts their creative output, that self-reported cognitive outsourcing (treating a partner as a passive resource rather than engaging collaboratively) predicts worse outcomes specifically in human-human pairs, and that prior exposure to highly creative ideas improves subsequent co-creation regardless of partner type — a tractable seeding intervention. The authors release the platform, code, and dataset as a shared experimental testbed for future controlled studies of human-AI co-creation.

What carries the argument

The central object is the two-player Alternate Uses Test platform: a browser-based chat application where two players generate unusual uses for a target object together under a fixed time limit, with conditions for human-human, human-GPT-4, and non-interactive (pre-rated seed ideas) pairings. The platform enforces strict turn-taking and matched instructions across conditions, and supports decomposition of co-creative performance into participant traits, partner perceptions, and content dynamics.

What would settle it

If a future, larger study using the same platform finds a robust originality difference between human and GPT-4 partners under matched conditions — or if removing the AI behavioral guardrails produces a significant asymmetry — the equivalence claim would be undermined.

Watch

Extended reading notes

Core claim

The paper's core contribution is a controlled experimental apparatus — a two-player, time-limited chat-based Alternate Uses Test — that can simultaneously vary partner identity (human vs GPT-4), interaction structure (interactive vs passive exposure to pre-rated ideas), participant traits, and content exposure, within a single within-subjects design. Using this apparatus, the authors demonstrate that originality with a GPT-4 partner is statistically equivalent to originality with a human partner when interaction conditions are matched, and that the determinants of co-creative success decompose into separable layers: trait-level factors (approach motivation, socioemotional sensitivity), attun

Load-bearing premise

The paper assumes that the behavioral guardrails imposed on GPT-4 (one idea per turn, no boilerplate, single-sentence responses, token-overlap filtering) achieve interaction parity with the unconstrained human partner condition. If these constraints channel the AI's contributions differently from how a freely responding human would contribute, the observed equivalence between human and AI partners could reflect the prompt design rather than a genuine property of human-AI co-.

Editorial extensions

If this is right

  • If the equivalence finding holds at scale, the widespread claim that AI is more creative than humans in divergent thinking may be an artifact of comparing independent agents rather than interactive partners — interaction structure, not just model capability, determines the conclusion.
  • The finding that approach motivation moderates who benefits from interactive partnership suggests that AI creativity tools should be matched to user motivational profiles rather than deployed uniformly.
  • The seeding effect — prior exposure to highly creative human ideas raising subsequent originality — offers a concrete, low-cost intervention that could be tested in educational and workplace brainstorming settings.
  • The platform's open release as a pre-registerable testbed could standardize how future studies isolate the mechanisms of human-AI co-creation, reducing the confounds that plague both independent-agent comparisons and field studies.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If interaction structure is the key variable explaining the gap between prior AI-surpasses-human findings and the equivalence found here, then studies comparing independent LLM output to independent human output may be measuring the wrong comparison entirely — the relevant question is not who is more creative alone, but who is more creative as a collaborator.
  • The reversal of the outsourcing effect across partner types (harmful with humans, possibly beneficial with GPT-4) hints at a deeper asymmetry: structured AI turn-taking may partially substitute for the motivational engagement that human partners require, which would mean AI partners serve a different cognitive function than human partners rather than an equivalent one.
  • The seeding effect could interact with the homogenization concern raised by prior work — if seeding with diverse highly creative human ideas counteracts the convergence tendency observed under passive AI exposure, then pre-exposure design may be a practical lever for preserving creative diversity in AI-assisted workflows.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. This paper introduces a two-player extension of the Alternate Uses Test (AUT) designed as a controlled testbed for comparing human-human and human-AI co-creation. The platform includes interactive conditions (human-human, human-GPT-4) and non-interactive baselines (pre-rated creative, uncreative, and GPT-4-only idea sets). A pilot study (N=62) demonstrates the testbed's utility by decomposing co-creative performance into participant traits (RMET, BIS/BAS), partner perceptions (outsourcing, warmth, competence), and content dynamics (seeding effects). The central empirical finding is that originality with a GPT-4 partner is statistically equivalent to that with a human partner under matched time limits, contrasting with prior reports of AI surpassing humans on independent divergent-thinking tasks. The authors release the platform, code, and dataset.

Significance. The paper addresses a genuine methodological gap between controlled but non-interactive AI creativity studies and ecologically valid but confounded field studies. The testbed design is thoughtful: within-subjects randomization, counterbalanced objects, TOST and Bayesian tests for the equivalence claim, rater random effects in the LMM, and calibrated non-interactive baselines sourced from external data (Stevenson et al.). The release of the platform, code, and dataset as a shared, pre-registerable testbed is a strong contribution that enables future replication and extension. The decomposition framework separating traits, perceptions, and content dynamics is a useful organizing principle for the field.

major comments (3)
  1. §3.1 and §4.1: The central equivalence claim (β_{GPT−HUM} = −0.11, BF₀₁ = 10.95) is conditional on interaction parity between the GPT-4 and human conditions. However, the GPT-4 guardrails (one idea per turn, single-sentence responses, no boilerplate, token-overlap filter) are applied only to the AI condition, creating a systematic asymmetry in expressive bandwidth. Human partners can elaborate, provide multi-sentence context, or offer conversational scaffolding that GPT-4 cannot. The paper acknowledges this in §6 ('the guardrails' influence invites future work'), but the abstract and §4.1 state the equivalence as an established finding. The claim should be explicitly scoped as 'GPT-4 under guardrails vs. unconstrained humans' rather than 'GPT-4 vs. humans,' or the authors should provide evidence that the guardrails do not materially affect originality ratings. Without this, the reader's—
  2. §3.2 and §4.3: The CON sub-arms were unevenly assigned (~21 participants each across three sub-arms), and the seeding effect analysis (F(2, 1963) = 4.68, p = .009) relies on this between-subjects variation. With N≈21 per sub-arm, the power to detect seeding effects is limited, and the authors correctly flag this as exploratory. However, the claim that 'prior exposure to highly creative ideas improves later performance' is presented as one of three main findings in the abstract and conclusion. Given the uneven assignment and modest N, this finding should be more cautiously framed throughout, or the analysis should include sensitivity checks (e.g., effect size stability under leave-one-subject-out).
  3. §4.2: The outsourcing × partner-type interaction is based on n=20 for the human-human correlation (r = −0.504, p = .024). This is a very small subsample, and the p-value is marginally significant. The reversal direction (Δβ ≈ +0.35) is described as 'anecdotal in Bayesian terms' for the GPT condition. While the authors note a confirmatory test is in progress, presenting this as a key finding (listed in the abstract and conclusion) risks overstating the evidence. The authors should either move this to a secondary/exploratory analysis or explicitly state the n and power limitations in the results section.
minor comments (7)
  1. §3.4: The abstract states 1,928 ideas, but §3.4 states 1,923 distinct ideas. Please reconcile.
  2. §3.4: The inter-rater reliability is mentioned only as 'r = .54' between crowdsourced and researcher curation ratings. A formal ICC or Krippendorff's alpha for the six-rater crowdsourced ratings would strengthen the originality measure validation.
  3. §4.1: The TOST equivalence bound of |β| ≤ 0.33 is not justified. Was this bound pre-registered or derived from a smallest-effect-size-of-interest rationale? A brief justification would strengthen the equivalence claim.
  4. Figure 1 caption: '225 Prolific raters' is mentioned, but §3.4 states '226 crowdsourced workers.' Please reconcile.
  5. §3.1: The GPT-4 model version is not specified beyond 'gpt-4.' Given the rapid iteration of model versions, specifying the exact snapshot (e.g., gpt-4-0613) would improve reproducibility.
  6. §4.2: The random forest variable importance (Figure 2) is based on 200 bootstrap replicates, but no out-of-bag error or cross-validated R² is reported. Including a measure of predictive accuracy would contextualize the variable importance rankings.
  7. §5: The discussion of RMET results could note the Oakley et al. [18] critique more prominently, as the paper uses RMET as a socioemotional sensitivity measure rather than a ToM measure. The current framing is appropriate but could be clearer about what RMET does and does not measure.

Simulated Author's Rebuttal

3 responses · 0 unresolved

We thank the referee for a careful and constructive review. The referee raises three major concerns, all of which are substantive and well-taken: (1) the equivalence claim should be explicitly scoped to acknowledge the guardrail asymmetry between GPT-4 and human partners; (2) the seeding-effect finding rests on modest per-sub-arm N and should be framed more cautiously; (3) the outsourcing × partner-type interaction is based on a very small subsample (n=20) and risks being overstated. We agree with all three points and will revise the manuscript accordingly. The core contribution—the platform and decomposition framework—remains intact; the revisions involve scoping claims more carefully and adding sensitivity checks.

read point-by-point responses
  1. Referee: §3.1 and §4.1: The central equivalence claim is conditional on interaction parity, but GPT-4 guardrails create a systematic asymmetry in expressive bandwidth. The claim should be explicitly scoped as 'GPT-4 under guardrails vs. unconstrained humans,' or the authors should provide evidence that the guardrails do not materially affect originality ratings.

    Authors: The referee is correct. The guardrails (one idea per turn, single-sentence responses, no boilerplate, token-overlap filter) are applied only to the AI condition, and this creates an asymmetry in expressive bandwidth. We acknowledge this in §6 but do not scope the claim in the abstract or §4.1. We will revise the abstract, §4.1, and the conclusion to explicitly state that the equivalence finding concerns 'GPT-4 under conversational guardrails vs. unconstrained human partners.' We will also add a sentence in §3.1 noting that the guardrails were designed to neutralize model-specific asymmetries (verbosity, meta-commentary) rather than to constrain creative output per se, but that we cannot rule out an effect of the guardrails on originality without a dedicated manipulation. We agree that the current wording overstates the generality of the finding. revision: yes

  2. Referee: §3.2 and §4.3: The CON sub-arms were unevenly assigned (~21 participants each), and the seeding effect analysis relies on this between-subjects variation. With N≈21 per sub-arm, the power to detect seeding effects is limited. The claim that 'prior exposure to highly creative ideas improves later performance' is presented as one of three main findings in the abstract and conclusion. This finding should be more cautiously framed throughout, or the analysis should include sensitivity checks.

    Authors: We agree that the seeding finding should be framed more cautiously given the modest per-sub-arm N. We will take two steps. First, we will add a leave-one-subject-out sensitivity analysis for the seeding effect (F(2, 1963) = 4.68, p = .009) and report the stability of the effect size and p-value across iterations. Second, we will revise the abstract and conclusion to describe this as an 'exploratory finding from a modest sample' rather than presenting it as an established result on equal footing with the equivalence claim. We already flag the finding as exploratory in §4.3 and §5, but the abstract and conclusion do not carry this qualification, and we will fix that. revision: yes

  3. Referee: §4.2: The outsourcing × partner-type interaction is based on n=20 for the human-human correlation (r = −0.504, p = .024). This is a very small subsample, and the p-value is marginally significant. The reversal direction (Δβ ≈ +0.35) is described as 'anecdotal in Bayesian terms' for the GPT condition. Presenting this as a key finding risks overstating the evidence. The authors should either move this to a secondary/exploratory analysis or explicitly state the n and power limitations in the results section.

    Authors: The referee is right that n=20 is too small to support the prominence this finding currently receives in the abstract and conclusion. We will make two changes. First, we will explicitly state the subsample size (n=20 for the human-human correlation, and the corresponding n for the GPT condition) and note the limited power in §4.2 itself, not only in the general limitations section. Second, we will reframe the outsourcing finding in the abstract and conclusion as a hypothesis-generating result that motivates a confirmatory test, rather than listing it as one of three main findings. The confirmatory study with direct self- and partner-effort items is already in progress, and we will note this. We believe the finding is worth reporting given the theoretically predicted direction reversal, but we agree it should not be presented as an established result. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found; derivation is self-contained against external data

full rationale

The paper's central claims are empirical findings tested against externally sourced data. The equivalence claim (β_{GPT−HUM} = −0.11, TOST, BF₀₁ = 10.95) is evaluated using crowdsourced originality ratings from 226 independent Prolific workers, not a measure defined by the authors. The CON baselines use ideas externally sourced from Stevenson et al. [25]. The trait measures (RMET [2], BIS/BAS [4]) are standard published instruments not authored by the present authors. The TOST equivalence test [15] and Bayesian re-estimation use standard statistical methods. The one self-citation, Hemmatian & Sloman [13], provides theoretical motivation for the outsourcing-vs-collaboration distinction, but it does not define any measured variable in terms of the outcome, and the predictions it motivates (BAS Drive moderation, outsourcing effects) are tested empirically against independent data rather than being true by construction. No step in the derivation chain reduces to its own inputs by definition or by self-citation. The paper is self-contained against external benchmarks.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

The paper introduces no new physical entities, particles, forces, or mathematical objects. The platform itself is a software artifact, not a theoretical entity. The free parameters are operational choices (temperature, token limits, session duration) rather than theoretical constructs fitted to data. The axioms are domain assumptions about measurement validity and interaction parity, with one ad-hoc-to-paper assumption about the prompt guardrails. The theoretical framework (outsourcing vs. collaboration) is cited from prior work [13] and not invented here.

free parameters (6)
  • GPT-4 temperature = 0.7
    Set for the AI partner's response generation; not derived from theory but chosen to balance determinism and variability.
  • GPT-4 max_tokens = 3000
    Operational parameter for the AI partner, chosen to avoid truncation.
  • Token-overlap filter threshold = 0.50
    Rejects candidate responses with >50% token overlap with prior messages; chosen ad hoc to prevent repetition.
  • Chat session duration = 4 minutes
    Fixed time limit for all interactive sessions; chosen for experimental feasibility.
  • Solo AUT duration = 2 minutes
    Personal baseline duration; chosen for experimental feasibility.
  • Number of CON seed ideas per set = 16
    Pre-rated ideas shown in non-interactive condition; chosen to match exposure levels.
assumptions (5)
  • domain assumption Crowdsourced Likert ratings (5-point scale) are a valid measure of originality for AUT ideas.
    Section 3.4 defines originality via crowdsourced ratings. The paper notes modest alignment with researcher curation (r=.54) and plans RA-rated replication, but the central analyses depend on this measurement assumption.
  • domain assumption RMET indexes socioemotional sensitivity relevant to text-based co-creation.
    Section 3.3 acknowledges RMET may measure emotion recognition rather than Theory of Mind, but assumes it indexes 'a trait-level tendency to attend to partners' states' relevant to textual co-creation.
  • domain assumption BIS/BAS scales capture motivational traits that moderate co-creative performance.
    Section 3.3 uses BIS/BAS as the motivational framework; the moderation finding depends on this construct validity.
  • ad hoc to paper The prompt guardrails for GPT-4 achieve interaction parity with human partners.
    Section 3.1 describes behavioral guardrails (one idea per turn, no boilerplate, single-sentence responses) added to neutralize model-side asymmetries. The equivalence claim depends on these guardrails not differentially constraining AI vs. human creative contributions.
  • domain assumption Four-minute chat sessions capture meaningful co-creative processes.
    The 4-minute limit is a practical constraint; whether it captures the dynamics of longer co-creative sessions is assumed, not tested.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Two-player Alternate Uses Test: A Controlled Testbed for Interactive Human-AI and Human-Human Co-Creation." pith.science (2026). https://pith.science/paper/KXGLSBMZ

@misc{pith2026260707522,
  author       = {Pith},
  title        = {Pith review of: Two-player Alternate Uses Test: A Controlled Testbed for Interactive Human-AI and Human-Human Co-Creation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KXGLSBMZ}},
  note         = {Machine review of arXiv:2607.07522}
}
read the original abstract

Controlled research on AI ideation typically compares independent agents, while field studies of human-AI collaboration sacrifice experimental control. We introduce a controlled, two-player extension of the Alternate Uses Test (AUT) that enables comparison of human-human and human-AI co-creation under matched interactive conditions, alongside calibrated non-interactive baselines. The platform supports decomposition of performance into three typically confounded factors: participant traits, partner perceptions, and content dynamics. An in-person pilot (N = 62) demonstrates its utility. Under matched time limits, originality with a GPT-4 partner is statistically equivalent to that with a human partner. Approach motivation (BAS Drive) moderates whether interactive partnership benefits originality, and self-reported cognitive outsourcing predicts lower originality specifically in human-human dyads. Prior exposure to highly creative ideas improves later performance, suggesting a "seeding" intervention. We release the platform, code, and dataset as a shared testbed for controlled studies of human-AI co-creation.

Figures

Figures reproduced from arXiv: 2607.07522 by the authors.

Figure 1
Figure 1. Experimental design and data pipeline. Each participant (N = 62) runs three randomized 4-minute co-creation sessions—once [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Random-forest variable importance for predicting originality, over 200 bootstrap replicates. Bars show mean importance; [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. BAS Drive moderates whether interactive partnership helps creativity. Bars show mean originality (1–5 scale) per BAS Drive [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

27 extracted references · 27 canonical work pages

  1. [1]

    Joshua Ashkinaze, Julia Mendelsohn, Li Qiwei, Ceren Budak, and Eric Gilbert. 2025. How AI ideas affect the creativity, diversity, and evolution of human ideas: Evidence from a large, dynamic experiment. In Proceedings of the ACM Collective Intelligence Conference (CI ‘25). ACM, New York, NY, USA. https://doi.org/10.1145/3715928.3737481

  2. [2]

    Reading the Mind in the Eyes

    Simon Baron-Cohen, Sally Wheelwright, Jacqueline Hill, Yogini Raste, and Ian Plumb. 2001. The “Reading the Mind in the Eyes” test revised version: A study with normal adults, and adults with Asperger syndrome or high-functioning autism. Journal of Child Psychology and Psychiatry 42, 2 (2001), 241–251

  3. [3]

    Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakçı, and Rei Mariman. 2025. Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences 122, 26 (2025), e2422633122. https://doi.org/10.1073/pnas.2422633122

  4. [4]

    Carver and Teri L

    Charles S. Carver and Teri L. White. 1994. Behavioral inhibition, behavioral activation, and affective responses to impending reward and punishment: The BIS/BAS scales. Journal of Personality and Social Psychology 67, 2 (1994), 319–333

  5. [5]

    Nicholas Davis, Chih-Pin Hsiao, Kunwar Yashraj Singh, Brenda Lin, and Brian Magerko. 2017. Creative Sense-Making: Quantifying Interaction Dynamics in Co-Creation. In Proceedings of the 2017 ACM SIGCHI Conference on Creativity and Cognition (C&C ‘17). ACM, New York, NY, USA, 356–366. https://doi.org/10.1145/3059454.3059478

  6. [6]

    Mollick, Hila Lifshitz-Assaf, Katherine Kellogg, Saran Rajendran, Lisa Krayer, François Candelon, and Karim R

    Fabrizio Dell’Acqua, Edward McFowland III, Ethan R. Mollick, Hila Lifshitz-Assaf, Katherine Kellogg, Saran Rajendran, Lisa Krayer, François Candelon, and Karim R. Lakhani. 2023. Navigating the jagged technological frontier: Field experimental evidence of the effects of artificial intelligence on knowledge worker productivity and quality. Harvard Business ...

  7. [7]

    Manoj Deshpande, Jisu Park, Supratim Pait, and Brian Magerko. 2024. Perceptions of Interaction Dynamics in Co-Creative AI: A Comparative Study of Interaction Modalities in Drawcto. In Proceedings of the 16th Conference on Creativity & Cognition (C&C ‘24). ACM, New York, NY, USA. https://doi.org/10.1145/3635636.3656202

  8. [8]

    Doshi and Oliver P

    Anil R. Doshi and Oliver P. Hauser. 2024. Generative AI enhances individual creativity but reduces the collective diversity of novel content. Science Advances 10, 28 (2024), eadn5290. https://doi.org/10.1126/sciadv.adn5290

Show all 27 references
  1. [9]

    Jing, Christopher F

    Daniel Engel, Anita Williams Woolley, Lisa X. Jing, Christopher F. Chabris, and Thomas W. Malone. 2014. Reading the Mind in the Eyes or Reading between the Lines? Theory of Mind Predicts Collective Intelligence Equally Well Online and Face-To-Face. PLoS ONE 9, 12 (2014), e1152...

  2. [10]

    Fiske, Amy J

    Susan T. Fiske, Amy J. C. Cuddy, and Peter Glick. 2007. Universal dimensions of social cognition: Warmth and competence. Trends in Cognitive Sciences 11, 2 (2007), 77–83

  3. [11]

    Fiske, Amy J

    Susan T. Fiske, Amy J. C. Cuddy, Peter Glick, and Jun Xu. 2002. A model of (often mixed) stereotype content: Competence and warmth respectively follow from perceived status and competition. Journal of Personality and Social Psychology 82, 6 (2002), 878–902

  4. [12]

    J. P. Guilford. 1967. The Nature of Human Intelligence. McGraw-Hill, New York, NY

  5. [13]

    Babak Hemmatian and Steven A. Sloman. 2020. Two systems for thinking with a community: Outsourcing versus collaboration. In Logic and Uncertainty in the Human Mind: A Tribute to David E. Over, Shira Elqayam, Igor Douven, Jonathan St. B. T. Evans, and Nicole Cruz (Eds.). Routle...

  6. [14]

    Hubert, Kim N

    Kent F. Hubert, Kim N. Awa, and Darya L. Zabelina. 2024. The current state of artificial intelligence generative language models is more creative than humans on divergent thinking tasks. Scientific Reports 14, 1 (2024), 3440. https://doi.org/10.1038/s41598-024-53303-w

  7. [15]

    Daniël Lakens. 2017. Equivalence tests: A practical primer for t tests, correlations, and meta-analyses. Social Psychological and Personality Science 8, 4 (2017), 355–362. https://doi.org/10.1177/1948550617697177

  8. [16]

    Duri Long, Mikhail Jacob, Nicholas Davis, and Brian Magerko. 2017. Designing for Socially Interactive Systems. In Proceedings of the 2017 ACM SIGCHI Conference on Creativity and Cognition (C&C ‘17). ACM, New York, NY, USA, 39–50. https://doi.org/10.1145/3059454.3059479

  9. [17]

    Simon Maier, Max Schneider, and Stefan Feuerriegel. 2026. Partnering with generative AI: Experimental evaluation of human-led and model-led interaction in human–AI co-creation. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI ‘26). ACM. arXi...

  10. [18]

    Bonnie F. M. Oakley, Rebecca Brewer, Geoffrey Bird, and Caroline Catmur. 2016. Theory of mind is not theory of emotion: A cautionary note on the Reading the Mind in the Eyes test. Journal of Abnormal Psychology 125, 6 (2016), 818–823. https://doi.org/10.1037/abn0000182

  11. [19]

    Jeba Rezwana and Corey Ford. 2025. Human-Centered AI Communication in Co-Creativity: An Initial Framework and Insights. In Proceedings of the 2025 ACM Creativity and Cognition Conference (C&C ‘25). ACM, New York, NY, USA, 651–665. https://doi.org/10.1145/3698061.3726932

  12. [20]

    Jeba Rezwana and Mary Lou Maher. 2023. Designing Creative AI Partners with COFI: A Framework for Modeling Interaction in Human-AI Co- Creative Systems. ACM Transactions on Computer-Human Interaction 30, 5, Article 67 (Sept. 2023), 28 pages. https://doi.org/10.1145/3519026

  13. [21]

    Malone, and Anita Williams Woolley

    Christoph Riedl, Young Ji Kim, Pranav Gupta, Thomas W. Malone, and Anita Williams Woolley. 2021. Quantifying collective intelligence in human groups. Proceedings of the National Academy of Sciences 118, 21 (2021), e2005737118. https://doi.org/10.1073/pnas.2005737118

  14. [22]

    Keith Sawyer and Stacy DeZutter

    R. Keith Sawyer and Stacy DeZutter. 2009. Distributed creativity: How collective creations emerge from collaboration. Psychology of Aesthetics, Creativity, and the Arts 3, 2 (2009), 81–92. https://doi.org/10.1037/a0013282

  15. [23]

    Kun, and Hagit Ben Shoshan

    Orit Shaer, Angelora Cooper, Osnat Mokryn, Andrew L. Kun, and Hagit Ben Shoshan. 2024. AI-augmented brainwriting: Investigating the use of LLMs in group ideation. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (CHI ‘24). ACM, New York, NY, USA....

  16. [24]

    Steven Sloman and Philip Fernbach. 2018. The Knowledge Illusion: Why We Never Think Alone. Riverhead Books, New York, NY

  17. [25]

    Stevenson, Iris Smal, Matthijs Baas, Raoul Grasman, and Han L

    Claire E. Stevenson, Iris Smal, Matthijs Baas, Raoul Grasman, and Han L. J. van der Maas. 2022. Putting GPT-3’s creativity to the (alternative uses) test. In Proceedings of the 13th International Conference on Computational Creativity (ICCC ‘22). arXiv:2206.08932. https://arxi...

  18. [26]

    Dingjue Wang, Dongjie Huang, Hanlu Shen, et al. 2025. A large-scale comparison of divergent creativity in humans and large language models. Nature Human Behaviour (2025). https://doi.org/10.1038/s41562-025-02331-1

  19. [27]

    Chabris, Alex Pentland, Nada Hashmi, and Thomas W

    Anita Williams Woolley, Christopher F. Chabris, Alex Pentland, Nada Hashmi, and Thomas W. Malone. 2010. Evidence for a Collective Intelligence Factor in the Performance of Human Groups. Science 330, 6004 (2010), 686–688. https://doi.org/10.1126/science.1193147

Pith tools

Reviewed July 9, 2026 · model on record in the stance chip above.