Pith. sign in

REVIEW 4 major objections 4 minor 33 references

Stop Writing for Me: Generative Refusal in AI Tools for Thought

T0 review · 4 major / 4 minor · reviewed 2026-08-02 · deepseek-v4-flash

Pith's one-line read This paper argues that generative AI tools for thought should refuse to write, asking users questions instead, to protect their own constructive cognition.

desk verdict Worth engaging for the design idea, but the field study does not isolate 'generative refusal' from generic question prompts, and the unverifiable citations need fixing. read the letter →

arxiv 2607.24751 v1 pith:32TNDUM4 submitted 2026-05-23 cs.HC

classification cs.HC
keywords generativerefusaltoolsforthoughtcognitiveoffloadingactortrainingreflectivewritingSocraticquestioninghuman-AIinteractioninternalization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that generative AI tools for thought should sometimes refuse to write. In actor training, where articulating a character's inner life is the core skill, a journaling tool that only asks context-aware questions—never drafting text—lowered the reported cognitive burden of writing while deepening reflection. The paper's 14-day field study found that actors using the question-asking tool wrote more lexically and emotionally varied entries and later reported internalizing the questioning style after the tool was removed. The authors propose treating 'generative refusal' as a core design feature and measuring success by how well users perform without the tool.

What carries the argument

The central mechanism is Generative Refusal: a deliberate design constraint in which the AI does not produce text, leaving a structured gap the user must fill. Actor's Note operationalizes this with three context-aware, second-person questions generated from the user's uploaded script, role, and current rehearsal stage (table work vs. run-through). This mechanism is what the paper says converts delegation into activation—the interaction model shifts from 'AI does it for me' to 'AI prompts me to do it.' The paper pairs this with the concept of 'desirable difficulties' and Vygotsky's zone of proximal development to argue that the friction is productive.

What would settle it

A three-arm experiment comparing AI that only asks questions, AI that drafts text, and no AI, measuring cognitive burden and post-intervention self-questioning, would settle it: if the draft-text arm matches the question-asking arm, the refusal mechanism is not doing the work.

Watch

Extended reading notes

Core claim

The central claim is that strategically withholding text generation—the paper calls it 'Generative Refusal'—turns a GenAI system from a co-author into a Maieutic Partner that forces users to do the cognitive work of articulation. In Actor's Note, the system reads the actor's script and rehearsal stage and generates three second-person questions daily, but never a draft. In a 14-day randomized crossover study with 29 actors, the question-asking condition significantly reduced cognitive burden and increased intrinsic motivation, acting confidence, lexical diversity, and emotional range compared with unassisted freewriting. The paper reports a residual effect: after the AI was removed, particip

Load-bearing premise

The paper attributes the benefits to the AI's refusal to generate text, but because the comparison was between question-asking AI and freewriting with no prompt, the improvements could equally be caused by having any structured question or by the tool's novelty, not by the withholding itself.

Editorial extensions

If this is right

  • Creativity support tools should treat 'refuse to write' as an intentional design feature, not a missing capability, and offer modes that ask questions instead of drafting.
  • Evaluation of AI tools for thought should include post-tool measures—how well users perform after the AI is removed—not just with-tool speed or output quality.
  • Scaffolding should adapt to the user's workflow phase: low-friction prompts early in a process to overcome inertia, high-friction challenges later to prevent fixation.
  • AI's lack of social judgment can be explicitly framed as a non-judgmental space, enabling raw, vulnerable reflection that users would hide from human peers or directors.
  • The success of question-driven journaling suggests a general pattern: for reflective-writing domains, ask-don't-write can protect the user's constructive cognitive process.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The study compared question-asking AI against unaided freewriting, so the observed gains could come from the presence of structured prompts rather than from the refusal itself; a three-arm trial that also includes a text-draft condition would isolate the mechanism.
  • The residual internalization score suggests a testable extension: repeated question-asking interactions should transfer to other reflective writing tasks and could be measured weeks after the intervention ends.
  • The same refusal mechanic may generalize to other domains, where 'refusal' would take domain-specific forms—for example, a coding assistant that refuses to complete a function and instead asks the programmer to state the intended behavior.
  • The reported safety of an unjudged AI space points to a boundary condition: private reflection may need a deliberate bridge to public, judged performance for skills to land, since the absence of social stakes may also remove performative energy.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper argues that generative AI tools for thought should sometimes withhold text generation and instead ask context-aware questions, a strategy the author terms 'Generative Refusal.' It presents Actor's Note, a journaling tool for actors that generates three daily questions but never draft text, and reports a 14-day field study with 29 actors comparing AI-assisted journaling to unassisted freewriting. Reported outcomes include reduced cognitive burden, increased intrinsic motivation and acting confidence, higher lexical diversity and emotion-word use, and a post-intervention tendency to self-generate similar questions. The discussion generalizes these findings into four design implications: Generative Refusal as a core mechanic, adaptive scaffolding, AI as a non-judgmental space, and internalization as the primary success metric.

Significance. If the empirical claims were methodologically established, the paper would make a valuable contribution to the Tools for Thought / human-AI interaction literature by articulating a design principle that runs against the default 'write-for-me' paradigm. The concept of Generative Refusal is provocative and generative, and the four workshop questions are well formed. The paper also deserves credit for grounding the argument in a real deployed tool and for reporting effect sizes. However, the current evidence cannot isolate the proposed mechanism from simpler alternatives, and the internalization claim—which is central to the paper's theoretical contribution—rests on a single self-report item. The design implications are therefore more credible as a research agenda than as conclusions from the reported study.

major comments (4)
  1. [§2.2, §3.1] The study cannot isolate 'Generative Refusal' from generic question scaffolding. The AI condition always supplies three context-specific questions and never generates draft text; the control condition supplies no questions and no text. These conditions differ in at least three ways simultaneously: presence of any structured prompt, context-specificity of the prompt, and absence of generated draft text. Any observed benefit could be due to the first or second factor alone. To support the paper's central claim, the design would need conditions that vary the refusal factor independently—for example, an AI condition that provides the same questions but also drafts a suggested response, or a non-AI condition that provides generic daily prompts. As it stands, the abstract's statement that 'this constraint significantly reduced cognitive burden' attributes causality to a factor that was never m
  2. [§2.2, §2.3] The crossover design is vulnerable to carryover, and the paper's own hypothesis makes carryover likely. If participants internalize the AI's questioning style, then those who receive AI first will bring that style into the subsequent unassisted phase, contaminating the 'unassisted' baseline. The reported subgroup difference (early AI as 'momentum starter' vs. late AI as 'deepener', p=.0128) may reflect order effects rather than rehearsal stage. The paper does not report an order-by-condition interaction or a washout analysis, so the early/late interpretation in §3.2 is not supported by the evidence presented.
  3. [§2.3, §3.4] The internalization result—a central claim of the paper—rests on a single post-intervention self-report item: 'participants reported a high tendency to recall the AI's questioning style or self-generate similar questions (M=4.87 on a 7-point scale)'. No baseline, control condition, distribution, confidence interval, or inferential test is reported, and no systematic qualitative coding of the interviews is provided. Given that Section 3.4 proposes 'internalization' as the primary success metric for Tools for Thought, this evidence is too thin to carry the paper's core argument. At minimum, the item wording, response scale, and pre-post comparison would need to be reported.
  4. [§2.3] Statistical reporting is insufficient for the strength of the claims. The section reports partial eta-squared and p-values (e.g., η²_p=.381, p<.001) without descriptive statistics, confidence intervals, or details of the repeated-measures models. The 'q' values used for lexical diversity and emotion-word analyses are undefined and no multiple-comparison adjustment is described. The subgroup analysis for Narrative Transportation (p=.0128) appears to be one of several subgroup or outcome tests, yet no correction is mentioned. Without these details, the reader cannot assess whether the observed effects are robust or whether the headline p-values survive proper control for multiplicity.
minor comments (4)
  1. [References [1]–[5]] Several citations are placeholders ('A. Author et al.', 'B. Author et al.', etc.), making the claimed prior work on Socratic agents and refusal strategies unverifiable. Please replace with real author names, titles, venues, and DOIs or remove them.
  2. [Reference [24]] The paper relies on the author's own companion paper [24] as the source of the study, but that paper is cited as 'to be presented at CHI 2026' rather than published. The relationship between this position paper and the companion paper should be stated explicitly, and the companion's availability should be clarified.
  3. [Figure 1] Figure 1 is not referenced in the body text. Please add an in-text callout or remove the figure.
  4. [Keywords] The keyword 'Tool for Thoughts' is inconsistent with the standard singular 'Tools for Thought' used throughout the paper. Please standardize.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: no derivation reduces to its own inputs; the main study-design confound is a causal-inference issue rather than a circular step.

full rationale

The paper contains no fitted parameters, equations, or prediction steps whose output is identical to an input by construction. 'Generative Refusal' is an intervention concept, and the reported outcomes—cognitive burden, motivation, lexical diversity, and internalized questioning—are empirical measurements that could have turned out differently. The central methodological weakness is a confound: Actor’s Note always provides structured context-aware questions in addition to withholding draft text (Section 2.2), so the comparison against unassisted freewriting does not isolate 'refusal' from generic question scaffolding. That threatens causal attribution, but confounding is not circularity. The empirical findings are sourced from the author's own companion paper [24], and the prior Socratic-AI citations [1–5] are unverifiable, but these are evidential dependencies rather than cases where the conclusion is presupposed by the premise. The internalization item (M=4.87 on a 7-point scale, Section 2.3) is a post-hoc self-report with no baseline, a reliability concern, not a tautology. Overall, the reasoning chain is not circular; the score reflects only the minor self-citation for empirical grounding and the unisolated mechanism, not derivational circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper makes no fitted mathematical claims. Its load-bearing assumptions are domain-specific: articulation is the learning mechanism in actor training; the observed benefits are attributable to the refusal mechanism rather than generic scaffolding; and self-reported internalization measures durable cognitive change. No invented entities are introduced.

assumptions (4)
  • domain assumption The labor of articulation in actor training is itself the mechanism of learning; bypassing it erodes constructive thought.
    Section 1, para 2; this is the value premise underlying Generative Refusal.
  • domain assumption Observed benefits are caused by withholding text generation, not by generic structured prompting, the AI's novelty, or increased attention.
    Sections 2.3 and 3.1; the design compares question-asking to freewriting, so the active ingredient is not isolated.
  • domain assumption A single self-report item on recalling the AI's questioning style (M=4.87) is a valid measure of internalized cognitive habit.
    Section 2.3; no behavioral transfer measure is reported.
  • domain assumption Bjork's desirable difficulties and Vygotsky's ZPD apply to adult professional actor journaling.
    Section 3.1; external learning theories imported as explanation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Stop Writing for Me: Generative Refusal in AI Tools for Thought." pith.science (2026). https://pith.science/paper/32TNDUM4

@misc{pith2026260724751,
  author       = {Pith},
  title        = {Pith review of: Stop Writing for Me: Generative Refusal in AI Tools for Thought},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/32TNDUM4}},
  note         = {Machine review of arXiv:2607.24751}
}
read the original abstract

In creative domains where the labor of articulation is central to the craft, how should we design Tools for Thought that enhance rather than bypass human cognition? Current GenAI paradigms often prioritize "cognitive offloading"-writing on behalf of users-risking the erosion of the constructive thought process essential to artistic training. In this position paper, we explore AI as a Maieutic Partner through "Generative Refusal"-strategically withholding text generation to demand user articulation. We discuss Actor's Note, a journaling tool that generates context-aware questions instead of draft text. Our field study suggests that this constraint significantly reduced cognitive burden while fostering a residual effect of internalized questioning habits. We use these findings to discuss broader design implications for protecting human cognition against the tendency of generative efficiency.

Figures

Figures reproduced from arXiv: 2607.24751 by the authors.

Figure 1
Figure 1. The Maieutic Interaction Framework. Instead of bypassing cognition, the system uses Generative Refusal to return [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

33 extracted references · 3 canonical work pages

  1. [1]

    Author et al

    A. Author et al. 2024. Enhancing Critical Thinking in Education by means of a Socratic Chatbot.arXiv preprint arXiv:2409.05511(2024)

  2. [5]

    Author et al

    E. Author et al. 2026. When AI only asks: how question-driven dialogue shapes prewriting in the classroom.Frontiers in Education(2026)

  3. [24]

    Sora Kang. 2026. Actor’s Note: Examining the Role of Al-Generated Questions in Character Journaling for Actor Training. InProceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI 26)(Barcelona, Spain). ACM, New York, NY, USA. doi:10.1145/3772318.3790370

  4. [2]

    Author et al

    B. Author et al. 2024. Using GenAI for Socratic Questioning: An Approach to Higher-Order Thinking for Nursing Education.PMC12557457 (2024)

  5. [3]

    Author et al

    C. Author et al. 2025. Exploring the Potential of Metacognitive Support Agents for Human-AI Co-Creation.arXiv preprint arXiv:2506.12879(2025)

  6. [4]

    Author et al

    D. Author et al. 2025. Investigating the effects of an LLM-based Socratic con- versational agent on students’ academic performance and reflective thinking in higher education.Computers & Education(2025)

  7. [6]

    J.S. Bascomb. 2019.Performing Arts and Performance Anxiety. Ph. D. Dissertation. Marshall University. Theses, Dissertations and Capstones. 1184

  8. [7]

    N. Begus. 2024. Experimental Narratives: A Comparison of Human Crowdsourced Storytelling and AI Storytelling.Humanities and Social Sciences Communications 11, 1 (2024), 1392. doi:10.1057/s41599-024-03868-8

Show all 33 references
  1. [8]

    Robert A Bjork. 1994. Memory and metamemory considerations in the training of human beings. InMetacognition: Knowing about knowing. MIT Press, 185–205

  2. [9]

    S.B. Blix. 2015. Professional Emotion Management as a Rehearsal Process.Pro- fessions and Professionalism5, 2 (2015). doi:10.7577/pp.1322

  3. [10]

    Branch, P

    B. Branch, P. Mirowski, and K.W. Mathewson. 2021. Collaborative Storytelling with Human Actors and AI Narrators.arXiv preprint(2021)

  4. [11]

    Bruder, L.M

    M. Bruder, L.M. Cohn, M. Olnek, N. Pollack, R. Previto, and S. Zigler. 2012.A Practical Handbook for the Actor. Knopf Doubleday Publishing Group

  5. [12]

    Vannevar Bush. 1945. As We May Think.The Atlantic Monthly176, 1 (1945), 101–108

  6. [13]

    Chakrabarty, V

    T. Chakrabarty, V. Padmakumar, F. Brahman, and S. Muresan. 2024. Creativity Support in the Age of Large Language Models: An Empirical Study Involving Emerging Writers.arXiv preprint(2024)

  7. [14]

    Andy Clark and David Chalmers. 1998. The Extended Mind.Analysis58, 1 (1998), 7–19

  8. [15]

    Coenen, L

    A. Coenen, L. Davis, D. Ippolito, E. Reif, and A. Yuan. 2021. Wordcraft: a Human- AI Collaborative Editor for Story Writing.arXiv preprint(2021)

  9. [16]

    Cremin, K

    T. Cremin, K. Goouch, L. Blakemore, E. Goff, and R. Macdonald. 2006. Connecting drama and writing: seizing the moment to write.Research in Drama Education: The Journal of Applied Theatre and Performance11, 3 (2006), 273–291. doi:10.1080/ 13569780600900636

  10. [17]

    Dayo, A.A

    F. Dayo, A.A. Memon, and N. Dharejo. 2023. Scriptwriting in the Age of AI: Revolutionizing Storytelling with Artificial Intelligence.Journal of Media & Communication4, 1 (2023), 24–38

  11. [18]

    Dharaniya, J

    R. Dharaniya, J. Indumathi, and V. Kaliraj. 2023. A design of movie script genera- tion based on natural language processing by optimized ensemble deep learning with heuristic algorithm.Data & Knowledge Engineering146 (2023), 102150. doi:10.1016/j.datak.2023.102150

  12. [19]

    Drama-Based Pedagogy. n.d.. Writing in Role. https://dbp.theatredance.utexas. edu/teaching-strategies/writing-role Accessed: 2025-08-20

  13. [20]

    Fleck and G

    R. Fleck and G. Fitzpatrick. 2010. Reflecting on reflection: framing a design land- scape. InProceedings of the 22nd Conference of the Computer-Human Interaction Special Interest Group of Australia on Computer-Human Interaction. 216–223

  14. [21]

    U. Hagen. 1991.Challenge For The Actor. Simon and Schuster

  15. [22]

    M.R. Hancock. 1993. Character Journals: Initiating Involvement and Identification through Literature.The Journal of Reading36 (1993)

  16. [23]

    C.K. Joyce. 2009.The Blank Page: Effects of Constraint on Creativity. Ph. D. Dissertation. University of California, Berkeley

  17. [25]

    Norrthon and A

    S. Norrthon and A. Schmidt. 2023. Knowledge Accumulation in Theatre Re- hearsals: The Emergence of a Gesture as a Solution for Embodying a Certain Aesthetic Concept.Human Studies46, 2 (2023), 337–369. doi:10.1007/s10746-022- 09654-2

  18. [26]

    Schultz and M

    L.M. Schultz and M. Rose. 1985. Writer’s Block: The Cognitive Dimension.College Composition and Communication36, 4 (1985), 497. doi:10.2307/357873

  19. [27]

    Stanislavski

    K. Stanislavski. 2009.An Actor’s Work on a Role. Routledge

  20. [28]

    Stanislavskij

    K.S. Stanislavskij. 1986.An actor prepares. Methuen

  21. [29]

    Stanley and P

    T. Stanley and P. Strandberg-Long. 2022.An Actor’s Research: Investigating Choices for Practice and Performance. Routledge

  22. [30]

    Always be relevant

    K.L. Sørensen, N. Hald, R.M. Olsen, and K.B. Ørjasæter. 2024. “Always be relevant”: a phenomenological study of the actor’s workday.Arts & Health(2024), 1–14. doi:10.1080/17533015.2024.2420822

  23. [31]

    Lev Tankelevitch, Advait Sarkar, Sean Rintel, Richard Banks, Mina Lee, et al. 2025. Understanding, Protecting, and Augmenting Human Cognition with Generative AI: A Synthesis of the CHI 2025 Tools for Thought Workshop.arXiv preprint arXiv:2508.21036(2025). https://arxiv.org/abs...

  24. [32]

    S.L. Taylor. 2016.Actor training and emotions: Finding a balance. Ph. D. Disserta- tion. Edith Cowan University

  25. [33]

    1978.Mind in society: The development of higher psychological processes

    Lev S Vygotsky. 1978.Mind in society: The development of higher psychological processes. Harvard University Press

Pith tools

Reviewed August 2, 2026 · model on record in the stance chip above.