Pith. sign in

REVIEW 4 major objections 7 minor 58 references

Plan Mode for spreadsheet agents cuts refinement and improves how people feel about the collaboration, without changing the final workbooks.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Plan Mode for spreadsheet agents shifts requirements to clarifying questions, reduces refinement, and improves perceived collaboration without changing final workbook quality.

T0 review reviewed 2026-07-30 challenge →

load-bearing objection Honest HCI study: Plan Mode shifts when requirements surface and improves comparative UX ratings, without changing final spreadsheets—and the “less refinement” headline is partly a coding artifact. the 4 major comments →

arxiv 2607.23670 v2 pith:JEKQWDRG submitted 2026-07-26 cs.HC cs.AIcs.SE

Plans Work in Mysterious Ways: Evaluating a Plan Mode for Spreadsheet Agents

classification cs.HC cs.AIcs.SE
keywords human-AI collaborationend-user programmingspreadsheet programminghuman-AI planningPlan Modeclarifying questionsagentic toolscreativity support
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Professional coding agents already let users co-write a plan before the agent edits anything. This paper asks whether that same Plan Mode helps ordinary spreadsheet users, who usually discover requirements as they go and care more about practical utility than technical correctness. The authors built a Plan Mode for an Excel agent—clarifying questions, an editable persistent plan, partial execution, and status updates—and compared it with a straight “Act Mode” baseline in a within-subjects study of 24 people on open-ended creation and analysis tasks. Final workbooks looked much the same across modes, yet Plan Mode shifted requirements from after-the-fact refinement into answers to clarifying questions, used fewer tokens, and earned stronger comparative ratings on creativity support and human–machine collaboration. The practical claim is that planning’s value for end-user spreadsheet work shows up mainly in the interaction and the sense of control, not in more diverse or higher-quality artifacts.

Core claim

In a within-subjects study (N=24), Plan Mode and Act Mode produced similar final spreadsheet features and outcomes, but Plan Mode reduced refinement-driven requirements and turns, lowered approximate token use, and produced better comparative perceptions across creativity-support and human–machine collaboration dimensions.

What carries the argument

Plan Mode: a separated planning agent that asks clarifying questions, produces a short editable persistent plan artifact users can partially execute, then hands off to an Act agent that updates step status—keeping planning read-only until the user explicitly executes.

Load-bearing premise

That short, time-capped tasks with mostly first-time agentic-spreadsheet users, plus a comparative preference scale and a more thorough Plan Mode tutorial, measure real lasting tool value rather than novelty or forced contrast—especially since non-comparative ratings of the outputs were nearly identical.

What would settle it

Repeat the same within-subjects design with longer real personal workbooks, matched tutorials, and both comparative and absolute instruments: if Plan Mode no longer reduces refine turns/tokens and loses the preference edge once novelty and tutorial asymmetry are removed, the central claim fails.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Plan Modes for end-user spreadsheet agents can be justified even when final artifact metrics do not improve, because interaction cost and perceived collaboration improve.
  • Clarifying questions become the main channel for requirements; designers should treat them as the lever for intent, not only the plan text.
  • Unused controls (edit plan, partial execute) can still supply design value as visible steerability.
  • Entry into Plan Mode should eventually be mixed-initiative: creation-heavy or longer tasks and less one-shot-prompt-skilled users benefited more.
  • Findings already fed a shipped Excel Copilot Plan Mode, so the interaction pattern is intended as a production design template.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If outputs stay homogeneous while clarifying questions supply many ‘new’ requirements, agents may be steering users onto a default feature checklist rather than expanding creative range—suggesting question design needs deliberate diversity or friction.
  • Selective/depth-first users who like live refine loops may need a lightweight ‘plan-on-demand mid-task’ path rather than mandatory upfront planning.
  • Token savings under similar requirement counts imply planning can be a cost-control feature for agentic spreadsheet products, independent of quality gains.
  • The creation-vs-analysis preference split suggests product defaults could auto-offer Plan Mode more aggressively for blank-workbook creation than for filter/sort analysis.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. The paper introduces a Plan Mode for agentic spreadsheet programming — a Plan agent that asks clarifying questions, produces a persistent, directly manipulable plan (≤7 steps), supports partial execution and dynamic step status, and hands off to a separate Act agent for execution — and evaluates it in a within-subjects study (N=24) against a matched non-planning Act Mode built on the same base LLM (Claude Sonnet 4.5) and spreadsheet tooling. Participants completed either two creation or two analysis tasks, one per mode, counterbalanced. The authors code user requirements by origin and iteration (κ=0.93 extraction, 0.85 origin, 0.62 iteration), code workbook features hierarchically, and administer a comparative adaptation of the Mixed-Initiative Creativity Support Scale. Findings: requirement elicitation shifts from post-edit refinement (Act) toward clarifying-question responses (Plan); Plan Mode shows fewer refine turns and lower approximate token use; final workbooks are similar across modes; participants prefer Plan Mode on the comparative instrument across creativity-support and collaboration dimensions, moderated by task type and interaction style. The authors discuss design implications, "control without action," and mixed-initiative entry; the work informed the shipped Plan Mode in Excel Copilot.

Significance. If the results hold, this is the first controlled evaluation of an explicit Plan Mode in an end-user programming domain, testing whether a programmer-tool affordance transfers to spreadsheet users whose iterative workflows were hypothesized to conflict with upfront planning. Strengths: a matched baseline sharing the same LLM and spreadsheet tools; counterbalanced task types and orderings; transparent qualitative coding with reported inter-rater reliability and a published codebook; hierarchical workbook-feature coding; a systematic feature mapping against four commercial plan modes (Table VIII); candid limitations, including the tutorial confound and Microsoft positionality; and documented product impact. The "control without action" observation (UI features valued despite near-zero use, §V-A3/§VI-B) is a genuinely useful design contribution, and the process instrumentation (requirement origins, turns, tokens) provides falsifiable behavioral measures beyond self-report.

major comments (4)
  1. §V-A2, Tables II, IV–V (and abstract): the 'reduction in refinement' claim is partially built into the coding scheme. A 'refine' origin requires that the agent has already edited the workbook, but the Act agent is instructed to execute immediately and asked almost no clarifying questions (0.2/session), so any requirement not stated upfront in Act Mode can only surface post-edit and be coded 'refine'; Plan Mode adds two pre-edit channels (clarification, plan) that reclassify the same latent requirements. The data fit this relocation reading: new requirements are identical (6.5 vs 6.4, Table V) and follow-ups are higher in Plan (2.0 vs 0.7). The interpretive claim that planning 'enabled participants to build spreadsheets more aligned with their goals' has no direct support — outputs are similar (§V-B) and Fig 2b ratings near-identical. Please reframe descriptively (the RQ1 phrasing in §I/§
  2. [§V-A2] §V-A2: the token result (11.0k vs 17.0k) is presented as 'a reduction in cost,' but Plan Mode sessions used more turns (5.3 vs 3.6) and more user time (10.4 vs 8.8 min), and the token figure is a characters/4 approximation that includes planning dialogue on one side and execute-then-revise rework on the other. As written, 'cost' silently shifts from model tokens to overall efficiency. Please scope the claim to model-side rework, state the approximation's limits, and present the user-time trade-off alongside it; the 15–20 minute time cap should also be acknowledged as compressing post-edit iteration differently across modes.
  3. [§IV-D3, §V-C, Fig 2] The perception half of the abstract rests on a comparatively rephrased, forced-choice adaptation of the MICSS ('I enjoyed using this tool more') administered after an asymmetric tutorial — Plan Mode's tutorial demonstrated the novel UI, Act Mode's did not. This measures contrast under unequal scaffolding, not durable tool quality; the paper's own non-comparative items (Fig 2b) are near-identical across modes, and the first author moderated all sessions and walked participants through their questionnaire answers, adding demand-characteristic risk. Limitations concedes all of this, but the abstract and §V-C state the preference result without qualification, and no inference (e.g., Wilcoxon/sign test against neutral) is reported for the 5-point comparative items. Please temper the headline claim and report basic inference or effect sizes.
  4. [§V-C, Figs 3, 4, 6, 7] The moderation claims — task type (n=12/arm, between-subjects since each participant saw only one task type), task length, upfront requirements (n=9), and refinement behavior (n=9 vs 15) — are stated as findings ('Participants preferred Plan Mode more for Creation tasks across all dimensions') but rest on very small subgroups with no inferential support across multiple post-hoc splits. These are interesting hypotheses, and §VI-C indeed treats them as such; §V-C should likewise label them exploratory and avoid dimension-level assertions, or report appropriate tests with multiplicity caveats.
minor comments (7)
  1. [§IV-D1, §V-A1] κ=0.62 for the iteration coding is the weakest IRR, yet it underlies the follow-up contrast (2.0 vs 0.7; Table XII counts are small: 49 vs 16). Please advise caution in interpretation and report per-category agreement.
  2. [Fig 2a] Neutral responses are omitted, which makes the strength of the preference hard to judge. Please include them or report full counts per item.
  3. [Tables II/XII, Fig 2] Table XII's header uses 'Fix' where Table II uses 'refine'; the text refers to 'Table 2b' though it is Fig 2(b). Please harmonize terminology and cross-references.
  4. [§V-A1, Table IV] §V-A1 says participants expressed a 'similar number of unique requirements,' but Table IV totals differ (9.2 vs 7.6). Clarify 'unique/new' vs total expressed requirements.
  5. [Table VII] The Budget tables median (Act 10.5 vs Plan 4.5) is a large artifact disparity that sits awkwardly next to the 'similar features' conclusion; a sentence on artifact bloat in Act Mode would help.
  6. [§VII] 'We continue to observe positive feedback' is unsubstantiated; please cite telemetry or remove.
  7. [Appendix B2, §IV-C] The Act Mode tutorial told participants the AI 'will not try to create a plan,' an additional priming/demand characteristic worth listing alongside the tutorial-asymmetry confound; also report how many participants hit the 20-minute cap in each mode.

Circularity Check

0 steps flagged

No circular derivation: empirical within-subjects HCI study; process and preference claims rest on coded logs, artifacts, and questionnaires, not quantities forced by definition or self-citation.

full rationale

This paper evaluates a Plan Mode prototype against an Act Mode baseline in a within-subjects user study (N=24). Load-bearing claims are empirical contrasts—requirement origins and iteration codes (Tables IV–VI), turn/time/token counts, spreadsheet feature distributions (Table VII, Appendix E), and comparative questionnaire preferences (Fig. 2, Table III)—not first-principles predictions or fitted parameters renamed as results. Related-work citations (including author-overlapping work such as TableTalk) motivate design features and interpret findings; they do not supply uniqueness theorems or force the Plan-vs-Act outcome differences. The design that Plan Mode asks clarifying questions before edits, so more requirements are coded origin=clarification rather than refine, is the intended mechanism under test and is reported transparently (shift from refinement toward clarifying questions; identical new-requirement counts 6.5 vs 6.4). That is construct/operationalization risk for the phrase “reduction in refinement,” not circularity by construction of a derivation chain. No self-definitional loop, fitted-input-as-prediction, load-bearing self-citation uniqueness, ansatz smuggling, or renaming of a known law appears. Score 0.

Axiom & Free-Parameter Ledger

2 free parameters · 4 axioms · 1 invented entities

Load-bearing content is empirical method choice plus standard end-user programming background assumptions, not free parameters in a model. The central contrast rests on treating Act Mode as an adequate non-planning baseline, on the validity of the requirement/feature coding ontologies, and on interpreting comparative MICSS-style ratings as collaboration quality under short lab tasks.

free parameters (2)
  • Plan length cap (≤7 steps, one sentence each) = 7 steps, one sentence per step
    Hand-chosen instruction constraint that shapes what users see as a ‘plan’ and may suppress implementation detail and diversity of plans.
  • Task time budget (~15–20 minutes) = ~15 minutes instructed, up to 20 allowed
    Experimenter-chosen cap that can truncate refinement and artifact complexity differences the authors themselves flag as a limitation.
axioms (4)
  • domain assumption End-user spreadsheet programmers work iteratively, discover requirements emergently, and prioritize practical utility over technical correctness (invoking Ko et al.; Pandita et al.).
    Motivates the research question that Plan Mode may not fit spreadsheet workflows; frames interpretation of refinement vs planning.
  • ad hoc to paper A matched agentic Act Mode with the same base LLM and spreadsheet tools is a fair non-planning baseline for attributing effects to planning interaction rather than model capability.
    Central causal contrast in §IV; both modes share Claude Sonnet 4.5 and Act execution stack.
  • ad hoc to paper Coded ‘requirements’ (new information types / filter-sort criteria) and hierarchical workbook ‘features’ are valid proxies for information exchange and output quality/diversity.
    RQ1–RQ2 measurement strategy in §IV-D; IRR reported but constructs are study-specific.
  • domain assumption Comparative preference items adapted from the Mixed-Initiative Creativity Support Scale measure differences in creativity support and human–machine collaboration between modes.
    RQ3 instrument (§IV-D3, Table III); authors modified immersiveness→attention and added capability.
invented entities (1)
  • Spreadsheet Plan Mode prototype (Plan agent + Act agent split, persistent manipulable plan pane, partial execution, dynamic step status, read-only planning) independent evidence
    purpose: Operationalize interactive human–AI planning for Excel-style agentic tasks as the experimental treatment.
    Bundle of design features assembled from coding-agent Plan Modes plus spreadsheet-specific partial execution; evaluated as one system rather than ablated.

reviewed 2026-07-30 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Plans Work in Mysterious Ways: Evaluating a Plan Mode for Spreadsheet Agents." pith.science (2026). https://pith.science/paper/JEKQWDRG

@misc{pith2026260723670,
  author       = {Pith},
  title        = {Pith review of: Plans Work in Mysterious Ways: Evaluating a Plan Mode for Spreadsheet Agents},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JEKQWDRG}},
  note         = {Machine review of arXiv:2607.23670}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Plan Modes have become standard features in agentic programming tools, allowing users to gain transparency and control by working with the agent to develop a plan before task execution. However, it remains unclear whether the benefits of this feature translate to end-user programming environments such as spreadsheets. Since spreadsheet programmers tend to work iteratively and care less about technical correctness, upfront planning may not fit into their workflows as easily. In this paper, we build a prototype of a Plan Mode for spreadsheet programming and evaluate it against a non-planning baseline through a within-subjects user study (N=24). We found that despite similar task outcomes with both tools, using Plan Mode led to a reduction in refinement and a better perception of the tool across dimensions of creativity support and human-machine collaboration. We discuss the implications of these results for the future design of Plan Modes, and for the broader role of human-AI planning in end-user programming.

Figures

Figures reproduced from arXiv: 2607.23670 by Aayush Kumar, Advait Sarkar, Avik Dutta, Emerson Murphy-Hill, Gustavo Soares, Sumit Gulwani.

Figure 1
Figure 1. Figure 1: Interacting With Plan Mode plan during execution affordances are often reified through shared artifacts editable by users and AI agents that represent actions taken/to be taken by the agent. For example, He et al. [20] and Mozannar et al. [21] introduce tools with an iterative loop where the agent proposes a plan that users can update before execution. Cocoa [22] maintains a similar shared plan, but furthe… view at source ↗
Figure 3
Figure 3. Figure 3: Participant preferences based on task type [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figure 2
Figure 2. Figure 2: Questionnaire responses by participants strongly preferred Plan Mode. These specific dimensions indicate that participants working on Creation tasks felt more like they were interacting with a collaborative partner when working with Plan Mode. Further, those working on Analysis tasks felt that they had to pay more attention to the tool rather than the task when working with Plan Mode (unlike for Creation t… view at source ↗
Figure 6
Figure 6. Figure 6: Participant preferences across task complexity. [PITH_FULL_IMAGE:figures/full_fig_p015_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Participant preferences based on upfront requirements. [PITH_FULL_IMAGE:figures/full_fig_p015_7.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

58 extracted references · 1 linked inside Pith

  1. [1]

    Challenges in Human-Agent Communication,

    G. Bansal, J. W. Vaughan, S. Amershi, E. Horvitz, A. Fourney, H. Mozannar, V . Dibia, and D. S. Weld, “Challenges in Human-Agent Communication,” 2024

  2. [2]

    Choose a Permission Mode – Claude Code Docs

    Anthropic, “Choose a Permission Mode – Claude Code Docs.” https:// code.claude.com/docs/en/permission-modes, 2026. Accessed: 2026-05- 01

  3. [3]

    Planning with Agents in VS Code

    Microsoft, “Planning with Agents in VS Code.” https://code.visualstudio. com/docs/copilot/agents/planning, 2026. Accessed: 2026-05-01

  4. [4]

    Plan Mode – Cursor Docs

    Cursor, “Plan Mode – Cursor Docs.” https://cursor.com/docs/agent/ plan-mode, 2026. Accessed: 2026-05-01

  5. [5]

    Plan & Act Mode – Cline Documentation

    Cline, “Plan & Act Mode – Cline Documentation.” https://docs.cline. bot/core-workflows/plan-and-act, 2026. Accessed: 2026-05-01

  6. [6]

    The state of the art in end-user software engineering,

    A. J. Ko, R. Abraham, L. Beckwith, A. Blackwell, M. Burnett, M. Erwig, C. Scaffidi, J. Lawrance, H. Lieberman, B. Myers, M. B. Rosson, G. Rothermel, M. Shaw, and S. Wiedenbeck, “The state of the art in end-user software engineering,”ACM Comput. Surv., vol. 43, Apr. 2011

  7. [7]

    What is it like to program with artificial intelligence?,

    A. Sarkar, A. D. Gordon, C. Negreanu, C. Poelitz, S. S. Ragavan, and B. Zorn, “What is it like to program with artificial intelligence?,” 2022

  8. [8]

    No half- measures: A study of manual and tool-assisted end-user programming tasks in Excel,

    R. Pandita, C. Parnin, F. Hermans, and E. Murphy-Hill, “No half- measures: A study of manual and tool-assisted end-user programming tasks in Excel,” in2018 IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC), pp. 95–103, 2018

  9. [9]

    AI, Help Me Think—but for Myself: Assisting People in Complex Decision-Making by Providing Different Kinds of Cognitive Support,

    L. Reicherts, Z. T. Zhang, E. von Oswald, Y . Liu, Y . Rogers, and M. Hassib, “AI, Help Me Think—but for Myself: Assisting People in Complex Decision-Making by Providing Different Kinds of Cognitive Support,” inProceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI ’25, (New York, NY , USA), Association for Computing Machinery, 2025

  10. [10]

    Improving Steering and Verification in AI-Assisted Data Analysis with Interactive Task Decomposition,

    M. Kazemitabaar, J. Williams, I. Drosos, T. Grossman, A. Z. Henley, C. Negreanu, and A. Sarkar, “Improving Steering and Verification in AI-Assisted Data Analysis with Interactive Task Decomposition,” in Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology, UIST ’24, (New York, NY , USA), Association for Computing Machinery, 2024

  11. [11]

    Dango: A Mixed- Initiative Data Wrangling System using Large Language Model,

    W.-H. Chen, W. Tong, A. Case, and T. Zhang, “Dango: A Mixed- Initiative Data Wrangling System using Large Language Model,” inPro- ceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI ’25, (New York, NY , USA), Association for Computing Machinery, 2025

  12. [12]

    Dynamic Prompt Middleware: Contextual Prompt Refinement Controls for Comprehension Tasks,

    I. Drosos, J. Williams, A. Sarkar, N. Wilson, S. Rintel, and P. Panda, “Dynamic Prompt Middleware: Contextual Prompt Refinement Controls for Comprehension Tasks,” inProceedings of the 4th Annual Symposium on Human-Computer Interaction for Work, CHIWORK ’25, (New York, NY , USA), Association for Computing Machinery, 2025

  13. [13]

    TableTalk: Scaffolding Spreadsheet Development with a Language Agent,

    J. T. Liang, A. Kumar, Y . Bajpai, S. Gulwani, V . Le, C. Parnin, A. Radhakrishna, A. Tiwari, E. Murphy-Hill, and G. Soares, “TableTalk: Scaffolding Spreadsheet Development with a Language Agent,”ACM Trans. Comput.-Hum. Interact., vol. 32, Dec. 2025

  14. [14]

    “What It Wants Me To Say

    M. X. Liu, A. Sarkar, C. Negreanu, B. Zorn, J. Williams, N. Toronto, and A. D. Gordon, ““What It Wants Me To Say”: Bridging the Abstraction Gap Between End-User Programmers and Code-Generating Large Language Models,” inProceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI ’23, (New York, NY , USA), Association for Computing Mac...

  15. [15]

    Beyond Code Generation: LLM-supported Exploration of the Program Design Space,

    J. Zamfirescu-Pereira, E. Jun, M. Terry, Q. Yang, and B. Hartmann, “Beyond Code Generation: LLM-supported Exploration of the Program Design Space,” inProceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI ’25, (New York, NY , USA), Association for Computing Machinery, 2025

  16. [16]

    Why AI Agents Still Need You: Findings from Developer-Agent Collaborations in the Wild,

    A. Kumar, Y . Bajpai, S. Gulwani, G. Soares, and E. Murphy-Hill, “Why AI Agents Still Need You: Findings from Developer-Agent Collaborations in the Wild,” in2025 40th IEEE/ACM International Conference on Automated Software Engineering (ASE), pp. 432–444, 2025

  17. [17]

    BISCUIT: Scaffolding LLM-Generated Code with Ephemeral UIs in Computational Notebooks,

    R. Cheng, T. Barik, A. Leung, F. Hohman, and J. Nichols, “BISCUIT: Scaffolding LLM-Generated Code with Ephemeral UIs in Computational Notebooks,” in2024 IEEE Symposium on Visual Languages and Human- Centric Computing (VL/HCC), pp. 13–23, 2024

  18. [18]

    DataSpeck: An AI-Driven Human-in-the-Loop System for Automating Transformations in Data Conversion Workflows,

    A. Rahman, K. Niinuma, and A. Gupta, “DataSpeck: An AI-Driven Human-in-the-Loop System for Automating Transformations in Data Conversion Workflows,” inProceedings of the 2026 CHI Conference on Human Factors in Computing Systems, CHI ’26, (New York, NY , USA), Association for Computing Machinery, 2026

  19. [19]

    ”It’s like a rubber duck that talks back

    I. Drosos, A. Sarkar, X. Xu, C. Negreanu, S. Rintel, and L. Tankelevitch, “”It’s like a rubber duck that talks back”: Understanding Generative AI- Assisted Data Analysis Workflows through a Participatory Prompting Study,” inProceedings of the 3rd Annual Meeting of the Symposium on Human-Computer Interaction for Work, CHIWORK ’24, (New York, NY , USA), Ass...

  20. [20]

    Plan-Then-Execute: An Em- pirical Study of User Trust and Team Performance When Using LLM Agents As A Daily Assistant,

    G. He, G. Demartini, and U. Gadiraju, “Plan-Then-Execute: An Em- pirical Study of User Trust and Team Performance When Using LLM Agents As A Daily Assistant,” inProceedings of the 2025 CHI Confer- ence on Human Factors in Computing Systems, CHI ’25, (New York, NY , USA), Association for Computing Machinery, 2025

  21. [21]

    Magentic-UI: Towards Human-in-the-loop Agentic Systems,

    H. Mozannar, G. Bansal, C. Tan, A. Fourney, V . Dibia, J. Chen, J. Gerrits, T. Payne, M. K. Maldaner, M. Grunde-McLaughlin,et al., “Magentic-UI: Towards Human-in-the-loop Agentic Systems,”arXiv preprint arXiv:2507.22358, 2025

  22. [22]

    Cocoa: Co-Planning and Co- Execution with AI Agents,

    K. J. K. Feng, K. Pu, M. Latzke, T. August, P. Siangliulue, J. Bragg, D. S. Weld, A. X. Zhang, and J. C. Chang, “Cocoa: Co-Planning and Co- Execution with AI Agents,” inProceedings of the 2026 CHI Conference on Human Factors in Computing Systems, CHI ’26, (New York, NY , USA), Association for Computing Machinery, 2026

  23. [23]

    From correctness to collaboration: Toward a human-centered framework for evaluating ai agent behavior in software engineering,

    T. Dong, H. Sampath, J. Y . Lee, S. Y . Shi, and A. Macvean, “From correctness to collaboration: Toward a human-centered framework for evaluating ai agent behavior in software engineering,” 2025

  24. [24]

    Codeplan: Repository-level coding using llms and planning,

    R. Bairi, A. Sonwane, A. Kanade, V . D. C., A. Iyer, S. Parthasarathy, S. Rajamani, B. Ashok, and S. Shet, “Codeplan: Repository-level coding using llms and planning,”Proc. ACM Softw. Eng., vol. 1, July 2024

  25. [25]

    Agentic AI has a Human Oversight Problem,

    S. Passi, “Agentic AI has a Human Oversight Problem,”Available at SSRN 5529058, 2025

  26. [26]

    SheetCopilot: Bringing Soft- ware Productivity to the Next Level through Large Language Models,

    H. Li, J. Su, Y . Chen, Q. Li, and Z. Zhang, “SheetCopilot: Bringing Soft- ware Productivity to the Next Level through Large Language Models,” 2023

  27. [27]

    SheetAgent: Towards A Generalist Agent for Spreadsheet Reasoning and Manipulation via Large Language Models,

    Y . Chen, Y . Yuan, Z. Zhang, Y . Zheng, J. Liu, F. Ni, J. Hao, H. Mao, and F. Zhang, “SheetAgent: Towards A Generalist Agent for Spreadsheet Reasoning and Manipulation via Large Language Models,” 2025

  28. [28]

    The Invisible Mentor: Inferring User Actions from Screen Record- ings to Recommend Better Workflows,

    L. Yan, A. Head, K. Milne, V . Le, S. Gulwani, C. Parnin, and E. Murphy- Hill, “The Invisible Mentor: Inferring User Actions from Screen Record- ings to Recommend Better Workflows,” inProceedings of the 2026 CHI Conference on Human Factors in Computing Systems, CHI ’26, (New York, NY , USA), Association for Computing Machinery, 2026

  29. [29]

    COLDECO: An End User Spread- sheet Inspection Tool for AI-Generated Code,

    K. Ferdowsi, J. Williams, I. Drosos, A. D. Gordon, C. Negreanu, N. Po- likarpova, A. Sarkar, and B. Zorn, “COLDECO: An End User Spread- sheet Inspection Tool for AI-Generated Code,” in2023 IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC), pp. 82– 91, 2023

  30. [30]

    Excel Copilot

    Microsoft, “Excel Copilot.” https://support.microsoft.com/en-US/excel/ copilot/get-started-with-copilot-in-excel, 2026. Accessed: 2026-06-21

  31. [31]

    Shortcut AI

    Shortcut, “Shortcut AI.” https://shortcut.ai/, 2026. Accessed: 2026-06- 21

  32. [32]

    Claude for Excel

    Anthropic, “Claude for Excel.” https://claude.com/claude-for-excel,

  33. [33]

    Spreadsheet Comprehension: Guesswork, Giving Up and Going Back to the Author,

    S. Srinivasa Ragavan, A. Sarkar, and A. D. Gordon, “Spreadsheet Comprehension: Guesswork, Giving Up and Going Back to the Author,” inProceedings of the 2021 CHI Conference on Human Factors in Computing Systems, CHI ’21, (New York, NY , USA), Association for Computing Machinery, 2021

  34. [34]

    Will Code Remain a Relevant User Interface for End- User Programming with Generative AI Models?,

    A. Sarkar, “Will Code Remain a Relevant User Interface for End- User Programming with Generative AI Models?,” inProceedings of the 2023 ACM SIGPLAN International Symposium on New Ideas, New Paradigms, and Reflections on Programming and Software, Onward! 2023, (New York, NY , USA), p. 153–167, Association for Computing Machinery, 2023

  35. [35]

    How Knowledge Workers Use and Want to Use LLMs in an Enterprise Context,

    M. Brachman, A. El-Ashry, C. Dugan, and W. Geyer, “How Knowledge Workers Use and Want to Use LLMs in an Enterprise Context,” inExtended Abstracts of the CHI Conference on Human Factors in Computing Systems, CHI EA ’24, (New York, NY , USA), Association for Computing Machinery, 2024

  36. [36]

    Reliability and Inter-rater Reliability in Qualitative Research: Norms and Guidelines for CSCW and HCI Practice,

    N. McDonald, S. Schoenebeck, and A. Forte, “Reliability and Inter-rater Reliability in Qualitative Research: Norms and Guidelines for CSCW and HCI Practice,”Proc. ACM Hum.-Comput. Interact., vol. 3, Nov. 2019

  37. [37]

    The measurement of observer agreement for categorical data,

    J. R. Landis and G. G. Koch, “The measurement of observer agreement for categorical data,”Biometrics, vol. 33, no. 1, pp. 159–174, 1977

  38. [38]

    Drawing with Reframer: Emergence and Control in Co-Creative AI,

    T. Lawton, F. J. Ibarrola, D. Ventura, and K. Grace, “Drawing with Reframer: Emergence and Control in Co-Creative AI,” inProceedings of the 28th International Conference on Intelligent User Interfaces, IUI ’23, (New York, NY , USA), p. 264–277, Association for Computing Machinery, 2023

  39. [39]

    Complacency and Bias in Human Use of Automation: An Attentional Integration,

    R. Parasuraman and D. Manzey, “Complacency and Bias in Human Use of Automation: An Attentional Integration,”Human factors, vol. 52, pp. 381–410, 06 2010

  40. [40]

    Does Writing with Language Models Reduce Content Diversity?,

    V . Padmakumar and H. He, “Does Writing with Language Models Reduce Content Diversity?,” inThe Twelfth International Conference on Learning Representations, 2024

  41. [41]

    AI Should Challenge, Not Obey,

    A. Sarkar, “AI Should Challenge, Not Obey,”Commun. ACM, vol. 67, p. 18–21, Sept. 2024

  42. [42]

    Understanding, Protecting, and Augmenting Human Cognition with Generative AI: A Synthesis of the CHI 2025 Tools for Thought Workshop,

    L. Tankelevitch, E. L. Glassman, J. He, A. Kittur, M. Lee, S. Palani, A. Sarkar, G. Ramos, Y . Rogers, and H. Subramonyam, “Understanding, Protecting, and Augmenting Human Cognition with Generative AI: A Synthesis of the CHI 2025 Tools for Thought Workshop,” 2025

  43. [43]

    How Do Users Discover New Tools in Software Development and Beyond?,

    E. Murphy-Hill, D. Y . Lee, G. C. Murphy, and J. Mcgrenere, “How Do Users Discover New Tools in Software Development and Beyond?,” Comput. Supported Coop. Work, vol. 24, p. 389–422, Oct. 2015

  44. [44]

    Don’t Just Tell Me, Ask Me: AI Systems that Intelligently Frame Explanations as Questions Improve Human Logical Discernment Accuracy over Causal AI explanations,

    V . Danry, P. Pataranutaporn, Y . Mao, and P. Maes, “Don’t Just Tell Me, Ask Me: AI Systems that Intelligently Frame Explanations as Questions Improve Human Logical Discernment Accuracy over Causal AI explanations,” inProceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI ’23, (New York, NY , USA), Association for Computing Mach...

  45. [45]

    Design Frictions for Mindful Interactions: The Case for Microbound- aries,

    A. L. Cox, S. J. Gould, M. E. Cecchinato, I. Iacovides, and I. Renfree, “Design Frictions for Mindful Interactions: The Case for Microbound- aries,” inProceedings of the 2016 CHI Conference Extended Abstracts on Human Factors in Computing Systems, CHI EA ’16, (New York, NY , USA), p. 1389–1397, Association for Computing Machinery, 2016

  46. [46]

    To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI- assisted Decision-making,

    Z. Buc ¸inca, M. B. Malaya, and K. Z. Gajos, “To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI- assisted Decision-making,”Proc. ACM Hum.-Comput. Interact., vol. 5, Apr. 2021

  47. [47]

    Measuring User Experience Inclusivity in Human-AI Interaction via Five User Problem-Solving Styles,

    A. Anderson, J. N. Guevara, F. Moussaoui, T. Li, M. V orvoreanu, and M. Burnett, “Measuring User Experience Inclusivity in Human-AI Interaction via Five User Problem-Solving Styles,”ACM Trans. Interact. Intell. Syst., vol. 14, Sept. 2024. APPENDIX A. Relation of Design Features to Existing Plan Modes Table VIII summarizes how each of our prototype’s desig...

  48. [49]

    Which of the following apps do you use for professional purposes more than once a week? Select all that apply. •Notion (May Select) •Spreadsheet.com (May Select) •Smartsheet (May Select) •Adobe Photoshop (May Select) •Google Docs (May Select) •Gmail (May Select) •Microsoft Word (May Select) •Microsoft Excel (Qualify) •Google Slides (May Select) •Microsoft...

  49. [50]

    How do you most frequently access productivity applica- tions like Word, Excel, or PowerPoint? •Smartphone app (Disqualify) •Internet browser (Disqualify) •Desktop or laptop application (Qualify) •Tablet (Android/iPad) app (Disqualify)

  50. [51]

    How frequently do you use spreadsheets? •Once a month (Disqualify) •Once a week (Qualify) •Multiple times a week (Qualify) •Never (Disqualify)

  51. [52]

    What do you primarily use spreadsheets for? Please select all that apply. •Viewing information and charts created by others (May Select) •Creating new spreadsheets with structured/unstructured data (Qualify) •Editing existing spreadsheets through data entry/visu- alization (Qualify) •Performing in-depth data modelling/analysis (Qualify) •None of the above...

  52. [53]

    What is your expertise level with Excel? •Beginner (Basic data wrangling, sheet formatting) (Disqualify) •Intermediate (Formulas like SUM, COUNT, MAX, charts, PivotTables) (Qualify) •Advanced (Power Query, Data Model, formulas like VLOOKUP) (Qualify) •Expert (Data Model, indirect references, Python for- mulas, formulas like LET and LAMBDA) (Qualify)

  53. [54]

    aModality of the clarifying-question interaction: NL = natural language in chat; MCQ = multiple-choice questions through UI cards

    Have you used/heard of VLOOKUP in Excel before? •I have used VLOOKUP before (Qualify) •I have not used VLOOKUP before, but I have heard of it and know what it does (Qualify) •I have not used VLOOKUP before and I do not know what it does, but I have heard of it (Disqualify) •I have not used or heard of VLOOKUP before (Dis- Design Feature Our Prototype Clau...

  54. [55]

    Which of the following would you use to describe your- self? Choose the answer that best fits your scenario. •A college, graduate, university, or a post-graduate student (Disqualify) •Employee or Owner of a business with less than 50 employees (Qualify) •Employee or Owners of a business with 50 or more employees (Qualify) •Freelance Worker, Gig Worker or ...

  55. [56]

    What is your experience with using AI? •Extensive – I use AI-powered tools frequently for work and/or personal tasks (Qualify) •Occasional – I have used AI tools but only in limited scenarios (Qualify) •Rare – I have only tried AI tools a few times (Qualify) •None – I have never used AI for either personal or work tasks (Disqualify)

  56. [57]

    I want to create a gradebook for my class of grade 10 math students

    Study Protocol:Sessions followed one of two orderings, Plan Mode FirstorAct Mode First, counterbalanced across participants. Both followed the same structure: an introduction, a warm-up/tutorial task with the first tool, two 15-minute study tasks (one per tool), and a post-study questionnaire. Here, we provide the full protocol for the ordering where part...

  57. [58]

    My attention was fully tuned to the activity, and I forgot about the system or tool that I was using

    Task Descriptions: •Creation Tasks –Personal Budget:Try and create your own personal budget, based on your actual needs and preferences. Your objective is to create a useful spreadsheet that you could use practically to track your finances. We will send you the spreadsheet after the session for you to use if you want to. –Personal Schedule:Create a person...

  58. [2026]

    Accessed: 2026-06-21

This paper was first reviewed by grok-4.5 on July 30, 2026.