REVIEW 4 major objections 4 minor 2 cited by
TableTalk: Scaffolding Spreadsheet Development with a Language Agent
T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A step-by-step AI guide builds better spreadsheets than a free-form assistant, with lower cognitive load.
desk verdict A solid system paper with a real confound between the model and the design principles; the tool-level preference claim is plausible, the causal claim is not yet supported. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the TableTalk language agent, which implements a plan based on Pirolli and Card's expert sensemaking process for creating knowledge products from data: gather requirements, define a data-table schema, and create insight tables. Its mechanism has three parts: a system prompt that walks the agent through this plan; a chat interface that presents three model-generated suggestion pills for the next step, so the human can steer the plan; and a set of pre-written OfficeScript tools (create_table, sort_rows, filter_rows, add_chart, highlight_cell) that let the agent build atomic spreadsheet components incrementally rather than generating a whole spreadsheet in one shot.
What would settle it
Run the same 20-participant study with the baseline replaced by a version that uses the same underlying model and identical warm-up instructions; if the 70% preference for TableTalk disappears or reverses, the design principles themselves are not what produced the benefit.
Extended reading notes
Core claim
The central claim is that a language agent that scaffolds spreadsheet development through a structured expert plan, gather requirements, define the data schema, and extract insights, while offering three adaptive next-step suggestions and building spreadsheets incrementally with atomic tools, produces higher-quality spreadsheets and a better programmer experience than a non-scaffolded baseline agent. The evidence is a 20-participant controlled study in which blinded evaluators preferred TableTalk's spreadsheets 42 of 60 times, a Bradley-Terry ability score of 0.85 (2.3x odds, p<0.001), and a 1.9-minute reduction in thinking or verifying time, which is 12.6% of the total task time. The paper further reports lower mental demand and better conversational relevance scores for TableTalk.
Load-bearing premise
The comparison assumes the only meaningful difference between TableTalk and the baseline is the three design principles, not the underlying model or the way each tool was explained to participants.
Editorial extensions
If this is right
- Spreadsheet agents that scaffold the process will be preferred over non-scaffolded ones for open-ended analysis tasks, even when the task is too hard to finish in the allotted time.
- Programmers using a scaffolded agent spend more effort on requirements and high-level commands and less on low-level implementation details and verification.
- Lower mental demand and thinking time could make spreadsheet tools usable by more people, especially those who struggle with formulas and schema design.
- The three design principles, scaffolding, flexibility, and incrementality, can be applied to other end-user programming domains such as debugging, data cleaning, and report generation.
- The tradeoff between proactivity and user control must be made adjustable; participants were split on whether the agent should act autonomously or wait for explicit confirmation.
Reading between the lines
- If the mechanism is human-in-the-loop planning, then removing the suggestion pills or making them generic should reduce the quality gap; this is a testable ablation.
- The positive results may hinge on TableTalk using GPT-4o while the baseline used a different model; a matched-model replication would clarify whether the design principles or the underlying model cause the effect.
- The 12.6% cognitive-load reduction likely combines two separate effects: the scaffold reduces planning burden, but the agent's latency adds waiting time, so net load may vary in settings with faster or slower models.
- The paper's design guidelines imply that future agents should expose adjustable proactivity and direct manipulation controls like undo and stop, to satisfy participants who felt the proactive agent was invasive.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents TableTalk, a language-agent system for spreadsheet programming that embodies three design principles: scaffolding, flexibility, and incrementality. These principles are derived from an analysis of 85 Excel templates and a formative study with 7 spreadsheet programmers. TableTalk guides users through an expert-derived plan, proposes three next-step suggestions, and builds spreadsheets incrementally using atomic OfficeScript-based tools. The evaluation is a controlled study with 20 spreadsheet programmers who each used TableTalk and a baseline (Excel Copilot version 2409) on two spreadsheet tasks. The reported headline results are that spreadsheet evaluators preferred TableTalk-created spreadsheets 42 out of 60 times (70%), with a Bradley-Terry ability score of 0.85 (2.3 times higher odds of preference, p<0.001), and that TableTalk reduced self-reported mental demand and time spent thinking about spreadsheet actions. The paper concludes with design guidelines for agentic spreadsheet tools and a discussion of implications for spreadsheet programming, end-user programming, and human-agent collaboration.
Significance. If the causal claims held, this would be a valuable contribution to HCI and software engineering: it provides a concrete, well-described system that operationalizes three design principles, and it offers an unusually rich mixed-methods evaluation with open materials, inter-rater reliability checks, and candid limitation statements. The qualitative analyses of chat logs, activity timelines, evaluator comments, and interviews are a genuine strength, as are the publicly available protocols and codebooks. However, the central empirical claim that the design principles themselves drive the observed quality and cognitive-load benefits is not supported by the current comparison, because the two tools differ not only in the principles but also in underlying model and in tutorial instructions. The paper is therefore best read as a tool-level comparison of TableTalk against Excel Copilot; the broader design-guideline contribution requires either a repaired experiment or a substantially more cautious framing.
major comments (4)
- [Section 6.3 and Section 7.3] The evaluation does not isolate the three design principles from the underlying model and the warm-up instructions. Section 6.3 states that TableTalk uses GPT-4o through the OpenAI Assistants API, while Section 7.3 identifies the baseline only as 'Excel Copilot available in Excel version 2409' without naming its model. The same section gives baseline users the extra instruction to 'include a table with headers in the spreadsheet for the best results,' a tutorial that TableTalk users did not receive. Consequently, the 70% preference result (Section 7.5.1) and the mental-demand reduction (Section 7.5.3) could be driven by model capability or by the tutorial asymmetry rather than by scaffolding, flexibility, and incrementality. Since Section 8 attributes the results to the design principles, the paper overreaches. Please either re-frame the central claims as a tool-level comparison, add an ablation or matched-model control, or provide a rigorous argument that the only systematic differences between the two tools are the implemented design principles.
- [Section 7.2 and Section 7.5] The quality comparison is of partial artifacts, not completed spreadsheets. Section 7.5 states that no participant completed any task, and Section 7.2 explains the tasks were intentionally difficult and time-limited. The paper anticipates this by focusing on spreadsheet quality rather than completion, but the interaction between the 15-minute limit and the tools' very different latencies is not addressed. Section 7.5.2 reports that TableTalk users collectively spent 61.0 additional minutes waiting, meaning the two conditions differ in the time available for manual work; the 'more polished' TableTalk artifacts could reflect automation advantages or the structure of the tool rather than the specific design principles. At minimum, the manuscript should discuss whether evaluator preference for partial artifacts is a valid proxy for the quality of a completed spreadsheet, and how the latency difference might affect the comparison.
- [Section 7.4] The Bradley-Terry analysis is reported as a point estimate and a p-value, but the paper does not provide a confidence interval, cluster-robust standard errors, or an account of the repeated-measures structure. The 60 pairwise comparisons are not independent: the spreadsheets come from 20 participants (each contributing one artifact per tool) and are judged by 6 evaluators. Please report the uncertainty in the ability score and test the sensitivity of the conclusion to a model that treats participants or evaluators as random effects, or to a bootstrap procedure that resamples by participant or evaluator.
- [Section 8.3] The limitations section discusses confirmation bias, the shared-computer environment, and external validity, but it does not acknowledge the most direct internal-validity threat identified above: the mismatch between the models used by TableTalk and the baseline, and the asymmetric warm-up instruction. Since Section 8.3 does enumerate internal validity threats, the omission of the model confound is a gap in the manuscript's own self-assessment. Please add an explicit limitation and state precisely which conclusions can and cannot be drawn from the current comparison.
minor comments (4)
- [Section 2.3] The heading '2.3 AI for spreasheet programming' contains a typo: 'spreasheet' should be 'spreadsheet'.
- [Section 8.1.5] In the paragraph about code comprehension, the participant is identified as P9 in the first sentence but the quotation 'if score is below 100, do not display excellent' is attributed to P10; this appears to be an inconsistency and should be corrected to P9.
- [Acknowledgments] The reference [61] lists an author as 'Arjun Rahakrishna' instead of 'Arjun Radhakrishna'.
- [Figure 7] The activity label 'T ableT alk' in the figure legend appears to be a formatting artifact; it should read 'TableTalk'.
Circularity Check
No significant circularity: the central claims are evaluated against external human preferences and an external baseline, not the paper's own definitions or fitted quantities.
full rationale
No circular step meets the quoted-reduction standard. TableTalk's central quality claim (Section 7.5.1) is assessed by external spreadsheet evaluators comparing spreadsheets produced by TableTalk and an external baseline, with the Bradley-Terry ability score (0.85; 42/60 preferences) fitted to those external pairwise preferences rather than to any quantity defined by the authors' design principles. The design principles (DP1-DP3) are derived from the template study (Section 3) and the 7-participant formative study (Section 4), not from the evaluation outcome, so the principles are not self-defined in terms of the predicted result. The cognitive-load and conversation-quality measures (Section 7.5.3) come from standard instruments (NASA TLX, SUS, Finch and Choi's conversation quality items) and participant self-reports, not from quantities that the paper constructs. The paper does cite prior work with overlapping authors, notably ROBIN [19], but only as related scaffolding context and design inspiration; the load-bearing evidence for the preference result is the external human evaluation. The paper's own limitations (Section 8.3) acknowledge confirmation bias, environment effects, and generalizability threats, and the reader's confound concern about the GPT-4o backend and asymmetric warm-up instructions is a validity threat to causal attribution rather than a circularity finding. Because the claimed outcome is measured against external human preference data and an external baseline rather than being equivalent to an input by construction, the appropriate circularity score is 0.
Assumptions & free parameters
free parameters (1)
- TableTalk Bradley-Terry ability score =
0.85 (log-odds)
assumptions (5)
- domain assumption Pirolli and Card's expert sensemaking process is the right scaffolding structure for spreadsheet creation.
- domain assumption Evaluator preference among unfinished spreadsheets is a valid proxy for spreadsheet quality.
- domain assumption Baseline and TableTalk are comparable except for design-principle variations.
- domain assumption Self-reported NASA-TLX, SUS, and conversation-quality items measure perceived load and experience without systematic bias.
- domain assumption GPT-4o's behavior is stable enough that study outcomes reflect the system design rather than random sampling of model outputs.
Cite this review
Pith. "Pith review of TableTalk: Scaffolding Spreadsheet Development with a Language Agent." pith.science (2026). https://pith.science/paper/ECR3WKWR
@misc{pith2026250209787,
author = {Pith},
title = {Pith review of: TableTalk: Scaffolding Spreadsheet Development with a Language Agent},
year = {2026},
howpublished = {\url{https://pith.science/paper/ECR3WKWR}},
note = {Machine review of arXiv:2502.09787}
}
read the original abstract
Spreadsheet programming is challenging. Programmers use spreadsheet programming knowledge (e.g., formulas) and problem-solving skills to combine actions into complex tasks. Advancements in large language models have introduced language agents that observe, plan, and perform tasks, showing promise for spreadsheet creation. We present TableTalk, a spreadsheet programming agent embodying three design principles -- scaffolding, flexibility, and incrementality -- derived from studies with seven spreadsheet programmers and 85 Excel templates. TableTalk guides programmers through structured plans based on professional workflows, generating three potential next steps to adapt plans to programmer needs. It uses pre-defined tools to generate spreadsheet components and incrementally build spreadsheets. In a study with 20 programmers, TableTalk produced higher-quality spreadsheets 2.3 times more likely to be preferred than the baseline. It reduced cognitive load and thinking time by 12.6%. From this, we derive design guidelines for agentic spreadsheet programming tools and discuss implications on spreadsheet programming, end-user programming, AI-assisted programming, and human-agent collaboration.
Figures
Figures from the paper (11 more)
Forward citations
Cited by 2 Pith papers
-
Understanding User Experiences of Computer Use Agents: Design Space and Opportunities for Building Agent UX Prototypes
Agent prototyping for non-experts requires scaffolds for scoping the agent, designing its chat/UI display, defining user interactions, running it, and debugging its runtime behavior.
-
SheetMind: An End-to-End LLM-Powered Multi-Agent Framework for Spreadsheet Automation
A multi-agent LLM framework with Manager, Action, and Reflection agents translates natural language into BNF-constrained spreadsheet commands, claiming about 80 percent success on simple tasks and 70 percent on multi ...
Reference graph
Works this paper leans on
-
[1]
2023 Skills Compass Report — The Burning Glass Institute
2024. 2023 Skills Compass Report — The Burning Glass Institute. Retrieved August 27, 2024 from https://static1.squarespace.com/static/ 6197797102be715f55c0e0a1/t/63ea41b5a9bd001d8061abe3/1676296630197/Skills+Compass+Report+2023_final.pdf
2024
-
[2]
Assistants Overview - OpenAI API
2024. Assistants Overview - OpenAI API. Retrieved August 27, 2024 from https://platform.openai.com/docs/assistants/overview
2024
-
[3]
ExcelScript Package
2024. ExcelScript Package. Retrieved August 27, 2024 from https://learn.microsoft.com/en-us/javascript/api/office-scripts/excelscript?view=office- scripts
2024
-
[4]
2025. Cursor. Retrieved July 28, 2025 from https://cursor.com/
2025
-
[5]
2025. Devin. Retrieved January 9, 2025 from https://devin.ai/
2025
-
[6]
GPT Pilot
2025. GPT Pilot. Retrieved July 28, 2025 from https://github.com/Pythagora-io/gpt-pilot
2025
-
[7]
VSCode Agent Mode
2025. VSCode Agent Mode. https://code.visualstudio.com/docs/copilot/chat/chat-agent-mode. Retrieved 30 May, 2025
2025
-
[8]
Robin Abraham, Margaret M. Burnett, and Martin Erwig. 2008. Spreadsheet programming. In Wiley Encyclopedia of Computer Science and Engineering. doi:10.1002/9780470050118.ecse415
Show all 107 references
-
[9]
Robin Abraham and Martin Erwig. 2004. Header and unit inference for spreadsheets through spatial analyses. In =IEEE Symposium on Visual Languages-Human Centric Computing (VL/HCC) . 165–172. doi:10.1109/VLHCC.2004.29
2004 doi
-
[10]
Robin Abraham and Martin Erwig. 2006. Inferring templates from spreadsheets. In International Conference on Software Engineering (ICSE) . 182–191. doi:10.1145/1134285.1134312
2006
-
[11]
Robin Abraham and Martin Erwig. 2007. GoalDebug: A spreadsheet debugger for end users. In International Conference on Software Engineering (ICSE). IEEE, 251–260. doi:10.1109/ICSE.2007.39
2007 doi
-
[12]
Robin Abraham, Martin Erwig, Steve Kollmansberger, and Ethan Seifert. 2005. Visual specifications of correct spreadsheets. In IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC) . 189–196. doi:10.1109/VLHCC.2005.70
2005 doi
-
[13]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023). https://arxiv.org/abs/2303.08774
2023 arXiv
-
[14]
Alan Agresti. 2012. Categorical data analysis. Vol. 792. John Wiley & Sons. 436–439 pages
2012
-
[15]
Saleema Amershi, Dan Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi Iqbal, Paul N Bennett, Kori Inkpen, et al. 2019. Guidelines for human-AI interaction. In CHI Conference on Human Factors in Computing Systems (CHI) . 1–13. doi:10.1145...
2019
-
[16]
Jacob Andreas. 2022. Language models as agent models. In Findings of Empirical Methods in Natural Language Processing (EMNLP) . Association for Computational Linguistics, 5769–5779. doi:10.18653/v1/2022.findings-emnlp.423
2022 doi
-
[17]
Maryam Arab, Thomas D LaToza, Jenny Liang, and Amy J Ko. 2022. An exploratory study of sharing strategic programming knowledge. In CHI Conference on Human Factors in Computing Systems (CHI) . 1–15. doi:10.1145/3491102.3502070
2022
-
[18]
Maryam Arab, Jenny Liang, Yang Yoo, Amy J Ko, and Thomas D LaToza. 2021. HowToo: A platform for sharing, finding, and using programming strategies. In IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC) . IEEE, 1–9. doi:10.1109/VL/HCC51201.2021.9576337
2021 arXiv
-
[19]
Yasharth Bajpai, Bhavya Chopra, Param Biyani, Cagri Aslan, Dustin Coleman, Sumit Gulwani, Chris Parnin, Arjun Radhakrishna, and Gustavo Soares. 2024. Let’s fix this together: Conversational debugging with GitHub Copilot. In2024 IEEE Symposium on Visual Languages and Human-Cent...
2024
-
[20]
Sebastian Baltes and Stephan Diehl. 2018. Towards a theory of software development expertise. In ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE) . 187–200. doi:10.1145/3236024.3236061
2018
-
[21]
Aaron Bangor, Philip Kortum, and James Miller. 2009. Determining what individual SUS scores mean: Adding an adjective rating scale. Journal of Usability Studies 4, 3 (2009), 114–123. https://uxpajournal.org/wp-content/uploads/sites/7/pdf/JUS_Bangor_May2009.pdf
2009
-
[22]
Gagan Bansal, Jennifer Wortman Vaughan, Saleema Amershi, Eric Horvitz, Adam Fourney, Hussein Mozannar, Victor Dibia, and Daniel S Weld
-
[23]
Shraddha Barke, Michael B James, and Nadia Polikarpova. 2023. Grounded copilot: How programmers interact with code-generating models. Proceedings of the ACM on Programming Languages 7, OOPSLA1 (2023), 85–111. doi:10.1145/3586030
2023 doi
-
[24]
Brett A Becker, Paul Denny, James Finnie-Ansley, Andrew Luxton-Reilly, James Prather, and Eddie Antonio Santos. 2023. Programming is hard-or at least it used to be: Educational opportunities and challenges of AI code generation. In ACM Technical Symposium on Computer Science E...
2023
-
[25]
Alan F Blackwell. 2024. Moral codes: Designing alternatives to AI . MIT Press
2024
-
[26]
quick and dirty
J Brooke. 1996. SUS: A “quick and dirty” Usability Scale. Usability Evaluation in INdustry/Taylor and Francis (1996). doi:chapters/edit/10.1201/ 9781498710411-35/sus-quick-dirty-usability-scale-john-brooke
1996
-
[27]
Margaret Burnett, Curtis Cook, Omkar Pendse, Gregg Rothermel, Jay Summet, and Chris Wallace. 2003. End-user software engineering with assertions in the spreadsheet paradigm. In International Conference on Software Engineering (ICSE) . 93–103. doi:10.1109/ICSE.2003.1201191
2003 arXiv
-
[28]
what you see is what you test
Margaret Burnett, Andrei Sheretov, Bing Ren, and Gregg Rothermel. 2002. Testing homogeneous spreadsheet grids with the "what you see is what you test" methodology. IEEE Transactions on Software Engineering (TOSEM) 28, 6 (2002), 576–594. doi:10.1109/TSE.2002.1010060
2002 arXiv
-
[29]
Stephen Casper, Luke Bailey, Rosco Hunter, Carson Ezell, Emma Cabalé, Michael Gerovitch, Stewart Slocum, Kevin Wei, Nikola Jurkovic, Ariba Khan, Phillip J. K. Christoffersen, A. Pinar Ozisik, Rakshit Trivedi, Dylan Hadfield-Menell, and Noam Kolt. 2025. The AI agent index. arXi...
2025 arXiv
-
[30]
It’s freedom to put things where my mind wants
George Chalhoub and Advait Sarkar. 2022. “It’s freedom to put things where my mind wants”: Understanding and improving the user experience of structuring data in spreadsheets. In CHI Conference on Human Factors in Computing Systems (CHI) . 1–24. doi:10.1145/3491102.3501833
2022
-
[31]
Valerie Chen, Alan Zhu, Sebastian Zhao, Hussein Mozannar, David Sontag, and Ameet Talwalkar. 2025. Need help? Designing proactive ai assistants for programming. In CHI Conference on Human Factors in Computing Systems (CHI) . 1–18. doi:10.1145/3706598.3714002
2025
-
[32]
Xinyun Chen, Petros Maniatis, Rishabh Singh, Charles Sutton, Hanjun Dai, Max Lin, and Denny Zhou. 2021. Spreadsheetcoder: Formula prediction from semi-structured context. In International Conference on Machine Learning (ICML) . PMLR, 1661–1672. https://proceedings.mlr.press/v1...
2021
-
[33]
Haoyu Dong, Jianbo Zhao, Yuzhang Tian, Junyu Xiong, Shiyu Xia, Mengyu Zhou, Yun Lin, José Cambronero, Yeye He, Shi Han, and Dongmei Zhang. 2025. SpreadsheetLLM: Encoding spreadsheets for large language models. arXiv preprint arXiv:2407.09025 (2025)
2025 arXiv
-
[34]
Lun Du, Fei Gao, Xu Chen, Ran Jia, Junshan Wang, Jiang Zhang, Shi Han, and Dongmei Zhang. 2021. TabularNet: A neural network architecture for understanding semantic structures of tabular data. In ACM SIGKDD Conference on Knowledge Discovery & Data Mining (KDD) . 322–331. doi:1...
2021
-
[35]
Martin Erwig, Robin Abraham, Steve Kollmansberger, and Irene Cooperstein. 2006. Gencel: A program generator for correct spreadsheets. Journal of Functional Programming 16, 3 (2006), 293–325. doi:10.1017/S0956796805005794
2006 doi
-
[36]
KJ Feng, Kevin Pu, Matt Latzke, Tal August, Pao Siangliulue, Jonathan Bragg, Daniel S Weld, Amy X Zhang, and Joseph Chee Chang. 2024. Cocoa: Co-planning and co-execution with ai agents. arXiv preprint arXiv:2412.10999 (2024)
2024
-
[37]
Kasra Ferdowsi, Jack Williams, Ian Drosos, Andrew D Gordon, Carina Negreanu, Nadia Polikarpova, Advait Sarkar, and Benjamin Zorn. 2023. ColDeco: An end user spreadsheet inspection tool for AI-generated code. In IEEE Symposium on Visual Languages and Human-Centric Computing (VL...
2023
-
[38]
Finch and Jinho D
Sarah E. Finch and Jinho D. Choi. 2020. Towards unified dialogue system evaluation: A comprehensive analysis of current evaluation protocols. In Annual Meeting of the Special Interest Group on Discourse and Dialogue (SIGDIAL) , Olivier Pietquin, Smaranda Muresan, Vivian Chen, ...
2020 doi
-
[39]
Diana Franklin, Paul Denny, David A Gonzalez-Maldonado, and Minh Tran. 2025. Generative AI in computer science education: Challenges and opportunities. Cambridge University Press
2025
-
[40]
Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023. Pal: Program-aided language models. In International Conference on Machine Learning (ICML) . PMLR, 10764–10799. https://proceedings.mlr.press/v202/gao23f
2023
-
[41]
Priyanshu Gupta, Shashank Kirtania, Ananya Singha, Sumit Gulwani, Arjun Radhakrishna, Sherry Shi, and Gustavo Soares. 2024. Metareflection: Learning instructions for language agents using past reflections. arXiv preprint arXiv:2405.13009 (2024). https://arxiv.org/abs/2405.13009
2024 arXiv
-
[42]
Sean N Halpin. 2024. Inter-coder agreement in qualitative coding: Considerations for its use. American Journal of Qualitative Research 8, 3 (2024), 23–43. doi:10.29333/ajqr/14887
2024 doi
-
[43]
Sandra G Hart and Lowell E Staveland. 1988. Development of NASA-TLX (Task Load Index): Results of empirical and theoretical research. In Advances in Psychology. Vol. 52. Elsevier, 139–183. doi:10.1016/S0166-4115(08)62386-9
1988 doi
-
[44]
Andrew Gary Darwin Holmes. 2020. Researcher positionality–A consideration of its influence and place in qualitative research–A new researcher guide. Shanlax International Journal of Education 8, 4 (2020), 1–10. https://eric.ed.gov/?id=EJ1268044 Manuscript submitted to ACM Tabl...
2020
-
[45]
Eric Horvitz. 1999. Principles of mixed-initiative user interfaces. In CHI Conference on Human Factors in Computing Systems (CHI) . 159–166. doi:10.1145/302979.303030
1999
-
[46]
Yanwei Huang, Yurun Yang, Xinhuan Shu, Ran Chen, Di Weng, and Yingcai Wu. 2024. Table Illustrator: Puzzle-based interactive authoring of plain tables. In CHI Conference on Human Factors in Computing Systems (CHI) . 1–18. doi:10.1145/3613904.3642415
2024
-
[47]
Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan
Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. 2024. SWE-bench: Can language models resolve real-world GitHub issues? arXiv preprint arXiv:2310.06770 (2024)
2024 arXiv
-
[48]
Harshit Joshi, Abishai Ebenezer, José Cambronero Sanchez, Sumit Gulwani, Aditya Kanade, Vu Le, Ivan Radiček, and Gust Verbruggen. 2024. Flame: A small language model for spreadsheet formulas. In AAAI Conference on Artificial Intelligence (AAAI) , Vol. 38. 12995–13003. doi:10.1...
2024 doi
-
[49]
Majeed Kazemitabaar, Xinying Hou, Austin Henley, Barbara Jane Ericson, David Weintrop, and Tovi Grossman. 2023. How novices use LLM-based code generators to solve CS1 coding tasks in a self-paced learning environment. In Koli Calling International Conference on Computing Educa...
2023
-
[50]
Amy J Ko, Robin Abraham, Laura Beckwith, Alan Blackwell, Margaret Burnett, Martin Erwig, Chris Scaffidi, Joseph Lawrance, Henry Lieberman, Brad Myers, et al. 2011. The state of the art in end-user software engineering. ACM Computing Surveys (CSUR) 43, 3 (2011), 1–44. doi:10.11...
2011
-
[51]
Amy J Ko, Thomas D LaToza, and Margaret M Burnett. 2015. A practical guide to controlled experiments of software engineering tools with human participants. Empirical Software Engineering (ESE) 20 (2015), 110–141. doi:10.1007/s10664-013-9279-3
2015 doi
-
[52]
Amy J Ko, Brad A Myers, and Htet Htet Aung. 2004. Six learning barriers in end-user programming systems. In IEEE Symposium on Visual Languages-Human Centric Computing (VL/HCC) . IEEE, 199–206. doi:10.1109/VLHCC.2004.47
2004 doi
-
[53]
Aayush Kumar, Yasharth Bajpai, Sumit Gulwani, Gustavo Soares, and Emerson Murphy-Hill. 2025. Sharp Tools: How developers wield agentic AI in real software engineering tasks. arXiv preprint arxiv:2506.12347 (2025)
2025
-
[54]
J Richard Landis and Gary G Koch. 1977. The measurement of observer agreement for categorical data.biometrics (1977), 159–174. doi:10.2307/2529310
1977 doi
-
[55]
Thomas D LaToza, Maryam Arab, Dastyni Loksa, and Amy J Ko. 2020. Explicit programming strategies. Empirical Software Engineering 25, 4 (2020), 2416–2449. doi:10.1007/s10664-020-09810-1
2020 doi
-
[56]
Hao-Ping Lee, Advait Sarkar, Lev Tankelevitch, Ian Drosos, Sean Rintel, Richard Banks, and Nicholas Wilson. 2025. The impact of generative AI on critical thinking: Self-reported reductions in cognitive effort and confidence effects from a survey of knowledge workers. In CHI Co...
2025
-
[57]
Juho Leinonen, Paul Denny, Stephen MacNeil, Sami Sarsa, Seth Bernstein, Joanne Kim, Andrew Tran, and Arto Hellas. 2023. Comparing code explanations created by students and large language models. In Conference on Innovation and Technology in Computer Science Education (ITiCSE) ...
2023
-
[58]
James R Lewis. 2018. The system usability scale: Past, present, and future. International Journal of Human–Computer Interaction 34, 7 (2018), 577–590. doi:10.1080/10447318.2018.1455307
2018
-
[59]
Hongxin Li, Jingran Su, Yuntao Chen, Qing Li, and Zhao-Xiang Zhang. 2024. SheetCopilot: Bringing software productivity to the next level through large language models. Advances in Neural Information Processing Systems (NeurIPs) 36 (2024). https://proceedings.neurips.cc/paper_f...
2024
-
[60]
Jenny T Liang, Maryam Arab, Minhyuk Ko, Amy J Ko, and Thomas D LaToza. 2023. A qualitative study on the implementation design decisions of developers. In IEEE/ACM International Conference on Software Engineering (ICSE) . IEEE, 435–447. doi:10.1109/ICSE48619.2023.00047
2023
-
[61]
TableTalk: Scaffolding spreadsheet development with a language agent
Jenny T. Liang, Aayush Kumar, Yasharth Bajpai, Sumit Gulwani, Vu Le, Chris Parnin, Arjun Rahakrishna, Ashish Tiwari, Emerson Murphy-Hill*, and Gustavo Soares*. 2025. Supplemental Materials to "TableTalk: Scaffolding spreadsheet development with a language agent". doi:10.6084/m...
2025 doi
-
[62]
Jenny T Liang, Chenyang Yang, and Brad A Myers. 2024. A large-scale survey on the usability of AI programming assistants: Successes and challenges. In IEEE/ACM International Conference on Software Engineering (ICSE) . 1–13. doi:10.1145/3597503.3608128
2024
-
[63]
Jenny T Liang, Thomas Zimmermann, and Denae Ford. 2022. Understanding skills for OSS communities on GitHub. In ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE) . 170–182. doi:10.1145/3540250.3549082
2022
-
[64]
Junwei Liu, Kaixin Wang, Yixuan Chen, Xin Peng, Zhenpeng Chen, Lingming Zhang, and Yiling Lou. 2024. Large language model-based agents for software engineering: A survey. arXiv preprint arXiv:2409.02977 (2024)
2024 arXiv
-
[65]
Michael Xieyang Liu, Jane Hsieh, Nathan Hahn, Angelina Zhou, Emily Deng, Shaun Burley, Cynthia Taylor, Aniket Kittur, and Brad A Myers. 2019. Unakite: Scaffolding developers’ decision-making using the web. In ACM Symposium on User Interface Software and Technology (UIST) . 67–...
2019
-
[66]
We need structured output
Michael Xieyang Liu, Frederick Liu, Alexander J Fiannaca, Terry Koo, Lucas Dixon, Michael Terry, and Carrie J Cai. 2024. "We need structured output": Towards user-centered constraints on large language model output. In CHI Conference on Human Factors in Computing Systems Exten...
2024
-
[67]
Michael Xieyang Liu, Tongshuang Wu, Tianying Chen, Franklin Mingzhe Li, Aniket Kittur, and Brad A Myers. 2024. Selenite: Scaffolding online sensemaking with comprehensive overviews elicited from large language models. In CHI Conference on Human Factors in Computing Systems (CH...
2024
-
[68]
Pan Lu, Liang Qiu, Wenhao Yu, Sean Welleck, and Kai-Wei Chang. 2023. A survey of deep learning for mathematical reasoning. In Annual Meeting of the Association for Computational Linguistics (ACL) , Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.). 14605–14631. doi:10...
2023 doi
-
[69]
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al. 2024. Self-refine: Iterative refinement with self-feedback. Advances in Neural Information Processing Systems (NeurIPS) 36 (2024...
2024
-
[70]
Hussein Mozannar, Gagan Bansal, Adam Fourney, and Eric Horvitz. 2024. Reading between the lines: Modeling user behavior and costs in AI-assisted programming. In CHI Conference on Human Factors in Computing Systems (CHI) . 1–16. doi:10.1145/3613904.3641936
2024
-
[71]
Brad A Myers. 1991. Graphical techniques in a spreadsheet for specifying user interfaces. In CHI Conference on Human Factors in Computing Systems (CHI). 243–249. doi:10.1145/108844.108903
1991
-
[72]
Brad A Myers, Amy J Ko, Thomas D LaToza, and YoungSeok Yoon. 2016. Programmers are users too: Human-centered methods for improving programming tools. Computer 49, 7 (2016), 44–52. doi:10.1109/MC.2016.200
2016 doi
-
[73]
Bonnie A Nardi and James R Miller. 1991. Twinkling lights and nested loops: Distributed problem solving and spreadsheet development.International Journal of Man-Machine Studies 34, 2 (1991), 161–184. doi:10.1016/0020-7373(91)90040-E
1991 doi
-
[74]
Rahul Pandita, Chris Parnin, Felienne Hermans, and Emerson Murphy-Hill. 2018. No half-measures: A study of manual and tool-assisted end-user programming tasks in Excel. In IEEE Symposium on Visual Languages and Human-centric Computing (VL/HCC) . IEEE, 95–103. doi:10.1109/VLHCC...
2018
-
[75]
Debjit Paul, Mete Ismayilzada, Maxime Peyrard, Beatriz Borges, Antoine Bosselut, Robert West, and Boi Faltings. 2024. REFINER: Reasoning feedback on intermediate representations. (March 2024), 1100–1126. https://aclanthology.org/2024.eacl-long.67/
2024
-
[76]
Justin Payan, Swaroop Mishra, Mukul Singh, Carina Negreanu, Christian Poelitz, Chitta Baral, Subhro Roy, Rasika Chakravarthy, Benjamin Van Durme, and Elnaz Nouri. 2023. InstructExcel: A benchmark for natural language instruction in Excel. In Conference on Empirical Methods in ...
2023 doi
-
[77]
Peter Pirolli and Stuart Card. 2005. The sensemaking process and leverage points for analyst technology as identified through cog- nitive task analysis. In International Conference on Intelligence Analysis , Vol. 5. 2–4. https://www.researchgate.net/profile/Peter- Pirolli/publ...
2005
-
[78]
It’s weird that it knows what i want
James Prather, Brent N Reeves, Paul Denny, Brett A Becker, Juho Leinonen, Andrew Luxton-Reilly, Garrett Powell, James Finnie-Ansley, and Eddie Antonio Santos. 2023. “It’s weird that it knows what i want”: Usability and interactions with copilot for novice programmers. ACM Tran...
2023 doi
-
[79]
Kevin Pu, Daniel Lazaro, Ian Arawjo, Haijun Xia, Ziang Xiao, Tovi Grossman, and Yan Chen. 2025. Assistance or disruption? Exploring and evaluating the design and trade-offs of proactive AI programming support. In CHI Conference on Human Factors in Computing Systems (CHI) . 1–2...
2025
-
[80]
Brian J Reiser. 2018. Scaffolding complex learning: The mechanisms of structuring and problematizing student work. In Scaffolding. Psychology Press, 273–304. https://www.taylorfrancis.com/chapters/edit/10.4324/9780203764411-2/scaffolding-complex-learning-mechanisms-structuring...
2018 doi
-
[81]
Thomas Reschenhofer and Florian Matthes. 2015. An empirical study on spreadsheet shortcomings from an information systems perspective. In Business Information Systems (BIS) . Springer, 50–61. doi:10.1007/978-3-319-19027-3_5
2015 doi
-
[82]
Gordon, Neil D
Diana Robinson, Christian Cabrera, Andrew D. Gordon, Neil D. Lawrence, and Lars Mennen. 2025. Requirements are all you need: The final frontier for end-user software engineering. ACM Transactions Software Engineering Methodology (TOSEM) , Article 141 (2025). doi:10.1145/3708524
2025 doi
-
[83]
Maggi Savin-Baden and Claire Major. 2023. Qualitative research: The essential guide to theory and practice . Routledge
2023
-
[84]
John W Saye and Thomas Brush. 2002. Scaffolding critical reasoning about history and social issues in multimedia-supported learning environments. Educational Technology Research and Development 50, 3 (2002), 77–96. doi:10.1007/BF02505026
2002 doi
-
[85]
Christopher Scaffidi. 2017. Workers who use spreadsheets and who program earn more than similar workers who do neither. In IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC) . 233–237. doi:10.1109/VLHCC.2017.8103472
2017
-
[86]
Christopher Scaffidi, Mary Shaw, and Brad Myers. 2005. Estimating the numbers of end users and end user programmers. In IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC) . 207–214. doi:10.1109/VLHCC.2005.34
2005 doi
-
[87]
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2024. Toolformer: Language models can teach themselves to use tools. Advances in Neural Information Processing Systems (NeurIPs) 36 ...
2024
-
[88]
Donald Sharpe. 2015. Your chi-square test is statistically significant: now what?. Practical Assessment, Research & Evaluation 20, 8 (2015), n8. https://eric.ed.gov/?id=EJ1059772
2015
-
[89]
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2024. Reflexion: Language agents with verbal rein- forcement learning. Advances in Neural Information Processing Systems (NeurIPS) 36 (2024). https://papers.nips.cc/paper_files/paper/2023/file/ ...
2024
-
[90]
Sruti Srinivasa Ragavan, Zhitao Hou, Yun Wang, Andrew D Gordon, Haidong Zhang, and Dongmei Zhang. 2022. Gridbook: Natural language formulas for the spreadsheet grid. In ACM International Conference on Intelligent User Interfaces (IUI) . 345–368. doi:10.1145/3490099.3511161
2022
-
[91]
Sangho Suh, Bryan Min, Srishti Palani, and Haijun Xia. 2023. Sensecape: Enabling multilevel exploration and sensemaking with large language models. In ACM Symposium on User Interface Software and Technology (UIST) . 1–18. doi:10.1145/3586183.3606756
2023
-
[92]
Lu Sun, Aaron Chan, Yun Seo Chang, and Steven P Dow. 2024. ReviewFlow: Intelligent scaffolding to support academic peer reviewing. In ACM International Conference on Intelligent User Interfaces (IUI) . 120–137. doi:10.1145/3640543.3645159
2024
-
[93]
Lev Tankelevitch, Viktor Kewenig, Auste Simkute, Ava Elizabeth Scott, Advait Sarkar, Abigail Sellen, and Sean Rintel. 2024. The metacognitive demands and opportunities of generative AI. In CHI Conference on Human Factors in Computing Systems (CHI) . 1–24. doi:10.1145/3613904.3642902
2024
-
[94]
Michele Tufano, Anisha Agarwal, Jinu Jang, Roshanak Zilouchian Moghaddam, and Neel Sundaresan. 2024. AutoDev: Automated AI-driven development. arXiv preprint arXiv:2403.08299 (2024). doi:abs/2403.08299
2024 arXiv
-
[95]
Priyan Vaithilingam, Tianyi Zhang, and Elena L Glassman. 2022. Expectation vs. experience: Evaluating the usability of code generation tools powered by large language models. InCHI Conference on Human Factors in Computing Systems Extended Abstracts . 1–7. doi:10.1145/3491101.3519665
2022
-
[96]
Sanidhya Vijayvargiya, Xuhui Zhou, Akhila Yerukola, Maarten Sap, and Graham Neubig. 2025. Interactive agents to overcome ambiguity in software engineering. arXiv preprint arXiv:2502.13069 (2025)
2025
-
[97]
Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, et al. 2024. A survey on large language model based autonomous agents. Frontiers of Computer Science 18, 6 (2024), 186345. doi:10.1007/s11704-024-40231-1
2024 doi
-
[98]
Ruotong Wang, Ruijia Cheng, Denae Ford, and Thomas Zimmermann. 2024. Investigating and designing for trust in ai-powered code generation tools. In ACM Conference on Fairness, Accountability, and Transparency (FAccT) . 1475–1493. doi:10.1145/3630106.3658984
2024
-
[99]
Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H
Xingyao Wang, Boxuan Li, Yufan Song, Frank F. Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H. Tran, Fuqiang Li, Ren Ma, Mingzhang Zheng, Bill Qian, Yanjun Shao, Niklas Muennighoff, Yizhe Zhang, Binyuan Hui, Junyang Lin, Robert Brenna...
2025 arXiv
-
[100]
Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, et al. 2023. The rise and potential of large language model based agents: A survey. arXiv preprint arXiv:2309.07864 (2023)
2023 arXiv
-
[101]
Wanli Xing, Nia Nixon, Scott Crossley, Paul Denny, Andrew Lan, John Stamper, and Zhou Yu. 2025. The use of large language models in education. International Journal of Artificial Intelligence in Education 35, 2 (2025), 439–443. doi:10.1007/s40593-025-00457-x
2025 doi
-
[102]
John Yang, Carlos E Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik R Narasimhan, and Ofir Press. 2024. SWE-agent: Agent-computer interfaces enable automated software engineering. In Advances in Neural Information Processing Systems (NeurIPs) . https://openreview...
2024
-
[103]
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2024. Tree of thoughts: Deliberate problem solving with large language models. Advances in Neural Information Processing Systems (NeurIPS) 36 (2024). doi:paper_files/paper/2023/f...
2024
-
[104]
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. 2023. ReAct: Synergizing reasoning and acting in language models. In International Conference on Learning Representations (ICLR) . https://openreview.net/forum?id=WE_vluYUL-X
2023
-
[105]
Linghao Zhang, Shilin He, Chaoyun Zhang, Yu Kang, Bowen Li, Chengxing Xie, Junhao Wang, Maoquan Wang, Yufan Huang, Shengyu Fu, Elsie Nallipogu, Qingwei Lin, Yingnong Dang, Saravan Rajmohan, and Dongmei Zhang. 2025. SWE-bench goes live! arXiv preprint arXiv:2505.23419 (2025)
2025 arXiv
-
[106]
Excellent
Yuntong Zhang, Haifeng Ruan, Zhiyu Fan, and Abhik Roychoudhury. 2024. AutoCodeRover: Autonomous program improvement. InACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA) . 1592–1604. doi:10.1145/3650212.3680384 Appendix Overview To provide context on ...
2024
-
[2024]
arXiv preprint arXiv:2412.10380 (2024)
Challenges in human-agent communication. arXiv preprint arXiv:2412.10380 (2024)
2024 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.