{"id":"3f686c72-bbdb-40b3-baf7-78d1dbb2148b","arxiv_id":"2502.09787","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A scaffolding spreadsheet agent produced partial spreadsheets that evaluators preferred 2.3 times more often than a standard AI assistant's in a 20-user controlled study.","lead":"Researchers built TableTalk, an AI agent that guides people step by step through creating Excel spreadsheets and suggests what to do next. In a study with 20 spreadsheet users, spreadsheets made with TableTalk were preferred 2.3 times as often as those made with a standard AI assistant, with lower reported mental demand.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline preference claim is confounded: TableTalk's GPT-4o backend and asymmetric warm-up instructions, not just the three design principles, could drive the 70% preference gap.","rationale":"The reader's weakest assumption correctly identifies the model and tutorial asymmetry between TableTalk and the baseline as the key threat to internal validity. My stress-test concurs: the paper's headline claim is a comparison of two systems that differ not only in the three design principles but also in the underlying LLM (GPT-4o vs. an unspecified Copilot model) and in the warm-up instructions given to participants. The observed 70% preference could plausibly reflect the stronger model or the differential tutorial guidance rather than scaffolding, flexibility, and incrementality. The paper itself acknowledges latency and other limitations, but it does not address this confound. A model-matched ablation is the direct test that would settle whether the design principles are causally responsible. I therefore keep the reader's CONDITIONAL verdict unchanged: the tool comparison is promising, but the design-principle attribution requires additional evidence. No ad hominem or theatrical framing is needed; this is a standard internal-validity concern that the authors should address in revision.","tokens_in":40626,"tokens_out":1764,"duration_ms":18668,"concrete_test":"Run a model-matched ablation with identical warm-up scripts for both arms: (a) TableTalk as-is, (b) the same GPT-4o backend with the three design features removed (no expert-process plan in the system prompt, no suggestion pills, no incremental tool-composition workflow; instead a single-pass chat that performs requested actions directly), and (c) the original Excel Copilot baseline. If preference between (a) and (b) drops to chance, the design principles are not the driver; if (a) still beats (b) by a similar margin, the design claim survives. The warm-up protocol must be identical across arms, removing the 'include a table with headers' instruction from the baseline condition.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that TableTalk's design principles (scaffolding, flexibility, incrementality) cause higher spreadsheet quality and lower cognitive load. The evaluation does not isolate these principles. Section 6.3 states TableTalk uses GPT-4o via the OpenAI Assistants API; the baseline (Section 7.3) is 'Excel Copilot available in Excel version 2409,' whose underlying model is unspecified and likely weaker. Section 7.3 also gives the baseline an asymmetric warm-up instruction: participants were told to 'include a table with headers in the spreadsheet for the best results,' which can disadvantage the baseline by constraining how users interact with it. The preference result (42/60, BT ability 0.85, p<0.001) is a tool-level comparison, but the paper's framing and Section 8 attribute the improvement to the design principles. Because model capability and tutorial framing are confounded with the design manipulation, the causal attribution is underdetermined. The reader's conditional verdict is appropriate; this is the load-bearing threat to the paper's broader design-guidelines contribution.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents TableTalk, a language-agent system for spreadsheet programming that embodies three design principles: scaffolding, flexibility, and incrementality. These principles are derived from an analysis of 85 Excel templates and a formative study with 7 spreadsheet programmers. TableTalk guides users through an expert-derived plan, proposes three next-step suggestions, and builds spreadsheets incrementally using atomic OfficeScript-based tools. The evaluation is a controlled study with 20 spreadsheet programmers who each used TableTalk and a baseline (Excel Copilot version 2409) on two spreadsheet tasks. The reported headline results are that spreadsheet evaluators preferred TableTalk-created spreadsheets 42 out of 60 times (70%), with a Bradley-Terry ability score of 0.85 (2.3 times higher odds of preference, p<0.001), and that TableTalk reduced self-reported mental demand and time spent thinking about spreadsheet actions. The paper concludes with design guidelines for agentic spreadsheet tools and a discussion of implications for spreadsheet programming, end-user programming, and human-agent collaboration.","tokens_in":40751,"tokens_out":4856,"duration_ms":51889,"significance":"If the causal claims held, this would be a valuable contribution to HCI and software engineering: it provides a concrete, well-described system that operationalizes three design principles, and it offers an unusually rich mixed-methods evaluation with open materials, inter-rater reliability checks, and candid limitation statements. The qualitative analyses of chat logs, activity timelines, evaluator comments, and interviews are a genuine strength, as are the publicly available protocols and codebooks. However, the central empirical claim that the design principles themselves drive the observed quality and cognitive-load benefits is not supported by the current comparison, because the two tools differ not only in the principles but also in underlying model and in tutorial instructions. The paper is therefore best read as a tool-level comparison of TableTalk against Excel Copilot; the broader design-guideline contribution requires either a repaired experiment or a substantially more cautious framing.","major_comments":[{"comment":"The evaluation does not isolate the three design principles from the underlying model and the warm-up instructions. Section 6.3 states that TableTalk uses GPT-4o through the OpenAI Assistants API, while Section 7.3 identifies the baseline only as 'Excel Copilot available in Excel version 2409' without naming its model. The same section gives baseline users the extra instruction to 'include a table with headers in the spreadsheet for the best results,' a tutorial that TableTalk users did not receive. Consequently, the 70% preference result (Section 7.5.1) and the mental-demand reduction (Section 7.5.3) could be driven by model capability or by the tutorial asymmetry rather than by scaffolding, flexibility, and incrementality. Since Section 8 attributes the results to the design principles, the paper overreaches. Please either re-frame the central claims as a tool-level comparison, add an ablation or matched-model control, or provide a rigorous argument that the only systematic differences between the two tools are the implemented design principles.","section":"Section 6.3 and Section 7.3"},{"comment":"The quality comparison is of partial artifacts, not completed spreadsheets. Section 7.5 states that no participant completed any task, and Section 7.2 explains the tasks were intentionally difficult and time-limited. The paper anticipates this by focusing on spreadsheet quality rather than completion, but the interaction between the 15-minute limit and the tools' very different latencies is not addressed. Section 7.5.2 reports that TableTalk users collectively spent 61.0 additional minutes waiting, meaning the two conditions differ in the time available for manual work; the 'more polished' TableTalk artifacts could reflect automation advantages or the structure of the tool rather than the specific design principles. At minimum, the manuscript should discuss whether evaluator preference for partial artifacts is a valid proxy for the quality of a completed spreadsheet, and how the latency difference might affect the comparison.","section":"Section 7.2 and Section 7.5"},{"comment":"The Bradley-Terry analysis is reported as a point estimate and a p-value, but the paper does not provide a confidence interval, cluster-robust standard errors, or an account of the repeated-measures structure. The 60 pairwise comparisons are not independent: the spreadsheets come from 20 participants (each contributing one artifact per tool) and are judged by 6 evaluators. Please report the uncertainty in the ability score and test the sensitivity of the conclusion to a model that treats participants or evaluators as random effects, or to a bootstrap procedure that resamples by participant or evaluator.","section":"Section 7.4"},{"comment":"The limitations section discusses confirmation bias, the shared-computer environment, and external validity, but it does not acknowledge the most direct internal-validity threat identified above: the mismatch between the models used by TableTalk and the baseline, and the asymmetric warm-up instruction. Since Section 8.3 does enumerate internal validity threats, the omission of the model confound is a gap in the manuscript's own self-assessment. Please add an explicit limitation and state precisely which conclusions can and cannot be drawn from the current comparison.","section":"Section 8.3"}],"minor_comments":[{"comment":"The heading '2.3 AI for spreasheet programming' contains a typo: 'spreasheet' should be 'spreadsheet'.","section":"Section 2.3"},{"comment":"In the paragraph about code comprehension, the participant is identified as P9 in the first sentence but the quotation 'if score is below 100, do not display excellent' is attributed to P10; this appears to be an inconsistency and should be corrected to P9.","section":"Section 8.1.5"},{"comment":"The reference [61] lists an author as 'Arjun Rahakrishna' instead of 'Arjun Radhakrishna'.","section":"Acknowledgments"},{"comment":"The activity label 'T ableT alk' in the figure legend appears to be a formatting artifact; it should read 'TableTalk'.","section":"Figure 7"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for a CHI-style venue and the supplementary materials are a definite strength. My main concern is the confound between model/tutorial differences and the design principles; the reader's conditional verdict identifies this correctly. The paper can be salvaged by re-framing the claims as tool-level rather than principle-level, or by adding a matched-model control, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear Colleague,\n\nThe punchline: TableTalk is a competent, well-reported HCI systems paper, and the headline result—spreadsheet evaluators preferred TableTalk's output 42 out of 60 times—is plausible at the tool level. But the broader claim, that the three design principles (scaffolding, flexibility, incrementality) cause that improvement, is not supported by the experiment as run. The stress-test note is right: the baseline is an unnamed Excel Copilot 2409, TableTalk uses GPT-4o, and the warm-up gave baseline users an asymmetric instruction about headers. That confound is load-bearing.\n\nWhat's genuinely good: the formative work is careful. The template study (85 spreadsheets, κ=0.87) and the 7-participant user study produce a believable design space, and the three principles are grounded in that data. The implementation is described in enough detail to be reproducible in spirit, and the supplemental materials—protocols, template dataset, codebooks, video demo—are a real asset. The evaluation itself is above the norm for an end-to-end system study: 6 evaluators, 60 pairwise comparisons, Bradley-Terry, mixed-effects models, chi-square with residual analysis, and qualitative coding with reported IRR. The paper also openly acknowledges limitations: no one finished a task, latency was higher, and the researchers adjusted their interview questions to separate the latency effect from preference.\n\nThe soft spots, in proportion: the biggest one is the attribution gap. Sections 7.5.1 and 8 read as if the design principles caused the preference, but the data only support a tool-level comparison. A model-matched baseline or an ablation would fix this, and I'd want that before trusting the design-guidelines contribution. Two smaller issues: the abstract's cognitive-load claim (12.6% reduction) drops the waiting-time tradeoff—the NASA-TLX mental demand item was significant, but other TLX items weren't after correction—and the Bradley-Terry ability score has no confidence interval. None of these are fatal to the system itself, but they should be tightened.\n\nWho is this for: anyone working on spreadsheet tools, end-user programming, or human-agent collaboration. It deserves a serious referee. I'd send it to peer review with the expectation that the authors will need to soften or re-test the causal framing.\n\nRecommendation: engage with it—read it as a system paper, not as a proof of the design principles.","headline":"A solid system paper with a real confound between the model and the design principles; the tool-level preference claim is plausible, the causal claim is not yet supported.","tokens_in":41370,"tokens_out":2664,"would_cite":true,"duration_ms":25412,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A step-by-step AI guide builds better spreadsheets than a free-form assistant, with lower cognitive load.","keywords":["spreadsheet programming","language agents","scaffolding","human-agent collaboration","LLM-assisted programming","Excel","incremental development","cognitive load"],"falsifier":"Run the same 20-participant study with the baseline replaced by a version that uses the same underlying model and identical warm-up instructions; if the 70% preference for TableTalk disappears or reverses, the design principles themselves are not what produced the benefit.","tokens_in":40385,"feed_emoji":"📊","tokens_out":4448,"duration_ms":39821,"temperature":0.7,"pith_summary":"The paper claims that a spreadsheet-programming agent should not hand the whole task to the user at once, nor perform it autonomously; it should guide the user through an expert process step by step. TableTalk embodies three design principles, scaffolding, flexibility, and incrementality, derived from a study of 85 Excel templates and 7 spreadsheet programmers. In a controlled study with 20 programmers, outside evaluators preferred spreadsheets built with TableTalk over those from a baseline language agent 42 out of 60 times (70%), and TableTalk's odds of being preferred were 2.3 times higher under a Bradley-Terry model. The paper also reports a 12.6% reduction in thinking time and lower self-reported mental demand. If the claim holds, the way to build useful spreadsheet agents is to keep the human in the loop, letting the agent handle atomic actions while the user steers the plan.","feed_headline":"Step-by-step AI agent beats free-form assistant on spreadsheets","feed_subtitle":"TableTalk won 70% of spreadsheet-quality comparisons and cut thinking time by 12.6%.","key_machinery":"The central object is the TableTalk language agent, which implements a plan based on Pirolli and Card's expert sensemaking process for creating knowledge products from data: gather requirements, define a data-table schema, and create insight tables. Its mechanism has three parts: a system prompt that walks the agent through this plan; a chat interface that presents three model-generated suggestion pills for the next step, so the human can steer the plan; and a set of pre-written OfficeScript tools (create_table, sort_rows, filter_rows, add_chart, highlight_cell) that let the agent build atomic spreadsheet components incrementally rather than generating a whole spreadsheet in one shot.","core_discovery":"The central claim is that a language agent that scaffolds spreadsheet development through a structured expert plan, gather requirements, define the data schema, and extract insights, while offering three adaptive next-step suggestions and building spreadsheets incrementally with atomic tools, produces higher-quality spreadsheets and a better programmer experience than a non-scaffolded baseline agent. The evidence is a 20-participant controlled study in which blinded evaluators preferred TableTalk's spreadsheets 42 of 60 times, a Bradley-Terry ability score of 0.85 (2.3x odds, p<0.001), and a 1.9-minute reduction in thinking or verifying time, which is 12.6% of the total task time. The paper further reports lower mental demand and better conversational relevance scores for TableTalk.","pith_inferences":["If the mechanism is human-in-the-loop planning, then removing the suggestion pills or making them generic should reduce the quality gap; this is a testable ablation.","The positive results may hinge on TableTalk using GPT-4o while the baseline used a different model; a matched-model replication would clarify whether the design principles or the underlying model cause the effect.","The 12.6% cognitive-load reduction likely combines two separate effects: the scaffold reduces planning burden, but the agent's latency adds waiting time, so net load may vary in settings with faster or slower models.","The paper's design guidelines imply that future agents should expose adjustable proactivity and direct manipulation controls like undo and stop, to satisfy participants who felt the proactive agent was invasive."],"forward_implications":["Spreadsheet agents that scaffold the process will be preferred over non-scaffolded ones for open-ended analysis tasks, even when the task is too hard to finish in the allotted time.","Programmers using a scaffolded agent spend more effort on requirements and high-level commands and less on low-level implementation details and verification.","Lower mental demand and thinking time could make spreadsheet tools usable by more people, especially those who struggle with formulas and schema design.","The three design principles, scaffolding, flexibility, and incrementality, can be applied to other end-user programming domains such as debugging, data cleaning, and report generation.","The tradeoff between proactivity and user control must be made adjustable; participants were split on whether the agent should act autonomously or wait for explicit confirmation."],"supporting_citations":[{"why":"Supplies the expert sensemaking process (gather information, develop schema, extract insights) that TableTalk follows as its plan.","marker":"[77]"},{"why":"SheetCopilot is the closest agent-based spreadsheet tool, and TableTalk is positioned against its non-interactive approach.","marker":"[59]"},{"why":"ROBIN is the prior LLM scaffolding tool whose follow-up suggestion mechanism TableTalk adapts.","marker":"[19]"},{"why":"Table Illustrator provides prior table-structure analysis and a cognitive-load comparison point for spreadsheet tools.","marker":"[46]"},{"why":"GridBook is a prior conversational spreadsheet tool, used as a comparison for conversational interaction and formula generation.","marker":"[90]"},{"why":"Chalhoub and Sarkar's study is the source of the two evaluation tasks: student grades analysis and hours-worked/client-payments tracking.","marker":"[30]"},{"why":"The NASA-TLX instrument is used to measure cognitive load in the evaluation.","marker":"[43]"},{"why":"Provides the conversational-quality measures (relevance and proactivity) used in the evaluation questionnaire.","marker":"[38]"}],"fun_headline_variants":["Scaffolded AI agent beats free-form in spreadsheet quality","Guided AI agent cuts spreadsheet thinking time by 12.6%","TableTalk's step-by-step plans win 70% of quality tests","AI agent with structured steps yields better spreadsheets"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison assumes the only meaningful difference between TableTalk and the baseline is the three design principles, not the underlying model or the way each tool was explained to participants.","fun_headline_variants_meta":{"raw":{"variants":["Scaffolded AI agent beats free-form in spreadsheet quality","Guided AI agent cuts spreadsheet thinking time by 12.6%","TableTalk's step-by-step plans win 70% of quality tests","AI agent with structured steps yields better spreadsheets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000672,"raw_usage":{"total_tokens":3026,"prompt_tokens":877,"completion_tokens":2149,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":493,"completion_tokens_details":{"reasoning_tokens":2078}},"tokens_in":493,"tokens_out":2149,"duration_ms":14904,"temperature":1.0,"reasoning_tokens":2078,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T20:28:02.843399+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 20-participant study with the baseline replaced by a version that uses the same underlying model and identical warm-up instructions; if the 70% preference for TableTalk disappears or reverses, the design principles themselves are not what produced the benefit.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the expert sensemaking process (gather information, develop schema, extract insights) that TableTalk follows as its plan."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"ROBIN is the prior LLM scaffolding tool whose follow-up suggestion mechanism TableTalk adapts."}],"review_version":1}