{"id":"de563aed-1e1c-4c95-89dc-2737952bd6c7","arxiv_id":"2502.05612","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"In a ten-day field study, twelve writers used an LLM-assisted dictation tool and reported that speaking their drafts helped productivity and emotional expression, with nine of twelve adapting to the new paradigm.","lead":"This paper reports a ten-day diary study in which twelve writers used Rambler, an app that lets people dictate text and then reorganize it with AI help. It found users adopted two main writing approaches and generally felt the tool made writing faster and more emotionally expressive, pointing to design directions for AI-assisted speech writing.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Productivity-gain conclusion rests on self-estimated time bins with no objective baseline; the collected interaction logs can test this directly.","rationale":"I read the paper as a modest qualitative case study whose central claim is that speech-based writing with LLM support is viable because users adapted quickly and reported productivity and psychological benefits. For that claim to hold, the reported productivity gains need to be more than an artifact of novelty, compensation, or self-presentation. The weakest point in the argument is the quantification of productivity in §4.4.1: the paper states that completion time was not measured and relies on participants' own estimates, with no baseline or comparison condition. This is exactly the kind of claim that can be checked against data the authors already collected, since §3.3 says all interactions were logged. The other limitations noted by the reader, such as small sample size and all-non-native participants, are real but less decisive for a qualitative case study, which does not claim statistical generalization. However, since the conclusion explicitly rests on 'productivity gain,' an unsupported productivity measure directly weakens the central claim. The proposed log re-analysis is a concrete, feasible check: if it fails, the paper should be revised to frame productivity as perceived rather than measured; if it passes, the conditional acceptance is strengthened. I therefore recommend keeping the reader's CONDITIONAL verdict unchanged.","tokens_in":11593,"tokens_out":5943,"duration_ms":65704,"concrete_test":"Re-analyze the existing Rambler server logs (§3.3) for the 36 articles: compute each article's active writing duration from timestamps of logged actions, excluding long idle gaps, and compare the resulting distribution with the self-reported time bins in §4.4.1. If logged durations are systematically longer than the reported bins, or if most tasks exceed 30 minutes, the productivity-gain claim is not supported by objective data and the conclusion should be softened. If logged durations corroborate the self-reports, the concern is mitigated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central conclusion (§6) is a positive outlook for speech-plus-LLM writing based on 'its productivity gain and psychological benefits.' The productivity component is stated strongly in §4.3.2 ('Speaking Boosts Productivity') and quantified in §4.4.1, but the only quantitative support is participants' retrospective placement of each 500-word task into time bins (e.g., 66.7% in 10–30 minutes). Section 4.4.1 explicitly acknowledges that task completion time was not measured effectively. There is no typing baseline, no within-subject comparison, and no reported use of the interaction logs that were collected (§3.3) to verify durations. In an incentivized diary study with a novel tool, self-estimated completion time is vulnerable to demand characteristics and novelty effects. Because the 'viable paradigm' claim is justified by productivity gain, this is load-bearing: if Rambler-assisted writing is not actually faster, or is no faster than the participants' usual keyboard process, the stated basis for the positive outlook is unsupported, even if qualitative acceptance and psychological benefits remain. The concern is not that productivity gains are impossible; it is that the current evidence cannot distinguish perceived gains from actual ones.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a ten-day diary study in which twelve academic and creative writers used Rambler, an LLM-assisted speech-to-text writing tool, for their own writing tasks. The authors identify two main composition strategies (outline-driven expansion and free-speaking followed by restructuring), describe affordances related to emotional expression, perceived productivity, and self-efficacy, and discuss user acceptance and remaining challenges. The paper concludes that speech-plus-LLM writing is a viable paradigm, with the positive outlook justified by reported productivity and psychological benefits.","tokens_in":11786,"tokens_out":4730,"duration_ms":45921,"significance":"If the findings hold, this is a useful real-world complement to the earlier lab study of Rambler, showing that at least some motivated writers can adopt speech as a primary writing modality and that LLM-based gist manipulation supports both structured and reflective writing. The study's strengths include the collection of thirty-six articles, interaction logs, and rich qualitative material; the authors are also transparent about the inability to measure task completion time directly. The strategy taxonomy and design suggestions are concrete and actionable. However, the quantitative productivity claim is not supported by the current evidence, and the sample is narrow, so the paper's central positive outlook is stronger than what the data can establish.","major_comments":[{"comment":"The productivity-gain component of the central claim rests on self-estimated completion-time bins (66.7% of tasks in 10–30 minutes) with no typing baseline, no within-subject comparison, and no analysis of the interaction logs described in §3.3 to verify durations. The text in §4.4.1 explicitly concedes that task completion time was not measured effectively, and the heading in §4.3.2 ('Speaking Boosts Productivity') asserts a causal productivity benefit. Because the conclusion in §6 is explicitly justified by 'productivity gain,' the current evidence supports only perceived productivity. Please either reframe the productivity claims as perceived or self-reported, or add a log-based duration analysis and a baseline or comparison condition.","section":"§4.3.2, §4.4.1, §6"},{"comment":"All twelve diary-study participants were non-native English speakers, were self-selected volunteers, and received supermarket coupons, and the entire observation period was seven to ten days. The statements that 'users could get used to writing with speech in a short period' (§5.4) and the generally positive outlook in §6 generalize beyond this sample. Please restrict the claims to the studied population of motivated, English-proficient non-native writers, and discuss how selection, compensation, and novelty effects may have shaped acceptance and reported benefits.","section":"§3.1, §5.4, §6"},{"comment":"The coding procedure reports that two researchers coded 25% of the data independently and reached consensus, after which one researcher coded the remainder, but no inter-rater reliability statistic or codebook is reported. Since the thematic findings are the core evidence for the psychological and strategy claims, please report agreement measures or provide a more complete audit trail to support the reliability of the themes.","section":"§3.3"}],"minor_comments":[{"comment":"Please clarify the relationship between the 14 focus-group respondents and the 12 diary-study participants, including why P1, P7, and P11 were absent from the diary phase and whether any recruitment or attrition issues affected the final sample.","section":"§3.1, Appendix Table 1"},{"comment":"Please include the exact survey question and response options used for the time estimates, since the current report gives only bins and percentages.","section":"§4.4.1"},{"comment":"The statement that the deviation of comfort scores decreases after the first use is based on visual inspection; please report the underlying distributions or a compact summary statistic.","section":"Figure 4"},{"comment":"The description of Custom Magic Prompt does not state whether it applies to a single Ramble or to all Rambles at once, which matters because §4.5.1 reports a request to apply prompts to all Rambles simultaneously.","section":"§2"},{"comment":"The phrase 'a English-taught degree program' should read 'an English-taught degree program.'","section":"§3.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is an extended abstract for CHI EA, and the exploratory diary-study design fits the venue. My main concern is the gap between the strong productivity language in §4.3.2 and §6 and the self-report evidence in §4.4.1; the requested revision is feasible within the paper's scope. The heavy reliance on the authors' own prior paper [10] for the tool description is acceptable because the present study is explicitly a follow-up, but the new paper should still make its independent evidence stand alone."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a legitimate follow-up to the authors' lab study. The two writing strategies (outline expansion vs. free-speaking) show up again in real-world use, and the new themes—emotional expression, reduced overthinking, self-efficacy—are genuinely informative for HCI design. The paper is transparent about its limits, which counts for something.\n\nWhat's new: the in-the-wild confirmation, the device/environment data, and the affordance themes. The analysis is grounded in quotes and interaction logs, and the coding process, though lightly reported, is standard for a CHI EA paper. The authors are honest that they did not measure task time directly.\n\nWhere it gets soft: the productivity gain theme rests on self-estimated time bins with no baseline. The stress-test worry is real: §4.3.2 and §4.4.1 frame 'Speaking Boosts Productivity' as a finding, but the only quantitative support is participants' retrospective estimates. That is a perceived productivity theme, not measured productivity. It would be easy for the authors to either soften the language or analyze the interaction logs they collected to at least triangulate duration. As written, the conclusion leans on that unsupported gain. I don't think it's load-bearing—the acceptance and psychological benefits stand on their own—but it should be fixed before the positive outlook is taken at face value.\n\nThe all-non-native, self-selected sample is another limitation, but for a qualitative case study it's not disqualifying. The strategies and themes are directional, and the authors say as much.\n\nBottom line: this is a small, honest contribution. It deserves a serious referee, and with modest revisions (tone down the productivity claim, report coding reliability, consider releasing anonymized logs) it would be a solid extended abstract. I'd cite it for the in-the-wild strategy confirmation, and I'd probably bring it to a reading group if we're discussing speech-based writing.","headline":"A modest but honest diary study that confirms lab findings in the wild; the productivity claim is softer than the framing suggests, but the qualitative core holds up.","tokens_in":12299,"tokens_out":2215,"would_cite":true,"duration_ms":21903,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A ten-day diary study finds speech plus LLM support is a viable writing paradigm.","keywords":["diary study","speech-to-text","LLM-assisted writing","dictation","writing strategies","user acceptance","qualitative study","gist manipulation"],"falsifier":"A larger, preregistered longitudinal study with a control group writing on keyboards, objective time-on-task logging, and quality ratings by blind readers would settle the claim; the paradigm is undermined if speech-assisted writers show no completion-time advantage, no quality difference, and elevated dropout after the first week.","tokens_in":11424,"feed_emoji":"🎙️","tokens_out":3478,"duration_ms":31463,"temperature":0.7,"pith_summary":"This case study tries to establish that writing with speech, supported by LLM-assisted gist manipulation and macro-revision, is a viable paradigm for real-world writing rather than a lab-only curiosity. Twelve academic and creative writers used the Rambler tool to write blog posts, diaries, screenplays, notes, and fiction over seven to ten days, filing three articles each. The paper reports that nine of twelve participants adapted within that short period, that users adopted one of two distinguishable strategies depending on audience, and that participants perceived productivity gains and psychological benefits such as emotional expression and reduced overthinking. If right, the result means speech can serve as a primary text-input modality for at least some motivated writers outside controlled settings.","feed_headline":"LLM-assisted dictation holds up in a ten-day real-world writing test","feed_subtitle":"Twelve writers used Rambler for blogs, diaries, and screenplays; nine adapted and two stable strategies emerged.","key_machinery":"The central mechanism is Rambler's pair of semantic operations on dictated text: gist extraction (keywords and multi-level Semantic Zoom summaries) and macro-revision (Semantic Split, Semantic Merge, and a Custom Magic Prompt for arbitrary LLM transformations), plus manual editing for fine polish. These operations convert an unstructured spoken monologue into a structured written artifact, which is what makes speech usable as a primary input modality.","core_discovery":"The paper claims that LLM-assisted dictation, as embodied in Rambler, supports a complete writing workflow in the wild: users dictate impromptu thoughts, have the LLM clean disfluencies, extract gists for semantic zoom, and use semantic split, merge, and custom prompts to restructure and polish the text. The central discovery is that this combination yields two stable real-world writing strategies—'outline first' for academic or communicative writing with an external audience, and 'free speaking' for personal or reflective writing—and that users report productivity, emotional, and self-efficacy benefits that make the paradigm acceptable. The authors conclude that the approach has a positive outlook as a primary writing method.","pith_inferences":["If the perceived gains replicate, the same gist-manipulation mechanism could extend to other 'messy' input modalities, such as handwritten notes or transcribed meetings, where the bottleneck is restructuring rather than text entry.","The paper's self-efficacy finding suggests a testable extension: longitudinal language-learning studies could measure whether LLM-polished dictation improves writing proficiency faster than traditional editing feedback.","The reliance on self-reported time estimates implies that an objective keystroke/audio-log replication might show smaller productivity gains, since novelty and compensation could inflate perceived speed.","A natural design follow-up implied but not explored by the authors is an LLM that proposes the outline expansion or merge actions automatically, rather than waiting for the user to invoke semantic split and merge."],"forward_implications":["Speech-based writing with LLM support can move from lab tasks to users' own projects, with most participants adapting within 7 to 10 days.","Two distinct writing strategies emerge: outline-first for audience-directed texts and free-speaking for personal reflection, each using different tool functions.","Perceived productivity gains are substantial: 66.7% of 500-word tasks were estimated to take 10 to 30 minutes, lowering the barrier to writing routines.","LLM features prevent overthinking and build self-efficacy, with users learning writing style from model-polished output.","Remaining challenges—cursor-level editing, inability to multitask, and noisy environments—define the next design targets for speech-based writing tools."],"supporting_citations":[{"why":"Supplies the Rambler tool itself and the lab-study findings this diary study extends.","marker":"[10]"},{"why":"Provides the prior example of LLM paragraph summarization and semantic manipulation that motivates the macro-revision approach.","marker":"[4]"},{"why":"Supplies the inductive thematic analysis method used to code the qualitative data.","marker":"[5]"},{"why":"Supports the idea that generative models can assist writers via macro-level revision without writing for them.","marker":"[1]"},{"why":"Grounds the premise that speech bridges thought and text, the core rationale for the paradigm.","marker":"[3]"}],"fun_headline_variants":["LLM dictation aids writing: two stable strategies emerge in diary study","Twelve writers, two strategies: LLM-assisted speech writing works in the wild","Diary study: speech plus LLM yields productive writing workflows","Rambler field test: outline-first and free-speaking dominate LLM dictation","Real-world test: LLM dictation supports complete writing workflows"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's optimistic conclusion rests on the assumption that twelve self-selected, compensated, non-native-English participants who volunteered for a dictation tool, and their self-reported time estimates and feelings over seven to ten days, stand in for broader writer populations and reflect real productivity rather than novelty or payment effects.","fun_headline_variants_meta":{"raw":{"variants":["LLM dictation aids writing: two stable strategies emerge in diary study","Twelve writers, two strategies: LLM-assisted speech writing works in the wild","Diary study: speech plus LLM yields productive writing workflows","Rambler field test: outline-first and free-speaking dominate LLM dictation","Real-world test: LLM dictation supports complete writing workflows"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000629,"raw_usage":{"total_tokens":2843,"prompt_tokens":820,"completion_tokens":2023,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":436,"completion_tokens_details":{"reasoning_tokens":1927}},"tokens_in":436,"tokens_out":2023,"duration_ms":13071,"temperature":1.0,"reasoning_tokens":1927,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T18:37:03.790308+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A larger, preregistered longitudinal study with a control group writing on keyboards, objective time-on-task logging, and quality ratings by blind readers would settle the claim; the paradigm is undermined if speech-assisted writers show no completion-time advantage, no quality difference, and elevated dropout after the first week.","supporting_citations":[{"cited_title":"Zamfirescu-Pereira, Matthew G Lee, Sauhard Jain, Shanqing Cai, Piyawat Lertvittayakumjorn, Michael Xuelin Huang, Shumin Zhai, Bjoern Hartmann, and Can Liu","cited_arxiv_id":null,"evidence_quote":"Supplies the Rambler tool itself and the lab-study findings this diary study extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the idea that generative models can assist writers via macro-level revision without writing for them."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Grounds the premise that speech bridges thought and text, the core rationale for the paradigm."}],"review_version":1}