REVIEW 3 major objections 5 minor 15 references
Rambler in the Wild: A Diary Study of LLM-Assisted Writing With Speech
T0 review · 3 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read A ten-day diary study finds speech plus LLM support is a viable writing paradigm.
desk verdict A modest but honest diary study that confirms lab findings in the wild; the productivity claim is softer than the framing suggests, but the qualitative core holds up. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is Rambler's pair of semantic operations on dictated text: gist extraction (keywords and multi-level Semantic Zoom summaries) and macro-revision (Semantic Split, Semantic Merge, and a Custom Magic Prompt for arbitrary LLM transformations), plus manual editing for fine polish. These operations convert an unstructured spoken monologue into a structured written artifact, which is what makes speech usable as a primary input modality.
What would settle it
A larger, preregistered longitudinal study with a control group writing on keyboards, objective time-on-task logging, and quality ratings by blind readers would settle the claim; the paradigm is undermined if speech-assisted writers show no completion-time advantage, no quality difference, and elevated dropout after the first week.
Extended reading notes
Core claim
The paper claims that LLM-assisted dictation, as embodied in Rambler, supports a complete writing workflow in the wild: users dictate impromptu thoughts, have the LLM clean disfluencies, extract gists for semantic zoom, and use semantic split, merge, and custom prompts to restructure and polish the text. The central discovery is that this combination yields two stable real-world writing strategies—'outline first' for academic or communicative writing with an external audience, and 'free speaking' for personal or reflective writing—and that users report productivity, emotional, and self-efficacy benefits that make the paradigm acceptable. The authors conclude that the approach has a positive outlook as a primary writing method.
Load-bearing premise
The paper's optimistic conclusion rests on the assumption that twelve self-selected, compensated, non-native-English participants who volunteered for a dictation tool, and their self-reported time estimates and feelings over seven to ten days, stand in for broader writer populations and reflect real productivity rather than novelty or payment effects.
Editorial extensions
If this is right
- Speech-based writing with LLM support can move from lab tasks to users' own projects, with most participants adapting within 7 to 10 days.
- Two distinct writing strategies emerge: outline-first for audience-directed texts and free-speaking for personal reflection, each using different tool functions.
- Perceived productivity gains are substantial: 66.7% of 500-word tasks were estimated to take 10 to 30 minutes, lowering the barrier to writing routines.
- LLM features prevent overthinking and build self-efficacy, with users learning writing style from model-polished output.
- Remaining challenges—cursor-level editing, inability to multitask, and noisy environments—define the next design targets for speech-based writing tools.
Reading between the lines
- If the perceived gains replicate, the same gist-manipulation mechanism could extend to other 'messy' input modalities, such as handwritten notes or transcribed meetings, where the bottleneck is restructuring rather than text entry.
- The paper's self-efficacy finding suggests a testable extension: longitudinal language-learning studies could measure whether LLM-polished dictation improves writing proficiency faster than traditional editing feedback.
- The reliance on self-reported time estimates implies that an objective keystroke/audio-log replication might show smaller productivity gains, since novelty and compensation could inflate perceived speed.
- A natural design follow-up implied but not explored by the authors is an LLM that proposes the outline expansion or merge actions automatically, rather than waiting for the user to invoke semantic split and merge.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a ten-day diary study in which twelve academic and creative writers used Rambler, an LLM-assisted speech-to-text writing tool, for their own writing tasks. The authors identify two main composition strategies (outline-driven expansion and free-speaking followed by restructuring), describe affordances related to emotional expression, perceived productivity, and self-efficacy, and discuss user acceptance and remaining challenges. The paper concludes that speech-plus-LLM writing is a viable paradigm, with the positive outlook justified by reported productivity and psychological benefits.
Significance. If the findings hold, this is a useful real-world complement to the earlier lab study of Rambler, showing that at least some motivated writers can adopt speech as a primary writing modality and that LLM-based gist manipulation supports both structured and reflective writing. The study's strengths include the collection of thirty-six articles, interaction logs, and rich qualitative material; the authors are also transparent about the inability to measure task completion time directly. The strategy taxonomy and design suggestions are concrete and actionable. However, the quantitative productivity claim is not supported by the current evidence, and the sample is narrow, so the paper's central positive outlook is stronger than what the data can establish.
major comments (3)
- [§4.3.2, §4.4.1, §6] The productivity-gain component of the central claim rests on self-estimated completion-time bins (66.7% of tasks in 10–30 minutes) with no typing baseline, no within-subject comparison, and no analysis of the interaction logs described in §3.3 to verify durations. The text in §4.4.1 explicitly concedes that task completion time was not measured effectively, and the heading in §4.3.2 ('Speaking Boosts Productivity') asserts a causal productivity benefit. Because the conclusion in §6 is explicitly justified by 'productivity gain,' the current evidence supports only perceived productivity. Please either reframe the productivity claims as perceived or self-reported, or add a log-based duration analysis and a baseline or comparison condition.
- [§3.1, §5.4, §6] All twelve diary-study participants were non-native English speakers, were self-selected volunteers, and received supermarket coupons, and the entire observation period was seven to ten days. The statements that 'users could get used to writing with speech in a short period' (§5.4) and the generally positive outlook in §6 generalize beyond this sample. Please restrict the claims to the studied population of motivated, English-proficient non-native writers, and discuss how selection, compensation, and novelty effects may have shaped acceptance and reported benefits.
- [§3.3] The coding procedure reports that two researchers coded 25% of the data independently and reached consensus, after which one researcher coded the remainder, but no inter-rater reliability statistic or codebook is reported. Since the thematic findings are the core evidence for the psychological and strategy claims, please report agreement measures or provide a more complete audit trail to support the reliability of the themes.
minor comments (5)
- [§3.1, Appendix Table 1] Please clarify the relationship between the 14 focus-group respondents and the 12 diary-study participants, including why P1, P7, and P11 were absent from the diary phase and whether any recruitment or attrition issues affected the final sample.
- [§4.4.1] Please include the exact survey question and response options used for the time estimates, since the current report gives only bins and percentages.
- [Figure 4] The statement that the deviation of comfort scores decreases after the first use is based on visual inspection; please report the underlying distributions or a compact summary statistic.
- [§2] The description of Custom Magic Prompt does not state whether it applies to a single Ramble or to all Rambles at once, which matters because §4.5.1 reports a request to apply prompts to all Rambles simultaneously.
- [§3.1] The phrase 'a English-taught degree program' should read 'an English-taught degree program.'
Circularity Check
No circularity: the diary-study findings derive from new empirical data, and the self-citation to the prior Rambler paper is descriptive context rather than the load-bearing argument.
full rationale
This is an empirical qualitative diary study. It contains no equations, fitted parameters, or derivation chain in which an output is computed from an input. The central conclusions—two writing strategies, productivity and psychological affordances, and user acceptance—are inductive themes and self-report summaries collected from twelve participants' diary surveys, interaction logs, and exit interviews. The only self-citation is to the authors' prior lab-study paper [10], which is used to describe the Rambler tool and to frame the follow-up; the diary study's findings are new observations of real-world usage and are not derived from [10]. Section 4.4.1's productivity discussion is based on participants' retrospective time-bin estimates and explicitly acknowledges that task completion time was not measured effectively; this is an evidentiary weakness about validity, not circularity, because the estimates are not a fitted parameter renamed as a prediction and no quantity is defined in terms of the conclusion. No load-bearing step reduces to its own inputs by construction, and no uniqueness claim or ansatz is imported via citation. Possible demand characteristics or self-report bias would be study-design concerns, not circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption Participants' self-reported time estimates and acceptance ratings reflect their actual writing experience.
- domain assumption Inductive thematic analysis by two researchers, with one coding 25% and then one coding the rest, yields stable themes.
- domain assumption The Rambler tool operates as described in the prior lab study [10].
Cite this review
Pith. "Pith review of Rambler in the Wild: A Diary Study of LLM-Assisted Writing With Speech." pith.science (2026). https://pith.science/paper/QCHB7WUV
@misc{pith2026250205612,
author = {Pith},
title = {Pith review of: Rambler in the Wild: A Diary Study of LLM-Assisted Writing With Speech},
year = {2026},
howpublished = {\url{https://pith.science/paper/QCHB7WUV}},
note = {Machine review of arXiv:2502.05612}
}
read the original abstract
Speech-to-text technologies have been shown to improve text input efficiency and potentially lower the barriers to writing. Recent LLM-assisted dictation tools aim to support writing with speech by bridging the gaps between speaking and traditional writing. This case study reports on the real-world writing experiences of twelve academic or creative writers using one such tool, Rambler, to write various pieces such as blog posts, diaries, screenplays, notes, or fictional stories, etc. Through a ten-day diary study, we identified the participants' in-context writing strategies using Rambler, such as how they expanded from an outline or organized their loose thoughts for different writing goals. The interviews uncovered the psychological and productivity affordances of writing with speech, pointing to future directions of designing for this writing modality and the utilization of AI support.
Figures
Reference graph
Works this paper leans on
-
[1]
Kenneth C Arnold, April M Volzer, and Noah G Madrid. 2021. Generative Models can Help Writers without Writing for Them. In Joint Proceedings of the IUI 2021 Workshops. CEUR-WS Team, College Station, TX, USA, 8 pages. http://ceur- ws.org/Vol-2903/
work page 2021
-
[2]
Saksham Bassi, Giulio Duregon, Siddhartha Jalagam, and David Roth. 2023. End- to-End Speech Recognition and Disfluency Removal with Acoustic Language Model Pretraining. arXiv:2309.04516 [eess.AS] https://arxiv.org/abs/2309.04516
work page Pith review arXiv 2023
-
[3]
David Crystal. 1995. Speaking of writing and writing of speaking. Longman Language Review 1 (1995), 5–8
work page 1995
-
[4]
Hai Dang, Karim Benharrak, Florian Lehmann, and Daniel Buschek. 2022. Beyond Text Generation: Supporting Writers with Continuous Automatic Text Summaries. In Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology (Bend, OR, USA) (UIST ’22). Association for Computing Machinery, New York, NY, USA, Article 98, 13 pages. https:...
arXiv 2022
-
[5]
Jennifer Fereday and Eimear Muir-Cochrane. 2006. Demonstrating Rigor Using Thematic Analysis: A Hybrid Approach of Inductive and Deductive Coding and Theme Development. International Journal of Qualitative Methods 5, 1 (2006), 80–92. https://doi.org/10.1177/160940690600500107
-
[6]
Clare-Marie Karat, Christine Halverson, Daniel Horn, and John Karat. 1999. Pat- terns of entry and correction in large vocabulary continuous speech recognition systems. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Pittsburgh, Pennsylvania, USA) (CHI ’99). Association for Computing Machinery, New York, NY, USA, 568–575. ht...
arXiv 1999
-
[7]
Daniel Li, Thomas Chen, Albert Tung, and Lydia B Chilton. 2021. Hierar- chical Summarization for Longform Spoken Dialog. In The 34th Annual ACM Symposium on User Interface Software and Technology (Virtual Event, USA) (UIST ’21). Association for Computing Machinery, New York, NY, USA, 582–597. https://doi.org/10.1145/3472749.3474771
-
[8]
Daniel Li, Thomas Chen, Alec Zadikian, Albert Tung, and Lydia B Chilton. 2023. Improving Automatic Summarization for Browsing Longform Spoken Dialog. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Association for Computing Machinery, New York, NY, USA, Article 106, 20 pages. https://doi.org/10...
Show all 15 references
-
[9]
Junwei Liao, Sefik Eskimez, Liyang Lu, Yu Shi, Ming Gong, Linjun Shou, Hong Qu, and Michael Zeng. 2023. Improving Readability for Automatic Speech Recognition Transcription. ACM Trans. Asian Low-Resour. Lang. Inf. Process. 22, 5, Article 142 (May 2023), 23 pages. https://doi.o...
2023 doi
-
[10]
Zamfirescu-Pereira, Matthew G Lee, Sauhard Jain, Shanqing Cai, Piyawat Lertvittayakumjorn, Michael Xuelin Huang, Shumin Zhai, Bjoern Hartmann, and Can Liu
Susan Lin, Jeremy Warner, J.D. Zamfirescu-Pereira, Matthew G Lee, Sauhard Jain, Shanqing Cai, Piyawat Lertvittayakumjorn, Michael Xuelin Huang, Shumin Zhai, Bjoern Hartmann, and Can Liu. 2024. Rambler: Supporting Writing With Speech via LLM-Assisted Gist Manipulation. In Proce...
2024
-
[11]
Brinda Mehra, Kejia Shen, Hen Chen Yen, and Can Liu. 2023. Gist and Verbatim: Understanding Speech to Inform New Interfaces for Verbal Text Composition. In Proceedings of the 5th International Conference on Conversational User Interfaces (Eindhoven, Netherlands) (CUI ’23). Ass...
2023
-
[12]
Wobbrock, Kenny Liou, Andrew Ng, and James A
Sherry Ruan, Jacob O. Wobbrock, Kenny Liou, Andrew Ng, and James A. Landay
-
[13]
Tomohiro Tanaka, Ryo Masumura, Hirokazu Masataki, and Yushi Aono. 2018. Neural Error Corrective Language Models for Automatic Speech Recognition. In INTERSPEECH. International Speech Communication Association, Hyderabad, India, 401–405. https://doi.org/10.21437/Interspeech.2018-1430
2018 doi
-
[14]
Daijin Yang, Yanpeng Zhou, Zhiyuan Zhang, Toby Jia-Jun Li, and Ray LC. 2022. AI as an Active Writer: Interaction strategies with generated text in human-AI collaborative fiction writing. In Joint Proceedings of the IUI 2022 Workshops (CEUR Workshop Proceedings), Alison Smith-R...
2022
-
[2018]
Comparing Speech and Keyboard Text Entry for Short Messages in Two Languages on Touchscreen Phones. Proc. ACM Interact. Mob. Wearable Ubiquitous CHI EA ’25, April 26-May 1, 2025, Yokohama, Japan Yang, Li, et al. Technol. 1, 4, Article 159 (Jan. 2018), 23 pages. https://doi.org...
2025 doi
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.