Pith. sign in

REVIEW 3 major objections 4 minor 51 references

Content-Driven Local Response: Supporting Sentence-Level and Message-Level Mobile Email Replies With and Without AI

T0 review · 3 major / 4 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read The paper proposes Content-Driven Local Response (CDLR), a mobile email reply UI where users tap sentences in the incoming email to insert local replies, and argues this fills a new design space between sentence-level and message-level AI…

desk verdict A well-run study with a genuinely new UI concept; the author-built MSG baseline is the main caveat, but the design-space claim holds. read the letter →

arxiv 2502.06430 v1 pith:FK2NSTSY submitted 2025-02-10 cs.HC cs.CL

classification cs.HCcs.CL
keywords Content-DrivenLocalResponsemobileemailAIwritingassistancesentence-levelsuggestionsmessage-levelgenerationhuman-AIinteractionmicrotaskinguserstudy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper is trying to establish that mobile email reply UIs do not have to choose between manual typing, sentence-level suggestions, and full-message AI generation: a UI built around responding directly inside the incoming email, sentence by sentence, lets users mix all three. It claims this Content-Driven Local Response (CDLR) concept occupies a distinct middle spot in the design space, giving people flexible control over how much AI they involve while still reducing typing and errors. The authors support this with a controlled study of 126 participants comparing CDLR with manual writing and message-level generation. The result matters because current email apps largely add AI as a pop-up on top of an empty draft view, hiding the email and forcing users into an editor role.

What carries the argument

The load-bearing mechanism is sentence selection as a dual-purpose interaction: tapping a sentence both inserts a local response and expresses intent that conditions the AI's suggestions. The local response widget offers six suggestions (two positive, two negative, two neutral), generated from the selected sentence, the incoming email, and all local replies so far, while the optional message-level improvement pass turns the collected local responses into a polished draft shown with tracked changes. This combination makes AI support skippable at every step, so workflows can range from fully manual to fully AI-generated within one UI.

What would settle it

Take the same nine emails and briefings and compare CDLR against the actual production AI reply flows in Gmail, Outlook, or Superhuman (or a closely matched prototype), measuring completion time, keystrokes, and briefing conformity; if CDLR no longer falls between manual typing and message-level generation on these metrics, for instance if its 66-second speed gap disappears or reverses, the central design-space claim would be refuted.

Watch

Extended reading notes

Core claim

The central claim is that redesigning the reply UI around the content of the incoming email, rather than adding full-reply generation on top of a draft view, lets users dynamically set the degree of AI involvement. In CDLR, tapping a sentence opens a local widget where users can type a response or accept AI suggestions conditioned on that sentence and on earlier local replies; a later "improve email" pass optionally revises the whole message with tracked changes. In the study, CDLR significantly reduced keystrokes (about 48 percent versus manual) and error rates, increased writing speed when the improvement pass was used, and produced replies with more lexical and semantic diversity than full message generation. It was slower than message-level generation by about 66 seconds per task, and participants rated it lower on speed but valued it for control and quality, with 43.7 percent choosing it as their favourite versus 49.2 percent for message-level generation.

Load-bearing premise

The comparative conclusions depend on the message-level generation condition faithfully representing the AI reply flow in current mobile email apps; if the authors' MSG prototype is not representative, the claim that CDLR sits between existing sentence-level and message-level designs is weakened.

Editorial extensions

If this is right

  • If CDLR is correct, mobile email clients can offer sentence-level and message-level AI support in one UI without forcing an "AI-first" workflow, letting each reply choose its own level of delegation.
  • Users can get roughly half the keystroke reduction of full generation (48 percent versus 58 percent) while keeping more content diversity and control, so the speed-agency tradeoff becomes a dial rather than a fork.
  • Longer incoming emails make the local-response step more valuable: each additional word lowers the chance of skipping it by about 2.46 percent, predicting that CDLR matters most for complex, multi-question email threads.
  • When users generate a full reply with no prompt or local input, half of the replies miss a key briefing point, so AI assistance without user input is the biggest risk to task adherence, larger than the choice of UI.
  • A production email app could combine CDLR with a message-level workflow, since CDLR already supports skipping to an "improve" pass that acts as full-reply generation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The authors do not test this, but the "selection-as-prompt" pattern likely transfers beyond email: any mobile task where the source document stays on screen (chat threads, forms, long articles) could use tapping a passage to both record a response and steer generation.
  • The paper's Co-Creative Chain-of-Thought reading suggests UI structure can replace explicit reasoning prompts; a direct test would compare CDLR against a message-level condition that first asks users to mark the key points they want addressed.
  • Because the study's paid setting may inflate acceptance of suggestions, real-world deployments could show even stronger user editing and briefing gaps; field studies with actual inboxes would resolve this.
  • The lower lexical diversity of CDLR than MSG is surprising given CDLR's greater control, and may reflect the balanced positive/negative/neutral prompt template; varying suggestion diversity could alter the measured tradeoff.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper introduces Content-Driven Local Response (CDLR), a mobile email reply UI that lets users tap sentences in the incoming email to insert local responses, optionally using LLM suggestions, and later finalize the draft with an optional message-level improvement pass. A controlled within-subject study (N=126) compares CDLR with manual typing and with a message-level reply generation design modeled on current apps, measuring interaction logs, perceived control/speed/quality, and email characteristics (length, error rate, diversity, briefing conformity). The authors report that CDLR occupies an intermediate design space between sentence-level and message-level support, supporting flexible workflows with reduced typing and errors while being slower than full generation.

Significance. If the findings hold, the paper makes a useful contribution to human-AI interaction for mobile email: it offers a concrete, implementable UI concept that combines sentence-level and message-level AI involvement with optionality. The empirical study is ambitious (nine emails, three UI modes, counterbalanced, with mixed-effects modeling and dual coding), and the prototype and study materials are released for reuse. Several analyses (diversity metrics, error rates, workflow clustering, briefing conformity) go beyond typical self-report-only evaluations. The main qualification is that the comparative claim of occupying a 'new distinct spot' depends on the representativeness of the author-built MSG baseline.

major comments (3)
  1. [Section 5.1.2; Appendix A.1.4] The message-level generation (MSG) condition is an author-built prototype whose prompt template explicitly asks for 'a well written email' with a greeting and sign-off and instructs the model not to make anything up. No validation is reported that this condition reproduces the behavior, prompt quality, or response style of production systems such as Gmail/Gemini, Outlook/Copilot, Superhuman, or Shortwave. Because the abstract and Section 7.4 claim that CDLR takes 'a new distinct spot in the design space between sentence-level and message-level support,' this claim is anchored to an unrepresentative baseline: the observed CDLR-vs-MSG differences (completion time, keystrokes, reply length, diversity, edit behavior) may reflect the specific prompt rather than message-level generation per se. The internal comparison between conditions remains valid, but the external positioning needs either a validation substudy (e.g., running the same email/prompt set on production systems) or a reframing that restricts the contribution to the authors' specific MSG implementation.
  2. [Section 6.1.6; Figure 6] The workflow analysis asserts 'three main workflows' from a Gaussian Mixture Model with 3 components, but the number of components is not justified (e.g., via BIC, silhouette, or stability analysis), and the model is only presented as point estimates of cluster sizes. Since flexible workflow support is a central claimed benefit of CDLR, the clustering result should be substantiated (or described as an illustrative visualization) and accompanied by robustness checks.
  3. [Section 5.2] The exclusion of 36 of 162 participants (22%) is reported, but the criteria are not pre-registered and the analysis is complete-case only. In a within-subject design with counterbalancing, this could bias the comparison if exclusions were uneven across modes or email orders. Please report which conditions/exclusions contributed to the 18 technical-issue removals and provide sensitivity analyses (e.g., mixed models using all available data or imputation).
minor comments (4)
  1. [Section 3.5] In the sentence 'we removed the second screen ... and direly offered the third one for free text editing and finalising,' 'direly' appears to be a typo for 'directly.'
  2. [Table 1, row 5] The 95% confidence interval for the MSG error-rate predictor is printed as '[-.0025, .0015]', which appears to be a typo for a negative upper bound; as printed it is inconsistent with the reported negative coefficient.
  3. [Section 5.5] It would be helpful to state how the fixed effects of 'UImode' were coded (treatment contrasts with NoAI as baseline are implied) and whether the random structure included random slopes; the current table does not fully describe the model specifications.
  4. [Figure 6] The figure's y-axis label 'Draft Progress [% final length]' with values above 100% is potentially confusing; the caption explains it, but consider adding a note directly on the axis.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the central claims are measured from interaction logs and user feedback, not derived from the paper's own assumptions.

full rationale

This is an empirical HCI paper rather than a derivation chain. The central claim that CDLR occupies a distinct spot between sentence-level and message-level support rests on a controlled within-subject study (N=126) with logged interaction metrics, Likert ratings, coded open feedback, and email-content analyses (Sections 6.1-6.3). These results are not equivalent by construction to the design choices: for example, the finding that MSG was faster than CDLR while CDLR elicited more control-related reasoning is a measured contrast, not an analytic consequence of how CDLR was defined. The authors' own MSG baseline is described as resembling the typical UI pattern (Section 5.1.2), and the skeptical concern that this baseline may not fully represent production systems is an external-validity limitation, not circularity. The paper contains two minor self-citations ([6] Buschek et al. 2021 and [10] Dang et al. 2023) used to motivate the number and type of suggestions (Section 2.4), but these are ancillary design rationale, not the load-bearing evidence for the main contribution, which comes from the new study. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from the authors' prior work, and no known empirical result is merely relabeled. Accordingly, no specific circular step can be exhibited; the score reflects only the presence of self-citations that are not load-bearing.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

This is an empirical HCI study, so there are no mathematical free parameters in a derivation sense. The closest items are the GMM component count (model selection) and the domain assumptions about baseline representativeness, LLM quality, and briefing validity. No new physical or conceptual entities are introduced.

free parameters (1)
  • Number of GMM components for workflow clustering = 3
    In Section 6.1.6, the authors fit a Gaussian Mixture Model to identify workflow patterns and chose 3 components. This is a model selection choice made by hand, and no comparison with other component counts is reported. The claim of 'three main workflows' depends on this choice.
assumptions (4)
  • domain assumption The MSG condition faithfully represents the industry-standard AI reply generation flow in mobile email apps.
    Section 5.1.2 describes the MSG implementation as designed 'similar to the typical UI pattern' shown in Figure 2, but the implementation is the authors' own, using their chosen prompt templates and model. The general conclusions about existing apps (Gmail, Outlook, Superhuman) assume this baseline is representative.
  • domain assumption Participants' behavior in a paid online study, including acceptance of AI suggestions, reflects real-world email response choices.
    Section 7.5 acknowledges that the willingness to accept suggestions may be inflated in the study due to satisficing, and that the study had no negative consequences for writing a poor email. The interaction metrics, especially suggestion acceptance rates, rely on this assumption to generalize to real contexts.
  • domain assumption Llama 3 8B Instruct produces suggestions of sufficient quality and neutrality to represent current AI assistance in email.
    Section 4.2 states the authors used Llama 3 8B Instruct after qualitative comparison with GPT-3.5 Turbo, but no systematic evaluation was conducted. The measured tradeoffs between control and efficiency for CDLR and MSG could change with a different model or prompt set.
  • domain assumption The briefing conformity coding accurately captures the key information participants should include in their replies.
    Section 5.4.2 describes binary coding by two researchers with consensus, but the briefings were created by the authors and may not represent all plausible intentions. The result that CDLR had the lowest conformity (23 percent) is interpreted as satisficing, which assumes the briefing is a valid ground truth.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Content-Driven Local Response: Supporting Sentence-Level and Message-Level Mobile Email Replies With and Without AI." pith.science (2026). https://pith.science/paper/FK2NSTSY

@misc{pith2026250206430,
  author       = {Pith},
  title        = {Pith review of: Content-Driven Local Response: Supporting Sentence-Level and Message-Level Mobile Email Replies With and Without AI},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FK2NSTSY}},
  note         = {Machine review of arXiv:2502.06430}
}
read the original abstract

Mobile emailing demands efficiency in diverse situations, which motivates the use of AI. However, generated text does not always reflect how people want to respond. This challenges users with AI involvement tradeoffs not yet considered in email UIs. We address this with a new UI concept called Content-Driven Local Response (CDLR), inspired by microtasking. This allows users to insert responses into the email by selecting sentences, which additionally serves to guide AI suggestions. The concept supports combining AI for local suggestions and message-level improvements. Our user study (N=126) compared CDLR with manual typing and full reply generation. We found that CDLR supports flexible workflows with varying degrees of AI involvement, while retaining the benefits of reduced typing and errors. This work contributes a new approach to integrating AI capabilities: By redesigning the UI for workflows with and without AI, we can empower users to dynamically adjust AI involvement.

Figures

Figures reproduced from arXiv: 2502.06430 by the authors.

Figure 1
Figure 1. Replying to an email with Content-Driven Local Response: (1) In the local response view, users can insert responses (A) directly while reading the email. (B) Tapping on a sentence opens a response widget, (C) with a text box where users enter a response or a prompt that affects (D) the sentence suggestions below. (2) After adding local responses, users go to the draft view, to turn their responses into a full reply … view at source ↗
Figure 2
Figure 2. A commonly used UI and interaction design for reply generation in mobile email apps: [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Replying to an email with our first prototype: [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (10 more)
Figure 4
Figure 4. Figure 4: The text suggestions in the local response widget are flexible: [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Three measures of interaction behaviour: Task completion time [PITH_FULL_IMAGE:figures/full_fig_p011_5.png]
Figure 7
Figure 7. Figure 7: Participants spent their time on different screens [PITH_FULL_IMAGE:figures/full_fig_p012_7.png]
Figure 8
Figure 8. Figure 8: Likert results on perception of the UIs and interaction, rated after each email task. Overall, participants rated the [PITH_FULL_IMAGE:figures/full_fig_p013_8.png]
Figure 9
Figure 9. Figure 9: Four measures of email characteristics from our study. The plots show email length [PITH_FULL_IMAGE:figures/full_fig_p014_9.png]
Figure 10
Figure 10. Figure 10: Likert results from the formative study (Section 3.3). The first three items were asked via an in-app feedback screen [PITH_FULL_IMAGE:figures/full_fig_p021_10.png]
Figure 11
Figure 11. Figure 11: SUS [5] results from the formative study (Section 3.3) [PITH_FULL_IMAGE:figures/full_fig_p021_11.png]
Figure 12
Figure 12. Figure 12: The study-specific UI screens in our prototype: A screen showing the briefing before each email task [PITH_FULL_IMAGE:figures/full_fig_p022_12.png]
Figure 13
Figure 13. Figure 13: The UI design for manual typing (NoAI) used in the study. It has one screen to show the incoming email (1) and one with a text box to type the reply manually (2). This figure shows the state after typing a reply, as an example. As is usual on mobile devices, the keybo…
Figure 14
Figure 14. Figure 14: The UI design for message-level reply generation ( [PITH_FULL_IMAGE:figures/full_fig_p023_14.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

51 extracted references · 13 canonical work pages

  1. [1]

    AI@Meta. 2024. Llama 3 Model Card. (2024). https://github.com/meta-llama/ llama3/blob/main/MODEL_CARD.md

  2. [2]

    Tal August, Shamsi Iqbal, Michael Gamon, and Mark Encarnación. 2020. Charac- terizing the Mobile Microtask Writing Process. In 22nd International Conference on Human-Computer Interaction with Mobile Devices and Services (Oldenburg, Germany) (MobileHCI ’20). Association for Computing Machinery, New York, NY, USA, Article 26, 12 pages. https://doi.org/10.11...

  3. [3]

    Patti Bao, Jeffrey Pierce, Stephen Whittaker, and Shumin Zhai. 2011. Smart phone use by non-mobile business users. In Proceedings of the 13th International Conference on Human Computer Interaction with Mobile Devices and Services (Stockholm, Sweden) (MobileHCI ’11). Association for Computing Machinery, New York, NY, USA, 445–454. https://doi.org/10.1145/2...

  4. [4]

    Douglas Bates, Martin Mächler, Ben Bolker, and Steve Walker. 2015. Fitting Linear Mixed-Effects Models Using lme4. Journal of Statistical Software 67, 1 (2015), 1–48. https://doi.org/10.18637/jss.v067.i01

  5. [5]

    John Brooke et al. 1996. SUS-A quick and dirty usability scale.Usability evaluation in industry 189, 194 (1996), 4–7

  6. [6]

    Daniel Buschek, Martin Zürn, and Malin Eiband. 2021. The Impact of Multiple Parallel Phrase Suggestions on Email Input and Composition Behaviour of Native and Non-Native English Writers. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokohama, Japan) (CHI ’21). Association for Computing Machinery, New York, NY, USA, Artic...

  7. [7]

    Lee, Gagan Bansal, Yuan Cao, Shuyuan Zhang, Justin Lu, Jackie Tsay, Yinan Wang, Andrew M

    Mia Xu Chen, Benjamin N. Lee, Gagan Bansal, Yuan Cao, Shuyuan Zhang, Justin Lu, Jackie Tsay, Yinan Wang, Andrew M. Dai, Zhifeng Chen, Timothy Sohn, and Yonghui Wu. 2019. Gmail Smart Compose: Real-Time Assisted Writing. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (Anchorage, AK, USA) (KDD ’19) . Assoc...

  8. [8]

    Iqbal, and Michael S

    Justin Cheng, Jaime Teevan, Shamsi T. Iqbal, and Michael S. Bernstein. 2015. Break It Down: A Comparison of Macro- and Microtasks. In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems (Seoul, Republic of Korea) (CHI ’15). Association for Computing Machinery, New York, NY, USA, 4061–4064. https://doi.org/10.1145/2702123.2702146

Show all 51 references
  1. [9]

    Juliet M Corbin. 1990. Basics of qualitative research: Grounded theory procedures and techniques. Sage

  2. [10]

    Hai Dang, Sven Goller, Florian Lehmann, and Daniel Buschek. 2023. Choice Over Control: How Users Write with Large Language Models using Diegetic and Non-Diegetic Prompting. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ...

  3. [11]

    Elkin, Matthew Kay, James J

    Lisa A. Elkin, Matthew Kay, James J. Higgins, and Jacob O. Wobbrock. 2021. An Aligned Rank Transform Procedure for Multifactor Contrast Tests. In The 34th Annual ACM Symposium on User Interface Software and Technology (Virtual Event, USA) (UIST ’21). Association for Computing ...

  4. [12]

    Liye Fu, Benjamin Newman, Maurice Jakesch, and Sarah Kreps. 2023. Compar- ing Sentence-Level Suggestions to Message-Level Suggestions in AI-Mediated Communication. In Proceedings of the 2023 CHI Conference on Human Fac- tors in Computing Systems (Hamburg, Germany) (CHI ’23) . ...

  5. [13]

    Yue Fu, Sami Foell, Xuhai Xu, and Alexis Hiniker. 2024. From Text to Self: Users’ Perception of AIMC Tools on Interpersonal Communication and Self. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Association for Computi...

  6. [14]

    Goodman, Erin Buehler, Patrick Clary, Andy Coenen, Aaron Donsbach, Tiffanie N

    Steven M. Goodman, Erin Buehler, Patrick Clary, Andy Coenen, Aaron Donsbach, Tiffanie N. Horne, Michal Lahav, Robert MacDonald, Rain Breaw Michaels, Ajit Narayanan, Mahima Pushkarna, Joel Riley, Alex Santana, Lei Shi, Rachel Sweeney, Phil Weaver, Ann Yuan, and Meredith Ringel ...

  7. [15]

    Google. 2024. Draft emails with Gemini in Gmail - Android - Gmail Help. https://support.google.com/mail/answer/13955415?co=GENIE.Platform% 3DAndroid&oco=1

  8. [16]

    Jeffrey T Hancock, Mor Naaman, and Karen Levy. 2020. AI-Mediated Communication: Definition, Research Agenda, and Ethical Considerations. Journal of Computer-Mediated Communication 25, 1 (01 2020), 89–100. https: //doi.org/10.1093/jcmc/zmz022 arXiv:https://academic.oup.com/jcmc...

  9. [17]

    Iqbal, Jaime Teevan, Dan Liebling, and Anne Loomis Thompson

    Shamsi T. Iqbal, Jaime Teevan, Dan Liebling, and Anne Loomis Thompson. 2018. Multitasking with Play Write, a Mobile Microproductivity Writing Tool. In Pro- ceedings of the 31st Annual ACM Symposium on User Interface Software and Tech- nology (Berlin, Germany) (UIST ’18). Assoc...

  10. [18]

    Anjuli Kannan, Karol Kurach, Sujith Ravi, Tobias Kaufmann, Andrew Tomkins, Balint Miklos, Greg Corrado, Laszlo Lukacs, Marina Ganea, Peter Young, and Vivek Ramavajjala. 2016. Smart Reply: Automated Response Suggestion for Email. In Proceedings of the 22nd ACM SIGKDD Internatio...

  11. [20]

    Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. Advances in neural information processing systems 35 (2022), 22199–22213

  12. [21]

    Max Kreminski. 2024. The Dearth of the Author in AI-Supported Writing. arXiv:2404.10289 [cs.HC] https://arxiv.org/abs/2404.10289

  13. [22]

    Per Ola Kristensson and Keith Vertanen. 2014. The inviscid text entry rate and its application as a grand goal for mobile text entry. InProceedings of the 16th Interna- tional Conference on Human-Computer Interaction with Mobile Devices & Services (Toronto, ON, Canada) (Mobile...

  14. [23]

    Brockhoff, and Rune H

    Alexandra Kuznetsova, Per B. Brockhoff, and Rune H. B. Christensen. 2017. lmerTest Package: Tests in Linear Mixed Effects Models. Journal of Statistical Software 82, 13 (2017), 1–26. https://doi.org/10.18637/jss.v082.i13

  15. [24]

    Luis Leiva, Matthias Böhmer, Sven Gehring, and Antonio Krüger. 2012. Back to the app: the costs of mobile application interruptions. In Proceedings of the 14th International Conference on Human-Computer Interaction with Mobile Devices and Services (San Francisco, California, U...

  16. [25]

    Jenny Lewin-Jones. 2014. Understanding style, language and etiquette in email communication in higher education: a survey. Research in Post-Compulsory Education 19 (01 2014), 75–90. https://doi.org/10.1080/13596748.2014.872934

  17. [27]

    Zamfirescu-Pereira, Matthew G Lee, Sauhard Jain, Shanqing Cai, Piyawat Lertvittayakumjorn, Michael Xuelin Huang, Shumin Zhai, Bjoern Hartmann, and Can Liu

    Susan Lin, Jeremy Warner, J.D. Zamfirescu-Pereira, Matthew G Lee, Sauhard Jain, Shanqing Cai, Piyawat Lertvittayakumjorn, Michael Xuelin Huang, Shumin Zhai, Bjoern Hartmann, and Can Liu. 2024. Rambler: Supporting Writing With Speech via LLM-Assisted Gist Manipulation. In Proce...

  18. [28]

    Yihe Liu, Anushk Mittal, Diyi Yang, and Amy Bruckman. 2022. Will AI Console Me when I Lose my Pet? Understanding Perceptions of AI-Mediated Email Writing. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA) (CHI ’22). Associat...

  19. [29]

    Li Lucy, Su Lin Blodgett, Milad Shokouhi, Hanna Wallach, and Alexandra Olteanu

  20. [30]

    Hancock, Mor Naaman, Malte Jung, and Jess Hohenstein

    Hannah Mieczkowski, Jeffrey T. Hancock, Mor Naaman, Malte Jung, and Jess Hohenstein. 2021. AI-Mediated Communication: Language Use and Interpersonal Effects in a Referential Communication Task. Proc. ACM Hum.-Comput. Interact. 5, CSCW1, Article 17 (April 2021), 14 pages. https...

  21. [31]

    Hannah Nicole Mieczkowski. 2022. AI-Mediated Communication: Examining Agency, Ownership, Expertise, and Roles of AI Systems. Ph. D. Dissertation. Stanford, CA, USA. Advisor(s) Jeremy, Bailenson, and Mor, Naaman, and Byron, Reeves,. AAI29756350

  22. [32]

    Miles, A.M

    M.B. Miles, A.M. Huberman, and J. Saldana. 2013. Qualitative Data Analysis: A Methods Sourcebook. SAGE Publications. https://books.google.de/books?id= p0wXBAAAQBAJ

  23. [33]

    Jakob Nielsen. 1994. Enhancing the explanatory power of usability heuristics. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Boston, Massachusetts, USA) (CHI ’94). Association for Computing Machinery, New York, NY, USA, 152–158. https://doi.org/...

  24. [34]

    Antti Oulasvirta, Tye Rattenbury, Lingyi Ma, and Eeva Raita. 2012. Habits make smartphone use more pervasive. Personal and Ubiquitous Computing 16, 1 (Jan. 2012), 105–114. https://doi.org/10.1007/s00779-011-0412-2

  25. [35]

    Antti Oulasvirta, Sakari Tamminen, Virpi Roto, and Jaana Kuorelahti. 2005. Inter- action in 4-second bursts: the fragmented nature of attentional resources in mo- bile HCI. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Portland, Oregon, USA)(CH...

  26. [36]

    Vishakh Padmakumar and He He. 2024. Does Writing with Language Models Reduce Content Diversity? arXiv:2309.05196 [cs.CL] https://arxiv.org/abs/2309. 05196

  27. [37]

    Kseniia Palin, Anna Maria Feit, Sunjun Kim, Per Ola Kristensson, and Antti Oulasvirta. 2019. How do People Type on Mobile Devices? Observations from a Study with 37,000 Volunteers. In Proceedings of the 21st International Conference on Human-Computer Interaction with Mobile De...

  28. [38]

    Zhang, Luke S

    Soya Park, Amy X. Zhang, Luke S. Murray, and David R. Karger. 2019. Opportu- nities for Automating Email Processing: A Need-Finding Study. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow, Scotland Uk) (CHI ’19). Association for Computin...

  29. [39]

    Philip Quinn and Shumin Zhai. 2016. A Cost-Benefit Study of Text Entry Sugges- tion Interaction. In Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems (San Jose, California, USA) (CHI ’16). Association for Com- puting Machinery, New York, NY, USA, 83–...

  30. [40]

    R Core Team. 2020. R: A Language and Environment for Statistical Computing . R Foundation for Statistical Computing, Vienna, Austria. https://www.R-project. org

  31. [41]

    Dimitrios Raptis, Nikolaos Tselios, Jesper Kjeldskov, and Mikael B. Skov. 2013. Does size matter? investigating the impact of mobile phone screen size on users’ perceived usability, effectiveness and efficiency.. InProceedings of the 15th Interna- tional Conference on Human-Co...

  32. [42]

    B. Reeves. 2008. Teach Yourself - The Internet and Email for the Over 50s . McGraw- Hill. https://books.google.de/books?id=gz04JwAACAAJ

  33. [43]

    Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Em- pirical Methods in Natural Language Processing . Association for Computational Linguistics. https://arxiv.org/abs/1908.10084

  34. [44]

    I Can’t Reply with That

    Ronald E Robertson, Alexandra Olteanu, Fernando Diaz, Milad Shokouhi, and Peter Bailey. 2021. “I Can’t Reply with That”: Characterizing Problematic Email Reply Suggestions. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokohama, Japan) (CHI ’...

  35. [45]

    Niloufar Salehi, Jaime Teevan, Shamsi Iqbal, and Ece Kamar. 2017. Communicating Context to the Crowd for Complex Writing Tasks. In Proceedings of the 2017 ACM Conference on Computer Supported Cooperative Work and Social Computing (Portland, Oregon, USA) (CSCW ’17). Association...

  36. [46]

    Superhuman. 2024. Superhuman AI for iPhone and iPad. https://www.youtube. com/watch?v=ijfi16JJguI

  37. [47]

    Lev Tankelevitch, Viktor Kewenig, Auste Simkute, Ava Elizabeth Scott, Advait Sarkar, Abigail Sellen, and Sean Rintel. 2024. The Metacognitive Demands and Opportunities of Generative AI. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, U...

  38. [48]

    The Copilot Connection. 2023. M365 Copilot demo - Outlook on mobile. https: //www.youtube.com/watch?v=Ru8OUhdKYpI

  39. [49]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35 (2022), 24824–24837

  40. [50]

    Wobbrock, Leah Findlater, Darren Gergle, and James J

    Jacob O. Wobbrock, Leah Findlater, Darren Gergle, and James J. Higgins. 2011. The aligned rank transform for nonparametric factorial analyses using only anova pro- cedures. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Vancouver, BC, Canada) (C...

  41. [51]

    Ann Yuan, Andy Coenen, Emily Reif, and Daphne Ippolito. 2022. Wordcraft: Story Writing With Large Language Models. In Proceedings of the 27th International Conference on Intelligent User Interfaces (Helsinki, Finland) (IUI ’22). Association for Computing Machinery, New York, N...

  42. [52]

    existing_text

    Jian Zheng and Ge Gao. 2024. Fragmented Moments, Balanced Choices: How Do People Make Use of Their Waiting Time?. InProceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Association for Computing Machinery, New York, NY, USA, Ar...

  43. [2024]

    One-Size-Fits-All

    “One-Size-Fits-All”? Examining Expectations around What Constitute “Fair” or “Good” NLG System Behaviors. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), ...

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.