Pith. sign in

REVIEW 3 major objections 6 minor 84 references

Understanding and Supporting Formal Email Exchange by Answering AI-Generated Questions

T0 review · 3 major / 6 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read Replying to formal email by answering short AI-generated questions is faster and less mentally demanding than writing prompts, with equal or better quality — at the price of a diminished sense of agency.

desk verdict A useful QA-based email drafting pattern, but the headline efficiency and workload results are undercut by a metric that counts AI output, a direct H1-d contradiction, and an order effect. read the letter →

arxiv 2502.03804 v2 pith:FIJB5GZX submitted 2025-02-06 cs.HC cs.AI

classification cs.HCcs.AI
keywords AI-MediatedCommunicationLargeLanguageModelsEmailQA-basedapproachquestiongenerationcognitiveloadsenseofagencyformalexchange
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper proposes that the best way to help people reply to formal emails is not to have them write prompts for an AI but to have the AI ask them questions. Its prototype system, ResQ, reads an incoming email, generates a short set of multiple-choice questions covering every request in it, and drafts the reply from the user's answers. Across a controlled experiment (12 participants) and a five-day field study (8 participants), the authors report that this question-answering approach outperforms both writing replies by hand and writing AI prompts: replies took less effort to produce, users felt less cognitive load, and independent evaluators rated the emails as high quality. The cost is that users feel less authorship and control over what they send, and some feel more psychological distance from the person they are writing to. The paper's contribution is to show that structuring the human's job as answering questions instead of engineering prompts reshapes the workload of AI-mediated communication.

What carries the argument

The mechanism that carries the argument is ResQ's question-generation step, driven by a structured prompt that instructs the LLM to act as a secretary: extract every request from the incoming email and emit the minimal sufficient set of multiple-choice questions in a structured format, each tied to a verbatim quoted passage of the email that the interface highlights when the user clicks the question. This converts the open-ended act of composing a reply into a series of small decisions, and it is what replaces prompt construction, what lowers comprehension load by pointing at the relevant parts of the sender's message, and what the authors credit with lowering the barrier to starting the task. The same scaffolding is hypothesized to produce the side effects: because the AI does more of the formulation, the user feels less like an author and more like an editor.

What would settle it

Re-run Study 1 with keystroke-level logging that separates characters typed or edited by the user from characters produced by the AI, or hold final message length equal across conditions; if the QA condition's per-second character advantage disappears under either measurement, the central efficiency claim is an artifact of AI output length. A second check is internal to the paper: Section 6.1.6 reports significantly higher self-reported barriers in the QA condition while also concluding the hypothesis was supported, so a corrected re-analysis of that item would settle whether initiation barriers actually dropped.

Watch

Extended reading notes

Core claim

The paper's central claim is that an LLM-powered question-and-answer workflow is a better way for a person to steer an email-drafting AI than writing a prompt is. ResQ, the prototype, reads the incoming message and generates a minimal set of multiple-choice questions covering every demand in the sender's email, highlighting the relevant passage for each question; the user answers the questions, optionally adds their own options, and adjusts tone, style, and length, after which the LLM writes a draft that the user reviews, edits, and sends. The authors report that, compared with a conventional prompt-based approach, this process improves reply efficiency (final response characters per second), reduces cognitive load measured with a standard workload questionnaire, makes the sender's requests easier to understand, increases satisfaction, and lowers the perceived barrier to starting a reply, while independent evaluators rate the resulting emails as equal or better in politeness, readability, and responsiveness. The same studies show a trade-off: users' sense of agency and control over the text decreases, and some users feel more psychological distance from their counterpart, even as others find that faster, more polished exchanges bring them closer.

Load-bearing premise

The load-bearing premise is that measuring reply efficiency as the character count of the final message divided by the time taken to send it captures how efficiently the user worked, rather than reflecting the AI's output length; if the QA advantage comes mostly from AI-generated text, the headline efficiency gain does not describe the user's own productive work.

Editorial extensions

If this is right

  • Email clients could add a reply-by-questions mode that turns the unstructured task of drafting a response into a structured set of decisions, reducing both the effort to understand the sender and the need for prompt-writing skill.
  • Users who postpone replying to long or demanding messages may start responding sooner and more often, because the system absorbs the initial steps of reading and deciding what to say.
  • Because the approach measurably lowers users' sense of agency and control, designers will need to make the level of AI intervention adjustable — the number and type of questions or the granularity of suggestions — to help users keep a sense of authorship.
  • Organizations that prize complete, polite, and consistent replies in formal settings could adopt the approach without sacrificing perceived email quality, since independent evaluators rated QA-produced replies at least as highly as prompted ones on politeness, readability, and meeting demands.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same question scaffolding could transfer to other high-stakes writing tasks beyond email — formal requests, progress reports, or performance reviews — where the bottleneck is also turning intentions into well-structured prose; a testable extension would compare question-based against prompt-based drafting in those genres.
  • If the efficiency gain partly reflects the AI emitting longer text, the practical benefit to the user is genuine but smaller than the per-second character metric suggests, because part of the measured efficiency is produced by the model rather than by the user's workflow.
  • The reduced sense of agency and the reduced barrier to task initiation may be two faces of the same mechanism: the AI that digests the email and asks the questions is also the AI that decides what the exchange is about, so preserving user control will require giving users a say in which questions get asked.
  • The field study hints at a behavioral spiral — faster, more polished replies draw faster responses and additional assignments — so longitudinal measures of task volume, reply latency, and relationship closeness would show whether closer-feeling relationships or workload growth dominate over time.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper introduces ResQ, an LLM-based question-and-answer (QA) system for replying to formal emails: instead of writing an open-ended prompt, users answer AI-generated multiple-choice questions and then edit the resulting draft. The authors report a within-subject controlled experiment (N=12) comparing QA, prompt-based AI, and no-AI conditions, plus a five-day field study (N=8 usable participants after excluding one informal-context user). The headline claims are that the QA condition improves efficiency, reduces workload, and maintains email quality relative to the prompt-based condition, while decreasing users' sense of agency and control. The paper also reports mixed effects on psychological distance and impressions between senders and recipients.

Significance. If the headline results held, the contribution would be a useful design pattern for AI-mediated communication: structured question answering can lower the input cost of email drafting compared to prompt crafting, with a documented trade-off in agency. The paper has several strengths: it uses realistic email scenarios, a within-subject design with Latin-square counterbalancing, an independent evaluator panel for email quality, effect sizes for key comparisons, a field deployment, and an open-source prototype. These features make the empirical core unusually complete for an HCI systems paper. However, the central efficiency and workload claims rest on a metric that conflates AI-generated text with user effort and on a workload measure with a significant order effect, and one quantitative result is reported in the opposite direction of its conclusion. The comparative advantage over the prompt-based condition is therefore not yet established.

major comments (3)
  1. [§5.6.1, §6.1.1] The efficiency metric used to support H1-a is defined as the character count of the final reply divided by task completion time (§5.6.1). Because the final reply in both AI conditions is largely generated by the LLM, this numerator credits AI-authored text to the user; no keystroke- or edit-level data are reported to separate user-authored from AI-authored characters. The QA vs. Prompt difference (d=0.65) could therefore reflect longer AI-generated drafts in the QA condition rather than a genuine improvement in the user's productive effort, and the paper does not report draft length by condition to rule this out. This concern is reinforced by §6.2.1 and Table 2, where QA and Prompt drafts are not shown to differ in quality, so longer output would not be rewarded by evaluators. I ask the authors to re-analyze with separate measures of completion time, user-authored characters, and draft length, or to reframe the claim as system-mediated throughput rather than user typing efficiency.
  2. [§6.1.6] Section 6.1.6 contains a direct contradiction: the post-hoc test shows that participants in the QA condition perceived significantly higher barriers to initiating email responses than both No-AI and Prompt conditions, yet the next sentence states 'Therefore, H1-d was supported' and concludes that the QA approach reduced difficulty. This same pattern is repeated in Table 5 and contradicts the qualitative quotes in §6.4.1. H1-d should be labeled as not supported (or the statistics corrected if the direction was misreported), and the RQ1 claims about task initiation should be revised accordingly.
  3. [Appendix B, Table 6] Table 6 reports a significant order effect for Raw TLX (p<0.05; the text gives p=0.041), and Raw TLX is the main dependent variable for H1-b. Because the order in which participants experienced conditions influenced the workload measure, the H1-b support should be interpreted as tentative; the discussion in §9.1.1 presents the workload reduction as established. Please either report an order-corrected analysis or explicitly qualify the workload conclusion.
minor comments (6)
  1. [Figure 5] The middle panel in Figure 5 is labeled 'H1-b Prompt Character Count [chars]' even though prompt character count is used to test H1-a (see §5.6.2); relabel the panel.
  2. [§5.3, §5.4] Section 5.3 says the experiment lasted approximately two hours, while §5.4 says approximately two and a half hours; make these descriptions consistent.
  3. [§6.2.1] The QA vs. Prompt post-hoc comparison for perceived email quality is not reported; since the hypothesis concerns the QA approach relative to the prompt baseline, report that pairwise result with its effect size or state explicitly that it was not significant.
  4. [Appendix B] Table 6 reports the Raw TLX order effect as 'p < 0.05' while the text gives p=0.041; unify the notation.
  5. [§6.4.1] The same P5 quote ('By saving the time needed to read the counterpart's text...') appears twice in Section 6.4.1; remove the duplicate.
  6. [§8.3.1] The P3 quote 'I've been assigned more tasks than before' appears to describe a workload consequence rather than a relational benefit; clarify how it supports the claim of positive self-presentation.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: empirical comparison with disclosed prompts, no fitted inputs, and no load-bearing self-citations.

full rationale

This is an empirical HCI systems evaluation rather than a derivation from theoretical first principles. The prototype ResQ is implemented with GPT-4o and its LLM prompts are fully disclosed in the appendix; Study 1 compares QA-based, prompt-based, and no-AI conditions on eighteen new email scenarios, with email quality assessed by independent third-party evaluators. No parameter is fitted to the outcome data, no prediction is evaluated on the same data from which it was derived, and no load-bearing claim rests on a self-citation: the reference list contains no prior work by this author team. The efficiency metric in §5.6.1 (final response character count divided by task completion time) does include AI-generated text, so the observed QA advantage may partly reflect longer drafts rather than user typing effort; however, this is a construct-validity limitation, not circularity, because the result is not true by construction and depends on measured completion times and on the drafts participants actually accepted. The paper's own appendix B reports a significant order effect for Raw TLX (p = 0.041) and an order-by-condition interaction for IOS (p = 0.043), which weakens some secondary workload and psychological-distance claims but does not make them circular. Therefore no circular step is identifiable.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper introduces no fitted numerical parameters; the central quantitative claims rest on the validity of self-report measures, the efficiency metric, and the LLM's question-generation reliability.

assumptions (4)
  • domain assumption Total final response character count divided by completion time is a valid measure of user efficiency in email replying.
    Section 5.6.1 defines efficiency this way; the validity of this proxy is load-bearing for the H1-a result.
  • domain assumption The 18 email scenarios used in Study 1 are representative of formal email exchanges that require detailed and polite replies.
    Section 5.1 describes scenario construction from real emails provided by 10 volunteers; representativeness affects generalizability of the efficiency and quality findings.
  • standard math Standard repeated-measures ANOVA and Friedman test assumptions hold for N=12 with the applied corrections.
    Sections 6.1 and 6.2; the paper reports Shapiro-Wilk tests but does not report sphericity for all tests, and Appendix B shows a significant order effect for Raw TLX.
  • domain assumption GPT-4o reliably generates questions that cover the sender's requirements without omission when given the structured prompt in Appendix A.1.
    Section 4 and Appendix A.1; the entire QA approach depends on the LLM question generation being comprehensive and accurate.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Understanding and Supporting Formal Email Exchange by Answering AI-Generated Questions." pith.science (2026). https://pith.science/paper/FIJB5GZX

@misc{pith2026250203804,
  author       = {Pith},
  title        = {Pith review of: Understanding and Supporting Formal Email Exchange by Answering AI-Generated Questions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FIJB5GZX}},
  note         = {Machine review of arXiv:2502.03804}
}
read the original abstract

Replying to formal emails is time-consuming and cognitively demanding, as it requires crafting polite phrasing and providing an adequate response to the sender's demands. Although systems with Large Language Models (LLMs) were designed to simplify the email replying process, users still need to provide detailed prompts to obtain the expected output. Therefore, we proposed and evaluated an LLM-powered question-and-answer (QA)-based approach for users to reply to emails by answering a set of simple and short questions generated from the incoming email. We developed a prototype system, ResQ, and conducted controlled and field experiments with 12 and 8 participants. Our results demonstrated that the QA-based approach improves the efficiency of replying to emails and reduces workload while maintaining email quality, compared to a conventional prompt-based approach that requires users to craft appropriate prompts to obtain email drafts. We discuss how the QA-based approach influences the email reply process and interpersonal relationship dynamics, as well as the opportunities and challenges associated with using a QA-based approach in AI-mediated communication.

Figures

Figures reproduced from arXiv: 2502.03804 by the authors.

Figure 1
Figure 1. In our system, (1) users receive an email, (2) communicate their intentions by answering AI-generated questions, (3) [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. The overview of the process of creating a reply message using ResQ. A) The LLM first generates multiple-choice [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Interface of ResQ. On the left, the content of the email is displayed, with an editor and a “Reply” button below for [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Inclusion of Other in the Self (IOS). The diagram above the x-axis is an example of what participants were shown [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]
Figure 5
Figure 5. Figure 5: Results of participants’ efficiency and cognitive load of replying to emails. Left: Efficiency for replying to emails. [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 6
Figure 6. Figure 6: Summary of Likert scale responses. Measurements H2 and H3-a were assessed by third-party evaluators rather than the participants themselves. The significant differences between conditions were from post-hoc analysis after one-way repeated measure ANOVA or the Friedman …
Figure 7
Figure 7. Figure 7: Participants’ future preferences in Study 1. Significant differences between conditions were identified through post-hoc [PITH_FULL_IMAGE:figures/full_fig_p020_7.png]
Figure 8
Figure 8. Figure 8: UI of the Gmail Reply Box with the “Reply with AI” Feature, used in Study 2. Pressing the “Reply with AI” button [PITH_FULL_IMAGE:figures/full_fig_p020_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

84 extracted references · 58 canonical work pages

  1. [1]

    Abdulkareem Al-Alwani. 2014. A novel email response algorithm for email management systems. Journal of Computer Science 10, 4 (2014), 689. https://www.researchgate.net/profile/Abdulkareem-Al- Alwani/publication/288108641_A_novel_email_response_algorithm_for_ email_management_systems/links/5702ae1508aedbac126f3c90/A-novel-email- response-algorithm-for-emai...

  2. [2]

    Anthropic. 2024. Claude: Next-Generation AI Assistant. Retrieved in July 10, 2024 from https://www.anthropic.com/claude

  3. [3]

    Arnold, Krysta Chauncey, and Krzysztof Z

    Kenneth C. Arnold, Krysta Chauncey, and Krzysztof Z. Gajos. 2020. Predictive text encourages predictable writing. InProceedings of the 25th International Conference on Intelligent User Interfaces (Cagliari, Italy) (IUI ’20). ACM, New York, NY, USA, 128–138. https://doi.org/10.1145/3377325.3377523

  4. [4]

    Aron, and Danny Smollan

    Arthur Aron, Elaine N. Aron, and Danny Smollan. 1992. Inclusion of other in the self scale and the structure of interpersonal closeness. Journal of personality and social psychology 63, 4 (1992), 596. https://psycnet.apa.org/journals/psp/63/4/596/ Publisher: American Psychological Association

  5. [5]

    Albert Bandura. 2001. Social Cognitive Theory: An Agentic Perspective. Annual Review of Psychology 52, 1 (Feb. 2001), 1–26. https://doi.org/10.1146/annurev. psych.52.1.1

  6. [6]

    Ashish Bastola, Hao Wang, Judsen Hembree, Pooja Yadav, Zihao Gong, Emma Dixon, Abolfazl Razi, and Nathan McNeese. 2024. LLM-based Smart Reply (LSR): Enhancing Collaborative Performance with ChatGPT-mediated Smart Reply System. https://doi.org/10.48550/arXiv.2306.11980 arXiv:2306.11980 [cs]

  7. [7]

    Victoria Bellotti, Nicolas Ducheneaut, Mark Howard, Ian Smith, and Rebecca E. Grinter. 2005. Quality versus quantity: E-mail-centric task management and its relation with overload. Human–Computer Interaction 20, 1-2 (2005), 89–138. https://www.tandfonline.com/doi/abs/10.1080/07370024.2005.9667362 Publisher: Taylor & Francis

  8. [8]

    Daniel E. Berlyne. 1960. Conflict, arousal, and curiosity

Show all 84 references
  1. [9]

    Boomerang. 2024. Respondable: Write Better Emails with AI Assistance. Retrieved in July 10, 2024 from https://www.boomeranggmail.com/respondable

  2. [10]

    Petter Bae Brandtzaeg and Asbjørn Følstad. 2017. Why People Use Chatbots. In Internet Science, Ioannis Kompatsiaris, Jonathan Cave, Anna Satsiou, Georg Understanding and Supporting Formal Email Exchange by Answering AI-Generated Questions CHI ’25, April 26-May 1, 2025, Yokoham...

  3. [11]

    Virginia Braun and Victoria Clarke. 2006. Using thematic analysis in psychology. Qualitative Research in Psychology 3, 2 (Jan. 2006), 77–101. https://doi.org/10. 1191/1478088706qp063oa

  4. [12]

    Sondos Mahmoud Bsharat, Aidar Myrzakhan, and Zhiqiang Shen. 2024. Prin- cipled Instructions Are All You Need for Questioning LLaMA-1/2, GPT-3.5/4. arXiv:2312.16171 [cs.CL] https://arxiv.org/abs/2312.16171

  5. [13]

    Daniel Buschek, Martin Zürn, and Malin Eiband. 2021. The Impact of Multiple Parallel Phrase Suggestions on Email Input and Composition Behaviour of Native and Non-Native English Writers. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokohama,...

  6. [14]

    Byers, A

    James C. Byers, A. C. Bittner, and Susan G. Hill. 1989. Traditional and raw task load index (TLX) correlations: Are paired comparisons neces- sary. Advances in industrial ergonomics and safety 1 (1989), 481–485. https://books.google.com/books?hl=en&lr=&id=xuV4Bb7vsvkC&oi=fnd& ...

  7. [15]

    Yu, and Lichao Sun

    Yihan Cao, Siyu Li, Yixin Liu, Zhiling Yan, Yutong Dai, Philip S. Yu, and Lichao Sun

  8. [16]

    Yu, Qiang Yang, and Xing Xie

    Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, Wei Ye, Yue Zhang, Yi Chang, Philip S. Yu, Qiang Yang, and Xing Xie. 2024. A Survey on Evaluation of Large Language Models. ACM Transactions on Intelligent ...

  9. [17]

    Lee, Gagan Bansal, Yuan Cao, Shuyuan Zhang, Justin Lu, Jackie Tsay, Yinan Wang, Andrew M

    Mia Xu Chen, Benjamin N. Lee, Gagan Bansal, Yuan Cao, Shuyuan Zhang, Justin Lu, Jackie Tsay, Yinan Wang, Andrew M. Dai, Zhifeng Chen, Timothy Sohn, and Yonghui Wu. 2019. Gmail Smart Compose: Real-Time Assisted Writing. In Proceedings of the 25th ACM SIGKDD International Confer...

  10. [18]

    Yuan-shan Chen. 2015. Chinese learners’ cognitive processes in writing email requests to faculty. System 52 (2015), 51–62. https://www.sciencedirect.com/ science/article/pii/S0346251X15000743 Publisher: Elsevier

  11. [19]

    Valdemar Danry, Pat Pataranutaporn, Yaoli Mao, and Pattie Maes. 2023. Don’t Just Tell Me, Ask Me: AI Systems that Intelligently Frame Explanations as Questions Improve Human Logical Discernment Accuracy over Causal AI explanations. In Proceedings of the 2023 CHI Conference on ...

  12. [20]

    Dhillon, Somayeh Molaei, Jiaqi Li, Maximilian Golub, Shaochun Zheng, and Lionel Peter Robert

    Paramveer S. Dhillon, Somayeh Molaei, Jiaqi Li, Maximilian Golub, Shaochun Zheng, and Lionel Peter Robert. 2024. Shaping Human-AI Collaboration: Varied Scaffolding Levels in Co-writing with Language Models. InProceedings of the CHI Conference on Human Factors in Computing Syst...

  13. [21]

    Dotan Di Castro, Zohar Karnin, Liane Lewin-Eytan, and Yoelle Maarek. 2016. You’ve got Mail, and Here is What you Could do With It!: Analyzing and Predict- ing Actions on Email Messages. InProceedings of the Ninth ACM International Con- ference on Web Search and Data Mining (Sa...

  14. [22]

    Fiona Draxler, Anna Werner, Florian Lehmann, Matthias Hoppe, Albrecht Schmidt, Daniel Buschek, and Robin Welsch. 2024. The AI Ghostwriter Effect: When Users do not Perceive Ownership of AI-Generated Text but Self-Declare as Authors. ACM Transactions on Computer-Human Interacti...

  15. [23]

    Mark Dredze, Tova Brooks, Josh Carroll, Joshua Magarick, John Blitzer, and Fernando Pereira. 2008. Intelligent email: reply and attachment prediction. In Proceedings of the 13th international conference on Intelligent user interfaces (Gran Canaria, Spain) (IUI ’08). ACM, Gran ...

  16. [24]

    Casey Dugan, Aabhas Sharma, Michael Muller, Di Lu, Michael Brenndoerfer, and Werner Geyer. 2017. RemindMe: Plugging a Reminder Manager into Email for En- hancing Workplace Responsiveness. InHuman-Computer Interaction - INTERACT 2017, Regina Bernhaupt, Girish Dalvi, Anirudha Jo...

  17. [25]

    Mark Dunlop and John Levine. 2012. Multidimensional pareto optimization of touchscreen keyboards for speed, familiarity and improved spell checking. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Austin, Texas, USA) (CHI ’12). ACM, New York, NY,...

  18. [26]

    Bj Fogg. 2009. A behavior model for persuasive design. In Proceedings of the 4th International Conference on Persuasive Technology (Claremont, California, USA) (Persuasive ’09). ACM, New York, NY, USA, 1–7. https://doi.org/10.1145/1541948. 1541999

  19. [27]

    Andrew Fowler, Kurt Partridge, Ciprian Chelba, Xiaojun Bi, Tom Ouyang, and Shumin Zhai. 2015. Effects of Language Modeling and its Personalization on Touchscreen Typing Performance. InProceedings of the 33rd Annual ACM Confer- ence on Human Factors in Computing Systems (Seoul,...

  20. [28]

    Holmvall, and Laura E

    Lori Francis, Camilla M. Holmvall, and Laura E. O’brien. 2015. The influence of workload and civility of treatment on the perpetration of email incivility. Computers in Human Behavior 46 (2015), 191–201. https://www.sciencedirect. com/science/article/pii/S0747563214007675 Publ...

  21. [29]

    Liye Fu, Benjamin Newman, Maurice Jakesch, and Sarah Kreps. 2023. Compar- ing Sentence-Level Suggestions to Message-Level Suggestions in AI-Mediated Communication. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). ACM...

  22. [31]

    Morrow Von O

    Daniel G. Morrow Von O. Leirer Jill M . An. 1998. The Influence of List Format and Category Headers on Age Differences in Understanding Medication Instructions. Experimental Aging Research 24, 3 (June 1998), 231–256. https://doi.org/10.1080/ 036107398244238

  23. [32]

    Goodman, Erin Buehler, Patrick Clary, Andy Coenen, Aaron Donsbach, Tiffanie N

    Steven M. Goodman, Erin Buehler, Patrick Clary, Andy Coenen, Aaron Donsbach, Tiffanie N. Horne, Michal Lahav, Robert MacDonald, Rain Breaw Michaels, Ajit Narayanan, Mahima Pushkarna, Joel Riley, Alex Santana, Lei Shi, Rachel Sweeney, Phil Weaver, Ann Yuan, and Meredith Ringel ...

  24. [33]

    Google. 2024. People + AI Guidebook

  25. [34]

    Grammarly. 2024. Grammarly: Free AI Writing Assistance. Retrieved in July 10, 2024 from https://www.grammarly.com

  26. [35]

    S. G. Hart. 1988. Development of NASA-TLX (Task Load Index): Results of empirical and theoretical research. Human mental workload/Elsevier (1988)

  27. [36]

    Jess Hohenstein and Malte Jung. 2018. AI-Supported Messaging: An Investigation of Human-Human Text Conversation with AI Support. In Extended Abstracts of the 2018 CHI Conference on Human Factors in Computing Systems (Montreal QC, Canada) (CHI EA ’18). ACM, New York, NY, USA, 1...

  28. [37]

    Kizilcec, Dominic DiFranzo, Zhila Aghajari, Hannah Mieczkowski, Karen Levy, Mor Naaman, Jeffrey Hancock, and Malte F

    Jess Hohenstein, Rene F. Kizilcec, Dominic DiFranzo, Zhila Aghajari, Hannah Mieczkowski, Karen Levy, Mor Naaman, Jeffrey Hancock, and Malte F. Jung

  29. [38]

    McKinsey Global Institute. 2012. The Social Economy: Unlocking Value and Productivity through Social Technologies. Retrieved in July 10, 2024 from https://www.mckinsey.com/industries/technology-media-and- telecommunications/our-insights/the-social-economy

  30. [39]

    Scientific Reports 13, 1 (2023), 5487

    Artificial intelligence in communication impacts language and social relationships. Scientific Reports 13, 1 (2023), 5487. https://www.nature.com/ articles/s41598-023-30938-9 Publisher: Nature Publishing Group UK London

  31. [40]

    Dietmar Jannach, Ahtsham Manzoor, Wanling Cai, and Li Chen. 2022. A Survey on Conversational Recommender Systems. Comput. Surveys 54, 5 (June 2022), 1–36. https://doi.org/10.1145/3453154

  32. [41]

    Hancock, and Mor Naaman

    Maurice Jakesch, Megan French, Xiao Ma, Jeffrey T. Hancock, and Mor Naaman

  33. [42]

    Anjuli Kannan, Karol Kurach, Sujith Ravi, Tobias Kaufmann, Andrew Tomkins, Balint Miklos, Greg Corrado, Laszlo Lukacs, Marina Ganea, Peter Young, and Vivek Ramavajjala. 2016. Smart Reply: Automated Response Suggestion for Email. In Proceedings of the 22nd ACM SIGKDD Internatio...

  34. [43]

    Dae Hyun Kim, Hyungyu Shin, Shakhnozakhon Yadgarova, Jinho Son, Hariharan Subramonyam, and Juho Kim. 2024. AINeedsPlanner: A Workbook to Support Effective Collaboration Between AI Experts and Clients. In Designing Interactive Systems Conference (Copenhagen, Denmark) (DIS ’24)....

  35. [44]

    Kalman and Sheizaf Rafaeli

    Yoram M. Kalman and Sheizaf Rafaeli. 2011. Online Pauses and Silence: Chrone- mic Expectancy Violations in Written Computer-Mediated Communication. Communication Research 38, 1 (Feb. 2011), 54–69. https://doi.org/10.1177/ 0093650210378229

  36. [45]

    If the Machine Is As Good As Me, Then What Use Am I?

    Charlotte Kobiella, Yarhy Said Flores López, Franz Waltenberger, Fiona Draxler, and Albrecht Schmidt. 2024. "If the Machine Is As Good As Me, Then What Use Am I?" – How the Use of ChatGPT Changes Young Professionals’ Perception of Productivity and Accomplishment. In Proceeding...

  37. [46]

    Farshad Kooti, Luca Maria Aiello, Mihajlo Grbovic, Kristina Lerman, and Amin Mantrach. 2015. Evolution of Conversations in the Age of Email Overload. InPro- ceedings of the 24th International Conference on World Wide Web (Florence, Italy) (WWW ’15). International World Wide We...

  38. [47]

    AI enhances our performance, I have no doubt this one will do the same

    Agnes Mercedes Kloft, Robin Welsch, Thomas Kosch, and Steeven Villa. 2024. "AI enhances our performance, I have no doubt this one will do the same": The Placebo effect is robust to negative descriptions of AI. In Proceedings of the 2024 CHI Conference on Human Factors in Compu...

  39. [48]

    Sebastian Linxen, Christian Sturm, Florian Brühlmann, Vincent Cassau, Klaus Opwis, and Katharina Reinecke. 2021. How WEIRD is CHI?. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokohama, Japan) (CHI ’21). ACM, New York, NY, USA, 1–14. https:...

  40. [49]

    Yihe Liu, Anushk Mittal, Diyi Yang, and Amy Bruckman. 2022. Will AI Console Me when I Lose my Pet? Understanding Perceptions of AI-Mediated Email Writing. In CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA) (CHI ’22). ACM, New York, NY, USA, 1–13. ht...

  41. [50]

    Erin Chao Ling, Iis Tussyadiah, Aarni Tuomi, Jason Stienmetz, and Athina Ioan- nou. 2021. Factors influencing users’ adoption and use of conversational agents: A systematic review. Psychology & Marketing 38, 7 (July 2021), 1031–1051. https://doi.org/10.1002/mar.21491

  42. [51]

    Microsoft. 2024. Microsoft Copilot: AI-Powered Assistance for Productivity. Retrieved in July 10, 2024 from https://www.microsoft.com/en-us/microsoft- 365/copilot

  43. [52]

    Hannah Mieczkowski and Jeffrey Hancock. 2022. Examining agency, ex- pertise, and roles of AI systems in AI-mediated communication. OSF Preprints 15 (2022). https://files.osf.io/v1/resources/asnv4/providers/osfstorage/ 62d1e312f66a94273e230192?action=download&direct&version=1

  44. [53]

    Iqbal, Mary Czerwinski, Paul Johns, Akane Sano, and Yuliya Lutchyn

    Gloria Mark, Shamsi T. Iqbal, Mary Czerwinski, Paul Johns, Akane Sano, and Yuliya Lutchyn. 2016. Email Duration, Batching and Self-interruption: Patterns of Email Use on Productivity and Stress. InProceedings of the 2016 CHI Conference on Human Factors in Computing Systems (Sa...

  45. [54]

    Moore and Paul C

    James W. Moore and Paul C. Fletcher. 2012. Sense of agency in health and disease: a review of cue integration approaches. Consciousness and cognition 21, 1 (2012), 59–68. https://www.sciencedirect.com/science/article/pii/S1053810011002005 Publisher: Elsevier

  46. [55]

    Uma Parthavi Moravapalle and Raghupathy Sivakumar. 2017. DejaVu: A case for assisted email replies on smartphones. In 2017 IEEE 13th International Conference on Wireless and Mobile Computing, Networking and Communications (WiMob) . IEEE, 1–8. https://ieeexplore.ieee.org/abstra...

  47. [56]

    Pashutan Modaresi, Philipp Gross, Siavash Sefidrodi, Mirja Eckhof, and Stefan Conrad. 2017. On (Commercial) Benefits of Automatic Text Summarization Systems in the News Domain: A Case of Media Monitoring and Media Response Analysis. https://doi.org/10.48550/arXiv.1701.00728 ar...

  48. [57]

    Asif Naeem, I

    M. Asif Naeem, I. Wayan S. Linggawa, Aftab A. Mughal, Christof Lutteroth, and Gerald Weber. 2018. A smart email client prototype for effective reuse of past replies. IEEE Access 6 (2018), 69453–69471. https://ieeexplore.ieee.org/abstract/ document/8517098/ Publisher: IEEE

  49. [58]

    Nandhini and S

    K. Nandhini and S. R. Balasundaram. 2013. Use of Genetic Algorithm for Cohesive Summary Extraction to Assist Reading Difficulties. Applied Computational Intelli- gence and Soft Computing 2013 (2013), 1–11. https://doi.org/10.1155/2013/945623

  50. [59]

    Qianqian Mu, Marcel Borowski, Jens Emil Sloth Grønbæk, Susanne Bødker, and Eve Hoggan. 2024. Whispering Through Walls: Towards Inclusive Backchannel Communication in Hybrid Meetings. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA)...

  51. [60]

    OpenAI. 2024. ChatGPT (4o) [Large language model]. Retrieved in July 10, 2024 from https://chatgpt.com/

  52. [61]

    OpenAI. 2024. Hello GPT-4o | OpenAI. Retrieved in September 8, 2024 from https://openai.com/index/hello-gpt-4o/

  53. [62]

    Chi, and Gregorio Convertino

    Les Nelson, Rowan Nairn, Ed H. Chi, and Gregorio Convertino. 2011. Mail2tag: Augmenting email for sharing with implicit tag-based categorization. In 2011 International Conference on Collaboration Technologies and Systems (CTS) . IEEE, 23–30. https://ieeexplore.ieee.org/abstrac...

  54. [63]

    PL Patrick Rau, Ye Li, and Dingjun Li. 2009. Effects of communication style and culture on ability to accept recommendations from robots. Computers in Human Behavior 25, 2 (2009), 587–595. https://www.sciencedirect.com/science/article/ pii/S0747563208002367 Publisher: Elsevier

  55. [64]

    Sarah Resendes, Thammi Ramanan, Angela Park, Brad Petrisor, and Mohit Bhandari. 2012. Send it: study of e-mail etiquette and notions from doc- tors in training. Journal of surgical education 69, 3 (2012), 393–403. https: //www.sciencedirect.com/science/article/pii/S19317204110...

  56. [65]

    Philip Quinn and Shumin Zhai. 2016. A Cost-Benefit Study of Text Entry Sugges- tion Interaction. In Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems (San Jose, California, USA) (CHI ’16). ACM, New York, NY, USA, 83–88. https://doi.org/10.1145/285803...

  57. [66]

    Supraja Sankaran, Chao Zhang, Henk Aarts, and Panos Markopoulos. 2021. Exploring peoples’ perception of autonomy and reactance in everyday ai inter- actions. Frontiers in psychology 12 (2021), 713074. https://www.frontiersin.org/ articles/10.3389/fpsyg.2021.713074/full Publish...

  58. [67]

    Schouwenburg

    Henri C. Schouwenburg. 1992. Procrastinators and fear of failure: an exploration of reasons for procrastination. European Journal of Personality 6, 3 (Sept. 1992), 225–236. https://doi.org/10.1002/per.2410060305

  59. [68]

    I Can’t Reply with That

    Ronald E Robertson, Alexandra Olteanu, Fernando Diaz, Milad Shokouhi, and Peter Bailey. 2021. “I Can’t Reply with That”: Characterizing Problematic Email Reply Suggestions. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokohama, Japan) (CHI ’...

  60. [69]

    Stephens, Renee L

    Keri K. Stephens, Renee L. Cowan, and Marian L. Houser. 2011. Organiza- tional norm congruency and interpersonal familiarity in e-mail: Examining messages from two different status perspectives. Journal of Computer-Mediated Communication 16, 2 (2011), 228–249. https://academic...

  61. [70]

    Sansiri Tarnpradab, Fei Liu, and Kien A. Hua. 2017. Toward extractive summa- rization of online forum discussions via hierarchical attention networks. In The Thirtieth International Flairs Conference . https://cdn.aaai.org/ocs/15500/15500- 68662-1-PB.pdf

  62. [71]

    Shirren and J.G

    S. Shirren and J.G. Phillips. 2011. Decisional style, mood and work communication: email diaries. Ergonomics 54, 10 (Oct. 2011), 891–903. https://doi.org/10.1080/ 00140139.2011.609283

  63. [72]

    Vignovic and Lori Foster Thompson

    Jane A. Vignovic and Lori Foster Thompson. 2010. Computer-mediated cross- cultural collaboration: Attributing communication errors to the person versus the situation. Journal of Applied Psychology 95, 2 (2010), 265. https://psycnet. apa.org/journals/apl/95/2/265/ Publisher: Am...

  64. [73]

    Steve Whittaker, Victoria Bellotti, and Jacek Gwizdka. 2006. Email in personal information management. Commun. ACM 49, 1 (Jan. 2006), 68–73. https://doi. org/10.1145/1107458.1107494

  65. [74]

    Rajan Vaish and Andrés Monroy-Hernández. 2017. CrowdTone: Crowd-powered tone feedback and improvement system for emails. https://doi.org/10.48550/ arXiv.1701.01793 arXiv:1701.01793 [cs]

  66. [75]

    Chen Zhou, Zihan Yan, Ashwin Ram, Yue Gu, Yan Xiang, Can Liu, Yun Huang, Wei Tsang Ooi, and Shengdong Zhao. 2024. GlassMail: Towards Personalised Wearable Assistant for On-the-Go Email Creation on Smart Glasses. InDesigning Interactive Systems Conference (Copenhagen, Denmark) ...

  67. [77]

    Sungjoon (Steve) Won and Laura A. Dabbish. 2009. Designing for email response management. In CHI ’09 Extended Abstracts on Human Factors in Computing Systems (Boston, MA, USA) (CHI EA ’09). ACM, New York, NY, USA, 3661–3666. https://doi.org/10.1145/1520340.1520551

  68. [79]

    You must create questions with choices for your audience and output the results in JSON format

  69. [80]

    The questions must be created in the native language of your audience

  70. [81]

    If necessary , your audience can write any free answers to your questions , so you will be penalized if you create an " other " option

  71. [82]

    That is , output c o rr e s po n di n g _p a r t = In co mi ng Mai l_ HT ML [ x : x + h ]

    In ' corresponding_part ', you must quote a part of the provided ' Incoming Mail ' verbatim . That is , output c o rr e s po n di n g _p a r t = In co mi ng Mai l_ HT ML [ x : x + h ]

  72. [83]

    You must quote spaces , `<br > `, periods , and commas exactly as in the provided ' Incoming Mail '

  73. [84]

    You will be penalized if you edit or combine multiple parts of the ' Incoming Mail ' for your questions

  74. [85]

    questions

    You will be penalized if you create unhelpful questions to compose a reply . You must keep the number of questions to a minimum . I 'm going to tip $100 for a better solution ! Ensure that your output is unbiased and avoids relying on stereotypes . ### Output JSON Format ### {...

  75. [2019]

    In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow, Scotland Uk) (CHI ’19)

    AI-Mediated Communication: How the Perception that Profile Text was Written by AI Affects Trustworthiness. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow, Scotland Uk) (CHI ’19). ACM, New York, NY, USA, 1–13. https://doi.org/10.1145/32...

  76. [2023]

    https://doi.org/10.48550/arXiv.2303.04226 arXiv:2303.04226 [cs]

    A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT. https://doi.org/10.48550/arXiv.2303.04226 arXiv:2303.04226 [cs]

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.