{"id":"e6f551ca-dd82-41d0-b01b-3b475f590804","arxiv_id":"2505.14370","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A study with 18 employees found that a generative AI Meeting Purpose Assistant can help people clarify meeting goals, anticipate challenges, and change how they prepare, with social and technical barriers to adoption.","lead":"The paper reports a study of a generative AI chat tool that helps workers reflect on the purpose, challenges, and success of upcoming meetings. It suggests that such \"prospective reflection\" can make people more intentional and better prepared, though adoption faces social and technical barriers.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Observed impacts are not cleanly attributable to the MPA because the researcher co-produced the reflective conversation; §6.4.2's own caveat leaves the AI-specific contribution unestablished.","rationale":"The reader's verdict is CONDITIONAL with high confidence, and I agree with the weakest-assumption identification: attribution of the observed impacts to the AI system is the most load-bearing unresolved point. The paper's own limitation statement (6.4.2) concedes that the parallel conversations 'risk the conflation of human and AI-driven reflection,' and §4.3.2 describes the researcher as actively eliciting reasoning and 'surfacing a wider range of topics than the MPA alone.' That admission directly undermines the central inference that the MPA's generative, adaptive questioning caused the clarification, prioritization, perspective shifts, and action plans reported in Findings. The design is a technology probe, so some conflation is acceptable for discovery, but the paper's framing and RQ2 ask about the 'impacts of GenAI-assisted prospective reflection,' which requires the AI to be the active ingredient. No current analysis in the paper separates prompt source from impact; quotes are presented as illustrations without attribution to speaker (MPA vs researcher). A concrete re-coding of the existing transcripts can settle this without new data collection. If the re-coding shows most impacts follow researcher questions, the central claim should be softened to 'researcher-guided reflection with an AI transcript tool,' and future field deployments without a human guide are necessary. If the re-coding shows MPA questions dominate, the current CONDITIONAL verdict could be strengthened. Given that the paper already frames itself as formative and lists this as a limitation, my read does not move the verdict; it stays CONDITIONAL, but the condition is precisely this attribution test. I find no internal inconsistency or overclaim beyond the attribution issue; the rich quotes and transparent analysis are genuine strengths, and the authors deserve credit for surfacing the limitation rather than hiding it.","tokens_in":48628,"tokens_out":4153,"duration_ms":40893,"concrete_test":"Re-code all 18 session transcripts and follow-up surveys: for each statement coded as an impact (e.g., §5.2.1 'making purpose explicit', §5.3.1 'sharing meeting intentions'), identify whether the immediately preceding question or prompt came from the MPA or from the researcher. Report the distribution and a representative trace. If the majority of impact-adjacent prompts are researcher-initiated, the findings cannot be attributed to the MPA; if they are MPA-initiated, the AI-specific role is substantially supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that generative AI (the MPA) catalyzes prospective reflection and meeting intentionality. The evidence for this is the observed impacts in §5.2–§5.3 (clarifying purpose, prioritization, perspective change, preparation, changed meeting plans). However, the participatory prompting methodology means every session had a researcher actively guiding the participant, answering questions, and 'surfacing a wider range of topics than the MPA alone' (§4.3.2). The analysis does not distinguish which prompts—MPA or researcher—preceded each reported impact, so the observed reflection and its outcomes could be largely the product of the human guide, with the MPA serving as a structured transcription surface. The authors explicitly note the 'risk [of] conflation of human and AI-driven reflection' (§6.4.2), and demand characteristics from the researcher's presence likely amplify positive self-report in the follow-up survey. Because the paper's stated contribution is specifically 'GenAI-assisted prospective reflection' (RQ2), not 'human-guided reflection with a chat tool,' the AI-specific attribution is load-bearing and currently unsupported. This is a validity threat, not an internal contradiction; the study remains a promising exploratory probe, but the central claim is conditional on future isolation of the AI's role.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that meeting technologies lack support for prospective reflection (thinking about why a meeting is needed and what might happen) and introduces a generative-AI technology probe, the Meeting Purpose Assistant (MPA), to coach users in articulating meeting purpose, success conditions, and challenges. In a participatory prompting study, 18 employees of a global technology company used the MPA on three real upcoming meetings each, with a researcher present to guide the interaction. Thematic analysis of session transcripts and follow-up surveys identified four groups of findings: the process of reflective interaction (including personalization and AI-generated reflection summaries); impacts on thinking (making purpose explicit, prioritization, reflecting on unknowns, perspective change, preparation); impacts on the meeting itself (sharing intentions, accountability, efficiency, changing meeting series); and barriers plus timing considerations. The paper concludes with design and workflow implications for AI-assisted prospective reflection.","tokens_in":48889,"tokens_out":3718,"duration_ms":42099,"significance":"If the central claim holds, this is a valuable exploratory contribution to meeting science and HCI: it opens a new design space for pre-meeting, GenAI-mediated reflection that goes beyond scheduling and in-meeting support, and it provides a richly documented set of participant responses, barriers, and design considerations. The paper is transparent about its method and limitations, includes a follow-up survey to probe whether effects outlasted the session, and provides full protocol and meta-prompt appendices, which supports reproducibility of the probe even if not of the specific model outputs. The study is appropriately framed as a technology probe rather than a controlled efficacy trial. However, the central claim that the observed impacts are attributable to the AI system is not yet established because the researcher co-produced every reflective session; the paper's own limitation statement (§6.4.2) acknowledges this risk.","major_comments":[{"comment":"The central claim of the paper is that GenAI-assisted prospective reflection produced the observed impacts, but the study design embeds a researcher in every session. As described in §4.3.2, the researcher \"could build a better understanding of the individual meeting, whilst surfacing a wider range of topics than the MPA alone,\" and the analysis does not code which prompts—MPA messages versus researcher utterances—preceded each reported impact. The authors themselves state in §6.4.2 that \"the parallel conversations between the participant and AI, and the participant and researcher, risk the conflation of human and AI-driven reflection.\" Because RQ2 is specifically about \"GenAI-assisted prospective reflection,\" this attribution gap is load-bearing. I recommend adding an attribution analysis: code each impact or reflective episode to the immediately preceding prompt source (MPA message, researcher question, participant-initiated turn) and report the distribution, or explicitly re-scope the claims to \"reflection with an AI probe in a researcher-mediated session.\" Without this, the reported impacts could largely reflect the human guide rather than the AI.","section":"§5.2–§5.3 and RQ2"},{"comment":"The follow-up survey contains a leading question: \"How did your interaction with the Meeting Purpose Assistant during the study influence the effectiveness of the meeting, if at all?\" This presupposes an influence and, combined with the researcher's active presence during the session, creates a clear demand-characteristic pathway to the positive self-reports quoted in §5.3, such as the claim that the interaction helped meetings stay focused and effective. Please acknowledge this wording in the limitations section and soften the strength of the causal language in the findings (e.g., in §5.3.3) to \"participants perceived, when asked, that the interaction influenced effectiveness,\" or provide neutral follow-up questions in any future iteration.","section":"§4.3.3 / Appendix C.3"},{"comment":"The finding \"Making Purpose Explicit\" is partly realized by construction: the MPA's meta-prompts instruct it to ask users to articulate purpose, success conditions, and challenges, and to probe until these are expressed. Observing that participants did articulate these elements is therefore to some extent a check that the probe followed its prompt, not an emergent effect. The more informative evidence for RQ2 is in the action-level impacts in §5.3 (changed plans, shared summaries, altered communication) and the follow-up survey reports. I suggest foregrounding those concrete action and follow-up findings in the abstract and in the answer to RQ2, and re-presenting §5.2.1 as evidence that the interaction elicited the designed reflection rather than as an independent impact.","section":"§5.2.1 and Appendix E"}],"minor_comments":[{"comment":"There is a typo in \"an obvious confabulation by the MPA, wbhich suggested\"—\"wbhich\" should be \"which.\"","section":"§5.4.1"},{"comment":"The transcript excerpts repeatedly contain the unexplained token \"TYPES:\" (e.g., §5.2.2, §5.3.1, §5.4.3). This looks like a speech-to-text artifact from the researcher/participant talk. Please explain this notation in a footnote or remove it, as it currently confuses the reader about whether these are MPA-chat messages or spoken remarks.","section":"§5 and quotes throughout"},{"comment":"The sentence \"Users of may end up talking to and reflecting with AI more than they do other people\" is missing a noun after \"of\"; it should be \"Users of such systems may end up...\"","section":"§6.4.1"},{"comment":"The findings table and section headings refer to \"Impact of Reflection: Change in Thinking\" and \"Impact of Reflection: Changing the Meeting,\" but the discussion and abstract use \"observed impacts.\" Consider adding a qualifier such as \"perceived\" or \"reported\" in these headings and in the summary of findings to match the self-report nature of the evidence.","section":"Table 1 and §5"}],"recommendation":"major_revision","confidential_remarks":"This is a carefully reported exploratory probe, and the authors are commendably explicit about their study's limitations. That transparency makes the central attribution problem easy to see: because the researcher co-produced every reflective session, the paper's headline claim about GenAI-assisted reflection is not yet supported. The good news is that fixable evidence exists in the transcripts—an attribution analysis of which prompt source preceded each reported impact would substantially strengthen the paper. The leading follow-up survey question should also be addressed. With those changes, this could be a solid contribution to CHIWORK. I do not see novelty or disclosure concerns; the paper fits the venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is the first study I know of that uses an adaptive GenAI conversation to get people to reflect on the purpose, success conditions, and challenges of upcoming meetings. That is a genuinely new application, and the empirical contribution is real: 18 employees, 54 meetings, rich quotes, and a follow-up survey after the meetings. The paper is well written, transparent about method, and the qualitative analysis looks careful. The design implications in §6.3 are useful and grounded in data. Credit where due.\n\nThe soft spot is exactly where the reader and stress test put it. The participatory prompting methodology means a researcher sat in every session, prompted thinking aloud, explained the system, and even surfaced topics beyond the MPA. The paper itself says in §6.4.2 that parallel conversations risk 'conflation of human and AI-driven reflection.' That caveat is honest but it is load-bearing for RQ2. The observed impacts—clarifying purpose, prioritization, perspective change, preparation, changed plans—are not disaggregated by whether the MPA or the researcher elicited them. Demand characteristics from the researcher's presence probably also inflate the positive self-report in the follow-up survey. So the central claim that GenAI-assisted reflection caused these impacts is supported only as a promising hypothesis. The stress test is right that this is a validity threat, not an internal contradiction. The paper would be stronger if it framed RQ2 as 'impacts of a researcher-assisted MPA session' or compared MPA-only sessions.\n\nOther limitations are minor and mostly acknowledged: self-selected sample from one company, no baseline, no control condition. For a technology probe at an exploratory stage that is acceptable. The barriers findings (§5.4) are actually a nice counterweight—participants resisted specifics, confidentiality, solutions, and social goals—so the paper is not a pure cheerleader.\n\nVerdict: this deserves peer review. It opens a new design space, reports the study with enough detail to replicate, and its limitations are stated rather than buried. The right outcome is publication with the AI-attribution framing tightened. I'd bring it to a reading group and cite it in related work.","headline":"A new, well-reported probe of GenAI for pre-meeting reflection, but the observed impacts are co-produced with a researcher, so the AI-specific claim is real but still conditional.","tokens_in":49372,"tokens_out":1387,"would_cite":true,"duration_ms":24766,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Meetings keep failing because the technology around them never asks why they exist; this paper shows that a generative-AI conversation before a meeting, one that probes purpose, success conditions, and challenges, can make intentions…","keywords":["meetings","meeting intentionality","prospective reflection","generative AI","conversational user interface","technology probe","participatory prompting","workplace reflection"],"falsifier":"Run the same protocol in three conditions — the full MPA, a static questionnaire asking the identical purpose, success, and challenge questions without adaptive follow-up, and no tool at all — and compare preparation actions, meeting changes, and post-meeting effectiveness ratings; if the static questionnaire matches the MPA's effects, adaptive AI probing is not the active ingredient. A second test compares MPA sessions with and without a researcher present: if reported impacts vanish when no researcher is in the room, demand characteristics or human guidance rather than the AI explain the effect.","tokens_in":48371,"feed_emoji":"📅","tokens_out":10358,"duration_ms":90909,"temperature":0.7,"pith_summary":"Most meeting technology optimizes the \"how\" of meetings: scheduling, coordination, in-meeting engagement, and post-meeting summaries. This paper argues the missing piece is the \"why\" — prospective reflection before a meeting, where participants think about why the meeting is needed, what success looks like, and what might go wrong. To test this, the authors built the Meeting Purpose Assistant (MPA), a generative-AI chatbot that walks users through exactly that reflection and then produces a structured Reflection Summary, and had 18 employees of a global technology company use it on three of their real upcoming meetings each. They report that the reflection clarified and prioritized purposes, shifted perspectives, improved preparation and communication, and in several cases led participants to change or even cancel meetings, with some effects confirmed after the meetings actually took place. The paper also catalogues why people resist such reflection — time cost, confidentiality concerns, the awkwardness of making social goals explicit, and the expectation that AI should give answers rather than ask questions.","feed_headline":"Pre-meeting AI reflection reshaped how 18 workers planned meetings","feed_subtitle":"A conversational assistant probing purpose, success, and challenges turned vague plans into concrete goals and actions.","key_machinery":"The carrying object is the Meeting Purpose Assistant (MPA), a technology probe built on GPT-4 Turbo behind an enterprise firewall, consisting of a chat interface and a one-click Reflection Summary pane. The MPA works through three linked components: a title assistant that extracts the meeting name, a purpose assistant that conducts the reflective conversation under a meta-prompt instructing it to use active listening, single open-ended questions, and a divergence-then-convergence arc (first exploring purpose and challenges broadly, then asking users to prioritize), and a summary assistant that distills the thread into bullet points under three fixed headings: why we are meeting, what success looks like, and what could prevent success. The Reflection Summary is the load-bearing artifact: it converts private reflection into a shareable work object, which is how reflection reaches the meeting itself through invitations, attendee preparation, and agenda building. The second piece of machinery is the participatory prompting methodology, in which a researcher sits beside each participant and guides their interaction with the AI, eliciting reasoning and resolving technical confusion, at the cost of confounding human and AI contributions to the reflection.","core_discovery":"Meeting intentionality — a clear, articulated sense of why a meeting is happening and what success looks like — is largely unsupported by current meeting technology, and the paper argues this absence cascades into inefficiency, fatigue, and derailed discussions. The central claim is that a generative-AI conversational assistant can fill this gap by coaching users through prospective reflection: asking open questions about a meeting's purpose, probing for challenges and uncertainties, prompting prioritization, and finally distilling the conversation into a structured Reflection Summary that can be shared with attendees or pasted into a meeting invitation. With 18 employees reflecting on three real upcoming meetings each, the paper reports impacts in thinking — implicit goals made explicit, priorities clarified, unknown variables surfaced, other attendees' perspectives taken, anxiety reduced — and in action: participants shared agendas ahead of time, asked attendees to prepare, communicated potential challenges tactfully, changed recurring meeting formats, and sometimes concluded that a meeting should be canceled or that their own attendance was unnecessary. Follow-up surveys after the meetings had occurred confirmed that several of these intended changes were realized in practice. The paper presents this as an early, exploratory demonstration that GenAI can act as a thought-provoking coach in workplace reflection, while also cataloguing the social, temporal, and technological barriers that any real deployment would have to navigate.","pith_inferences":["If the MPA's effects replicate without a researcher present, the strongest product form suggested by this study is not a standalone chat but reflection embedded in the calendar: context-aware prompts at invitation time, with private provocative questioning and a shareable summary offered only when the user opts to go public; the authors gesture at this direction but do not test it.","The therapy- and coaching-like responses, including anxiety reduction and feeling validated, hint that part of the measured benefit may be emotional regulation rather than planning quality; a controlled study measuring stress and self-efficacy alongside plan changes could separate these two channels, which the paper does not do.","The confidentiality resistance observed in one-on-one and sensitive meetings predicts that privacy-preserving local or on-device reflection, or explicit guarantees about who can access reflective input, is a precondition for adoption in exactly the meetings where reflection is most needed; this is a testable deployment hypothesis the paper leaves open.","The finding that attendees with little control over a meeting found reflection less useful suggests the highest-leverage target for such tools is the meeting organizer, or any role with agency over format and agenda; a deployment study randomizing organizer-focused versus attendee-focused prompts would test this directly."],"forward_implications":["If prospective reflection works as described, meeting tools should add a pre-meeting reflection step, with the natural implementation point being the invitation: a prompt to articulate purpose and success when scheduling, and a nudge to attendees for important or uncertain meetings.","Reflection Summaries are the mechanism that makes reflection actionable; participants pasted them into invitations and chat threads, and several reported in follow-up surveys that sharing them improved attendance, engagement, and meeting focus.","Reflection can change whether a meeting happens at all: participants concluded that some meetings were unnecessary, that their own presence was dispensable, or that a recurring series needed format changes such as attendance policies or timed agendas.","The optimal timing of reflection tracks the meeting's role and routine: soon after scheduling for organizers, close to the meeting for attendees and for instances of recurring meetings, and periodically at the series level to renew intentionality for the series as a whole.","GenAI's value in this context is as a persistent question-asker rather than an answer-giver; participants who expected instant solutions were frustrated, while those who engaged with the questions reported the lasting effects on clarity, preparation, and plans."],"supporting_citations":[{"why":"Supplies the framing of meeting intentionality and the gap in meeting technology that this paper sets out to address.","marker":"[139]"},{"why":"Provides the participatory prompting methodology that the study adopts as its core protocol.","marker":"[134]"},{"why":"Provides the technology probe methodology that justifies the exploratory, non-validated probe design.","marker":"[68]"},{"why":"Demonstrates qualitative interviewing with generative AI, the direct template for the MPA's conversational design.","marker":"[36]"},{"why":"Supplies design resources for reflection technologies, used to interpret the MPA's reframing, provocation, and question-asking behaviors.","marker":"[17]"},{"why":"Shows that reflective goal-setting in the workplace produces perceived behavioral change, the nearest prior demonstration this work extends to meetings.","marker":"[94]"},{"why":"Provides evidence that contingency planning sustains motivation and performance, motivating the MPA's probing for challenges and unknowns.","marker":"[109]"},{"why":"Provides the premortem strategy that the MPA's challenge-elicitation mimics.","marker":"[75]"},{"why":"Represents prior retrospective in-meeting reflection support, the contrast that situates this paper's prospective, purpose-focused contribution.","marker":"[132]"}],"fun_headline_variants":["AI coach makes workers rethink why they meet","Before the meeting, AI asks 'why' and changes plans","Prospective reflection with AI shifts meeting outcomes","Meeting Purpose Assistant: AI that asks the tough questions","AI-assisted pre-meeting reflection makes plans concrete"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the observed reflection benefits came from the AI's questioning rather than from the researcher who sat alongside each participant guiding the session and eliciting reasoning; if the human guide did the work, the AI's specific contribution is overstated, a conflation risk the authors themselves acknowledge.","fun_headline_variants_meta":{"raw":{"variants":["AI coach makes workers rethink why they meet","Before the meeting, AI asks 'why' and changes plans","Prospective reflection with AI shifts meeting outcomes","Meeting Purpose Assistant: AI that asks the tough questions","AI-assisted pre-meeting reflection makes plans concrete"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000974,"raw_usage":{"total_tokens":4149,"prompt_tokens":961,"completion_tokens":3188,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":577,"completion_tokens_details":{"reasoning_tokens":3125}},"tokens_in":577,"tokens_out":3188,"duration_ms":25251,"temperature":1.0,"reasoning_tokens":3125,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T15:34:23.488055+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same protocol in three conditions — the full MPA, a static questionnaire asking the identical purpose, success, and challenge questions without adaptive follow-up, and no tool at all — and compare preparation actions, meeting changes, and post-meeting effectiveness ratings; if the static questionnaire matches the MPA's effects, adaptive AI probing is not the active ingredient. A second test compares MPA sessions with and without a researcher present: if reported impacts vanish when no researcher is in the room, demand characteristics or human guidance rather than the AI explain the effect.","supporting_citations":[{"cited_title":"Participatory prompting: a user-centric research method for eliciting AI assistance opportunities in knowledge workflows","cited_arxiv_id":"2312.16633","evidence_quote":"Provides the participatory prompting methodology that the study adopts as its core protocol."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the premortem strategy that the MPA's challenge-elicitation mimics."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Represents prior retrospective in-meeting reflection support, the contrast that situates this paper's prospective, purpose-focused contribution."}],"review_version":1}