{"id":"f2816f79-e0c6-47a6-be9b-07e89b150bc5","arxiv_id":"2505.01678","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An AI-powered speaking assistant for non-native speakers in live multilingual teams did not improve measured speaking competence but improved perceived logical flow and depth, while adding anxiety and workload.","lead":"This paper tests an AI assistant that generates speaking suggestions for non-native English speakers during live group discussions. In a 31-team experiment, the tool did not measurably improve speaking skill, but interviews showed it helped structure and deepen speech while raising workload and anxiety.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"ANOVA on 76 clustered queries may overstate RQ1; the central RQ2 claim is not affected.","rationale":"The reader's weakest_assumption exactly matches my concern: Section 4.1.5 ANOVA treats 76 queries as independent despite clustering under 31 participants. This is structural for the RQ1 effort claim. My independent reading of the manuscript confirms the concern is real and that no mixed-effects or cluster-robust analysis is reported. However, the central RQ2 claim — the null speaking-competence result and the qualitative benefits — is not undermined: the Wilcoxon test in Section 4.2 is participant-level (31 pairs), and the interviews are qualitative. The RQ3 workload/anxiety claims are also based on participant-level Wilcoxon tests and qualitative evidence, so they remain credible. Even if the ANOVA concern lands, the paper's central message about real-time AI content support having both benefits and interaction costs still stands, because the interaction-cost story is independently supported by 20 NNSs describing inputting information as challenging (Section 4.3.2) and by the honest reporting of null quantitative results. Thus the verdict should remain CONDITIONAL; the concern warrants a statistical re-analysis but not a change of the overall verdict.","tokens_in":26150,"tokens_out":1469,"duration_ms":12561,"concrete_test":"Re-analyze the 76 input queries with a linear mixed-effects model: fixed effect of input pattern (four levels), random intercept for NNS participant (31 participants), and the same Brown-Forsythe-appropriate variance structure if needed. Report the pattern effect on input duration and modification count with participant-level random effects. If the rationalizing-decisions contrast remains significant at p<0.05 with a random intercept, the RQ1 effort claim survives; if the effect shrinks below significance or the variance components show most variance is between participants, Section 5.1 recommendations need softening.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is two-pronged: AISA did not improve speaking competence (Wilcoxon z=0.171, p=0.875, Section 4.2) but qualitative reports indicate benefits in logical flow/depth, and multitasking raised workload/anxiety (Section 4.3). The weakest point is the Section 4.1.5 one-way ANOVA (Brown-Forsythe, F[3,42.268]=20.382, p<0.001) that treats each of 76 AISA input queries as independent observations. These queries are clustered within 31 NNS participants, with some participants contributing many queries (e.g., rationalizing decisions n=36 across 31 participants). If a few participants dominate the rationalizing-decisions category, the significant differences in input duration and modification count could be driven by participant-level traits like typing speed or verbosity rather than by the pattern itself. The paper reports no mixed-effects model, no cluster-robust standard errors, and no intra-class correlation analysis. This matters because the RQ1 claim that 'rationalizing decisions' imposes a significant burden is used to motivate design recommendations in Section 5.1. The RQ2 null result and the qualitative findings on workload/anxiety rest on interviews and Wilcoxon tests at the participant level, so they are not damaged by this clustering issue. The paper's own limitations are honestly reported, but this particular statistical structural issue is not flagged.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a mixed-methods study of an AI-based speaking assistant (AISA) that provides real-time speaking references to non-native speakers (NNSs) during multilingual team communication. In a within-subjects experiment, 31 teams consisting of two native speakers and one NNS completed two collaborative survival tasks, one with and one without AISA access. The authors identify four input patterns from 76 AISA queries (seeking word translation, rationalizing decisions, stating viewpoints, only keywords), report an ANOVA suggesting that the rationalizing-decisions pattern requires significantly more input duration and character modifications, and find no significant quantitative effects of AISA on self-rated speaking competence (p=0.875), anxiety (p=0.235), or workload (p=0.383). Follow-up interviews indicate that AISA improved perceived logical flow and depth of speech but also introduced multitasking demands that could increase anxiety and workload. The paper concludes with design recommendations for reducing input effort, preserving user autonomy, and mitigating workload and anxiety.","tokens_in":26497,"tokens_out":3916,"duration_ms":41970,"significance":"If the findings hold, this is a useful exploratory contribution to CSCW and AI-mediated communication: it is one of the first studies to examine real-time AI content support for NNS speaking, and it reports null results honestly with participant-level nonparametric tests. The detailed system and prompt design are described well enough to be adapted by other researchers, and the qualitative analysis is grounded in concrete interview quotes. The main value lies in the balanced claim that real-time AI assistance can improve perceived speech content while introducing measurable interaction costs. However, the quantitative support for the RQ1 claim about the burden of the rationalizing-decisions pattern is weakened by the clustering issue discussed below, and the RQ2 null result concerns self-rated competence rather than objectively measured speaking competence.","major_comments":[{"comment":"The one-way ANOVA with Brown-Forsythe correction and the subsequent Games-Howell post-hoc tests treat each of the 76 AISA queries as an independent observation, but these queries are nested within 31 NNS participants. Participants differ in typing speed, verbosity, and engagement with the tool, so the significant differences in input duration (F[3,42.268]=20.382, p<0.001) and modification count (F[3,42.856]=34.908, p<0.001) could be driven by participant-level traits rather than by the input pattern itself. The paper reports no mixed-effects model, no cluster-robust standard errors, and no intra-class correlation. Because the claim that the rationalizing-decisions pattern is significantly more effortful is load-bearing for RQ1 and for the design recommendations in Section 5.1, the authors should reanalyze these data with participant-level random effects or cluster-robust inference and report whether the pattern differences remain.","section":"Section 4.1.5"},{"comment":"The abstract and conclusion state that AISA 'did not improve NNSs' speaking competence,' but the measure used is the NNSs' self-rated speaking competence on the Duran scale, not an objective or observer-rated measure. Section 3.5 correctly labels this as 'self-rated speaking competence,' but the RQ2 results and the central claim are repeatedly phrased without that qualification. The claim should be reframed as 'AISA did not significantly change self-perceived speaking competence' to avoid overstating the null result, especially since the interview data suggest benefits in dimensions not captured by this self-report scale.","section":"Section 4.2"}],"minor_comments":[{"comment":"There is a numerical inconsistency in the reported frequency of the 'only keywords' pattern: Section 4.1 states it occurred 12 times, while Section 4.1.6 reports sixteen instances of Only English for this category, which would exceed the total. Please reconcile these counts and ensure Figure 3(a) matches them.","section":"Section 4.1 and Section 4.1.6"},{"comment":"The speaking competence scale is reported in Section 3.6.2 with M=4.27 and SD=1.15, but Section 4.2 reports M=3.357 (with AISA) and M=3.226 (without AISA). Please clarify whether the Section 3.6.2 statistics are pooled across conditions or whether one of these values is a typo.","section":"Section 3.6.2 and Section 4.2"},{"comment":"There is a typo: 'analyis' should be 'analysis' in the sentence describing the Games-Howell post-hoc analysis.","section":"Section 4.1.5"},{"comment":"The Wilcoxon signed-rank tests are reported with z and p values but without effect sizes or confidence intervals; adding a standardized effect size such as r or matched-pairs rank-biserial correlation would strengthen the interpretation of the null results.","section":"Section 4.2 and Section 4.3"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for a CSCW venue and the qualitative findings are interesting. The main technical concern is the clustering issue in Section 4.1.5; if the reanalysis changes the significance of the pattern differences, the RQ1 claim and Section 5.1 recommendations will need to be adjusted accordingly. I did not find evidence of circularity or problematic self-citation; the related work by the same group is used for motivation and task design. The manuscript would also benefit from a careful pass to fix the count inconsistency for the 'only keywords' pattern before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a competent exploratory CSCW/HCI study of a real-time AI speaking assistant for non-native speakers. What's actually new: it provides the first empirical look I know of at NNSs using an LLM to generate speaking references during live group discussion, rather than practicing in advance or managing the floor. The four-category taxonomy of input patterns — word translation, rationalizing decisions, stating viewpoints, keywords — is a genuine contribution. The within-subjects design (31 teams) is plausible, and the reporting of null results for speaking competence, anxiety, and workload is honest and not overclaimed. The interviews give credible evidence that the tool improved logical flow and depth of argument while adding multitasking burden. The limitations section is unusually candid.\n\nThe soft spots are real but not fatal. Section 4.1.5 runs an ANOVA on 76 queries nested within 31 participants without accounting for the clustering. The stress-test note is correct: a mixed model or cluster-robust errors could change the effort-related claim about rationalizing decisions. The RQ2 and RQ3 conclusions rest on participant-level Wilcoxon tests and interviews, so they are not damaged by this issue. I also found internal inconsistencies in the reported counts: the abstract and Section 4.1 give 12 instances of \"only keywords,\" but Section 4.1.6 reports 16; the sub-counts for rationalizing decisions and stating viewpoints do not sum to their stated totals. These look like reporting errors rather than anything sinister, but they need correction before publication. The system itself is described at a high level, with no evaluation of output quality or latency, and the speaking-competence measure is self-report only. None of this sinks the paper; it is a solid exploratory study that belongs in the review process.\n\nWho is this for: HCI and CSCW researchers working on AI-mediated communication, language barriers in collaboration, and real-time assistive tools. A serious referee can engage with it profitably. My recommendation: send it to peer review. Ask the authors to redo the clustered analysis, reconcile the numbers, and clarify what the self-report measure can and cannot show. With those revisions, it is a useful contribution to a rapidly growing design space.","headline":"A solid exploratory HCI study with honestly reported null results and a useful input-pattern taxonomy, held back by a clustered ANOVA and internal count discrepancies that a serious revision can fix.","tokens_in":26924,"tokens_out":1728,"would_cite":true,"duration_ms":18999,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An AI speech aid did not improve non-native speakers' measured speaking competence, but interviews found clearer, better-structured speech.","keywords":["AI-mediated communication","non-native speakers","real-time multilingual communication","speaking assistance","large language models","input patterns","cognitive workload","speaking anxiety"],"falsifier":"Reanalyze the 76 logged queries with a mixed-effects model that adds a random intercept per non-native participant and tests whether the input-pattern differences in duration and modification count survive; the claim about the extra burden of rationalizing decisions would be undercut if the pattern effect no longer reaches significance, or if the per-participant variance dominates the pattern variance.","tokens_in":25936,"feed_emoji":"🗣️","tokens_out":5989,"duration_ms":59008,"temperature":0.7,"pith_summary":"This paper tries to establish what happens when a real-time AI assistant supplies non-native speakers with ready-made English sentences during live multilingual group discussions. It shows that the assistant's benefits are real but narrow: self-rated speaking competence, anxiety, and workload showed no statistically significant change, yet interviews found that speech became more logical, more on-topic, and better argued. It also identifies four ways non-native speakers asked the assistant for help, and argues that one of them—asking it to rationalize a decision—costs noticeably more time and editing effort. A sympathetic reader would care because the result separates content support from interaction cost: helping someone say something is not the same as helping them speak better under real-time pressure.","feed_headline":"AI speech aid: no fluency gain, but clearer arguments","feed_subtitle":"In 31 mixed teams, self-rated competence stayed flat while interviews showed better-structured speech and extra multitasking load.","key_machinery":"The load-bearing object is AISA itself, a system built on a large language model that turns a non-native speaker's short input—keywords, phrases, or partial sentences, often mixing English with their native language—into complete, conversational English sentences, conditioned on the task background and the live transcript. The prompt template, which includes background, conversation history, self-introduction, and output requirements, is the mechanism that makes generated references context-aware and first-person. Around that generation step sits the interaction loop the paper calls the cost: type a query, wait for output, review and adapt it, then speak. That loop, and the four input patterns it produces, carries the argument that assistance can enhance content while the overhead of requesting it competes for attention.","core_discovery":"The central claim is that a tool which generates speaking references in real time can improve the substance of what non-native speakers say without improving the linguistic competence with which they say it, and that the act of using the tool carries its own attention cost. In a within-subjects experiment with 31 teams of two native speakers and one non-native speaker completing survival tasks with and without AISA, the assistant produced no significant change in self-reported speaking competence (Wilcoxon signed-rank test, z = 0.171, p = 0.875), speaking anxiety (z = 1.201, p = 0.235), or workload (z = 0.882, p = 0.383). The qualitative data instead report clearer logical flow, stronger arguments, and better alignment with the topic, alongside feelings of reduced agency, extra anxiety from entering and reviewing queries, and a workload that increased for some users while decreasing for others. The paper reads these together as evidence that the added multitasking of using the tool can offset the content-level gains, and it grounds the null competence result in this trade-off rather than in the tool failing to help at all.","pith_inferences":["Beyond the paper, the same content-benefit/interaction-cost pattern likely applies to any synchronous AI writing or reply assistant: measured output quality can rise while fluency or agency perceptions stay flat, so evaluations should separate what the tool produced from what the user had to do to get it.","The paper's own design hints at a testable fork the authors did not run: offering discrete words or phrases instead of complete sentences may preserve autonomy and reduce the 'it is controlling me' effect; this could be tested head-to-head against full-sentence output.","A longitudinal extension would be to let non-native speakers use AISA across several sessions; the input-language mix strategy (native language for logic, English for precision) may improve with practice, potentially converting the qualitative content gains into measurable competence gains."],"forward_implications":["If the trade-off claim is right, future speaking-assistance tools should be judged on content quality (logic, relevance, depth) and on interaction cost, not only on grammar and vocabulary scores.","Future designs that reduce the cost of entering queries—for example voice input or native-language vocal queries—could preserve the content benefits without raising workload.","If the extra multitasking is the reason competence did not improve, then designs that offload attention, such as streaming partial suggestions or having an agent ask the group to wait, may unlock the benefits the interviews observed.","The four input patterns, if stable, give designers a direct menu: support word lookup, decision rationalization, viewpoint completion, and keyword expansion in one tool."],"supporting_citations":[{"why":"Supplies the speaking-competence scale (grammar, tense, vocabulary) that yields the paper's null result for RQ2.","marker":"[22]"},{"why":"Supplies the foreign-language state anxiety scale used to test whether AISA changed anxiety.","marker":"[9]"},{"why":"Supplies the NASA-TLX workload items used to test whether AISA changed workload.","marker":"[19]"},{"why":"Defines AI-mediated communication, the conceptual frame that positions AISA as an AIMC tool.","marker":"[26]"},{"why":"Describes the prior automatic agent that increased NNS speaking opportunities, the approach this paper extends from opportunity to content support.","marker":"[45]"},{"why":"Presents EnglishBot, an AI practice partner for English learners, the main prior system the paper contrasts with real-time speaking reference generation.","marker":"[72]"},{"why":"Supports the real-time transcript feature by showing automated transcripts help non-native speakers keep up with missed conversation.","marker":"[27]"},{"why":"Provides evidence that real-time transcription improves non-native speakers' comprehension, justifying the transcript panel in the system design.","marker":"[64]"}],"fun_headline_variants":["AI speaking aid: more depth, no fluency gain","Real-time AI voice help: better ideas, same competence","Speaking assistant refines arguments, not language skills","AI improves speech content, at cost of extra attention"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The quantitative comparison of input effort treats each of the 76 queries as an independent measurement, even though they come from only 31 participants; if the clustering by participant were taken into account, the reported differences between input patterns might weaken or disappear.","fun_headline_variants_meta":{"raw":{"variants":["AI speaking aid: more depth, no fluency gain","Real-time AI voice help: better ideas, same competence","Speaking assistant refines arguments, not language skills","AI improves speech content, at cost of extra attention"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000246,"raw_usage":{"total_tokens":1572,"prompt_tokens":1012,"completion_tokens":560,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":628,"completion_tokens_details":{"reasoning_tokens":497}},"tokens_in":628,"tokens_out":560,"duration_ms":6155,"temperature":1.0,"reasoning_tokens":497,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:12:19.945224+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Reanalyze the 76 logged queries with a mixed-effects model that adds a random intercept per non-native participant and tests whether the input-pattern differences in duration and modification count survive; the claim about the extra burden of rationalizing decisions would be undercut if the pattern effect no longer reaches significance, or if the per-participant variance dominates the pattern variance.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the speaking-competence scale (grammar, tense, vocabulary) that yields the paper's null result for RQ2."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes the prior automatic agent that increased NNS speaking opportunities, the approach this paper extends from opportunity to content support."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Presents EnglishBot, an AI practice partner for English learners, the main prior system the paper contrasts with real-time speaking reference generation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the real-time transcript feature by showing automated transcripts help non-native speakers keep up with missed conversation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides evidence that real-time transcription improves non-native speakers' comprehension, justifying the transcript panel in the system design."}],"review_version":1}