Pith. sign in

REVIEW 2 major objections 4 minor 99 references

AI-Based Speaking Assistant: Supporting Non-Native Speakers' Speaking in Real-Time Multilingual Communication

T0 review · 2 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read An AI speech aid did not improve non-native speakers' measured speaking competence, but interviews found clearer, better-structured speech.

desk verdict A solid exploratory HCI study with honestly reported null results and a useful input-pattern taxonomy, held back by a clustered ANOVA and internal count discrepancies that a serious revision can fix. read the letter →

arxiv 2505.01678 v1 pith:6EFZRJBV submitted 2025-05-03 cs.HC

classification cs.HC
keywords AI-mediatedcommunicationnon-nativespeakersreal-timemultilingualspeakingassistancelargelanguagemodelsinputpatternscognitiveworkloadanxiety
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish what happens when a real-time AI assistant supplies non-native speakers with ready-made English sentences during live multilingual group discussions. It shows that the assistant's benefits are real but narrow: self-rated speaking competence, anxiety, and workload showed no statistically significant change, yet interviews found that speech became more logical, more on-topic, and better argued. It also identifies four ways non-native speakers asked the assistant for help, and argues that one of them—asking it to rationalize a decision—costs noticeably more time and editing effort. A sympathetic reader would care because the result separates content support from interaction cost: helping someone say something is not the same as helping them speak better under real-time pressure.

What carries the argument

The load-bearing object is AISA itself, a system built on a large language model that turns a non-native speaker's short input—keywords, phrases, or partial sentences, often mixing English with their native language—into complete, conversational English sentences, conditioned on the task background and the live transcript. The prompt template, which includes background, conversation history, self-introduction, and output requirements, is the mechanism that makes generated references context-aware and first-person. Around that generation step sits the interaction loop the paper calls the cost: type a query, wait for output, review and adapt it, then speak. That loop, and the four input patterns it produces, carries the argument that assistance can enhance content while the overhead of requesting it competes for attention.

What would settle it

Reanalyze the 76 logged queries with a mixed-effects model that adds a random intercept per non-native participant and tests whether the input-pattern differences in duration and modification count survive; the claim about the extra burden of rationalizing decisions would be undercut if the pattern effect no longer reaches significance, or if the per-participant variance dominates the pattern variance.

Watch

Extended reading notes

Core claim

The central claim is that a tool which generates speaking references in real time can improve the substance of what non-native speakers say without improving the linguistic competence with which they say it, and that the act of using the tool carries its own attention cost. In a within-subjects experiment with 31 teams of two native speakers and one non-native speaker completing survival tasks with and without AISA, the assistant produced no significant change in self-reported speaking competence (Wilcoxon signed-rank test, z = 0.171, p = 0.875), speaking anxiety (z = 1.201, p = 0.235), or workload (z = 0.882, p = 0.383). The qualitative data instead report clearer logical flow, stronger arguments, and better alignment with the topic, alongside feelings of reduced agency, extra anxiety from entering and reviewing queries, and a workload that increased for some users while decreasing for others. The paper reads these together as evidence that the added multitasking of using the tool can offset the content-level gains, and it grounds the null competence result in this trade-off rather than in the tool failing to help at all.

Load-bearing premise

The quantitative comparison of input effort treats each of the 76 queries as an independent measurement, even though they come from only 31 participants; if the clustering by participant were taken into account, the reported differences between input patterns might weaken or disappear.

Editorial extensions

If this is right

  • If the trade-off claim is right, future speaking-assistance tools should be judged on content quality (logic, relevance, depth) and on interaction cost, not only on grammar and vocabulary scores.
  • Future designs that reduce the cost of entering queries—for example voice input or native-language vocal queries—could preserve the content benefits without raising workload.
  • If the extra multitasking is the reason competence did not improve, then designs that offload attention, such as streaming partial suggestions or having an agent ask the group to wait, may unlock the benefits the interviews observed.
  • The four input patterns, if stable, give designers a direct menu: support word lookup, decision rationalization, viewpoint completion, and keyword expansion in one tool.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper, the same content-benefit/interaction-cost pattern likely applies to any synchronous AI writing or reply assistant: measured output quality can rise while fluency or agency perceptions stay flat, so evaluations should separate what the tool produced from what the user had to do to get it.
  • The paper's own design hints at a testable fork the authors did not run: offering discrete words or phrases instead of complete sentences may preserve autonomy and reduce the 'it is controlling me' effect; this could be tested head-to-head against full-sentence output.
  • A longitudinal extension would be to let non-native speakers use AISA across several sessions; the input-language mix strategy (native language for logic, English for precision) may improve with practice, potentially converting the qualitative content gains into measurable competence gains.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. This paper presents a mixed-methods study of an AI-based speaking assistant (AISA) that provides real-time speaking references to non-native speakers (NNSs) during multilingual team communication. In a within-subjects experiment, 31 teams consisting of two native speakers and one NNS completed two collaborative survival tasks, one with and one without AISA access. The authors identify four input patterns from 76 AISA queries (seeking word translation, rationalizing decisions, stating viewpoints, only keywords), report an ANOVA suggesting that the rationalizing-decisions pattern requires significantly more input duration and character modifications, and find no significant quantitative effects of AISA on self-rated speaking competence (p=0.875), anxiety (p=0.235), or workload (p=0.383). Follow-up interviews indicate that AISA improved perceived logical flow and depth of speech but also introduced multitasking demands that could increase anxiety and workload. The paper concludes with design recommendations for reducing input effort, preserving user autonomy, and mitigating workload and anxiety.

Significance. If the findings hold, this is a useful exploratory contribution to CSCW and AI-mediated communication: it is one of the first studies to examine real-time AI content support for NNS speaking, and it reports null results honestly with participant-level nonparametric tests. The detailed system and prompt design are described well enough to be adapted by other researchers, and the qualitative analysis is grounded in concrete interview quotes. The main value lies in the balanced claim that real-time AI assistance can improve perceived speech content while introducing measurable interaction costs. However, the quantitative support for the RQ1 claim about the burden of the rationalizing-decisions pattern is weakened by the clustering issue discussed below, and the RQ2 null result concerns self-rated competence rather than objectively measured speaking competence.

major comments (2)
  1. [Section 4.1.5] The one-way ANOVA with Brown-Forsythe correction and the subsequent Games-Howell post-hoc tests treat each of the 76 AISA queries as an independent observation, but these queries are nested within 31 NNS participants. Participants differ in typing speed, verbosity, and engagement with the tool, so the significant differences in input duration (F[3,42.268]=20.382, p<0.001) and modification count (F[3,42.856]=34.908, p<0.001) could be driven by participant-level traits rather than by the input pattern itself. The paper reports no mixed-effects model, no cluster-robust standard errors, and no intra-class correlation. Because the claim that the rationalizing-decisions pattern is significantly more effortful is load-bearing for RQ1 and for the design recommendations in Section 5.1, the authors should reanalyze these data with participant-level random effects or cluster-robust inference and report whether the pattern differences remain.
  2. [Section 4.2] The abstract and conclusion state that AISA 'did not improve NNSs' speaking competence,' but the measure used is the NNSs' self-rated speaking competence on the Duran scale, not an objective or observer-rated measure. Section 3.5 correctly labels this as 'self-rated speaking competence,' but the RQ2 results and the central claim are repeatedly phrased without that qualification. The claim should be reframed as 'AISA did not significantly change self-perceived speaking competence' to avoid overstating the null result, especially since the interview data suggest benefits in dimensions not captured by this self-report scale.
minor comments (4)
  1. [Section 4.1 and Section 4.1.6] There is a numerical inconsistency in the reported frequency of the 'only keywords' pattern: Section 4.1 states it occurred 12 times, while Section 4.1.6 reports sixteen instances of Only English for this category, which would exceed the total. Please reconcile these counts and ensure Figure 3(a) matches them.
  2. [Section 3.6.2 and Section 4.2] The speaking competence scale is reported in Section 3.6.2 with M=4.27 and SD=1.15, but Section 4.2 reports M=3.357 (with AISA) and M=3.226 (without AISA). Please clarify whether the Section 3.6.2 statistics are pooled across conditions or whether one of these values is a typo.
  3. [Section 4.1.5] There is a typo: 'analyis' should be 'analysis' in the sentence describing the Games-Howell post-hoc analysis.
  4. [Section 4.2 and Section 4.3] The Wilcoxon signed-rank tests are reported with z and p values but without effect sizes or confidence intervals; adding a standardized effect size such as r or matched-pairs rank-biserial correlation would strengthen the interpretation of the null results.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: central claims derive from new experimental data, interviews, and external scales; self-citations are motivational only.

full rationale

The paper's central claims are empirical findings from a newly conducted within-subject experiment (31 teams) and follow-up interviews: the identification of four AISA input patterns from log data (Section 4.1), the null result on speaking competence (Wilcoxon z = 0.171, p = 0.875, Section 4.2), and qualitative reports on workload and anxiety (Section 4.3). No quantity used as an input is re-labeled as a prediction: speaking competence is measured with an external scale by Duran [22], anxiety with the Foreign Language State Anxiety Scale [9], and workload with NASA-TLX items [19]. The input patterns are derived from the collected query logs rather than being used to define the outcome. The paper does cite prior work co-authored by one of the present authors (e.g., He et al. [29] for the survival-task design and Li et al. [45] for an agent that opens speaking floors), but these citations provide task materials, motivation, and comparison with prior systems; they do not supply the paper's conclusions or forbid alternative interpretations. The statistical concern raised by the skeptical reader, that the one-way ANOVA in Section 4.1.5 treats 76 queries as independent despite clustering within 31 participants, is a validity or robustness issue, not a circularity issue, because the RQ2 null result and the qualitative workload/anxiety findings rest on participant-level tests and interviews. The paper's stated limitations are honest and do not conceal a derivation that reduces to its own inputs. Overall, the derivation chain is self-contained relative to the new data, so no circular step is present.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The study's claims rest on established self-report scales, a specific participant population, and a custom LLM pipeline; no numeric free parameters are fitted and no new physical or conceptual entities are postulated.

assumptions (3)
  • domain assumption Self-report scales for speaking competence, anxiety, and workload are valid proxies for the constructs.
    Used in Section 3.6 to answer RQ2 and RQ3; the scales have acceptable Cronbach alpha values (0.86, 0.86, 0.85), but self-report may diverge from objective speech quality.
  • domain assumption The team composition of two native speakers and one non-native speaker, plus survival tasks, elicits realistic real-time multilingual communication challenges.
    This underlies the experimental design in Section 3.3; the authors note in Section 6 that survival tasks and Chinese IELTS 5.5-6.5 participants limit generalizability.
  • standard math The 76 AISA input queries are independent for the purpose of one-way ANOVA.
    Section 4.1.5 applies one-way ANOVA without accounting for clustering by participant; this is a load-bearing assumption for the significant effort differences.

how reviews work

0 comments
Cite this review

Pith. "Pith review of AI-Based Speaking Assistant: Supporting Non-Native Speakers' Speaking in Real-Time Multilingual Communication." pith.science (2026). https://pith.science/paper/6EFZRJBV

@misc{pith2026250501678,
  author       = {Pith},
  title        = {Pith review of: AI-Based Speaking Assistant: Supporting Non-Native Speakers' Speaking in Real-Time Multilingual Communication},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6EFZRJBV}},
  note         = {Machine review of arXiv:2505.01678}
}
read the original abstract

Non-native speakers (NNSs) often face speaking challenges in real-time multilingual communication, such as struggling to articulate their thoughts. To address this issue, we developed an AI-based speaking assistant (AISA) that provides speaking references for NNSs based on their input queries, task background, and conversation history. To explore NNSs' interaction with AISA and its impact on NNSs' speaking during real-time multilingual communication, we conducted a mixed-method study involving a within-subject experiment and follow-up interviews. In the experiment, two native speakers (NSs) and one NNS formed a team (31 teams in total) and completed two collaborative tasks--one with access to the AISA and one without. Overall, our study revealed four types of AISA input patterns among NNSs, each reflecting different levels of effort and language preferences. Although AISA did not improve NNSs' speaking competence, follow-up interviews revealed that it helped improve the logical flow and depth of their speech. Moreover, the additional multitasking introduced by AISA, such as entering and reviewing system output, potentially elevated NNSs' workload and anxiety. Based on these observations, we discuss the pros and cons of implementing tools to assist NNS in real-time multilingual communication and offer design recommendations.

Figures

Figures reproduced from arXiv: 2505.01678 by the authors.

Figure 1
Figure 1. Web interface with AISA. Where to represents the components of the interface; to represents NNSs’ operation steps when using AISA. Specifically, basic information header; game instruction button. When users click this button, detailed instructions for the survival game will be expanded; transcript panel; AISA input box; AISA output panel; items panel. Here, John and Jessie are NSs, while Xin is a NNS. As for the dem… view at source ↗
Figure 2
Figure 2. Input duration and the number of modified characters by input pattern. The asterisks denote levels of statistical significance, [PITH_FULL_IMAGE:figures/full_fig_p014_2.png] view at source ↗
Figure 3
Figure 3. Distribution of the language used across input patterns and input lengths by each language used. [PITH_FULL_IMAGE:figures/full_fig_p015_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

99 extracted references · 58 canonical work pages

  1. [1]

    Samira Al Hosni. 2014. Speaking difficulties encountered by young EFL learners. International Journal on Studies in English Language and Literature (IJSELL) 2, 6 (2014), 22–30

  2. [2]

    Ali Ali. 2024. Improving trust-building through more transparent conversational agent communication, in the context of medical decision support. B.S. thesis. University of Twente

  3. [3]

    Jason L Anthony, Emily J Solari, Jeffrey M Williams, Kimberly D Schoger, Zhou Zhang, Lee Branum-Martin, and David J Francis. 2009. Development of bilingual phonological awareness in Spanish-speaking English language learners: The roles of vocabulary, letter knowledge, and prior phonological awareness. Scientific Studies of Reading 13, 6 (2009), 535–564

  4. [4]

    Stella Aririguzoh. 2022. Communication competencies, culture and SDGs: effective processes to cross-cultural communication. Humanities and Social Sciences Communications 9, 1 (2022), 1–11

  5. [5]

    Ariyanti Ariyanti. 2016. Psychological factors affecting EFL students’ speaking performance. ASIAN TEFL Journal of Language Teaching and Applied Linguistics 1, 1 (2016)

  6. [6]

    Solomon E Asch. 2016. Effects of group pressure upon the modification and distortion of judgments. In Organizational influence processes. Routledge, 295–303

  7. [7]

    Alan D Baddeley and Graham James Hitch. 1974. Working memory (Vol. 8). New York: GA Bower (ed), Recent advances in learning and motivation (1974)

  8. [8]

    J Baker. 2003. Essential Speaking Skills/Joanna Baker. Westrup Heather: Continuum International Publishing Group (2003)

Show all 99 references
  1. [9]

    Melissa Baralt and Laura Gurzynski-Weiss. 2011. Comparing learners’ state anxiety during task-based interaction in computer-mediated and face-to-face communication. Language Teaching Research 15, 2 (2011), 201–229

  2. [10]

    Ashish Bastola, Hao Wang, Judsen Hembree, Pooja Yadav, Nathan McNeese, and Abolfazl Razi. 2023. LLM-based Smart Reply (LSR): Enhancing Collaborative Performance with ChatGPT-mediated Smart Reply System. arXiv preprint (2023)

  3. [11]

    Patrice Béchard and Orlando Marquez Ayala. 2024. Reducing hallucination in structured outputs via Retrieval-Augmented Generation. arXiv preprint arXiv:2404.08189 (2024)

  4. [12]

    Boomerang. 2018. Respondable: Write Better Email. https://www.boomeranggmail.com/respondable/

  5. [13]

    Virginia Braun and Victoria Clarke. 2006. Using thematic analysis in psychology. Qualitative research in psychology 3, 2 (2006), 77–101

  6. [14]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020), 1877–1901

  7. [15]

    Matt A Casado and Mary I Dershiwsky. 2004. EFFECT OF EDUCATIONAL STRATEGIES ON ANXIETY IN THE SECONDLANGUAGE CLASSROOM: AN EXPLORATORY COMPARATIVE STUDY BETWEENU. S. AND SPANISH FIRST-SEMESTER UNIVERSITY STUDENTS. College Student Journal 38, 1 (2004)

  8. [16]

    Contextual

    Anastasia Chan. 2023. GPT-3 and InstructGPT: Technological dystopianism, utopianism, and “Contextual” perspectives in AI ethics and industry. AI and Ethics 3, 1 (2023), 53–64

  9. [17]

    Robert B Cialdini and Noah J Goldstein. 2004. Social influence: Compliance and conformity. Annu. Rev. Psychol. 55 (2004), 591–621

  10. [18]

    Mark Coeckelbergh. 2020. Artificial intelligence, responsibility attribution, and a relational justification of explainability. Science and engineering ethics 26, 4 (2020), 2051–2068

  11. [19]

    Lacey Colligan, Henry WW Potts, Chelsea T Finn, and Robert A Sinkin. 2015. Cognitive workload changes for nurses transitioning from a legacy system with paper documentation to a commercial electronic health record. International journal of medical informatics 84, 7 (2015), 469–476

  12. [20]

    Jim Cummins. 2007. Rethinking monolingual instructional strategies in multilingual classrooms. Canadian journal of applied linguistics 10, 2 (2007), 221–240

  13. [21]

    Wen Duan, Naomi Yamashita, and Susan R Fussell. 2019. Increasing native speakers’ awareness of the need to slow down in multilingual conversations using a real-time speech speedometer. Proceedings of the ACM on Human-Computer Interaction 3, CSCW (2019), 1–25

  14. [22]

    Robert L Duran. 1992. Communicative adaptability: A review of conceptualization and measurement. Communication Quarterly 40, 3 (1992), 253–268

  15. [23]

    Liye Fu, Benjamin Newman, Maurice Jakesch, and Sarah Kreps. 2023. Comparing Sentence-Level Suggestions to Message-Level Suggestions in AI-Mediated Communication. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–13. Manuscript submitted to ACM...

  16. [24]

    Yue Fu, Sami Foell, Xuhai Xu, and Alexis Hiniker. 2024. From Text to Self: Users’ Perception of AIMC Tools on Interpersonal Communication and Self. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–17

  17. [25]

    Youssouf Haidara. 2016. Psychological Factor Affecting English Speaking Performance for the English Learners in Indonesia. Universal Journal of Educational Research 4, 7 (2016), 1501–1505

  18. [26]

    Jeffrey T Hancock, Mor Naaman, and Karen Levy. 2020. AI-Mediated Communication: Definition, Research Agenda, and Ethical Considerations.Jour- nal of Computer-Mediated Communication 25, 1 (01 2020), 89–100. https://doi.org/10.1093/jcmc/zmz022 arXiv:https://academic.oup.com/jcmc...

  19. [27]

    Ari Hautasaari and Naomi Yamashita. 2014. Do automated transcripts help non-native speakers catch up on missed conversation in audio conferences?. In Proceedings of the 5th ACM international conference on Collaboration across boundaries: culture, distance & technology . 65–72

  20. [28]

    An E He. 2012. Systematic use of mother tongue as learning/teaching resources in target language instruction.Multilingual Education 2 (2012), 1–15

  21. [29]

    Helen Ai He, Naomi Yamashita, Ari Hautasaari, Xun Cao, and Elaine M. Huang. 2017. Why Did They Do That? Exploring Attribution Mismatches Between Native and Non-Native Speakers Using Videoconferencing. In Proceedings of the 2017 ACM Conference on Computer Supported Cooperative ...

  22. [30]

    Helen Ai He, Naomi Yamashita, Ari Hautasaari, Xun Cao, and Elaine M Huang. 2017. Why did they do that? Exploring attribution mismatches between native and non-native speakers using videoconferencing. In Proceedings of the 2017 ACM Conference on Computer Supported Cooperative W...

  23. [31]

    Matthew Henderson, Rami Al-Rfou, Brian Strope, Yun-Hsuan Sung, László Lukács, Ruiqi Guo, Sanjiv Kumar, Balint Miklos, and Ray Kurzweil. 2017. Efficient natural language response suggestion for smart reply. arXiv preprint arXiv:1705.00652 (2017)

  24. [32]

    Suzanne Hidi and Valerie Anderson. 1986. Producing written summaries: Task demands, cognitive operations, and implications for instruction. Review of educational research 56, 4 (1986), 473–493

  25. [33]

    Pamela J Hinds, Tsedal B Neeley, and Catherine Durnell Cramton. 2014. Language as a lightning rod: Power contests, emotion regulation, and subgroup dynamics in global teams. Journal of International Business Studies 45 (2014), 536–561

  26. [34]

    Jess Hohenstein, Rene F Kizilcec, Dominic DiFranzo, Zhila Aghajari, Hannah Mieczkowski, Karen Levy, Mor Naaman, Jeffrey Hancock, and Malte F Jung. 2023. Artificial intelligence in communication impacts language and social relationships. Scientific Reports 13, 1 (2023), 5487

  27. [35]

    Elaine K Horwitz, Michael B Horwitz, and Joann Cope. 1986. Foreign language classroom anxiety.The Modern language journal 70, 2 (1986), 125–132

  28. [36]

    Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al. 2023. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Info...

  29. [37]

    David G Jansson and Steven M Smith. 1991. Design fixation. Design studies 12, 1 (1991), 3–11

  30. [38]

    Jaeho Jeon. 2021. Chatbot-assisted dynamic assessment (CA-DA) for L2 vocabulary learning and diagnosis. Computer Assisted Language Learning (2021), 1–27

  31. [39]

    Junaidi Junaidi. 2020. Artificial intelligence in EFL context: rising students’ speaking performance with Lyra virtual assistance. International Journal of Advanced Science and Technology Rehabilitation 29, 5 (2020), 6735–6741

  32. [40]

    Hyunjin Kang and Chen Lou. 2022. AI agency vs. human agency: understanding human–AI interactions on TikTok and their implica- tions for user engagement. Journal of Computer-Mediated Communication 27, 5 (08 2022), zmac014. https://doi.org/10.1093/jcmc/zmac014 arXiv:https://acad...

  33. [41]

    Anjuli Kannan, Karol Kurach, Sujith Ravi, Tobias Kaufmann, Andrew Tomkins, Balint Miklos, Greg Corrado, Laszlo Lukacs, Marina Ganea, Peter Young, et al. 2016. Smart reply: Automated response suggestion for email. In Proceedings of the 22nd ACM SIGKDD international conference o...

  34. [42]

    Raja Muhammad Ishtiaq Khan, Noor Raha Mohd Radzuan, Muhammad Shahbaz, Ainol Haryati Ibrahim, and Ghulam Mustafa. 2018. The role of vocabulary knowledge in speaking development of Saudi EFL learners. Arab World English Journal (A WEJ) Volume9 (2018)

  35. [43]

    Paul E King and Amber N Finn. 2017. A test of attention control theory in public speaking: Cognitive load influences the relationship between state anxiety and verbal production. Communication Education 66, 2 (2017), 168–182

  36. [44]

    J Clayton Lafferty, Patrick M Eady, and J Elmers. 1974. The desert survival problem. Experimental learning methods (1974)

  37. [45]

    Xiaoyan Li, Naomi Yamashita, Wen Duan, Yoshinari Shirai, and Susan R Fussell. 2023. Improving Non-Native Speakers’ Participation with an Automatic Agent in Multilingual Groups. Proceedings of the ACM on Human-Computer Interaction 7, GROUP (2023), 1–28

  38. [46]

    John Lim and Yin Ping Yang. 2008. Exploring computer-based multilingual negotiation support for English–Chinese dyads: can we negotiate in our native languages? Behaviour & Information Technology 27, 2 (2008), 139–151

  39. [47]

    Stephanie Lindemann. 2002. Listening with an attitude: A model of native-speaker comprehension of non-native speakers in the United States. Language in Society 31, 3 (2002), 419–441

  40. [48]

    Julie S Linsey, Ian Tseng, Katherine Fu, Jonathan Cagan, Kristin L Wood, and Christian Schunn. 2010. A study of design fixation, its mitigation and perception in engineering design faculty. (2010)

  41. [49]

    Yihe Liu, Anushk Mittal, Diyi Yang, and Amy Bruckman. 2022. Will AI console me when I lose my pet? Understanding perceptions of AI-mediated email writing. In Proceedings of the 2022 CHI conference on human factors in computing systems . 1–13. Manuscript submitted to ACM 26 Pei...

  42. [50]

    Amama Mahmood, Junxiang Wang, Bingsheng Yao, Dakuo Wang, and Chien-Ming Huang. 2023. LLM-Powered Conversational Voice Assistants: Interaction Patterns, Opportunities, Challenges, and Design Guidelines. arXiv preprint arXiv:2309.13879 (2023)

  43. [51]

    Potsawee Manakul, Adian Liusie, and Mark JF Gales. 2023. Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models. arXiv preprint arXiv:2303.08896 (2023)

  44. [52]

    Y Manurung and S Izar. 2019. Challenging factors affecting students’ speaking performance. In Proceedings of the 2nd International Conference on Language, Literature and Education, ICLLE 2019, 22-23 August, Padang, West Sumatra, Indonesia

  45. [53]

    Aaron Marcus and Emilie West Gould. 2000. Crosscurrents: cultural dimensions and global Web user-interface design.interactions 7, 4 (2000), 32–46

  46. [54]

    Jon McCormack, Toby Gifford, and Patrick Hutchings. 2019. Autonomy, authenticity, authorship and intention in computer generated art. In International conference on computational intelligence in music, sound, art and design (part of EvoStar) . Springer, 35–50

  47. [55]

    Sarnoff Mednick. 1962. The associative basis of the creative process. Psychological review 69, 3 (1962), 220

  48. [56]

    Sheetal Mehrotra. 2022. Decision-Making Quality: A Cognitive Load Perspective. 16 (01 2022)

  49. [57]

    Catalin Mitelut, Ben Smith, and Peter Vamplew. 2023. Intent-aligned AI systems deplete human agency: the need for agency foundations research in AI safety. arXiv preprint arXiv:2305.19223 (2023)

  50. [58]

    Tsedal B Neeley. 2013. Language matters: Status loss and achieved status distinctions in global organizations. Organization Science 24, 2 (2013), 476–497

  51. [59]

    Tsedal B Neeley and Tracy L Dumas. 2016. Unearned status gain: Evidence from a global language mandate. Academy of Management Journal 59, 1 (2016), 14–43

  52. [60]

    Tran Tin Nghi, Luu Quy Khuong, et al. 2021. A study on communication breakdowns between native and non-native speakers in English speaking classes. Journal of English Language Teaching and Applied Linguistics 3, 6 (2021), 01–06

  53. [61]

    Bernard A Nijstad and Wolfgang Stroebe. 2006. How the group affects the mind: A cognitive model of idea generation in groups. Personality and social psychology review 10, 3 (2006), 186–213

  54. [62]

    Richard Oelschlager. 2024. Evaluating the impact of hallucinations on user trust and satisfaction in llm-based systems

  55. [63]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems 35...

  56. [64]

    Yingxin Pan, Danning Jiang, Michael Picheny, and Yong Qin. 2009. Effects of real-time transcription on non-native speaker’s comprehension in computer-mediated communications. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems . 2353–2356

  57. [65]

    Yingxin Pan, Danning Jiang, Lin Yao, Michael Picheny, and Yong Qin. 2010. Effects of automated transcription quality on non-native speakers’ comprehension in real-time computer-mediated communication. In Proceedings of the SIGCHI Conference on Human Factors in Computing System...

  58. [66]

    N Eleni Pappamihiel. 2002. English as a second language students and English language anxiety: Issues in the mainstream classroom. Research in the Teaching of English (2002), 327–355

  59. [67]

    Harold Pashler. 1994. Dual-task interference in simple tasks: data and theory. Psychological bulletin 116, 2 (1994), 220

  60. [68]

    Kaisa Pekkala and Ward van Zoonen. 2022. Work-related social media use: The mediating role of social media communication self-efficacy.European Management Journal 40, 1 (2022), 67–76

  61. [69]

    A Terry Purcell and John S Gero. 1996. Design and other types of fixation. Design studies 17, 4 (1996), 363–383

  62. [70]

    Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. 2018. Improving language understanding by generative pre-training. (2018)

  63. [71]

    Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog 1, 8 (2019), 9

  64. [72]

    Sherry Ruan, Liwei Jiang, Qianyao Xu, Zhiyuan Liu, Glenn M Davis, Emma Brunskill, and James A Landay. 2021. Englishbot: An ai-powered conversational system for second language learning. In 26th international conference on intelligent user interfaces . 434–444

  65. [73]

    Sherry Ruan, Jacob O Wobbrock, Kenny Liou, Andrew Ng, and James A Landay. 2018. Comparing speech and keyboard text entry for short messages in two languages on touchscreen phones. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 1, 4 (2018), 1–23

  66. [74]

    Emanuel A Schegloff. 1982. Discourse as an interactional achievement: Some uses of ‘uh huh’and other things that come between sentences. Analyzing discourse: Text and talk 71 (1982), 71–93

  67. [75]

    Jim Scivener. 2011. Learning teaching: A guidebook for English language teachers

  68. [76]

    Lennart Seitz, Sigrid Bekmeier-Feuerhahn, and Krutika Gohil. 2022. Can we trust a chatbot like a physician? A qualitative study on understanding the emergence of trust toward diagnostic chatbots. International Journal of Human-Computer Studies 165 (2022), 102848

  69. [77]

    Ashish Sharma, Inna W Lin, Adam S Miner, David C Atkins, and Tim Althoff. 2023. Human–AI collaboration enables more empathic conversations in text-based peer-to-peer mental health support. Nature Machine Intelligence 5, 1 (2023), 46–57

  70. [78]

    Nobuhiro Shimogori, Tomoo Ikeda, and Sougo Tsuboi. 2010. Automatically generated captions: will they help non-native speakers communicate in english?. In Proceedings of the 3rd international conference on Intercultural collaboration . 79–86

  71. [79]

    Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. 2021. Retrieval augmentation reduces hallucination in conversation. arXiv preprint arXiv:2104.07567 (2021)

  72. [80]

    Pao Siangliulue, Joel Chan, Krzysztof Z Gajos, and Steven P Dow. 2015. Providing timely examples improves the quantity and quality of generated ideas. In Proceedings of the 2015 ACM SIGCHI Conference on Creativity and Cognition . 83–92. Manuscript submitted to ACM AI-Based Spe...

  73. [81]

    S Shyam Sundar and Sampada S Marathe. 2010. Personalization versus customization: The importance of agency, privacy, and power usage. Human communication research 36, 3 (2010), 298–322

  74. [82]

    Merrill Swain, Thomas Andrew KIRKPATRICK, and Jim Cummins. 2011. How to have a guilt-free life using Cantonese in the English class: A handbook for the English language teacher in Hong Kong. (2011)

  75. [83]

    John Sweller. 1988. Cognitive load during problem solving: Effects on learning. Cognitive science 12, 2 (1988), 257–285

  76. [84]

    Yohtaro Takano and Akiko Noda. 1993. A temporary decline of thinking ability during foreign language processing. Journal of Cross-Cultural Psychology 24, 4 (1993), 445–462

  77. [85]

    Noparat Tananuraksakul. 2011. Non-native English students’ linguistic and cultural challenges in Australia. Journal of International Students 2012 Vol 2 Issue 1 (2011), 107

  78. [86]

    M Iftekhar Tanveer, Emy Lin, and Mohammed Hoque. 2015. Rhema: A real-time in-situ intelligent interface to help people with public speaking. In Proceedings of the 20th international conference on intelligent user interfaces . 286–295

  79. [87]

    Helene Tenzer, Markus Pudelko, and Anne-Wil Harzing. 2014. The impact of language barriers on trust formation in multinational teams. Journal of International Business Studies 45 (2014), 508–535

  80. [88]

    Ha Trinh, Reza Asadi, Darren Edge, and T Bickmore. 2017. Robocop: A robotic coach for oral presentations. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 1, 2 (2017), 1–24

  81. [89]

    Takahiro Tsumura and Seiji Yamada. 2023. Influence of agent’s self-disclosure on human empathy. PLoS One 18, 5 (2023), e0283955

  82. [90]

    Penny Ur. 1999. A course in language teaching

  83. [91]

    Lauren Jodi Van Scoy, Allison M Scott, Michael J Green, Pamela D Witt, Emily Wasserman, Vernon M Chinchilli, and Benjamin H Levi. 2022. Communication Quality Analysis: A User-friendly Observational Measure of Patient–Clinician Communication. Communication methods and measures ...

  84. [92]

    Luis A Vasconcelos and Nathan Crilly. 2016. Inspiration and fixation: Questions, methods, findings, and challenges. Design Studies 42 (2016), 1–32

  85. [93]

    Viswanath Venkatesh and Xiaojun Zhang. 2010. Unified theory of acceptance and use of technology: US vs. China. Journal of global information technology management 13, 1 (2010), 5–27

  86. [94]

    VIETNAM Vietnam. 2015. FACTORS AFFECTING STUDENTS’SPEAKING PERFORMANCE AT LE THANH HIEN HIGH SCHOOL. Asian Journal of Educational Research Vol 3, 2 (2015), 8–23

  87. [95]

    Phuong Quyen Vo, Thi My Nga Pham, and Thao Nguyen Ho. 2018. Challenges to speaking skills encountered by English-majored students: A story of one Vietnamese university in the Mekong Delta. CTU Journal of Innovation and Sustainable Development 54, 5 (2018), 38–44

  88. [96]

    Ching-Lin Wu, Shih-Yuan Huang, Pei-Zhen Chen, and Hsueh-Chih Chen. 2020. A systematic review of creativity-related studies applying the remote associates test from 2000 to 2019. Frontiers in psychology 11 (2020), 573432

  89. [97]

    Qian Yang, Aaron Steinfeld, Carolyn Rosé, and John Zimmerman. 2020. Re-examining whether, why, and how human-AI interaction is uniquely difficult to design. In Proceedings of the 2020 chi conference on human factors in computing systems . 1–13

  90. [98]

    Yuze Zeng, Junze Xiao, Danfeng Li, Jiaxiu Sun, Qingqi Zhang, Ai Ma, Ke Qi, Bin Zuo, and Xiaoqian Liu. 2023. The influence of victim self-disclosure on bystander intervention in cyberbullying. Behavioral Sciences 13, 10 (2023), 829

  91. [99]

    Ruichen Zhang, Hongyang Du, Yinqiu Liu, Dusit Niyato, Jiawen Kang, Sumei Sun, Xuemin Shen, and H Vincent Poor. 2024. Interactive AI with retrieval-augmented generation for next generation networking. IEEE Network (2024). Manuscript submitted to ACM

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.