Pith. sign in

REVIEW 3 major objections 5 minor 143 references

A mixed-reality AI agent that watches a live group conversation and privately coaches one participant on what to say and how to behave can improve that person's engagement and contribution, and can learn from its own suggestions.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 15:01 UTC pith:5XV2IZA5

load-bearing objection ChatMuse is a competent and candid design exploration of MR support for small-group conversation, but its self-improvement claim rests on an unvalidated gaze proxy; RQ1 mostly holds, RQ2 does not. the 3 major comments →

arxiv 2607.18556 v2 pith:5XV2IZA5 submitted 2026-07-20 cs.HC

ChatMuse: Supporting In-Person Small-Group Conversation Experience with a Proactive Assistive AI Agent in Mixed Reality

classification cs.HC
keywords Mixed realityproactive AI agentsmall-group conversationconversational supportnonverbal cuesgaze-based feedbackLLM agentparticipation inequality
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tries to establish that a proactive AI agent in mixed reality can help a participant hold their own in an in-person small-group conversation. ChatMuse listens to what everyone says and watches where they look, then privately suggests example phrases, non-verbal behaviors, and background facts at moments the agent chooses. The system also scores each suggestion by how long the user's gaze rests on it, and uses that score to reshape future support. In six group sessions with 18 participants, the authors report seven benefits and measurable decreases in speaking-time inequality with the feedback loop enabled. The claim matters because it moves conversational AI support from dyadic settings into the harder "many minds" territory of live group interaction.

Core claim

ChatMuse is a mixed-reality system for in-person small-group conversation support. It continuously builds a "group conversation context" containing each participant's speech, tone, speaker turns, and a directed graph of where people are looking. A pre-trained LLM reads this context and decides whether and what to suggest — example phrases, non-verbal behavior cues, background facts — rendered privately as an MR overlay. After a suggestion is shown, the system estimates how useful it was by dividing the user's gaze dwell time on the overlay by the expected reading time of the text, and feeds that score back into the context to improve later suggestions. In a within-subject study with six grou

What carries the argument

The central mechanism is a two-pipeline agentic loop. The feedforward inference pipeline maintains a "group conversation context" — speaker identity, transcribed speech, inferred speech emotion, and a directed attention graph of who looks at whom — serialized into a prompt for a pre-trained LLM, which outputs a proactive-decision flag and support text (max 50 words). The feedback enhancement pipeline then computes a usefulness score u = t / T, where t is the MR user's gaze dwell time on the overlay and T is the expected reading time of the text (60*N/238 seconds), and appends that score to the episode so the LLM can later reference which kinds of support were useful.

Load-bearing premise

ChatMuse's self-improvement relies on how long the user looks at the suggested text being a true measure of how useful that suggestion was; if people stare out of curiosity or confusion instead, the learning loop is chasing the wrong signal.

What would settle it

Run the system with deliberately unhelpful suggestions (e.g., off-topic or random tips) and measure gaze dwell time; if dwell time remains comparable to helpful suggestions, the usefulness score is not measuring usefulness.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • With the feedback loop active, participation inequality fell in all six groups compared to the no-feedback condition.
  • Support adoption — how much the user's later speech resembles the suggested example — rose significantly in four of six groups (Tau-U, p<.05).
  • Most conversation partners could not tell the supported participant was receiving help, although some described the MR user as leading or less genuine.
  • Users reported seven benefits: less friction, summaries, background info, topic-starting, context emphasis, topic extension, and inspiration.
  • No reliable workload or usability differences were found across conditions, so the support came at no measured cognitive cost.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The gaze-based usefulness score is a proxy that could be misled by curiosity or confusion; future work could validate it against self-reported usefulness or behavioral uptake of the suggested content.
  • The proactive rate of roughly 88–90% suggests the agent nearly always chooses to intervene; testing an explicit cost for unwanted interruptions could clarify how to preserve conversation ownership.
  • The same pipeline could generalize beyond 3–4 person, 15-minute sessions if attention graphs scale, but longer conversations would likely require memory mechanisms to keep context coherent.
  • Placing overlays above the person being addressed reduces occlusion but may lag fast gaze shifts; an optimization-based placement algorithm could better handle multi-speaker moments.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. ChatMuse is an MR system that supports one participant in a three-to-four-person in-person conversation by analyzing all participants' verbal and non-verbal cues and proactively rendering private suggestions on speech, non-verbal behavior, and background knowledge. The system is designed from a formative provotype/focus-group study that yields three design considerations, and is evaluated in a counterbalanced within-subject study (six groups, N=18; one MR user per group) with post-session questionnaires, focus groups, and logged agent rationales. The paper reports seven key benefits for engagement/contribution, reduced participation inequality, and significant Tau-U support-adoption effects for four of six groups when the feedback enhancement pipeline is enabled. The contribution is framed as a design exploration of proactive MR conversational support, with the novelty concentrated in the feedback pipeline that uses gaze-derived usefulness to update the LLM's context.

Significance. If the results hold, ChatMuse is a credible demonstration that a proactive LLM-based MR overlay can help a less-engaged participant enter and contribute to a small-group conversation without disrupting partners, and the design considerations from the formative study are practically useful. The paper is commendably transparent: counterbalanced ordering, null SUS/TLX results reported, rationales logged and analyzed (209 vs 271), calibration benchmarking in Appendix C, and a frank limitations section. The qualitative material is rich and triangulated. However, the system's most distinctive claim — closed-loop improvement through the feedback pipeline — is not yet supported, because the sole learning signal (usefulness from gaze duration) is unvalidated and confounded by overlay placement, and the outcome measure for adoption is itself an unvalidated proxy. The effectiveness claim is further limited by a treatment sample of six MR users and null objective usability results. Overall, this is a promising design exploration whose central qualitative findings are plausible, but whose headline mechanism needs reframing or further validation.

major comments (3)
  1. [§4.4–4.5, Fig. 5, Table 1] The usefulness score u = t/T (§4.4) is the only learning signal in the feedback pipeline, but it is unvalidated and structurally confounded by the rendering logic in §4.5: in the common placement the overlay is anchored above the head of the partner the MR user is looking at, so gaze dwell on the overlay can reflect eye-contact/attention to the partner rather than reading of the suggestion; dwell time is also inflated by confusion or readability problems. Because u updates the episode context (Fig. 5) and conditions all subsequent LLM generations, the RQ2 findings (Table 1 Tau-U for G3–G6; rationales such as 'Prior support with facts was useful') do not establish genuine self-improvement; the agent's rationales merely restate the unvalidated signal. No calibration against independent usefulness ratings is presented. Fix: validate u on a sample of episodes with human ratings, or re-analyz
  2. [§5.1, §5.3, Table 1, Appendix F] The central effectiveness claim rests on a much smaller treatment sample than the abstract suggests: one MR user per group (orange highlight in Table 4), i.e., six supported participants; the seven benefit themes (RQ1) are drawn largely from these six participants' focus-group remarks, while the other twelve contributed peripheral observations. The quantitative backbone is also weak: SUS and NASA-TLX showed no significant differences (Appendix F), the participation-inequality reduction in Table 1 is purely descriptive and not monotonic (Baseline already has lower inequality than C_NoFeedback for G1, G4, G5), and Q4/Q5 had a majority reporting no perceived difference. The N=18 framing in the abstract should be corrected and the qualitative N=6 basis acknowledged; the descriptive inequality claims should be supported by an appropriate test or discarded.
  3. [§5.2] The support adoption score (max cosine similarity between suggested example speech and the MR user's subsequent utterances, all-MiniLM-L6-v2) is itself an unvalidated proxy for 'adoption.' Since the feedback pipeline's generation is conditioned on the same conversational context — and on the possibly invalid u — an increase in Tau-U (Table 1) could reflect the suggestions increasingly echoing the user's own phrasing or the metric's properties rather than improved support quality. The paper should validate the measure (e.g., against manual coding of whether the user used the suggestion) or discuss this confound explicitly. Also specify how episodes with no user utterance before the next support rendering are scored.
minor comments (5)
  1. [§6.2] The limitations list omits the main validity gap: the gaze-based usefulness proxy of §4.4 is never calibrated against any independent usefulness rating. Given its load-bearing role in RQ2, it should be acknowledged here.
  2. [§5.1] State explicitly that the supported condition was experienced by one participant per group (six total, orange in Table 4); the N=18 framing in the abstract and conclusion obscures the treatment sample size.
  3. [§5.3] Proactive rates of 87.98%/89.91% are reported without a test or discussion. With the agent providing support in ~90% of invocations, the 'proactive decision' rarely withholds, which sits uneasily with DC1 and participants' reports of both unnecessary and missing support; a brief analysis would help.
  4. [Table 1] Clarify the definition of participation inequality (%) and report effect sizes or confidence intervals for the Tau-U indices; also specify how the adoption score is computed when the MR user produced no utterance before the next support rendering.
  5. [§5.3, RQ1] Partner perceptions ('presenter,' 'less genuine,' 'artificial while transitioning topics') are reported but not analyzed; a short paragraph on the risk of over-reliance/inauthenticity would balance the claimed benefits.

Circularity Check

0 steps flagged

No significant circularity: the feedback-pipeline claims are measured with independent speech-based metrics, and the lone self-citation is background, not load-bearing.

full rationale

The paper's derivation chain is not circular. The feedback loop defines usefulness as u = t/T (Sec. 4.4) and feeds u back into the episode context, but the RQ2 claims are not evaluated with u. RQ2 uses (i) participation inequality, the standard deviation of speaking-time proportions, and (ii) support adoption score, the maximum cosine similarity between the generated example speech and the MR user's subsequent utterances, tested with Tau-U (Sec. 5.2, Table 1). These metrics are computed from speech/transcripts, not from gaze duration, so improvement in them is not equivalent to u by construction. The only self-citation ([133], Related Work: "Zhou et al. [133] demonstrated the need for and value of providing private information support to facilitate in-person group conversations") is a background statement about prior needs research and is not load-bearing; the present paper independently runs a formative study (Sec. 3) that produces the design considerations that drive the system. No uniqueness theorem, no fitted-parameter-renamed-as-prediction, and no renaming pattern is present. The unvalidated nature of the gaze proxy is an internal-validity threat, and the paper itself acknowledges related sensing limitations in Sec. 6.2, but that is a correctness/validity concern, not circularity.

Axiom & Free-Parameter Ledger

3 free parameters · 5 axioms · 0 invented entities

No new physical or theoretical entities are introduced. The invented 'entity' is a software agent and a usefulness metric, both of which are configurations of existing components. The ledger instead captures the modeling assumptions that the system's effectiveness depends on.

free parameters (3)
  • Gaze/attention detection threshold = 30 degrees
    Section 4.3 sets a 30-degree horizontal angle as the condition for one participant looking at another; this hand-set threshold defines the attention graph that feeds the LLM context.
  • Reading-speed constant in usefulness score = wpm = 238
    Section 4.4 uses 238 words per minute from Brysbaert [20] to compute expected reading time T; not fitted to study data, but a model choice that scales the usefulness score.
  • Maximum support output length = 50 words
    Section 4.2 limits generated support to 50 words as an empirical design constraint; it affects readability and cognitive load but is not a core statistical parameter.
axioms (5)
  • domain assumption Headpose direction approximates gaze direction for non-MR participants.
    Section 4.1 uses forward-facing headpose as a proxy for gaze, citing Jha and Busso [61] for horizontal-accuracy; the attention graph is built on this proxy.
  • domain assumption Gaze duration on the support overlay reflects how useful the support was.
    Section 4.4 defines usefulness u = t/T and justifies it with the general claim that longer viewing reflects sustained attention; no task-specific validation is provided.
  • domain assumption Embedding cosine similarity between suggested example speech and subsequent user speech measures actual adoption of the support.
    Section 5.2 defines support adoption score as the maximum cosine similarity using all-MiniLM-L6-v2; this is a proxy, not a verified measure of behavioral adoption.
  • domain assumption Speech-emotion labels from a pretrained classifier are accurate enough for the agent's context.
    Section 4.3 relies on a pretrained speech emotion classification model with reported benchmark accuracy 83.79%, and on Ekman's six basic emotions; no in-situ accuracy check is reported.
  • domain assumption LLM-generated proactive decisions and textual rationales are meaningful evidence of conversational grounding.
    Section 5.3 treats the agent's rationales as supporting evidence for RQ2; this assumes the LLM's stated reasons reflect the actual inputs rather than confabulation.

pith-pipeline@v1.3.0-alltime-deepseek · 31818 in / 11081 out tokens · 131195 ms · 2026-08-01T15:01:27.755846+00:00 · methodology

0 comments
read the original abstract

In-person small-group conversations occur across nearly every aspect of daily life and play a crucial role in social interaction. However, achieving effective in-person group conversations can be challenging and cognitively demanding. While recent Mixed Reality (MR) headsets show promise as a conversational support system by presenting relevant information through overlays, it remains unclear how such supporting information should be designed and generated for in-person group conversations. We propose ChatMuse, a novel MR-based proactive assistive system for in-person small-group conversation experience. ChatMuse analyzes verbal and non-verbal cues from all conversation participants and proactively provides real-time guidance on the user's verbal and non-verbal behaviors. The behavioral responses of the supported users are then used to improve ChatMuse's support capabilities in subsequent interactions. We conducted a within-subject study to evaluate and demonstrate the feasibility and effectiveness of ChatMuse in assisting users to engage in and contribute to in-person small-group conversations. Our research around ChatMuse represents a design exploration of a new interaction space that investigates the feasibility of supporting in-person small-group conversations through a proactive assistive AI agent in MR.

Figures

Figures reproduced from arXiv: 2607.18556 by Chen Chen, Christine Lisetti, Diana Nelly Rivera Rodriguez, Janet G. Johnson, Joaquin Frangi, Lingyao Li, Rawan Alghofaili, Renkai Ma, Shaoze Zhou.

Figure 1
Figure 1. Figure 1: ChatMuse provides real-time, MR information support for in-person small-group conversations; (a) a third-person [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: (a) FPV of provotypes in the formative study. (b) [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: ChatMuse’s agentic pipeline design. DC2 Supporting information should preserve conversa￾tional ownership while integrating verbal and nonverbal contexts to offer summaries, clarifications, overlooked as￾pects, and relevant others’ perspectives. The suggestive infor￾mation should be context-aware and depend on the conversation participants. The type of information may vary during different phases of the con… view at source ↗
Figure 4
Figure 4. Figure 4: Setup of ChatMuse. (a) TPV with three participants; [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figure 6
Figure 6. Figure 6: Placements of virtually rendered support within [PITH_FULL_IMAGE:figures/full_fig_p007_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Post-study survey responses from (a) MR users and [PITH_FULL_IMAGE:figures/full_fig_p009_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Tracking performance of the RGB-D camera and [PITH_FULL_IMAGE:figures/full_fig_p016_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Survey responses of SUS and NASA TLX. (a) Responses of SUS from MR users. (b) Responses of NASA TLX from [PITH_FULL_IMAGE:figures/full_fig_p019_9.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

143 extracted references · 21 canonical work pages · 1 internal anchor

  1. [1]

    http://fac-staff.seattleu

    2002.Interpersonal-Communication-Skills-Inventory. http://fac-staff.seattleu. edu/thompson/web/communication/communication_skills_inventory.pdf Ac- cessed on January 8, 2026

  2. [2]

    https://openai.com/index/gpt-4-1 Accessed on March 16, 2026

    2024.GPT 4.1. https://openai.com/index/gpt-4-1 Accessed on March 16, 2026

  3. [3]

    https://platform.openai.com/docs/models/gpt-4o-mini Accessed on January 21, 2026

    2024.GPT-4o-mini. https://platform.openai.com/docs/models/gpt-4o-mini Accessed on January 21, 2026

  4. [4]

    https://openai.com/index/introducing-gpt-live/ Accessed on July 14, 2026

    2026.GPT-Live. https://openai.com/index/introducing-gpt-live/ Accessed on July 14, 2026

  5. [5]

    https://optitrack.com Accessed on March 18, 2026

    2026.OptiTrack. https://optitrack.com Accessed on March 18, 2026

  6. [6]

    Yuvraj Agarwal, Christopher Harrison, Gierad Laput, Sudershan Boovaragha- van, Chen Chen, Abhijit Hota, Bo Robert Xiao, and Yang Zhang. 2019. Virtual Sensor System. https://patents.google.com/patent/US20200033163A1/en US Patent 10436615

  7. [7]

    Yuvraj Agarwal, Christopher Harrison, Gierad Laput, Sudershan Boovaragha- van, Chen Chen, Abhijit Hota, Bo Robert Xiao, and Yang Zhang. 2020. Virtual Sensor System. https://patents.google.com/patent/US20200033163A1/en US Patent Application 16591987

  8. [8]

    Allen, Curry I

    James F. Allen, Curry I. Guinn, and Eric Horvtz. 1999. Mixed-Initiative In- teraction.IEEE Intelligent Systems and their Applications14, 5 (1999), 14–23. https://doi.org/10.1109/5254.796083

  9. [9]

    Sumit Asthana, Sagi Hilleli, Pengcheng He, and Aaron Halfaker. 2025. Sum- maries, Highlights, and Action Items: Design, Implementation and Evaluation of an LLM-Powered Meeting Recap System.Proceedings of the ACM on Human- Computer Interaction9, 2 (may 2025), 1–29. https://doi.org/10.1145/3711074

  10. [10]

    Robert F Bales. 1950. Interaction Process Analysis: A Method for the Study of Small Groups. (1950). https://doi.org/10.2307/2572427

  11. [11]

    Jad Bendarkawi, Ashley Ponce, Sean Chidozie Mata, Aminah Aliu, Yuhan Liu, Lei Zhang, Amna Liaqat, Varun Nagaraj Rao, and Andrés Monroy-Hernández

  12. [12]

    Indrani Bhattacharya, Michael Foley, Ni Zhang, Tongtao Zhang, Christine Ku, Cameron Mine, Heng Ji, Christoph Riedl, Brooke Foucault Welles, and Richard J. Radke. 2018. A Multimodal-Sensor-Enabled Room for Unobtrusive Group Meet- ing Analysis. InProceedings of the 20th ACM International Conference on Multi- modal Interaction(Boulder, CO, USA)(ICMI ’18). As...

  13. [13]

    Millard J Bienvenu Sr. 1971. An Interpersonal Communication Inventory.Jour- nal of communication21, 4 (1971), 381–388. https://doi.org/j.1460-2466.1971. tb02937.x

  14. [14]

    Laurens Boer and Jared Donovan. 2012. Provotypes for Participatory Innovation. InProceedings of the Designing Interactive Systems Conference(Newcastle Upon Tyne, United Kingdom)(DIS ’12). Association for Computing Machinery, New York, NY, USA, 388–397. https://doi.org/10.1145/2317956.2318014

  15. [15]

    Sudershan Boovaraghavan, Chen Chen, Anurag Maravi, Mike Czapik, Yang Zhang, Chris Harrison, and Yuvraj Agarwal. 2023. Mites: Design and Deploy- ment of a General-Purpose Sensing Infrastructure for Buildings.Proc. ACM Interact. Mob. Wearable Ubiquitous Technol.7, 1, Article 2 (mar 2023), 32 pages. https://doi.org/10.1145/3580865

  16. [16]

    Lenore A Boyd and Arthur J Roach. 1977. Interpersonal Communication Skills Differentiating More Satisfying From Less Satisfying Marital Relationships. Journal of Counseling Psychology24, 6 (1977), 540. https://doi.org/10.1037/0022- 0167.24.6.540

  17. [17]

    Virginia Braun and Victoria Clarke. 2006. Using Thematic Analysis in Psy- chology.Qualitative research in psychology3, 2 (2006), 77–101. https: //doi.org/10.1191/1478088706qp063oa

  18. [18]

    2012.Thematic Analysis

    Virginia Braun and Victoria Clarke. 2012.Thematic Analysis. American Psy- chological Association. https://doi.org/10.1037/13620-004

  19. [19]

    John Brooke. 1996. SUS - A Quick and Dirty Usability Scale.Usability Evaluation in Industry189, 194 (1996), 4–7

  20. [20]

    Marc Brysbaert. 2019. How Many Words Do We Read Per Minute? A Review and Meta-Analysis of Reading Rate.Journal of memory and language109 (2019), 104047. https://doi.org/10.1016/j.jml.2019.104047

  21. [21]

    Runze Cai, Nuwan Nanayakkarawasam Peru Kandage Janaka, Shengdong Zhao, and Minghui Sun. 2023. ParaGlassMenu: Towards Social-Friendly Subtle In- teractions in Conversations. InProceedings of the 2023 CHI Conference on Hu- man Factors in Computing Systems(Hamburg, Germany)(CHI ’23). Associa- tion for Computing Machinery, New York, NY, USA, Article 721, 21 p...

  22. [22]

    Marco Caliendo, Deborah A Cobb-Clark, Juliana Silva-Goncalves, and Arne Uhlendorff. 2024. Locus of Control and the Preference for Agency.European Economic Review165 (2024), 104737. https://doi.org/10.1016/j.euroecorev.2024. 104737

  23. [23]

    Chen Chen, Cuong Nguyen, Jane Hoffswell, Jennifer Healey, Trung Bui, and Nadir Weibel. 2023. PaperToPlace: Transforming Instruction Documents into Spatialized and Context-Aware Mixed Reality Experiences. InProceedings of the 36th Annual ACM Symposium on User Interface Software and Technology(San Francisco, CA, USA)(UIST ’23). Association for Computing Mac...

  24. [24]

    Chen Chen, Ke Sun, and Xinyu Zhang. 2020. CapTag: Toward Printable Ubiq- uitous Internet of Things: Poster Abstract. InProceedings of the 18th Confer- ence on Embedded Networked Sensor Systems(Virtual Event, Japan)(SenSys ’20). Association for Computing Machinery, New York, NY, USA, 669–670. https://doi.org/10.1145/3384419.3430410

  25. [25]

    Xinyue Chen, Lev Tankelevitch, Rishi Vanukuru, Ava Elizabeth Scott, Payod Panda, and Sean Rintel. 2025. Are We On Track? AI-Assisted Active and Passive Goal Reflection During Meetings. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). Association for Computing Machinery, New York, NY, USA, Article 705, 22 pages. htt...

  26. [26]

    Yifei Cheng, Yukang Yan, Xin Yi, Yuanchun Shi, and David Lindlbauer. 2021. SemanticAdapt: Optimization-Based Adaptation of Mixed Reality Layouts Lever- aging Virtual-Physical Semantic Connections. InThe 34th Annual ACM Sympo- sium on User Interface Software and Technology(Virtual Event, USA)(UIST ’21). Association for Computing Machinery, New York, NY, US...

  27. [27]

    Jun-Ho Choi, Marios Constantinides, Sagar Joglekar, and Daniele Quercia. 2021. KAIROS: Talking Heads and Moving Bodies for Successful Meetings. InPro- ceedings of the 22nd International Workshop on Mobile Computing Systems and Applications(Virtual, United Kingdom)(HotMobile ’21). Association for Com- puting Machinery, New York, NY, USA, 30–36. https://doi...

  28. [28]

    Marios Constantinides, Sanja Šćepanović, Daniele Quercia, Hongwei Li, Ugo Sassi, and Michael Eggleston. 2020. ComFeel: Productivity is a Matter of the Senses Too.Proc. ACM Interact. Mob. Wearable Ubiquitous Technol.4, 4, Article 123 (dec 2020), 21 pages. https://doi.org/10.1145/3432234

  29. [29]

    Gus Cooney, Adam M Mastroianni, Nicole Abi-Esber, and Alison Wood Brooks

  30. [30]

    Daft and Robert H

    Richard L. Daft and Robert H. Lengel. 1986. Organizational Information Re- quirements, Media Richness and Structural Design.Manage. Sci.32, 5 (May 1986), 554–571. https://doi.org/10.1287/mnsc.32.5.554

  31. [31]

    Ionut Damian, Chiew Seng (Sean) Tan, Tobias Baur, Johannes Schöning, Kris Luyten, and Elisabeth André. 2015. Augmenting Social Interactions: Realtime Behavioural Feedback using Social Signal Processing Techniques. InProceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems (Seoul, Republic of Korea)(CHI ’15). Association for Comp...

  32. [32]

    Shakiba Davari and Doug A. Bowman. 2024. Towards Context-Aware Adap- tation in Extended Reality: A Design Space for XR Interfaces and an Adap- tive Placement Strategy.arXiv preprint arXiv: 2411.02607(2024). https: //doi.org/10.48550/arXiv.2411.02607

  33. [33]

    Duchowski

    Andrew T. Duchowski. 2017.Eye Tracking Methodology. Springer International Publishing, Cham. https://doi.org/10.1007/978-3-319-57883-5

  34. [34]

    Paul Ekman. 1992. Are There Basic Emotions?Psychological Review99, 3 (1992), 550–553. https://doi.org/10.1037/0033-295X.99.3.550

  35. [35]

    Paul Ekman and Wallace V Friesen. 1971. Constants Across Cultures in the Face and Emotion.Journal of personality and social psychology17, 2 (1971), 124. https://doi.org/10.1037/h0030377

  36. [36]

    Satu Elo and Helvi Kyngäs. 2008. The Qualitative Content Analysis Process. Journal of Advanced Nursing62, 1 (2008), 107–115. https://doi.org/10.1111/j. 1365-2648.2007.04569.x

  37. [37]

    Sayed M Elsayed-Elkhouly, Harold Lazarus, and Volville Forsythe. 1997. Why is a Third of Your Time Wasted in Meetings?Journal of Management Development 16, 9 (1997), 672–676. https://doi.org/10.1108/02621719710190185

  38. [38]

    Gerhard Fischer. 2012. Context-Aware Systems: The ‘Right’ Information, at the ‘Right’ Time, in the ‘Right’ Place, in the ‘Right’ Way, to the ‘Right’ Person. In Proceedings of the International Working Conference on Advanced Visual Interfaces (Capri Island, Italy)(A VI ’12). Association for Computing Machinery, New York, NY, USA, 287–294. https://doi.org/1...

  39. [39]

    1950.Statistical Methods for Research Workers

    Ronald Aylmer Fisher. 1950.Statistical Methods for Research Workers. Oliver and Boyd, Edinburgh. xv + 354 pp. pages. Issue llth ed. revised. https://doi.org/ 10.1038/139737d0

  40. [40]

    2025.FIU-ICA VE: Integrated Computer Augmented Virtual Environment

    Florida International University (FIU). 2025.FIU-ICA VE: Integrated Computer Augmented Virtual Environment. https://icave.fiu.edu/ Accessed on December 3, 2025

  41. [41]

    Milton Friedman. 1937. The Use of Ranks to Avoid the Assumption of Normality Implicit in the Analysis of Variance.J. Amer. Statist. Assoc.32, 200 (1937), 675–

  42. [42]

    Jenny Xiyu Fu, Brennan Antone, Kowe Kadoma, and Malte Jung. 2025. Large Language Model Use Impact Locus of Control.arXiv preprint arXiv: 2505.11406 ChatMuse (2025). https://doi.org/10.48550/arXiv.2505.11406

  43. [43]

    Szu-Wei Fu, Yaran Fan, Yasaman Hosseinkashi, Jayant Gupchup, and Ross Cutler

  44. [44]

    Yuichiro Fujimoto. 2025. ChatAR: Conversation Support using Large Language Model and Augmented Reality.arXiv preprint arXiv: 2506.16008(2025). https: //doi.org/10.48550/arXiv.2506.16008

  45. [45]

    Daniel Gatica-Perez. 2006. Analyzing Group Interactions in Conversations: A Review. In2006 IEEE International Conference on Multisensor Fusion and Integration for Intelligent Systems. 41–46. https://doi.org/10.1109/MFI.2006. 265658

  46. [46]

    Rahul Gavas, Debatri Chatterjee, and Aniruddha Sinha. 2017. Estimation of Cognitive Load Based on the Pupil Size Dilation. In2017 IEEE international conference on systems, man, and cybernetics (SMC). IEEE, 1499–1504. https: //doi.org/10.1109/SMC.2017.8122826

  47. [47]

    Zixuan Guo and Tomoo Inoue. 2019. Using a Conversational Agent to Facili- tate Non-Native Speaker’s Active Participation in Conversation. InExtended Abstracts of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow, Scotland Uk)(CHI EA ’19). Association for Computing Machinery, New York, NY, USA, 1–6. https://doi.org/10.1145/3290607.3313075

  48. [48]

    Ki-Won Haan, Christoph Riedl, and Anita Woolley. 2021. Discovering Where We Excel: How Inclusive Turn-Taking in Conversation Improves Team Per- formance. InCompanion Publication of the 2021 International Conference on Multimodal Interaction(Montreal, QC, Canada)(ICMI ’21 Companion). As- sociation for Computing Machinery, New York, NY, USA, 278–283. https:...

  49. [49]

    Hart and Lowell E

    Sandra G. Hart and Lowell E. Staveland. 1988. Development of NASA-TLX (Task Load Index): Results of Empirical and Theoretical Research. InHuman Mental Workload, Peter A. Hancock and Najmedin Meshkati (Eds.). Advances in Psychology, Vol. 52. North-Holland, 139–183. https://doi.org/10.1016/S0166- 4115(08)62386-9

  50. [50]

    Uri Hasson, Asif A Ghazanfar, Bruno Galantucci, Simon Garrod, and Christian Keysers. 2012. Brain-to-Brain Coupling: A Mechanism for Creating and Sharing a Social World.Trends in cognitive sciences16, 2 (2012), 114–121. https: //doi.org/10.1016/j.tics.2011.12.007

  51. [51]

    Jess Hohenstein and Malte Jung. 2018. AI-Supported Messaging: An Inves- tigation of Human-Human Text Conversation with AI Support. InExtended Abstracts of the 2018 CHI Conference on Human Factors in Computing Systems (Montreal QC, Canada)(CHI EA ’18). Association for Computing Machinery, New York, NY, USA, 1–6. https://doi.org/10.1145/3170427.3188487

  52. [52]

    Eric Horvitz. 1999. Principles of Mixed-Initiative User Interfaces. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems(Pittsburgh, Pennsylvania, USA)(CHI ’99). Association for Computing Machinery, New York, NY, USA, 159–166. https://doi.org/10.1145/302979.303030

  53. [53]

    Ross, Dario An- dres Silva Moran, Gabriel Enrique Gonzalez, Siya Kunde, Morgan A

    Stephanie Houde, Kristina Brimijoin, Michael Muller, Steven I. Ross, Dario An- dres Silva Moran, Gabriel Enrique Gonzalez, Siya Kunde, Morgan A. Foreman, and Justin D. Weisz. 2025. Controlling AI Agent Participation in Group Conver- sations: A Human-Centered Approach. InProceedings of the 30th International Conference on Intelligent User Interfaces (IUI ’...

  54. [54]

    Wei-Chieh Huang, Weizhi Zhang, Yueqing Liang, Yuanchen Bei, Yankai Chen, Tao Feng, Xinyu Pan, Zhen Tan, Yu Wang, Tianxin Wei, Shanglin Wu, Ruiyao Xu, Liangwei Yang, Rui Yang, Wooseong Yang, Chin-Yuan Yeh, Hanrong Zhang, Haozhen Zhang, Siqi Zhu, Henry Peng Zou, Wanjia Zhao, Song Wang, Wu- jiang Xu, Zixuan Ke, Zheng Hui, Dawei Li, Yaozu Wu, Langzhou He, Che...

  55. [55]

    2023.Learn about Eye Tracking on Meta Quest Pro

    Meta Inc. 2023.Learn about Eye Tracking on Meta Quest Pro. https://www.meta. com/help/quest/8107387169303764 Accessed on February 18, 2026

  56. [56]

    2023.Meta Quest Pro

    Meta Inc. 2023.Meta Quest Pro. https://www.meta.com/quest/quest-pro Ac- cessed on October 9, 2025

  57. [57]

    2024.Meta Quest 3S

    Meta Inc. 2024.Meta Quest 3S. https://www.meta.com/quest/quest-3s Accessed on January 18, 2026

  58. [58]

    Shivesh Jadon, Mehrad Faridan, Edward Mah, Rajan Vaish, Wesley Willett, and Ryo Suzuki. 2024. Augmented Conversation with Embedded Speech-Driven On-the-Fly Referencing in AR.arXiv preprint arXiv: 2405.18537(2024). https: //doi.org/10.48550/arXiv.2405.18537

  59. [59]

    Maurice Jakesch, Advait Bhat, Daniel Buschek, Lior Zalmanson, and Mor Naa- man. 2023. Co-Writing with Opinionated Language Models Affects Users’ Views. InProceedings of the 2023 CHI Conference on Human Factors in Computing Sys- tems(Hamburg, Germany)(CHI ’23). Association for Computing Machinery, New York, NY, USA, Article 111, 15 pages. https://doi.org/1...

  60. [60]

    Nuwan Janaka, Chloe Haigh, Hyeongcheol Kim, Shan Zhang, and Sheng- dong Zhao. 2022. Paracentral and Near-Peripheral Visualizations: Towards Attention-Maintaining Secondary Information Presentation on OHMDs Dur- ing In-Person Social Interactions. InProceedings of the 2022 CHI Conference on Human Factors in Computing Systems(New Orleans, LA, USA)(CHI ’22). ...

  61. [61]

    Sumit Jha and Carlos Busso. 2016. Analyzing the Relationship between Head Pose and Gaze to Model Driver Visual Attention. In2016 IEEE 19th International Conference on Intelligent Transportation Systems (ITSC). 2157–2162. https: //doi.org/10.1109/ITSC.2016.7795905

  62. [62]

    Johnson, Macarena Peralta, Mansanjam Kaur, Ruijie Sophia Huang, Sheng Zhao, Ruijia Guan, Shwetha Rajaram, and Michael Nebeling

    Janet G. Johnson, Macarena Peralta, Mansanjam Kaur, Ruijie Sophia Huang, Sheng Zhao, Ruijia Guan, Shwetha Rajaram, and Michael Nebeling. 2025. Ex- ploring Collaborative GenAI Agents in Synchronous Group Settings: Elic- iting Team Perceptions and Design Considerations for the Future of Work. Proc. ACM Hum.-Comput. Interact.9, 7, Article CSCW414 (oct 2025),...

  63. [63]

    Just and Patricia A

    Marcel A. Just and Patricia A. Carpenter. 1980. A Theory of Reading: From Eye Fixations to Comprehension.Psychological Review87, 4 (1980), 329–354. https://doi.org/10.1037/0033-295X.87.4.329

  64. [64]

    Vaiva Kalnikaitundefined, Patrick Ehlen, and Steve Whittaker. 2012. Markup as You Talk: Establishing Effective Memory Cues While Still Contributing to a Meeting. InProceedings of the ACM 2012 Conference on Computer Supported Cooperative Work(Seattle, Washington, USA)(CSCW ’12). Association for Com- puting Machinery, New York, NY, USA, 349–358. https://doi...

  65. [65]

    Carolyn M Keeler and R Kirk Steinhorst. 1995. Using Small Groups to Promote Active Learning in the Introductory Statistics Course: A Report from the Field. Journal of Statistics Education3, 2 (1995). https://doi.org/10.1080/10691898.1995. 11910485

  66. [66]

    1990.Conducting Interaction: Patterns of Behavior in Focused Encounters

    Adam Kendon. 1990.Conducting Interaction: Patterns of Behavior in Focused Encounters. Vol. 7. CUP Archive. https://doi.org/10.1017/S0272263100011700

  67. [67]

    Adam Kendon. 1990. Spatial Organization in Social Encounters: The F- Formation System.Conducting interaction: Patterns of behavior in focused encounters(1990)

  68. [68]

    Soomin Kim, Jinsu Eun, Changhoon Oh, Bongwon Suh, and Joonhwan Lee

  69. [69]

    Steffi Kohl, André Calero Valdez, and Kay Schröder. 2023. Using Speech Con- tribution Visualization to Improve Team Performance of Divergent Thinking Tasks. InProceedings of the 15th Conference on Creativity and Cognition(Virtual Event, USA)(C&C ’23). Association for Computing Machinery, New York, NY, USA, 319–324. https://doi.org/10.1145/3591196.3596824

  70. [70]

    Robert M Krauss, Robert A Dushay, Yihsiu Chen, and Frances Rauscher. 1995. The Communicative Value of Conversational Hand Gesture.Journal of exper- imental social psychology31, 6 (1995), 533–552. https://doi.org/10.1006/jesp. 1995.1024

  71. [71]

    Robert E Kraut, Robert S Fish, Robert W Root, Barbara L Chalfonte, et al. 1990. Informal Communication in Organizations: Form, Function, and Technology. InHuman reactions to technology: Claremont symposium on applied social psy- chology, Vol. 145. 1–20

  72. [72]

    2014.Focus Groups: A Practical Guide for Applied Research (5th ed.)

    Richard A Krueger. 2014.Focus Groups: A Practical Guide for Applied Research (5th ed.). Sage publications, Thousand Oaks, CA. https://doi.org/10.1080/ 03098260600927575

  73. [73]

    InProceedings of the 2020 CHI Conference on Human Factors in Computing Systems(Honolulu, HI, USA)(CHI ’20)

    Bot in the Bunch: Facilitating Group Chat Discussion by Improving Efficiency and Participation with a Chatbot. InProceedings of the 2020 CHI Conference on Human Factors in Computing Systems(Honolulu, HI, USA)(CHI ’20). Association for Computing Machinery, New York, NY, USA, 1–13. https: //doi.org/10.1145/3313831.3376785

  74. [74]

    2010.Research Methods in Human-Computer Interaction

    Jonathan Lazar, Jinjuan Heidi Feng, and Harry Hochheiser. 2010.Research Methods in Human-Computer Interaction. Wiley Publishing. https://doi.org/10. 5555/1841406

  75. [75]

    Soohwan Lee, Seoyeong Hwang, Dajung Kim, and Kyungho Lee. 2025. Conver- sational Agents as Catalysts for Critical Thinking: Challenging Social Influence in Group Decision-Making. InProceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems (CHI EA ’25). Associa- tion for Computing Machinery, New York, NY, USA, Articl...

  76. [76]

    SooHwan Lee, Mingyu Kim, Seoyeong Hwang, Dajung Kim, and Kyungho Lee. 2025. Amplifying Minority Voices: AI-Mediated Devil’s Advocate Sys- tem for Inclusive Group Decision-Making. InCompanion Proceedings of the 30th International Conference on Intelligent User Interfaces (IUI ’25 Companion). Association for Computing Machinery, New York, NY, USA, 17–21. ht...

  77. [77]

    Joanne Leong, John Tang, Edward Cutrell, Sasa Junuzovic, Gregory Paul Barib- ault, and Kori Inkpen. 2024. Dittos: Personalized, Embodied Agents That Partici- pate in Meetings When You Are Unavailable.Proc. ACM Hum.-Comput. Interact. 8, CSCW2, Article 494 (Nov. 2024), 28 pages. https://doi.org/10.1145/3687033

  78. [78]

    Yining Lang, Wei Liang, and Lap-Fai Yu. 2019. Virtual Agent Positioning Driven by Scene Semantics in Mixed Reality. In2019 IEEE Conference on Virtual Reality and 3D User Interfaces (VR). 767–775. https://doi.org/10.1109/VR.2019.8798018

  79. [79]

    Lizi Liao, Grace Hui Yang, and Chirag Shah. 2023. Proactive Conversational Agents in the Post-ChatGPT World. InProceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval (Taipei, Taiwan)(SIGIR ’23). Association for Computing Machinery, New York, NY, USA, 3452–3455. https://doi.org/10.1145/3539618.3594250

  80. [80]

    Tianjian Liu, Hongzheng Zhao, Yuheng Liu, Xingbo Wang, and Zhenhui Peng

Showing first 80 references.