Pith. sign in

REVIEW 3 major objections 5 minor 84 references

SocialMind: LLM-based Proactive AR Social Assistive System with Human-like Perception for In-situ Live Interactions

T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read SocialMind claims that a proactive AR assistant that reads facial expressions, gestures, and personas in real time can lift live-conversation engagement by 38.3%.

desk verdict The system is real, but the headline engagement gain is an artifact of mismatched evaluation inputs. read the letter →

arxiv 2412.04036 v1 pith:53T7XVCX submitted 2024-12-05 cs.AI

classification cs.AI
keywords socialassistivesystemsaugmentedrealitylargelanguagemodelsnonverbalcueperceptionproactiveassistanceimplicitpersonaadaptationliveinteractionsmartglasses
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

SocialMind is a proposed system that uses AR glasses and large language models to give a person live suggestions during face-to-face conversations. The paper's central claim is that by perceiving nonverbal cues, social context, and the implicit interests of both speakers, the system can generate timely, personalized social suggestions that raise conversational engagement. It reports that on three public dialogue datasets, SocialMind achieves 38.3% higher engagement and 38.7% higher personalization than reactive text-only baselines, and that 95% of 20 user-study participants would use it in live interactions. A sympathetic reader would take the paper as showing that proactive, context-aware assistance during live conversation is both technically feasible and genuinely useful.

What carries the argument

The central mechanism is the multi-tier collaborative suggestion generation strategy: a social factor-aware cache stores pairs of conversational utterances and corresponding suggestions, grouped by social factors, so that familiar situations are answered in about 50 ms, while cache misses trigger deeper LLM reasoning. An intention-infer-based strategy processes partial utterances every two seconds to prepare suggestions before the partner finishes speaking, and a proactive update mechanism refreshes the display only when the conversation's meaning changes. The perception stack that feeds this machinery combines pose and facial tracking for nonverbal cues, vibration-based primary-user identification, and LLM-based extraction of implicit personas from historical conversations.

What would settle it

Record a set of live conversations with the glasses' camera and microphone, have human coders label the partner's facial expression, gesture, and distance every few seconds, and compare those labels to the system's automatically produced cues; if agreement falls substantially below the level needed to support the suggested responses, the central claim of human-like perception fails.

Watch

Extended reading notes

Core claim

The paper argues that live social interactions can be assisted in situ by a proactive system that combines human-like perception with LLM reasoning. SocialMind extracts verbal and nonverbal cues from glasses-mounted sensors, parses social factors such as relation, formality, and location, and adapts to the implicit personas of both parties learned from prior conversations. These cues are fed into a multi-tier generation strategy that returns short bullet-point suggestions with example sentences on AR glasses. The reported outcome is that this pipeline produces more personalized, engaging, and nonverbal-aware suggestions than zero-shot prompting, chain-of-thought prompting, and a specialized social-assistant retrieval baseline, while keeping latency low enough not to disrupt the natural flow of conversation.

Load-bearing premise

The load-bearing premise is that the on-glasses perception stack reliably extracts facial expressions, gestures, and proximity from real noisy sensor streams; the evaluation feeds hand-selected cues in simulation and reports only latency and power in the live test.

Editorial extensions

If this is right

  • If the reported numbers hold, a proactive glasses-based assistant can raise conversational engagement by 38.3% and personalization by 38.7% over reactive text-only assistants, as measured by an LLM-based judge on three public datasets.
  • The social factor-aware cache improves matching accuracy by 4.6% over a generic semantic cache, while cutting LLM input tokens by 26.8% and output tokens by 31.4% at a cache size of 300 and threshold of 0.95, making live suggestions affordable in practice.
  • The vibration-based primary-user identification yields a 32.3% lower false reject rate and 12.1% higher success rate than a volume-based approach, supporting a privacy-preserving way to know who is speaking.
  • The system runs under 2 W on off-the-shelf AR glasses, supports about 70 minutes of use, and keeps cache latency near 50 ms and LLM latency near 2.8 s, which the paper argues is acceptable for real-time in-situ assistance.
  • In a 20-participant user study, 95% expressed willingness to use SocialMind in live interactions, and roughly 80% were willing to converse with someone else using the same system.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: If this works as claimed, the same architecture could extend to multi-party conversations by tracking multiple speakers through the camera view, a direction the paper lists as future work.
  • Editorial inference: The paper's strongest untested assumption is the real-world accuracy of nonverbal cue perception, which could be checked directly by comparing automatic cue extraction against human-coded labels on recorded live conversations.
  • Editorial inference: A natural next test is whether the 38.3% engagement gain holds when the conversation partner is not an LLM agent but an unscripted human, since the current quantitative evaluation uses simulated partners.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. SocialMind is an AR-glasses-based proactive social assistive system that extracts verbal/nonverbal cues, social factors, and implicit personas from multi-modal sensors, and uses an LLM with a social-factor-aware cache and intention-infer reasoning to generate and display in-situ social suggestions. The authors report a 38.3% engagement improvement over baselines based on an LLM-as-a-Judge evaluation on three public dialogue datasets with simulated role-play, and a 20-participant user study reporting 95% willingness to use the system.

Significance. If the engagement improvements were credible, SocialMind would be a valuable step toward practical in-situ social AI assistance: it addresses a real gap, is motivated by a 60-person survey, and includes a functional prototype with measured system latency (<70 ms perception, 2.8 s LLM, 50 ms cache) and power (<2 W). The vibration-based primary-user identification is a clever and well-measured contribution. However, the current evaluation does not establish the central engagement claim, so the paper's significance is conditional on a revised, bias-controlled evaluation.

major comments (3)
  1. [§5.1.2 / §5.1.4 / Figure 28] The headline 38.3% engagement gain is an artifact of the evaluation setup. Section 5.1.4 defines Engagement as whether suggestions 'consider the conversational partner's implicit personas,' and the LLM judge prompt (Figure 28) supplies ground-truth partner personas and instructs the judge to score exactly this dimension. In the Section 5.1.2 simulation, only SocialMind receives persona cues and randomly selected nonverbal cues (Figure 24); the Zero-shot, CoT, and Tianji baselines receive only dialogue context (Figures 26-27) and are not instructed to use personas or nonverbal cues. Thus the judge rewards SocialMind for information that the baselines are never given, so the comparison measures input asymmetry rather than engagement. Additionally, the cache is initialized from LLM-simulated conversations resembling the test distribution, and the judge (GPT-4o) is from the same model family as the generator, so the evaluation may further favor SocialMind. The authors should either give baselines the same persona/nonverbal information, or run a matched ablation where SocialMind also lacks that information, and report the results under those conditions.
  2. [§5.4.2] The 20-person user study provides no control condition and no behavioral outcome measure; questionnaire items Q1-Q6 ask about prior experience, satisfaction, latency acceptance, willingness to use, willingness to interact with another user, and perceived innovativeness. It therefore cannot support the 38.3% engagement claim, which rests entirely on the simulated LLM-as-a-Judge result. To support the claim of higher engagement in live interactions, the authors need a controlled comparison and a behavioral or partner-rated engagement measure.
  3. [§5.1.2 / §5.4.1] The perception pipeline is not validated end-to-end. In Section 5.1.2, nonverbal cues are randomly selected subcategories from Table 3 and fed as clean inputs to SocialMind, bypassing the MediaPipe-based perception stack. Section 5.4.1 reports only latency and power for the real system, with no recognition accuracy for facial expressions, gestures, or proximity. The system's core premise is that it can reliably extract these cues from noisy real-world sensors; without such accuracy data, the practical value of the generated suggestions is unestablished. The authors should provide per-cue accuracy on real data or explain why clean-input simulation suffices.
minor comments (5)
  1. [Table 1] The legend ' means included' is incomplete; the symbols used in the table (checkmarks and crosses) are undefined and should be explicitly listed in the caption.
  2. [§4.2.3] There is a typo: 'utlizes' should be 'utilizes'.
  3. [§5.4.2 / Figure 21] The percentages in Figure 21 are presented without any statistical significance testing or confidence intervals; please report effect sizes or at least descriptive statistics for the questionnaire responses.
  4. [Figures 22 and 23] The captions for Figures 22 and 23 are identical, and the prompt templates appear duplicated; please differentiate the dialogue-based and social-factor-based role-play prompts or merge the figures.
  5. [§4.4.3] The statement that N=70 is 'optimal for full display on the eye screen' cites 'measurement experiments' but no details are given; please provide the measurement procedure or a reference.

Circularity Check

2 steps flagged · score 8.0 of 10

The headline 38.3% engagement gain is built into the evaluation: the LLM judge scores whether partner personas appear in suggestions, and only SocialMind is given those personas.

  1. self definitional [Section 5.1.4 (Evaluation Metrics), Fig. 24 (SocialMind prompt), Fig. 28 (LLM evaluation prompt)]
    "Engagement. ... assessing whether the social suggestions consider the conversational partner’s implicit personas. / Fig. 24: 'The Conversation Partner’s Persona Clues $[ Partner Persona Clues].' / Fig. 28: 'Engagement. this metric assesses whether the social suggestions consider the conversational partner implicit personas.'"

    The engagement metric is defined as whether suggestions mention the partner's implicit personas. SocialMind's prompt is the only assistant prompt that supplies 'The Conversation Partner’s Persona Clues' and instructs the model to generate suggestions considering both parties' persona cues. The Zero-shot and CoT baselines (Figs. 26-27) receive only the generic instruction to help the user maintain engaging conversations, with no persona fields. The judge is then given ground-truth partner personas and told to score engagement on that exact dimension. Therefore the reported 38.3% engagement advantage is entailed by which system received the partner-persona input; it is not evidence about real engagement.

  2. self definitional [Section 5.1.4 (Evaluation Metrics), Fig. 24 (SocialMind prompt), Section 5.1.5 (Baselines)]
    "Personalization. This metric evaluates ... whether the social suggestions incorporate users’ implicit personas, including personal interests and backgrounds. / Nonverbal Cues Utilization. ... assess whether social suggestions take into account the conversational partner’s nonverbal cues. / Fig. 24: 'Current Nonverbal Cues: $[ Nonverbal Cues].'"

    Personalization and Nonverbal Cues Utilization are defined as the presence, in the suggestion, of the user personas and partner nonverbal cues. Those cues are inputs that only SocialMind's prompt contains: Fig. 24 lists 'The User’s Persona Clues' and 'Current Nonverbal Cues,' while the baseline prompts (Figs. 26-27) contain no persona or nonverbal-cue fields. The LLM judge is instructed to reward exactly these dimensions. Hence the reported 38.7% personalization and 61.7% nonverbal-cue gains are by-construction consequences of input asymmetry rather than measured improvements in suggestion quality.

full rationale

The main quantitative claim (38.3% higher engagement, 38.7% higher personalization, 61.7% higher nonverbal-cue utilization) is not an independent measurement of social-suggestion quality. The paper defines Engagement as 'assessing whether the social suggestions consider the conversational partner’s implicit personas,' Personalization as whether suggestions 'incorporate users’ implicit personas,' and Nonverbal Cues Utilization as whether suggestions 'take into account the conversational partner’s nonverbal cues.' SocialMind's runtime prompt is the only condition that receives the user persona clues, partner persona clues, and current nonverbal cues, and it is explicitly instructed to include them. The LLM judge is given ground-truth personas and instructed to score exactly those dimensions. The baselines receive only a generic instruction to help the user maintain engaging conversations, so the comparison measures input asymmetry, not engagement. The user study is self-reported satisfaction and willingness with no control condition and no behavioral engagement measure, so it does not independently support the engagement claim. The remaining engineering evaluations (primary-user detection FAR/FRR, cache hit ratio and accuracy, latency and power) are tied to external measurements and are not circular. I found no load-bearing self-citation or imported-uniqueness circularity. Because the headline engagement result is forced by the metric definition and prompt design, the circularity score is 8.

Assumptions & free parameters 6 free parameters · 4 assumptions · 0 invented entities

The system has several hand-tuned thresholds and design parameters, and it depends on unverified assumptions about the perception pipeline and the validity of LLM-based evaluation. There are no invented physical entities. The most consequential free parameters are the vibration threshold (fit via grid search) and the cache similarity threshold (hand-tuned).

free parameters (6)
  • Vibration energy threshold for primary user detection = 1.1
    Chosen via grid search on real-world measurements (Section 4.2.2, Section 5.3.3); the reported success rate depends on this threshold.
  • Cache semantic similarity threshold = 0.95
    Hand-set high to avoid logically inconsistent cache hits (Section 4.5.1); impact analyzed with cache size in Figures 15 and 19.
  • Speech offloading interval = 2 seconds
    Set based on average speaking speed of 150 words per minute (Section 4.5.2).
  • Suggestion refresh interval = 3 seconds
    Set based on reading speed of 200 words per minute and the 70-word limit (Section 4.5.3).
  • Maximum suggestion length N = 70 words
    Tuned for full display on the glasses (Section 4.4.3).
  • Vibration sample rate = 466 Hz
    Chosen to reduce bandwidth (Section 4.2.2).
assumptions (4)
  • domain assumption LLM-as-a-Judge with GPT-4o yields valid scores for open-ended social suggestions.
    The paper adopts LLM-as-a-Judge (Section 5.1.4) without human validation of the scores; this is a load-bearing evaluation assumption.
  • domain assumption Two LLM agents role-playing as user and partner faithfully reproduce live social interaction dynamics.
    The simulated evaluation (Section 5.1.2) relies on LLM role-play to generate conversations because public datasets are fixed and lack cues; if agent behavior diverges from real humans, the quantitative results may not transfer.
  • ad hoc to paper The randomly selected nonverbal cue subcategories are representative of real-world nonverbal behavior and are available as clean inputs.
    In the simulation, cues are randomly assigned (Section 5.1.2), bypassing the perception pipeline; this masks sensor noise and detection errors.
  • domain assumption The lightweight on-glasses perception models accurately extract facial expressions, gestures, and proximity in real environments.
    The system's human-like perception claim (Section 4.2.1) assumes MediaPipe and specialized models work reliably, but no accuracy evaluation is provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SocialMind: LLM-based Proactive AR Social Assistive System with Human-like Perception for In-situ Live Interactions." pith.science (2026). https://pith.science/paper/53T7XVCX

@misc{pith2026241204036,
  author       = {Pith},
  title        = {Pith review of: SocialMind: LLM-based Proactive AR Social Assistive System with Human-like Perception for In-situ Live Interactions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/53T7XVCX}},
  note         = {Machine review of arXiv:2412.04036}
}
read the original abstract

Social interactions are fundamental to human life. The recent emergence of large language models (LLMs)-based virtual assistants has demonstrated their potential to revolutionize human interactions and lifestyles. However, existing assistive systems mainly provide reactive services to individual users, rather than offering in-situ assistance during live social interactions with conversational partners. In this study, we introduce SocialMind, the first LLM-based proactive AR social assistive system that provides users with in-situ social assistance. SocialMind employs human-like perception leveraging multi-modal sensors to extract both verbal and nonverbal cues, social factors, and implicit personas, incorporating these social cues into LLM reasoning for social suggestion generation. Additionally, SocialMind employs a multi-tier collaborative generation strategy and proactive update mechanism to display social suggestions on Augmented Reality (AR) glasses, ensuring that suggestions are timely provided to users without disrupting the natural flow of conversation. Evaluations on three public datasets and a user study with 20 participants show that SocialMind achieves 38.3% higher engagement compared to baselines, and 95% of participants are willing to use SocialMind in their live social interactions.

Figures

Figures reproduced from arXiv: 2412.04036 by the authors.

Figure 1
Figure 1. Overview of SocialMind. SocialMind provides in-situ social assistance to the user to help the user during live social [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Survey results for social experience and assistance demand. [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. System overview of SocialMind. SocialMind leverages the multi-modal sensor data to achieve human-like perception. [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (19 more)
Figure 4
Figure 4. Figure 4: SocialMind’s primary user detection [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]
Figure 5
Figure 5. Figure 5: Implicit persona adaptation in SocialMind. [PITH_FULL_IMAGE:figures/full_fig_p010_5.png]
Figure 7
Figure 7. Figure 7: Real-world test settings. Participants engage in live face-to-face social interactions with other parties, wearing the [PITH_FULL_IMAGE:figures/full_fig_p013_7.png]
Figure 8
Figure 8. Figure 8: Overall performance of the social suggestions generated by SocialMind and baselines across three datasets. Personal. [PITH_FULL_IMAGE:figures/full_fig_p016_8.png]
Figure 9
Figure 9. Figure 9: Social suggestion performance across different types of social scenarios. Two LLM agents are prompted to engage in [PITH_FULL_IMAGE:figures/full_fig_p017_9.png]
Figure 10
Figure 10. Figure 10: Examples of social conversations and social suggestions SocialMind and baselines. Words highlighted in red and blue [PITH_FULL_IMAGE:figures/full_fig_p018_10.png]
Figure 11
Figure 11. Figure 11: Example of using LLMs for scoring the social suggestions. [PITH_FULL_IMAGE:figures/full_fig_p018_11.png]
Figure 12
Figure 12. Figure 12: Example of the intention infer-based suggestion generation in SocialMind. Words highlighted in red indicate that [PITH_FULL_IMAGE:figures/full_fig_p019_12.png]
Figure 13
Figure 13. Figure 13: Examples of the extracted implicit persona cues [PITH_FULL_IMAGE:figures/full_fig_p020_13.png]
Figure 16
Figure 16. Figure 16: Effectiveness of social factor-aware cache. [PITH_FULL_IMAGE:figures/full_fig_p020_16.png]
Figure 17
Figure 17. Figure 17: Effectiveness of the personas and the nonverbal cues [PITH_FULL_IMAGE:figures/full_fig_p021_17.png]
Figure 19
Figure 19. Figure 19: Impact of the threshold and catch size on the LLM token saving ratio. [PITH_FULL_IMAGE:figures/full_fig_p021_19.png]
Figure 20
Figure 20. Figure 20: The primary user detection performance of SocialMind and the impact of parameters on different solutions . [PITH_FULL_IMAGE:figures/full_fig_p022_20.png]
Figure 21
Figure 21. Figure 21: SocialMind’s user study results. system, such as balancing using their own cognitive abilities versus referring to the text displayed on the glasses, ensuring a natural conversation flow. Additionally, some participants find the system helpful when they lose focus or …
Figure 22
Figure 22. Figure 22: Prompt templates of the user agent and the conver [PITH_FULL_IMAGE:figures/full_fig_p028_22.png]
Figure 24
Figure 24. Figure 24: Prompt template of SocialMind. 28 [PITH_FULL_IMAGE:figures/full_fig_p028_24.png]
Figure 25
Figure 25. Figure 25: Prompt of implicit persona extraction in SocialMind. [PITH_FULL_IMAGE:figures/full_fig_p029_25.png]
Figure 26
Figure 26. Figure 26: Prompt of Zero-shot baseline approach. Prompt Template of Zero-shot # OVERALL INSTRUCTIONS You are playing the role of a social assistant, helping a user during live social conversations. Your user is currently engaged in a conversation with another individual. Your g…
Figure 28
Figure 28. Figure 28: Prompt template of the LLM evaluation. 29 [PITH_FULL_IMAGE:figures/full_fig_p029_28.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

84 extracted references · 56 canonical work pages

  1. [1]

    Apple Siri

    2024. Apple Siri. https://www.apple.com/siri/

  2. [2]

    New in Gemini: Gemini Live and connected Google apps in more languages

    2024. New in Gemini: Gemini Live and connected Google apps in more languages. https://blog.google/products/gemini/gemini-live- extensions-language-expansion/

  3. [3]

    Quality of life indicators - social interactions

    2024. Quality of life indicators - social interactions. https://ec.europa.eu/eurostat/statistics-explained/index.php?title=Quality_of_life_ indicators_-_social_interactions

  4. [4]

    RayNeo X2

    2024. RayNeo X2. https://rayneo.cn/product/x2/specs/

  5. [5]

    Social Anxiety Disorder

    2024. Social Anxiety Disorder. https://adaa.org/understanding-anxiety/social-anxiety-disorder

  6. [6]

    2024. Tianji. https://github.com/SocialAI-tianji

  7. [7]

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)

  8. [8]

    Shashank Ahire, Benjamin Simon, and Michael Rohs. 2024. WorkFit: Designing Proactive Voice Assistance for the Health and Well-Being of Knowledge Workers. In Proceedings of the 6th ACM Conference on Conversational User Interfaces . 1–14

Show all 84 references
  1. [9]

    Michael Argyle and Janet Dean. 1965. Eye-contact, distance and affiliation. Sociometry (1965), 289–304

  2. [10]

    Michael C Ashton. 2022. Individual differences and personality . Academic Press

  3. [11]

    Fu Bang. 2023. GPTCache: An open-source semantic cache for LLM applications enabling faster answers and cost savings. InProceedings of the 3rd Workshop for Natural Language Processing Open Source Software (NLP-OSS 2023) . 212–218

  4. [12]

    Federico Bianchi, Patrick John Chia, Mert Yuksekgonul, Jacopo Tagliabue, Dan Jurafsky, and James Zou. 2024. How well can llms negotiate? negotiationarena platform and analysis. arXiv preprint arXiv:2402.05863 (2024)

  5. [13]

    Marc Brysbaert. 2019. How many words do we read per minute? A review and meta-analysis of reading rate. Journal of memory and language 109 (2019), 104047

  6. [14]

    Renato AC Capuruço and Luiz F Capretz. 2009. Building social-aware software applications for the interactive learning age. Interactive Learning Environments 17, 3 (2009), 241–255

  7. [15]

    Harrison Chase. 2022. LangChain. https://github.com/langchain-ai/langchain

  8. [16]

    Tao Chen, Yongjie Yang, Chonghao Qiu, Xiaoran Fan, Xiuzhen Guo, and Longfei Shangguan. 2024. Enabling Hands-Free Voice Assistant Activation on Earphones. In Proceedings of the 22nd Annual International Conference on Mobile Systems, Applications and Services . 155–168

  9. [17]

    Yuqi Chu, Lizi Liao, Zhiyuan Zhou, Chong-Wah Ngo, and Richang Hong. 2024. Towards Multimodal Emotional Support Conversation Systems. arXiv preprint arXiv:2408.03650 (2024)

  10. [18]

    Yang Deng, Lizi Liao, Zhonghua Zheng, Grace Hui Yang, and Tat-Seng Chua. 2024. Towards human-centered proactive conversational agents. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval . 807–818

  11. [19]

    Utsav Drolia, Katherine Guo, Jiaqi Tan, Rajeev Gandhi, and Priya Narasimhan. 2017. Cachier: Edge-caching for recognition applications. In 2017 IEEE 37th international conference on distributed computing systems (ICDCS) . IEEE, 276–286

  12. [20]

    Starkey Duncan Jr. 1969. Nonverbal communication. Psychological bulletin 72, 2 (1969), 118

  13. [21]

    Zachary Englhardt, Richard Li, Dilini Nissanka, Zhihan Zhang, Girish Narayanswamy, Joseph Breda, Xin Liu, Shwetak Patel, and Vikram Iyer. 2024. Exploring and characterizing large language models for embedded system development and debugging. In Extended Abstracts of the CHI Co...

  14. [22]

    Chris Frith. 2009. Role of facial expressions in social interactions. Philosophical Transactions of the Royal Society B: Biological Sciences 364, 1535 (2009), 3453–3458

  15. [23]

    Bin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang, Djordje Jevdjic, Junbo Deng, Xingkun Yang, Zhou Yu, and Pengfei Zuo. 2024. {Cost-Efficient} Large Language Model Serving for Multi-turn Conversations with{CachedAttention}. In 2024 USENIX Annual Technical Conference (USENIX ATC ...

  16. [24]

    Ge Gao, Alexey Taymanov, Eduardo Salinas, Paul Mineiro, and Dipendra Misra. 2024. Aligning llm agents by learning latent preference from user edits. arXiv preprint arXiv:2404.15269 (2024)

  17. [25]

    Weiwei Gao, Kexin Du, Yujia Luo, Weinan Shi, Chun Yu, and Yuanchun Shi. 2024. EasyAsk: An In-App Contextual Tutorial Search Assistant for Older Adults with Voice and Touch Inputs. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 8, 3 (2024), 1–27

  18. [26]

    Sonia Garcia-Salicetti, Charles Beumier, Gérard Chollet, Bernadette Dorizzi, Jean Leroux les Jardins, Jan Lunter, Yang Ni, and Dijana Petrovska-Delacrétaz. 2003. BIOMET: A multimodal person authentication database including face, voice, fingerprint, hand and signature modaliti...

  19. [27]

    In Gim, Guojun Chen, Seung-seob Lee, Nikhil Sarda, Anurag Khandelwal, and Lin Zhong. 2024. Prompt cache: Modular attention reuse for low-latency inference. Proceedings of Machine Learning and Systems 6 (2024), 325–338

  20. [28]

    Google. 2024. Google Assistant. https://assistant.google.com/learn/. 24

  21. [29]

    Google. 2024. MediaPipe. https://github.com/google/mediapipe Accessed: 2024-10-30

  22. [30]

    Peizhen Guo, Bo Hu, Rui Li, and Wenjun Hu. 2018. Foggycache: Cross-device approximate computation reuse. In Proceedings of the 24th annual international conference on mobile computing and networking . 19–34

  23. [31]

    Yunqi Guo, Jinghao Zhao, Boyan Ding, Congkai Tan, Weichong Ling, Zhaowei Tan, Jennifer Miyaki, Hongzhe Du, and Songwu Lu. 2023. Sign-to-911: Emergency Call Service for Sign Language Users with Assistive AR Glasses. In Proceedings of the 29th Annual International Conference on ...

  24. [32]

    Judith A Hall, Terrence G Horgan, and Nora A Murphy. 2019. Nonverbal communication. Annual review of psychology 70, 1 (2019), 271–294

  25. [33]

    Dirk Hovy and Diyi Yang. 2021. The importance of modeling social factors of language: Theory and practice. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human language technologies . 588–602

  26. [34]

    Yuncheng Hua, Zhuang Li, Linhao Luo, Kadek Ananta Satriadi, Tao Feng, Haolan Zhan, Lizhen Qu, Suraj Sharma, Ingrid Zukerman, Zhaleh Semnani-Azad, et al. 2024. Sadas: A dialogue assistant system towards remediating norm violations in bilingual socio-cultural conversations. arXi...

  27. [35]

    Yuncheng Hua, Lizhen Qu, and Gholamreza Haffari. 2024. Assistive Large Language Model Agents for Socially-Aware Negotiation Dialogues. arXiv preprint arXiv:2402.01737 (2024)

  28. [36]

    Apple Inc. 2024. Apple Vision Pro. https://www.apple.com/apple-vision-pro/ Mixed-reality headset by Apple Inc., announced in 2023, offering advanced augmented and virtual reality experiences

  29. [37]

    INMO Glass. 2024. INMO Air2 - Next-Gen Wireless AR Glasses. https://air2.inmoglass.com/ Accessed: 2024-10-31

  30. [38]

    Pegah Jandaghi, Xianghai Sheng, Xinyi Bai, Jay Pujara, and Hakim Sidahmed. 2024. Faithful Persona-based Conversational Dataset Generation with Large Language Models. In Findings of the Association for Computational Linguistics ACL 2024 , Lun-Wei Ku, Andre Martins, and Vivek Sr...

  31. [39]

    It’s the only thing I can trust

    JiWoong Jang, Sanika Moharana, Patrick Carrington, and Andrew Begel. 2024. “It’s the only thing I can trust”: Envisioning Large Language Model Use by Autistic Workers for Communication Assistance. InProceedings of the CHI Conference on Human Factors in Computing Systems. 1–18

  32. [40]

    Julie Jiang and Emilio Ferrara. 2023. Social-LLM: Modeling User Behavior at Scale using Language Models and Social Network Data. arXiv preprint arXiv:2401.00893 (2023)

  33. [41]

    Xiaoqing Jing, Chun Yu, Kun Yue, Liangyou Lu, Nan Gao, Weinan Shi, Mingshan Zhang, Ruolin Wang, and Yuanchun Shi. 2024. AngleSizer: Enhancing Spatial Scale Perception for the Visually Impaired with an Interactive Smartphone Assistant. Proceedings of the ACM on Interactive, Mob...

  34. [42]

    Neha U Keshav, Joseph P Salisbury, Arshya Vahabzadeh, and Ned T Sahin. 2017. Social communication coaching smartglasses: Well tolerated in a diverse sample of children and adults with autism. JMIR mHealth and uHealth 5, 9 (2017), e8534

  35. [43]

    Hyunwoo Kim, Youngjae Yu, Liwei Jiang, Ximing Lu, Daniel Khashabi, Gunhee Kim, Yejin Choi, and Maarten Sap. 2022. Prosocialdialog: A prosocial backbone for conversational agents. arXiv preprint arXiv:2205.12688 (2022)

  36. [44]

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing S...

  37. [45]

    Huining Li, Chenhan Xu, Aditya Singh Rathore, Zhengxiong Li, Hanbin Zhang, Chen Song, Kun Wang, Lu Su, Feng Lin, Kui Ren, et al

  38. [46]

    Jiaxing Li, Chi Xu, Feng Wang, Isaac M von Riedemann, Cong Zhang, and Jiangchuan Liu. 2024. SCALM: Towards Semantic Caching for Automated Chat Services with Large Language Models. arXiv preprint arXiv:2406.00025 (2024)

  39. [47]

    Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017. Dailydialog: A manually labelled multi-turn dialogue dataset. arXiv preprint arXiv:1710.03957 (2017)

  40. [48]

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024. Visual instruction tuning. Advances in neural information processing systems 36 (2024)

  41. [49]

    Jiachen Liu, Zhiyu Wu, Jae-Won Chung, Fan Lai, Myungjin Lee, and Mosharaf Chowdhury. 2024. Andes: Defining and Enhancing Quality-of-Experience in LLM-Based Text Streaming Services. arXiv preprint arXiv:2404.16283 (2024)

  42. [50]

    Kaiwei Liu, Bufang Yang, Lilin Xu, Yunqi Guo, Neiwen Ling, Zhihe Zhao, Guoliang Xing, Xian Shuai, Xiaozhe Ren, Xin Jiang, et al. 2024. Tasking Heterogeneous Sensor Systems with LLMs. In Proceedings of the 22nd ACM Conference on Embedded Networked Sensor Systems . 901–902

  43. [51]

    Chaquopy LLC. 2024. Chaquopy: the Python SDK for Android. https://github.com/chaquo/chaquopy. Accessed: 2024-10-30

  44. [52]

    Cheng Charles Ma, Kevin Hyekang Joo, Alexandria K Vail, Sunreeta Bhattacharya, Álvaro Fernández García, Kailana Baker-Matsuoka, Sheryl Mathew, Lori L Holt, and Fernando De la Torre. 2024. Multimodal Fusion with LLMs for Engagement Prediction in Natural Conversation. arXiv prep...

  45. [53]

    Meta. 2024. Introducing Orion, Our First True Augmented Reality Glasses. https://about.fb.com/news/2024/09/introducing-orion-our- first-true-augmented-reality-glasses/ Accessed: 2024-10-31

  46. [54]

    Antje S Meyer. 2023. Timing in conversation. Journal of Cognition 6, 1 (2023)

  47. [55]

    Microsoft. 2024. Limited Access to Speaker Recognition. https://learn.microsoft.com/en-us/legal/cognitive-services/speech-service/ speaker-recognition/limited-access-speaker-recognition#registration-process

  48. [56]

    Microsoft Learn. 2024. Speaker Recognition Overview. https://learn.microsoft.com/en-us/azure/ai-services/speech-service/speaker- recognition-overview Accessed: 2024-10-31

  49. [57]

    Sheshera Mysore, Zhuoran Lu, Mengting Wan, Longqi Yang, Steve Menezes, Tina Baghaee, Emmanuel Barajas Gonzalez, Jennifer Neville, and Tara Safavi. 2023. Pearl: Personalizing large language model writing assistants with generation-calibrated retrievers. arXiv preprint arXiv:231...

  50. [58]

    Pedregosa, G

    F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. 2011. Scikit-learn: Machine Learning in Python. Journal of Machine Le...

  51. [59]

    Even Realities. 2024. G1: Next-Gen Smart Glasses with Display. https://www.evenrealities.com/g1 Accessed: 2024-10-31

  52. [60]

    N Reimers. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. arXiv preprint arXiv:1908.10084 (2019)

  53. [61]

    Frank Seide, Morrie Doulaty, Yangyang Shi, Yashesh Gaur, Junteng Jia, and Chunyang Wu. 2024. Speech ReaLLM–Real-time Streaming Speech Recognition with Multimodal LLMs by Teaching the Flow of Time. arXiv preprint arXiv:2406.09569 (2024)

  54. [62]

    WIRED Staff. 2024. XRAI Glass Caption AR Glasses: First Look. https://www.wired.com/story/xrai-glass-caption-ar-glasses-first-look/ Accessed: 2024-10-31

  55. [63]

    Emma M Templeton, Luke J Chang, Elizabeth A Reynolds, Marie D Cone LeBeaumont, and Thalia Wheatley. 2022. Fast response times signal social connection in conversation. Proceedings of the National Academy of Sciences 119, 4 (2022), e2116915119

  56. [64]

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288 (2023)

  57. [65]

    Chongyang Wang, Yuan Feng, Lingxiao Zhong, Siyi Zhu, Chi Zhang, Siqi Zheng, Chen Liang, Yuntao Wang, Chengqi He, Chun Yu, et al. 2024. UbiPhysio: Support Daily Functioning, Fitness, and Rehabilitation with Action Understanding and Feedback in Natural Language. Proceedings of t...

  58. [66]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35 (2022), 24824–24837

  59. [67]

    Yu Wu, Zhoujun Li, Wei Wu, and Ming Zhou. 2018. Response selection with topic clues for retrieval-based chatbots. Neurocomputing 316 (2018), 251–261

  60. [68]

    Zhifei Xie and Changqiao Wu. 2024. Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming. arXiv preprint arXiv:2408.16725 (2024)

  61. [69]

    Mengwei Xu, Mengze Zhu, Yunxin Liu, Felix Xiaozhu Lin, and Xuanzhe Liu. 2018. Deepcache: Principled cache for mobile deep vision. In Proceedings of the 24th annual international conference on mobile computing and networking . 129–144

  62. [70]

    Qianli Xu, Shue Ching Chia, Bappaditya Mandal, Liyuan Li, Joo-Hwee Lim, Michal Akira Mukawa, and Cheston Tan. 2016. SocioGlass: social interaction assistance with face recognition on google glass. Scientific Phone Apps and Mobile Devices 2 (2016), 1–4

  63. [71]

    Zhenyu Xu, Hailin Xu, Zhouyang Lu, Yingying Zhao, Rui Zhu, Yujiang Wang, Mingzhi Dong, Yuhu Chang, Qin Lv, Robert P Dick, et al

  64. [72]

    Bufang Yang, Lixing He, Neiwen Ling, Zhenyu Yan, Guoliang Xing, Xian Shuai, Xiaozhe Ren, and Xin Jiang. 2023. Edgefm: Leveraging foundation model for open-set learning on the edge. In Proceedings of the 21st ACM Conference on Embedded Networked Sensor Systems . 111–124

  65. [73]

    Bufang Yang, Lixing He, Kaiwei Liu, and Zhenyu Yan. 2024. VIAssist: Adapting Multi-modal Large Language Models for Users with Visual Impairments. arXiv preprint arXiv:2404.02508 (2024)

  66. [74]

    Bufang Yang, Siyang Jiang, Lilin Xu, Kaiwei Liu, Hai Li, Guoliang Xing, Hongkai Chen, Xiaofan Jiang, and Zhenyu Yan. 2024. DrHouse: An LLM-empowered diagnostic reasoning system through harnessing outcomes from sensor data and expert knowledge. Proceedings of the ACM on Interac...

  67. [75]

    Ziqi Yang, Xuhai Xu, Bingsheng Yao, Ethan Rogers, Shao Zhang, Stephen Intille, Nawar Shara, Guodong Gordon Gao, and Dakuo Wang

  68. [76]

    Yao Yao, Zuchao Li, and Hai Zhao. 2024. SirLLM: Streaming infinite retentive LLM. arXiv preprint arXiv:2405.12528 (2024)

  69. [77]

    Haolan Zhan, Zhuang Li, Xiaoxi Kang, Tao Feng, Yuncheng Hua, Lizhen Qu, Yi Ying, Mei Rianto Chandra, Kelly Rosalin, Jureynolds Jureynolds, et al. 2024. RENOVI: A Benchmark Towards Remediating Norm Violations in Socio-Cultural Conversations. In Findings of the Association for C...

  70. [78]

    Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 8, 2 (2024), 1–35

    Talk2Care: An LLM-based Voice Assistant for Communication between Healthcare Providers and Older Adults. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 8, 2 (2024), 1–35

  71. [79]

    Haolan Zhan, Yufei Wang, Zhuang Li, Tao Feng, Yuncheng Hua, Suraj Sharma, Lizhen Qu, Zhaleh Semnani Azad, Ingrid Zukerman, and Reza Haf. 2024. Let’s Negotiate! A Survey of Negotiation Dialogue Systems. In Findings of the Association for Computational Linguistics: EACL 2024. 2019–2031

  72. [80]

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems 36 (2023), 46595–46623. 27 A...

  73. [81]

    Haolan Zhan, Zhuang Li, Yufei Wang, Linhao Luo, Tao Feng, Xiaoxi Kang, Yuncheng Hua, Lizhen Qu, Lay-Ki Soon, Suraj Sharma, et al

  74. [2020]

    In Proceedings of the 18th Conference on Embedded Networked Sensor Systems

    Vocalprint: exploring a resilient and secure voice authentication via mmwave biometric interrogation. In Proceedings of the 18th Conference on Embedded Networked Sensor Systems . 312–325

  75. [2023]

    In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval

    Socialdial: A benchmark for socially-aware dialogue systems. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval . 2712–2722

  76. [2024]

    Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 8, 2 (2024), 1–41

    Can Large Language Models Be Good Companions? An LLM-Based Eyewear System with Conversational Common Ground. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 8, 2 (2024), 1–41

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.