Pith. sign in

REVIEW 2 major objections 4 minor 53 references

An exploratory study with 9 autistic trainees and 7 job coaches finds that adding visible, graspable objects to an LLM-driven VR role-play scenario yields the preferred training mode, with usability and workload unchanged from conversation-

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 06:43 UTC pith:72YZZF6N

load-bearing objection Honest exploratory study with a real confound between prompt schema and modality; worth refereeing but the preference claim should be softened. the 2 major comments →

arxiv 2607.21769 v1 pith:72YZZF6N submitted 2026-07-23 cs.HC

From Grasping to Speaking: Generative AI-Based Environment-Grounded VR Communication Training for Autistic Individuals

classification cs.HC
keywords communication trainingautism spectrum disordervirtual realitylarge language modelsworkplace simulationenvironment groundinggrasp interaction
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper tries to establish that job-communication training for autistic individuals can be improved by embedding conversation in the trainee's physical task environment rather than keeping it as isolated verbal role-play. The authors built an LLM-driven virtual customer whose dialogue is conditioned on the VR scene's objects and on the trainee's grasp actions, and compared three levels of grounding — chat-only, chat with visible objects, and chat with objects plus grasping — across café, butcher-shop, and fast-food scenarios. In a study with 9 autistic trainees and 7 job coaches, usability and workload were statistically indistinguishable across the three conditions, yet both groups preferred the most grounded condition (C+O+G), citing sustained engagement, more realistic task-embedded conversation, and the chance to practise multi-tasking. The authors present this as exploratory evidence for flexible, adjustable grounding in VR communication training rather than a one-size-fits-all modality.

Core claim

The paper's central claim is that environment-grounded VR communication training can raise engagement and perceived workplace relevance for autistic trainees without measurable usability or workload costs. Concretely, when the virtual customer's prompt includes the list of objects in the room and the trainee's grasp/release actions (rendered as descriptive cues such as 'left hand grabs Coffee from Counter'), both trainees and job coaches preferred that condition over conversation-only practice and over a condition with objects but no interaction. The preference was accompanied by qualitative reports that physical interaction gave trainees conversational anchors, encouraged longer and more ob

What carries the argument

The carrying mechanism is a prompting structure that feeds the LLM-driven avatar three layers of context: the scenario and role definition, the catalog of environmental objects and containers, and the user's action cues describing which hand grabs or releases which object from which parent surface. This turns a physical grasp into a text event, letting the avatar produce dialogue that references the shared spatial situation and the work in progress. The three conditions are incremental unlocks of this schema: the conversation-only baseline omits the object list and action cues, C+O adds the object catalog but disables grasping, and C+O+G adds the grasp cues, enabling the avatar to respond to

Load-bearing premise

The load-bearing premise is that the three conditions differ only in environmental grounding: the conversation-only baseline receives no object or grasp information, and the richer conditions do not change how the avatar talks (response length, number of questions per turn, or pacing). If the prompt schema leaks object context into the baseline, or if the grounded conditions happen to produce friendlier, slower, or simpler dialogue, the observed preference could come from con

What would settle it

Take the exact system prompts from the three conditions and run the same scripted dialogue with an empty object list in C. Measure the avatar's average number of questions per turn and per-turn length; if the object list appears anywhere in C's system context, or if C+O+G's turns are systematically different in structure or pace from C's, the central comparison is confounded and the preference could be explained by dialogue style rather than grounding.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Environment-grounded VR role-play can embed communication practice inside task performance, matching the situated nature of real workplace conversation that job coaches emphasized.
  • Adding visible, graspable objects did not degrade SUS or raise NASA-TLX scores for this sample, suggesting richer interaction need not trade off against usability or cognitive load.
  • The preference for the C+O+G condition was tied to reports of sustained engagement and more object-referenced talk, implying that grounding can serve as a conversational scaffold for trainees who find open-ended small talk effortful.
  • Because participants valued different modalities for different trainees and goals, training systems should offer configurable grounding levels rather than a single fixed mode.
  • The avatar's occasional multi-question turns overwhelmed some trainees, so prompts should constrain the agent to one request or question per turn even when grounding enriches context.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If this preference generalizes beyond the small sample, VR vocational training for autistic individuals may shift from abstract social-skills rehearsal toward task-embedded conversation; the paper gestures at this (calling the hands-on activity a possible anchor) but does not test it directly.
  • A next experiment could separate task-relevant grounding from mere motor activity, comparing C+O+G with a condition where the physical task is unrelated to the conversation, to see which component drives engagement.
  • The study stops at experience and preference; the key practical promise — that skills practiced in the grounded condition transfer to real workplaces — remains untested and is a natural target for a longitudinal follow-up.
  • The coaches' request for per-session flexibility implies an authoring tool that lets them toggle object visibility and grasp interactivity per trainee; the paper lists this as future work, so it is our inference that such a tool is a direct design consequence.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper presents an exploratory within-subjects study of three VR-based communication training modalities for autistic individuals: conversation-only (C), conversation with environmental objects (C+O), and conversation with objects plus grasp interactions (C+O+G). Nine autistic trainees and seven job coaches used the system; the authors measured System Usability Scale (SUS), raw NASA-TLX, preference ratings, and collected qualitative observations and interviews. The central claim is that usability and workload were comparable across the three modalities, while both trainees and job coaches preferred the most environment-grounded condition (C+O+G), which was perceived as more engaging and as better integrating communication practice into task performance. The paper also offers design considerations for future environment-grounded VR communication training.

Significance. If the finding holds, the paper makes a useful contribution to accessible VR communication training by showing that adding environmental objects and grasp interactions does not necessarily increase workload or reduce usability for autistic trainees, and that such features may be preferred. The study is unusual in including both autistic trainees and job coaches, and the qualitative data are rich. However, the causal interpretation is undermined by a confound: the three conditions differ not only in the user's interaction modality but also in the prompt schema given to the LLM-driven agent, so the agent's dialogue is generated from different contextual inputs across conditions. This unaddressed confound weakens the central preference claim and needs to be resolved before the paper can be accepted.

major comments (2)
  1. [§3.1, §4.1.1, §5.2.2, §5.3.2] The independent variable is not limited to the trainee's environmental grounding; it also changes the agent's prompt context. Section 4.1.1 states that the C condition uses a 'basic system context prompt' while C+O and C+O+G receive the full prompting schema, with C+O+G additionally including grasp/release cues. Consequently, the LLM-generated avatar responses differ across conditions by construction. Section 5.2.2 acknowledges that the avatar's responses in C+O and C+O+G were 'constrained by the virtual objects defined in the schema,' and Section 5.3.2 reports that the agent sometimes asked multi-part questions that overwhelmed trainees. No transcript-level analysis of agent behavior (e.g., turn length, number of questions per turn, object reference frequency, response delays) is reported. The preference for C+O+G (Section 5.3.4, Figure 4c) could therefore be driven by prompt-driven dif
  2. [§5.1.1, Abstract, §7] The quantitative support for the preference claim is marginal. The Friedman test for overall preference yields χ²(2)=6.00, p=.050, Kendall's W=.33, and the authors appropriately avoid pairwise comparisons. Yet the abstract and conclusion state unqualifiedly that 'both trainees and coaches preferred' C+O+G. This rests on descriptive means (4.89 vs. 4.67 vs. 4.44) and on qualitative reports (6/7 coaches, all trainees). With n=9 and a borderline p-value, the claim should be more carefully qualified, or supported by a pre-planned comparison or effect-size measures, so that readers can judge the strength of the evidence.
minor comments (4)
  1. [§5.2] Please clarify how ChatGPT was used to 'assist with organizing the observation notes' and how the first author's coding was verified against the original notes. This transparency is valuable in qualitative work, especially given the importance of establishing that the AI tool did not influence theme development.
  2. [§4.3, §5.1.2] Because J01 and J06 each supported two trainees, coach ratings are nested (9 sessions but only 7 coaches). The descriptive analysis is acceptable, but the non-independence should be acknowledged when presenting coach ratings to avoid implying fully independent observations.
  3. [Throughout] Minor notation inconsistencies: 'Raw NASA-TLX' versus 'raw NASA-TLX'; 'χ²(2)' is used but the degree-of-freedom formatting varies; 'Kendall's W' is sometimes typeset as W=.03, .07, .33 with inconsistent punctuation. A consistent notation style would improve readability.
  4. [§4.1.2] The scenarios differ in object placement (all front, front/back split, all back). Although scenario-to-condition assignment was randomized, with only 9 participants the randomization may not balance scenario difficulty across conditions. Please report the actual scenario-condition mapping or discuss this as a potential carryover/confound.

Circularity Check

0 steps flagged

No significant circularity found: empirical user-study results are not derived from their inputs.

full rationale

The paper makes no mathematical derivation or fitted prediction. Its central claims—comparable SUS/TLX scores and preference for C+O+G—are empirical measurements from 9 autistic trainees and 7 job coaches. The system is built on the authors' prior prompting schema [25,27], and self-citations are used to credit design precedents, not to prove the outcome. The adapted schema is an input to the independent variable, not a hidden reproduction of the dependent variable. The one passage that could look like a problem, §5.2.2 ('the LLM-driven avatar’s responses in the C+O and C+O+G modalities were constrained by the virtual objects defined in the schema'), is an admitted potential confound between modality and agent behavior; this is a validity threat, not a circular reduction, and the paper also acknowledges limitations in §6.4 (small single-site sample, no transfer measure, avatar latency/expressiveness). No equation equates the claimed result to the input, and no fitted parameter is relabeled as a prediction. The preference, usability, and workload findings stand on independent participant ratings and interviews.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

The findings rest primarily on survey and interview self-report plus researcher coding; there are no fitted model parameters. The key domain assumptions are prompt fidelity across conditions, measurement validity for the population, single-coder qualitative reliability, and sample representativeness.

axioms (4)
  • domain assumption GPT-4.1 with the adapted environment-grounded prompt schema produces consistently different agent behavior across C, C+O, and C+O+G, realizing only the intended grounding differences.
    Section 3.1 and 4.1.1: the condition contrast is the independent variable; if prompts leaked object/action info into C or failed to produce object-aware responses, preference comparisons would be confounded.
  • domain assumption SUS, raw NASA-TLX, and the custom 5-point Likert items are valid and interpretable measures for this autistic trainee population.
    Section 4.4: used without validated adaptation for autistic adults; the paper cites prior use [23] but no psychometric validation for this specific population is reported.
  • domain assumption The first author's thematic coding, assisted by ChatGPT and verified against notes, is a reliable account of participant experience.
    Sections 5.2 and 5.3: no second coder or inter-rater reliability is reported; themes are investigator-derived and may reflect single-analyst interpretation.
  • domain assumption Nine trainees and seven coaches from one vocational organization are sufficient to support the qualitative preference claims.
    Sections 4.3 and 6.4: recruitment was from a single nonprofit, and participants did not have notable fine-motor or verbal difficulties; the authors acknowledge limited generalizability.

pith-pipeline@v1.3.0-alltime-deepseek · 21689 in / 11738 out tokens · 134748 ms · 2026-08-01T06:43:41.545636+00:00 · methodology

0 comments
read the original abstract

Autistic individuals often face barriers in workplace communication, where soft skills are embedded within ongoing tasks and surrounding environment context, not in isolated verbal exchange. Recent work has introduced LLM-driven agents into VR-based communication training and proposed prompting schemas that let agents generate dialogue grounded in the VR environment and the user's hand-based interactions. Building on this work, we explore how different levels of environmental grounding influence the training experience of autistic trainees and job coaches. We conducted an exploratory study with 9 autistic trainees and 7 job coaches across three modalities: conversation-only (C), conversation with environmental objects (C+O), and conversation with objects and grasp interactions (C+O+G). Usability and workload were comparable across modalities, while both trainees and coaches preferred the more interactive and environment-grounded condition (C+O+G). Participants viewed C+O+G as helpful for sustaining engagement and integrating communication practice into task performance. We discuss design considerations for more flexible, interactive, and workplace-relevant VR communication training for autistic individuals.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

53 extracted references · 14 canonical work pages

  1. [1]

    Vogus, Joshua Wade, and Nilanjan Sarkar

    Deeksha Adiani, Aaron Itzkovitz, Dayi Bian, Harrison Katz, Michael Breen, Spencer Hunt, Amy Swanson, Timothy J. Vogus, Joshua Wade, and Nilanjan Sarkar. 2022. Career Interview Readiness in Virtual Reality (CIRVR): A Plat- form for Simulated Interview Training for Autistic Individuals and Their Em- ployers. ACM Trans. Access. Comput. 15, 1, Article 2 (mar ...

  2. [2]

    Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Haus- man, Alex Herzog, Daniel Ho, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Eric Jang, Rosario Jauregui Ruano, Kyle Jeffrey, Sally Jesmonth, Nikhil J Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang,...

  3. [3]

    Pinaki Prasanna Babar, Mike Barry, and Roshan L Peiris. 2023. Understanding Job Coaches’ Perspectives on Using Virtual Reality as a Job Training Tool for Training People with Intellectual Disabilities. In Extended Abstracts of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI EA ’23). Association for Computing Machinery...

  4. [4]

    Aaron Bangor, Philip Kortum, and James Miller. 2009. Determining what individ- ual SUS scores mean: adding an adjective rating scale. J. Usability Studies 4, 3 (May 2009), 114–123

  5. [5]

    Lal Bozgeyikli, Evren Bozgeyikli, Andrew Raij, Redwan Alqasemi, Srinivas Katkoori, and Rajiv Dubey. 2017. Vocational Rehabilitation of Individuals with Autism Spectrum Disorder with Virtual Reality. ACM Transactions on Accessible Computing 10, 2 (April 2017), 5:1–5:25. doi:10.1145/3046786

  6. [6]

    Virginia Braun and Victoria Clarke. 2006. Using thematic analysis in psy- chology. Qualitative Research in Psychology 3, 2 (2006), 77–101. doi:10.1191/ 1478088706qp063oa

  7. [7]

    John Brooke. 1996. SUS – a quick and dirty usability scale. 189–194

  8. [8]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020), 1877–1901

  9. [9]

    Burke, Tammy Bresnahan, Tan Li, Katrina Epnere, Albert Rizzo, Mary Partin, Robert M

    Shanna L. Burke, Tammy Bresnahan, Tan Li, Katrina Epnere, Albert Rizzo, Mary Partin, Robert M. Ahlness, and Matthew Trimmer. 2018. Using Virtual Interactive Training Agents (ViTA) with Adults with Autism and Other Developmental Disabilities. Journal of Autism and Developmental Disorders 48, 3 (March 2018), 905–912. doi:10.1007/s10803-017-3374-z

  10. [10]

    Dasom Choi, Sunok Lee, Sung-In Kim, Kyungah Lee, Hee Jeong Yoo, Sangsu Lee, and Hwajung Hong. 2024. Unlock Life with a Chat(GPT): Integrating Conversa- tional AI with Large Language Models into Everyday Lives of Autistic Individuals. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Association for C...

  11. [11]

    Victoria Clarke and Virginia Braun. 2017. Thematic analysis. The journal of positive psychology 12, 3 (2017), 297–298

  12. [12]

    Nyaz Didehbani, Tandra Allen, Michelle Kandalaft, Daniel Krawczyk, and Sandra Chapman. 2016. Virtual Reality Social Cognition Training for children with high functioning autism. Computers in Human Behavior 62 (2016), 703–711. doi:10.1016/j.chb.2016.04.033

  13. [13]

    Ricardo Eiris, Masoud Gheisari, and Behzad Esmaeili. 2018. PARS: Using Aug- mented 360-Degree Panoramas of Reality for Construction Safety Training. International Journal of Environmental Research and Public Health 15, 11 (Nov. 2018), 2452. doi:10.3390/ijerph15112452

  14. [14]

    Sandra G. Hart. 2006. Nasa-Task Load Index (NASA-TLX); 20 Years Later. Proceedings of the Human Factors and Ergonomics Society Annual Meeting 50, 9 (2006), 904–908. doi:10.1177/154193120605000909

  15. [15]

    Hart and Lowell E

    Sandra G. Hart and Lowell E. Staveland. 1988. Development of NASA-TLX (Task Load Index): Results of Empirical and Theoretical Research. In Human Mental Workload, Peter A. Hancock and Najmedin Meshkati (Eds.). Advances in Psychology, Vol. 52. North-Holland, 139–183. doi:10.1016/S0166-4115(08)62386-9

  16. [16]

    Arno Hartholt, Sharon Mozgai, Ed Fast, Matt Liewer, Adam Reilly, Wendy Whitcup, and Albert "Skip" Rizzo. 2019. Virtual Humans in Augmented Reality: A First Step towards Real-World Embedded Virtual Roleplayers. In Proceedings of the 7th International Conference on Human-Agent Interaction (Kyoto, Japan) (HAI ’19). Association for Computing Machinery, New Yo...

  17. [17]

    Darren Hedley, Mirko Uljarević, Lauren Cameron, Santoshi Halder, Amanda Richdale, and Cheryl Dissanayake. 2017. Employment programmes and inter- ventions targeting adults with autism spectrum disorder: A systematic review of the literature. Autism 21, 8 (2017), 929–941. PMID: 27542395. doi:10.1177/ 1362361316661855

  18. [18]

    Dawn Hendricks. 2010. Employment and adults with autism spectrum disorders: Challenges and strategies for success. Journal of Vocational Rehabilitation 32, 2 (Jan. 2010), 125–134. Publisher: IOS Press. doi:10.3233/JVR-2010-0502

  19. [19]

    Wenlong Huang, Fei Xia, Ted Xiao, Harris Chan, Jacky Liang, Pete Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, Pierre Sermanet, Noah Brown, Tomas Jackson, Linda Luu, Sergey Levine, Karol Hausman, and Brian Ichter. 2022. Inner Monologue: Embodied Reasoning through Planning with Language Models. arXiv:2207.05608 [cs.RO] https://arxi...

  20. [20]

    It’s the only thing I can trust

    JiWoong Jang, Sanika Moharana, Patrick Carrington, and Andrew Begel. 2024. “It’s the only thing I can trust”: Envisioning Large Language Model Use by Autistic Workers for Communication Assistance. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Association for Computing Machinery, New York, NY...

  21. [21]

    Kandalaft, Nyaz Didehbani, Daniel C

    Michelle R. Kandalaft, Nyaz Didehbani, Daniel C. Krawczyk, Tandra T. Allen, and Sandra B. Chapman. 2012. Virtual Reality Social Cognition Training for Young Adults with High-Functioning Autism. Journal of Autism and Developmental Disorders 43, 1 (May 2012), 34–44. doi:10.1007/s10803-012-1544-6

  22. [22]

    Sung-In Kim, So-Youn Jang, Taewan Kim, Bogoan Kim, Dayoung Jeong, Taehyung Noh, Mingon Jeong, Kaely Hall, Meelim Kim, Hee Jeong Yoo, Kyungsik Han, Hwajung Hong, and Jennifer G Kim. 2024. Promoting Self-Efficacy of Individuals With Autism in Practicing Social Skills in the Workplace Using Virtual Reality and Physiological Sensors: Mixed Methods Study. JMIR...

  23. [23]

    Panagiotis Kourtesis, Evangelia-Chrysanthi Kouklari, Petros Roussos, Vasileios Mantas, Katerina Papanikolaou, Christos Skaloumbakas, and Artemios Pehli- vanidis. 2023. Virtual Reality Training of Social Skills in Adults with Autism Spectrum Disorder: An Examination of Acceptability, Usability, User Experience, Social Skills, and Executive Functions. Behav...

  24. [24]

    Hannah R Lawrence, Renee A Schneider, Susan B Rubin, Maja J Matarić, Daniel J McDuff, and Megan Jones Bell. 2024. The Opportunities and Risks of Large Language Models in Mental Health. JMIR Ment Health 11 (July 2024), e59479

  25. [25]

    Ziming Li, Pinaki Prasanna Babar, Mike Barry, and Roshan L Peiris. 2024. Ex- ploring the Use of Large Language Model-Driven Chatbots in Virtual Reality to Train Autistic Individuals in Job Communication Skills. In Extended Abstracts of the 2024 CHI Conference on Human Factors in Computing Systems (CHI EA ’24). Association for Computing Machinery, New York...

  26. [26]

    Ziming Li, Pinaki Prasanna Babar, and Roshan L Peiris. 2025. Generative Role-Play Communication Training in Virtual Reality for Autistic Individuals: A Study on Job Coach Experiences in Vocational Training Programs. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). Association for Computing Machinery, New York, NY,...

  27. [27]

    Ziming Li, Huadong Zhang, Chao Peng, and Roshan Peiris. 2025. Exploring Large Language Model-Driven Agents for Environment-Aware Spatial Interactions and Conversations in Virtual Reality Role-Play Scenarios. In 2025 IEEE Conference Virtual Reality and 3D User Interfaces (VR). 1–11. doi:10.1109/VR59515.2025.00025 ASSETS ’26, October 25–28, 2026, Vila Nova ...

  28. [28]

    MARCO, LEIGHTON B.N

    ELYSA J. MARCO, LEIGHTON B.N. HINKLEY, SUSANNA S. HILL, and SRIKAN- TAN S. NAGARAJAN. 2011. Sensory Processing in Autism: A Review of Neu- rophysiologic Findings. Pediatric Research 69, 5 Part 2 (May 2011), 48R–54R. doi:10.1203/pdr.0b013e3182130c54

  29. [29]

    Anne Masi, Marilena M DeMayo, Nicholas Glozier, and Adam J Guastella. 2017. An Overview of Autism Spectrum Disorder, Heterogeneity and Treatment Op- tions. Neurosci Bull 33, 2 (Feb. 2017), 183–193

  30. [30]

    Masullo and M

    H. Masullo and M. Galloro. 2025. Development of Virtual Reality Systems for Social Skills Training in Individuals with Autism Spectrum Disorder: A Systematic Literature Review. European Psychiatry 68, S1 (April 2025), S706–S707. doi:10. 1192/j.eurpsy.2025.1436

  31. [31]

    Mesibov and Victoria Shea

    Gary B. Mesibov and Victoria Shea. 2010. The TEACCH Program in the Era of Evidence-Based Practice. Journal of Autism and Developmental Disorders 40, 5 (01 May 2010), 570–579. doi:10.1007/s10803-009-0901-6

  32. [32]

    Leahy, Chien-Chun Lin, and Boyang Tong

    Nigel Newbutt, Connie Sung, Hung-Jen Kuo, Michael J. Leahy, Chien-Chun Lin, and Boyang Tong. 2016. Brief Report: A Pilot Study of the Use of a Virtual Reality Headset in Autism Populations. Journal of Autism and Developmental Disorders 46, 9 (2016), 3166–3176. doi:10.1007/s10803-016-2830-5

  33. [33]

    David B Nicholas, Mark Attridge, Lonnie Zwaigenbaum, and Margaret Clarke

  34. [34]

    Opal Ousley and Tracy Cermak. 2014. Autism Spectrum Disorder: Defining Dimensions and Subgroups. Current Developmental Disorders Reports 1, 1 (01 Mar 2014), 20–28. doi:10.1007/s40474-013-0003-1

  35. [35]

    Mengxu Pan, Alexandra Kitson, Hongyu Wan, and Mirjana Prpa. 2025. ELLMA-T: an Embodied LLM-agent for Supporting English Language Learning in Social VR. In Proceedings of the 2025 ACM Designing Interactive Systems Conference (DIS ’25). Association for Computing Machinery, New York, NY, USA, 576–594. doi:10.1145/3715336.3735786

  36. [36]

    Chris Papadopoulos. 2024. Large language models for autistic and neurodivergent individuals: Concerns, benefits and the path forward. Neurodiversity 2 (2024), 27546330241301938. doi:10.1177/27546330241301938

  37. [37]

    Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 1–22

  38. [38]

    As an Autistic Person Myself:

    Sohyeon Park, Aehong Min, Jesus Armando Beltran, and Gillian R Hayes. 2025. "As an Autistic Person Myself:" The Bias Paradox Around Autism in LLMs. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). Association for Computing Machinery, New York, NY, USA, Article 774, 17 pages. doi:10.1145/3706598.3713420

  39. [39]

    Sarah Parsons and Peter Mitchell. 2002. The potential of virtual reality in social skills training for people with autistic spectrum disorders. Journal of Intellectual Disability Research 46, Pt 5 (June 2002), 430–443. doi:10.1046/j.1365-2788.2002. 00425.x

  40. [40]

    K. A. Ritter and Terrence L. Chambers. 2021. Three-dimensional modeled envi- ronments versus 360 degree panoramas for mobile virtual reality training. Virtual Reality 26, 2 (March 2021), 571–581. doi:10.1007/s10055-021-00502-9

  41. [41]

    Melissa Scott, Ben Milbourn, Marita Falkmer, Melissa Black, Sven Bölte, Aly- cia Halladay, Matthew Lerner, Julie Lounds Taylor, and Sonya Girdler. 2019. Factors impacting employment for people with autism spectrum disorder: A scoping review. Autism 23, 4 (2019), 869–901. PMID: 30073870. doi:10.1177/ 1362361318787789

  42. [42]

    Murray Shanahan, Kyle McDonell, and Laria Reynolds. 2023. Role play with large language models. Nature 623, 7987 (2023), 493–498

  43. [43]

    Matthew J Smith, Emily J Ginger, Katherine Wright, Michael A Wright, Julie Lounds Taylor, Laura Boteler Humm, Dale E Olsen, Morris D Bell, and Michael F Fleming. 2014. Virtual reality job interview training in adults with autism spectrum disorder. Journal of autism and developmental disorders 44 (2014), 2450–2463

  44. [44]

    Smith, Rogério M

    Matthew J. Smith, Rogério M. Pinto, Leann Dawalt, J.D. Smith, Kari Sherwood, Rashun Miles, Julie Taylor, Kara Hume, Tamara Dawkins, Mary Baker-Ericzén, Thomas Frazier, Laura Humm, and Chris Steacy. 2020. Using community-engaged methods to adapt virtual reality job-interview training for transition-age youth on the autism spectrum. Research in Autism Spect...

  45. [45]

    Iulia Stanica, Maria-Iuliana Dascalu, Constanta Nicoleta Bodea, and Alin Dragos Bogdan Moldoveanu. 2018. VR Job Interview Simulator: Where Virtual Real- ity Meets Artificial Intelligence for Education. In 2018 Zooming Innovation in Consumer Technologies Conference (ZINC). 9–12. doi:10.1109/ZINC.2018.8448645

  46. [46]

    Cheryl Y Trepagnier, Dale E Olsen, Laura Boteler, and Corinne A Bell. 2011. Virtual conversation partner for adults with autism. Cyberpsychology, Behavior, and Social Networking 14, 1-2 (2011), 21–27

  47. [47]

    Sai Vemprala, Rogerio Bonatti, Arthur Bucker, and Ashish Kapoor. 2023. Chatgpt for robotics: Design principles and model abilities. arXiv preprint arXiv:2306.17582 (2023)

  48. [48]

    Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. 2023. Voyager: An Open-Ended Embodied Agent with Large Language Models. arXiv:2305.16291 [cs.AI] https://arxiv.org/ abs/2305.16291

  49. [49]

    Paul Wehman, Valerie Brooke, Alissa Molinelli Brooke, Whitney Ham, Carol Schall, Jennifer McDonough, Stephanie Lau, Hannah Seward, and Lauren Avel- lone. 2016. Employment for adults with autism spectrum disorders: A retrospec- tive review of a customized employment approach. Research in Developmental Disabilities 53-54 (2016), 61–72. doi:10.1016/j.ridd.20...

  50. [50]

    Paul Wehman, Fong Chan, Nicole Ditchman, and Hyun-Ju Kang. 2014. Effect of Supported Employment on Vocational Rehabilitation Outcomes of Transition-Age Youth With Intellectual and Developmental Disabilities: A Case Control Study. Intellectual and developmental disabilities 52 (08 2014), 296–310. doi:10.1352/1934- 9556-52.4.296

  51. [51]

    Wobbrock, Shaun K

    Jacob O. Wobbrock, Shaun K. Kane, Krzysztof Z. Gajos, Susumu Harada, and Jon Froehlich. 2011. Ability-Based Design: Concept, Principles and Examples. ACM Trans. Access. Comput. 3, 3, Article 9 (April 2011), 27 pages. doi:10.1145/1952383. 1952384

  52. [52]

    Xiuqi Tommy Zhu, Heidi Cheerman, Minxin Cheng, Sheri R Kiami, Leanne Chukoskie, and Eileen McGivney. 2025. Designing VR Simulation System for Clinical Communication Training with LLMs-Based Embodied Conversational Agents. In Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems (CHI EA ’25). Association for Comp...

  53. [2015]

    Autism 19, 2 (2015), 235–245

    Vocational support approaches in autism spectrum disorder: A synthesis review of the literature. Autism 19, 2 (2015), 235–245. PMID: 24449603. doi:10. 1177/1362361313516548