Pith. sign in

REVIEW 4 major objections 4 minor 71 references

Data-Driven and Participatory Approaches toward Neuro-Inclusive AI

T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The dissertation argues that AI systems built to mimic human communication systematically reproduce anti-autistic ableism, and that the path forward is to de-center neuronormative 'humanness' as the benchmark for machine intelligence.

desk verdict A worthwhile dissertation with a real contribution in AUTALIC, but the abstract's 90% claim outruns the evidence in Chapter 2. read the letter →

arxiv 2507.21077 v1 pith:34XOGXDJ submitted 2025-06-12 cs.HC cs.AIcs.CY

classification cs.HCcs.AIcs.CY
keywords neuro-inclusiveAIanti-autisticableismhuman-robotinteractionLLMbiashatespeechdetectionparticipatorydesignneurodiversityAUTALIC
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The dissertation argues that AI systems built to mimic human communication systematically reproduce anti-autistic ableism: the medical model treats autism as a deficit of neurotypical skills, autistic people are excluded from designing the technologies aimed at them, and large language models both censor autistic speech and fail to catch ableist speech. The author offers a definition of Neuro-Inclusive AI as an explicit alternative to benchmarks that equate intelligence with mimicking humanness. Across five studies—an education intervention, a 142-paper critical review, interviews with 16 AI creators, an annotation study, and a new benchmark—the dissertation makes the case that the exclusion is structural, not incidental, and provides tools to correct it. A reader should care because if the claim is right, the default evaluation criteria for human-like AI need to change, and content-moderation systems built on current models would systematically harm autistic users.

What carries the argument

The machinery is threefold: (1) the concept of Neuro-Inclusive AI, defined as design that de-centers neuronormative benchmarks; (2) the AUTALIC benchmark, a dataset of Reddit posts annotated for anti-autistic ableist language in context, with a binary labeling scheme developed through co-design with annotators; and (3) the critical-analytic categories of pathologizing, essentialism, and power imbalance used to code the human-robot interaction corpus. The Turing Test functions as the negative benchmark: the dissertation treats 'mimicking humanness' as the assumption to be dismantled.

What would settle it

An independent replication of the Chapter 2 corpus selection with a preregistered search protocol, full-text screening, multiple blinded coders, and inter-coder reliability statistics would settle whether the 90% exclusion figure holds; a substantially lower rate, or the discovery that many excluded studies reported participatory design outside the searched text, would undercut the quantitative centerpiece. For the AUTALIC claim, re-annotating a random sample with a larger autistic community panel and comparing LLM agreement would test whether the misclassification results generalize.

Watch

Extended reading notes

Core claim

On its own terms, the dissertation establishes the following: current AI systems designed to imitate human communication are built on a neuronormative standard of humanness, and this reproduces anti-autistic ableism. In support, it reports a critical review of 142 human-robot interaction studies, finding that nearly 90% did not include autistic input in design and that a large majority applied the medical model; interviews with 16 creators of human-like agents, most of whom did not treat accessibility or neurodiversity as their responsibility; and an annotation study in which a binary ableist/not-ableist label captured annotator nuances. It then introduces AUTALIC, a dataset of Reddit sentences labeled in context for anti-autistic ableist language, and shows that four open-source LLMs frequently misclassify autistic community speech and miss ableist speech. The positive claim is the definition of Neuro-Inclusive AI: systems that de-center neuronormative benchmarks and do not take 'mimicking humanness' as the goal.

Load-bearing premise

The dissertation's headline statistic—that nearly 90% of human-like AI agents exclude autistic perspectives—rests on a manually assembled corpus of 142 human-robot interaction papers, selected from venue lists and the first 50 Google Scholar results per venue, with thematic codes generated by the authors and no reported inter-coder reliability.

Editorial extensions

If this is right

  • If the dissertation is right, any AI system that treats 'human-like communication' as its goal should specify whose communication style is being modeled; otherwise it will inherit neuronormative assumptions.
  • Human-robot interaction research on autism should be reoriented from diagnosis and treatment toward support and participation, with autistic people as co-designers rather than subjects.
  • AUTALIC gives a concrete tool for fine-tuning or evaluating LLMs on anti-autistic ableist language; current popular models do poorly on it, so content-moderation deployments should not rely on them without such tuning.
  • A simplified binary annotation scheme can replace more complex labeling schemes for this task, making it cheaper and more consistent to build datasets for anti-autistic language.
  • Because many makers of human-like AI do not see ethics and accessibility as their responsibility, education, funding priorities, and organizational standards need to embed neuro-inclusion rather than treating it as an add-on.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A broader implication not stated in the dissertation: the same critique should apply to other neurodivergent and disabled populations, so any benchmark that defines 'human-like' behavior needs a representativeness check before deployment.
  • The AUTALIC findings imply keyword-based toxicity filters will be brittle in general; a concrete extension is stress-testing moderation pipelines on context-heavy examples from other marginalized groups.
  • The binary-label result suggests that finer-grained annotation schemes are not automatically better; a testable hypothesis is that co-designed binary schemes improve agreement in other hate-speech labeling tasks.
  • The dissertation's use of the contact hypothesis points to an untested opportunity: if human-like agents displayed neurodiverse communication styles, they might reduce real-world dehumanization of autistic people.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The dissertation, arXiv:2507.21077, argues that current AI systems built to mimic human communication reproduce anti-autistic ableism. It defines "Neuro-Inclusive AI" as an approach that de-centers neuronormative benchmarks, then presents five studies: a formative pilot of an autism-inclusive communication course for IT workers (Chapter 1); a critical review of 142 human-robot interaction papers quantifying the exclusion of autistic perspectives (Chapter 2); interviews with 16 creators of human-like AI about ethics, accessibility, and neurodiversity (Chapter 3); a participatory annotation study with six annotators comparing labeling schemes for anti-autistic language (Chapter 4); and AUTALIC, a benchmark dataset of Reddit posts annotated for anti-autistic ableist language, used to evaluate four LLMs (Chapter 5). The abstract claims "90% of human-like AI agents exclude autistic perspectives," a headline that generalizes beyond the HRI corpus in Chapter 2, and the dissertation also claims that LLMs frequently misclassify autistic community speech while failing to identify ableist speech. The work is framed as a participatory, community-oriented contribution with practical resources.

Significance. If the claims hold, the dissertation makes a valuable intervention: it names a concrete mechanism by which human-like AI can encode neuronormative assumptions, provides AUTALIC as a reusable resource for fine-tuning and evaluating models on anti-autistic language in context, and centers the perspectives of autistic and neurodivergent annotators and researchers throughout. The qualitative themes—pathologization, essentialism, and power imbalance in HRI research—are internally consistent and align with critical disability studies and prior work. The author also consistently acknowledges limitations: small samples, snowball recruitment, formative assessment, and US-centric perspectives. However, the most quantified headline, the "90%" exclusion statistic, rests on a hand-coded convenience corpus with no reported inter-coder reliability, and the abstract extends this from HRI papers to all human-like AI agents. That overreach is load-bearing for the abstract's central claim and must be fixed before the quantitative conclusions can be relied upon.

major comments (4)
  1. [Abstract; §2.2.2; §2.3.3; Figure 2.11] The abstract's claim that "90% of human-like AI agents exclude autistic perspectives" is not supported by the evidence in Chapter 2. What Chapter 2 actually reports is that, in a hand-assembled corpus of 142 HRI papers from selected venues between 2016 and 2022, "nearly 90% of the papers did not include the input of autistic people in the design process." The leap from a venue-specific, keyword-searched HRI corpus to all human-like AI agents (chatbots, voice assistants, embodied agents beyond robots) has no sampling frame. The recommendation should be to soften the abstract to state exactly what was measured, e.g., "in our corpus of HRI studies," or to provide additional evidence covering other classes of human-like agents.
  2. [§2.2.2; §2.2.4; §2.3.3] The headline 90% exclusion statistic inherits two unaddressed methodological fragilities. First, the corpus is built from venue lists plus the first 50 Google Scholar results per venue, which is a relevance-ranked convenience sample; no recall/precision audit, exclusion log, or inter-coder reliability statistic is reported for the manual coding of the "design input" variable. Second, the codebook for what counts as "input in the design process" is not provided, and no denominator rule is given for survey papers, editorials, or papers where no design process is described. If even 10–15 of the 142 papers were recoded, the headline would fall below 80%. The qualitative themes may be robust, but the quantitative headline is not; the manuscript should either report reliability and a clear coding protocol or demote this statistic from headline to context.
  3. [§5.4.2; Table 5.3; Figure 5.4] The claim that LLMs "frequently misclassify autistic community speech" and "fail to identify ableist speech due to reliance on simplistic keyword-based methods" needs a stronger quantitative basis. As presented, Table 5.3 reports F1 scores across prompts and in-context learning examples, and Figure 5.4 reports mean Cohen's Kappa values, but there is no baseline comparison (random classifier, keyword-matching classifier, or majority-class classifier), no confidence intervals or variance estimates across prompt repetitions, and no statistical test comparing human-LLM agreement against human-human agreement. Without these, the conclusion that LLM performance is deficient rather than merely moderate is not fully established. The AUTALIC resource remains useful, but the evaluation section should be explicit about the threshold for "frequent misclassification."
  4. [§4.3–§4.5] The claim that a binary annotation scheme "sufficiently captures the nuances" of labeling anti-autistic language is based on a study with six annotators and a co-design session. The manuscript acknowledges the small sample, but the generalizability of this conclusion to other annotator pools, other platforms, and other social media contexts is limited. Since this claim is used to justify the binary labeling in AUTALIC, a brief statement tying the binary scheme's sufficiency to the annotators' own reported preferences, rather than to a general claim about all annotation tasks, would make the inference more precise.
minor comments (4)
  1. [Throughout] There are multiple typographical errors, including "Pyschology" (Figure 2.5 caption), "contexualize" (§2.3.3), "immitation" (§3.2.1), "neuornormative" and "neuronomativity" (§3.1 and §3.2), and "AUTALICdataset" (Table 5.1). These should be corrected in a careful copyedit.
  2. [Table 1 and Table 1 (page 4 vs. page 9)] Two tables are both labeled "Table 1" (the key-terms glossary appears twice with identical numbering). The duplicated numbering should be fixed, and the glossary table should appear once or be renumbered consistently.
  3. [Figure 2.2] The sentence "There is a notable correlation between thehumanness" of robots and a power imbalance" contains a formatting error (missing space and an unmatched quote mark). Also, "correlation" here appears to mean an observed association in categorical codes, not a statistical correlation; the wording should be clarified.
  4. [§2.2.3; Figure 2.6] The claim that the majority of referenced works were published before Critical Autism Studies introduced inclusive theories is based on an annualized comparison, but no base-rate comparison is given for the overall publication volume in those fields. Without such a baseline, the figure may overstate the shift. A brief caveat would help.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the dissertation's claims are empirical measurements from corpus coding and annotator labels, not definitional or fitted outputs.

full rationale

The dissertation contains no equation-level derivation loop. The central quantified claim in Chapter 2 (that 'nearly 90% of the papers did not include the input of autistic people in the design process') is a descriptive count produced by manual thematic coding of a 142-paper corpus (Ch. 2.2.4, Table 2.2). The coding variable is described in the paper's own definition of autism inclusion, but the 90% figure is an empirical result rather than a restatement of the definition; its fragilities are convenience sampling and unreported inter-coder reliability, which are external-validity and measurement concerns, not circularity. The abstract's wording '90% of human-like AI agents' broadens the HRI-specific finding, but that is an overgeneralization, not a constructional equivalence. In Chapter 5, AUTALIC labels come from human annotators and are treated as external ground truth for evaluating LLMs; the LLM agreement scores (Cohen's Kappa) compare model outputs to those human labels, so no model output defines the benchmark or the outcome. The self-citations (Rizvi et al. 2021, 2024) provide framing and are cited as prior work, but the load-bearing theoretical claims about double empathy, dehumanization, and the medical model are also anchored in non-overlapping sources such as Milton 2012, Kapp et al. 2013, Baron-Cohen 1997, Williams 2021b, Woods et al. 2018, and Spiel et al. 2022. Thus the central claims do not reduce to the paper's own inputs by construction.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No numeric free parameters are fitted to data. The central claims depend on qualitative coding assumptions, corpus representativeness, and the decision to treat annotator majority labels as ground truth for the AUTALIC benchmark.

assumptions (3)
  • domain assumption The 142-paper HRI corpus, assembled from venue searches and the first 50 Google Scholar results per venue, is representative of human-robot interaction research on autism.
    Chapter 2.2.2. The 90% exclusion statistic is computed from this corpus; selection bias in venue choice or relevance filtering propagates into the headline claim.
  • domain assumption Author-generated thematic codes such as medical model, pathologizing, essentialism, and power imbalance reliably capture the intended constructs.
    Chapter 2.2.4 and 2.3. The central percentages depend on manual coding by three authors without reported inter-coder reliability metrics.
  • domain assumption Anti-autistic ableist language can be operationalized by keyword-filtered Reddit sentences labeled by recruited annotators, and majority annotator labels are an appropriate ground truth for LLM evaluation.
    Chapter 5.3 and 5.4. Keyword selection and annotator ground truth define the benchmark; external validation is not provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Data-Driven and Participatory Approaches toward Neuro-Inclusive AI." pith.science (2026). https://pith.science/paper/34XOGXDJ

@misc{pith2026250721077,
  author       = {Pith},
  title        = {Pith review of: Data-Driven and Participatory Approaches toward Neuro-Inclusive AI},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/34XOGXDJ}},
  note         = {Machine review of arXiv:2507.21077}
}
read the original abstract

Biased data representation in AI marginalizes up to 75 million autistic people worldwide through medical applications viewing autism as a deficit of neurotypical social skills rather than an aspect of human diversity, and this perspective is grounded in research questioning the humanity of autistic people. Turing defined artificial intelligence as the ability to mimic human communication, and as AI development increasingly focuses on human-like agents, this benchmark remains popular. In contrast, we define Neuro-Inclusive AI as datasets and systems that move away from mimicking humanness as a benchmark for machine intelligence. Then, we explore the origins, prevalence, and impact of anti-autistic biases in current research. Our work finds that 90% of human-like AI agents exclude autistic perspectives, and AI creators continue to believe ethical considerations are beyond the scope of their work. To improve the autistic representation in data, we conduct empirical experiments with annotators and LLMs, finding that binary labeling schemes sufficiently capture the nuances of labeling anti-autistic hate speech. Our benchmark, AUTALIC, can be used to evaluate or fine-tune models, and was developed to serve as a foundation for more neuro-inclusive future work.

Figures

Figures reproduced from arXiv: 2507.21077 by the authors.

Figure 1.1
Figure 1.1. The introduction slide to the online course. On the left is the course outline. 17 [PITH_FULL_IMAGE:figures/full_fig_p010_1_1.png] view at source ↗
Figure 2.3
Figure 2.3. An overview of the disability models applied in our main corpus reveals the medical model is the most commonly used model in HRI research. Papers were categorized as ’other’ if they did not exclusively fall under the medical or social model. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42 [PITH_FULL_IMAGE:figures/full_fig_p010_2_3.png] view at source ↗
Figure 2.6
Figure 2.6. The majority of our referenced works corpus was published when deficit￾based theories dominated autism research. The red area represents the papers that were published before the introduction of Critical Autism Studies which seeks to challenge dominant misunderstandings of autism and have a positive impact on the lives of autistic people (O’Dell et al., 2016). This data has been annualized to allow for a fair compar… view at source ↗
Figures from the paper (30 more)
Figure 2.7
Figure 2.7. Figure 2.7: An overview of the models applied in our referenced works corpus reveals that the medical model of autism is the most pervasive in these works. . . . 45 x [PITH_FULL_IMAGE:figures/full_fig_p010_2_7.png]
Figure 2.8
Figure 2.8. Figure 2.8: The majority of studies collecting eye contact data from the users focused on clinical applications such as therapy, skills training, or diagnosis. . . . . . 46 [PITH_FULL_IMAGE:figures/full_fig_p011_2_8.png]
Figure 4.1
Figure 4.1. Figure 4.1: Presenting the labeling schemes (a) used in all the rounds for score-based [PITH_FULL_IMAGE:figures/full_fig_p012_4_1.png]
Figure 4.2
Figure 4.2. Figure 4.2: The collaborative re-labeling activities our annotators complete in our co￾design session include assigning labels based on a static scale (a), assigning labels based on our scoring algorithm (b), and arranging sentences in order (c). These activities are designed to…
Figure 4.3
Figure 4.3. Figure 4.3: Comparing our annotator agreement using different labeling schemes (a) and annotation techniques (b). Although the granularity of the labels varies as discussed in Section 4.3.4, these scores have been converted to a binary classification of ableist or not ableist to…
Figure 5.2
Figure 5.2. Figure 5.2: An example of a sentence from our dataset. The search keyword is shown in red, while the word in blue is an example of a word defined in our glossary. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 119 […
Figure 1.1
Figure 1.1. Figure 1.1: The introduction slide to the online course. On the left is the course outline. The 15-minute course, divided into 4 units, employs video skits, quizzes, audio narrations, and charts to explain core concepts. The title screen of the course is shown in [PITH_FULL_IMA…
Figure 1.2
Figure 1.2. Figure 1.2: A picture of two coworkers talking to one another in an education video about stimming. The autistic coworker on the right is overwhelmed after her colleague on the left calls out her hair twirling. 1.5.1 Course Outline Our communication skills course is divided into…
Figure 2.1
Figure 2.1. Figure 2.1: Distribution of publication years of the papers included in our main corpus. Venue # Res. # Sel. ACM/IEEE Int. Conf. on Human-Robot Interaction (HRI)* 50 36 IEEE Intl. Symp. on Rob. & Human Interactive Comm. (RO-MAN)* 48 36 Int. Journal of Social Robotics (JSR)* 50 2…
Figure 2.2
Figure 2.2. Figure 2.2: There is a notable correlation between thehumanness” of robots and a power imbalance in user interactions with autistic people in our main corpus. More human-like robots were placed in mentor roles for diagnosis or helping autistic people move toward ’humanness’ (Wil…
Figure 2.4
Figure 2.4. Figure 2.4: The majority of papers in our main corpus focused on providing treatment to the end-users. Support is shown in orange as the codes for this theme are not exclusive to a particular disability model, while the codes for autism inclusion (light blue) exclusively do not …
Figure 2.5
Figure 2.5. Figure 2.5: It is important to note both psychology and medicine have studied autism using a [PITH_FULL_IMAGE:figures/full_fig_p064_2_5.png]
Figure 2.5
Figure 2.5. Figure 2.5: Psychology, medicine, and computer science are the most frequently referenced fields by our main corpus. Pyschology and medicine are shown in red as they have historically applied a medi￾cal model approach by viewing autism as a disorder (Golt and Kana, 2022). The fo…
Figure 2.7
Figure 2.7. Figure 2.7: An overview of the models applied in our referenced works corpus reveals that the medical model of autism is the most pervasive in these works. have identified barriers to effective communication that may arise due to neurotypical individuals misunderstanding, respon…
Figure 2.8
Figure 2.8. Figure 2.8: The majority of studies collecting eye contact data from the users focused on clinical applications such as therapy, skills training, or diagnosis. the double empathy problem entirely on autistic people. Prior work has pathologized autistic people’s unique communicat…
Figure 2.9
Figure 2.9. Figure 2.9: The majority of studies in our main corpus (shown in orange) did not report their gender data. Among the papers that did report, only 12 studies had a gender ratio that is representative of the autistic population (shown in blue) (Maenner et al., 2023). et al., 2019)…
Figure 2.10
Figure 2.10. Figure 2.10: An overview of the ages of the participants reported in our main corpus. The participant ages were reported in n=139 of the papers in our main corpus. in n=55 papers applying a binary definition of gender. The majority of these studies did not have a representation …
Figure 2.9
Figure 2.9. Figure 2.9: Future work should consider reporting this data to contextualize their results without [PITH_FULL_IMAGE:figures/full_fig_p072_2_9.png]
Figure 2.11
Figure 2.11. Figure 2.11: The majority of HRI studies did not include autistic people in the design stage black, and nearly one-fifth of them did not include the perspectives of autistic people at any stage of their study. Cohen, 1997). Yet, we found a correlation between the user’s eye cont…
Figure 3.1
Figure 3.1. Figure 3.1: An overview of our participants’ knowledge and acceptance of neurodi￾vergence. 3.4.1 Desirable Traits Through a qualitative thematic analysis of the participants’ responses during our interviews, we uncovered 12 desirable traits they implement to make their technolog…
Figure 3.2
Figure 3.2. Figure 3.2: A list of undesir￾able traits discussed by our par￾ticipants that may upset or frus￾trate their users [PITH_FULL_IMAGE:figures/full_fig_p093_3_2.png]
Figure 4.1
Figure 4.1. Figure 4.1: Presenting the labeling schemes (a) used in all the rounds for score-based labeling and the logic of our black-box algorithm (b), used in round 4 to dynamically assign scores by asking the annotators a series of questions about each sentence. In our co-design session…
Figure 4.2
Figure 4.2. Figure 4.2: The collaborative re-labeling activities our annotators complete in our co-design session include assigning labels based on a static scale (a), assigning labels based on our scoring algorithm (b), and arranging sentences in order (c). These activities are designed to…
Figure 4
Figure 4. Figure 4: figure 4.3, we share the improvement in our annotator the scores throughout each iteration [PITH_FULL_IMAGE:figures/full_fig_p123_4.png]
Figure 4.3
Figure 4.3. Figure 4.3: Comparing our annotator agreement using different labeling schemes (a) and annotation techniques (b). Although the granularity of the labels varies as discussed in Section 4.3.4, these scores have been converted to a binary classification of ableist or not ableist to…
Figure 5.1
Figure 5.1. Figure 5.1: The example illustrates the importance of labeling sentences in context. The target sentence alone, shown on the left, is difficult to classify as ableist toward autistic people. Adding the surrounding sentences, as shown on the right, provides context revealing the …
Figure 5.2
Figure 5.2. Figure 5.2: An example of a sentence from our dataset. The search keyword is shown in red, while the word in blue is an example of a word defined in our glossary. keywords using the default search settings, which filters posts based on relevancy by prioritizing rare words in the…
Figure 5.3
Figure 5.3. Figure 5.3: An example of a sentence from our dataset, shown in red, with a high level of disagreement among annotators. including a quantitative assessment of tag confusions that found the majority of disagreements are due to linguistically debatable cases rather than errors in…
Figure 5
Figure 5. Figure 5: contains an example of a sentence with a high disagreement among our [PITH_FULL_IMAGE:figures/full_fig_p146_5.png]
Figure 5.4
Figure 5.4. Figure 5.4: The mean Cohen’s Kappa scores of each LLM comparing the agreement with human annotators and other LLMs. In-Context Learning After providing the in-context learning examples, Llama3 (+22.96%) and Gemma2 (+12.68%) display the biggest relative improvement in F-1 scores,…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

71 extracted references · 61 canonical work pages

  1. [1]

    How do you feel about this new policy? • Strongly Disagree • Somewhat Disagree 180 • Neither agree nor disagree • Somewhat Agree • Strongly Agree

    Imagine your team switches to a remote-optional work environment and now allows your team members to have their videos off during meetings. How do you feel about this new policy? • Strongly Disagree • Somewhat Disagree 180 • Neither agree nor disagree • Somewhat Agree • Strongly Agree

  2. [2]

    Imagine your team does not allow people to have conversations in the open work spaces, and requires them to use designated spaces. How do you feel about this change in work policies? • Strongly Disagree • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree • Strongly Agree 3.Intersectionality impacts the work that I do. • Strongly Disagree • S...

  3. [3]

    How did you know that your work was effective? (ask if they created a solution to a problem, otherwise skip). a. Can you tell me about a time when a user interacted with your system and got frustrated or upset? b. What are some things that may make a user frustrated or not want to interact with your system?

  4. [4]

    Can you walk me through how or why you chose this specific mode of user interaction to reach your objective or solve your problem?

  5. [5]

    Autism research is in crisis

    doi:10.1145/3385980.3385985 Matthew Bennett, Amanda A Webster, Emma Goodall, and Susannah Rowland. 2019.Life on the autism spectrum: Translating myths and misconceptions into positive futures. Springer, Singapore. Katy Johanna Benson. 2023. Perplexing presentations: Compulsory neuronormativity and cognitive marginalisation in social work practice with aut...

  6. [6]

    • Strongly Disagree • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree • Strongly Agree

    People who work in computer science-related fields should participate in educational and training programs that help advance cultural competence within the profession. • Strongly Disagree • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree • Strongly Agree

  7. [7]

    People who work in computer science-related fields should understand the culture of computer science and its functions in human behavior and society, recognizing the strengths that exist in all cultures. • Strongly Disagree • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree • Strongly Agree 8.Because we live in the U.S., everyone should spe...

  8. [8]

    #DisabledOnIndianTwitter

    “#DisabledOnIndianTwitter” : A Dataset towards Understanding the Expression of People with Disabilities on Indian Twitter. InFindings of the Association for Computational Linguistics: AACL-IJCNLP 2022, Yulan He, Heng Ji, Sujian Li, Yang Liu, and Chua-Hui Chang (Eds.). Association for Computational Linguistics, Online only, 375–386. https: //aclanthology.o...

Show all 71 references
  1. [9]

    Is there anything else you’d like to say or add in regards to your experience building AI entities? A.2.2 Interview 2

  2. [11]

    theory of mind

    Annotators with attitudes: How annotator beliefs and identities bias toxic language detection.arXiv preprint arXiv:2111.07997(2021). Amanda Saxe. 2017. The theory of intersectionality: A new lens for understanding the barriers faced by autistic women.Canadian Journal of Disabi...

  3. [12]

    ACM Hum.-Comput

    The Potential of Diverse Youth as Stakeholders in Identifying and Mitigating Algorith- mic Bias for a Future of Fairer AI.Proc. ACM Hum.-Comput. Interact.7, CSCW2, Article 364 (Oct. 2023), 27 pages. doi:10.1145/3610213 Katta Spiel, Christopher Frauenberger, Eva Hornecker, and ...

  4. [13]

    Kaggle Team

    Gemma 2: Improving open language models at a practical size.arXiv preprint arXiv:2408.00118(2024). Kaggle Team. 2019. May 2015 Reddit Comments. https://www.kaggle.com/datasets/kaggle/ reddit-comments-may-2015 [Accessed 20-06-2024]. Priyanka Rebecca Tharian, Sadie Henderson, Na...

  5. [15]

    arXiv:2401.16292 Luke J

    Visual Stereotypes of Autism Spectrum in DALL-E, Stable Diffusion, SDXL, and Midjourney.arXiv preprint arXiv:2401.16292(jan 2024). arXiv:2401.16292 Luke J. Wood, Abolfazl Zaraki, Ben Robins, and Kerstin Dautenhahn. 2021. Developing Kaspar: a humanoid robot for children with au...

  6. [16]

    I Am Just Terrified of My Future

    Bias and Fairness in Chatbots: An Overview.APSIPA Transactions on Signal and Information Processing13, Article e102 (2024). doi:10.1561/116.00000064 Anon Ymous, Katta Spiel, Os Keyes, Rua M. Williams, Judith Good, Eva Hornecker, and Cynthia L. Bennett. 2020. "I Am Just Terrifi...

  7. [17]

    Which one of the following includes your total HOUSEHOLD income for last year, before taxes? • Less than $10,000 • $10,000 to under $20,000 • $20,000 to under $30,000 • $30,000 to under $40,000 • $40,000 to under $50,000 • $50,000 to under $65,000 • $65,000 to under $80,000 • ...

  8. [20]

    • Strongly Disagree 181 • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree • Strongly Agree

    It is my responsibility to support and advocate for recruitment and retention efforts in programs and agencies that ensure diversity. • Strongly Disagree 181 • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree • Strongly Agree

  9. [23]

    • Strongly Disagree • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree • Strongly Agree

    Membership in a minority group significantly increases risk factors for exposure to discrimination, economic deprivation, and oppression. • Strongly Disagree • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree • Strongly Agree

  10. [24]

    • Strongly Disagree • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree • Strongly Agree 14.Being lesbian, bisexual, or gay is a choice

    In the U.S., some people are often physically attacked because of their minority status. • Strongly Disagree • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree • Strongly Agree 14.Being lesbian, bisexual, or gay is a choice. • Strongly Disagree • Somewhat Disagr...

  11. [25]

    • True • False 18.In traffic, I am always polite and considerate of others

    I always admit my mistakes openly and face the potential for negative consequences. • True • False 18.In traffic, I am always polite and considerate of others. • True • False 19.I have tried illegal drugs (for example, marijuana, cocaine, etc.). • True • False 20.I always acce...

  12. [28]

    • True • False 29.During arguments I always stay objective and matter-of-fact

    I always stay friendly and courteous with other people, even when I am stressed out. • True • False 29.During arguments I always stay objective and matter-of-fact. • True • False

  13. [30]

    • True • False 31.I always eat a healthy diet

    There has been at least one occasion when I failed to return an item that I borrowed. • True • False 31.I always eat a healthy diet. • True • False 32.Sometimes I only help because I expect something in return. • True • False

  14. [33]

    • Strongly Disagree • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree 187 • Strongly Agree

    I am able to develop programs and services that reflect an understanding of the diversity between and within cultures. • Strongly Disagree • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree 187 • Strongly Agree

  15. [34]

    • Strongly Disagree • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree • Strongly Agree

    I feel confident about my knowledge and understanding of people with disabilities, needs, traditions, values, family systems, and artistic expressions. • Strongly Disagree • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree • Strongly Agree

  16. [35]

    • Strongly Disagree • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree • Strongly Agree

    I feel confident about my knowledge and understanding of African American and African history, traditions, values, family systems, and artistic expressions. • Strongly Disagree • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree • Strongly Agree

  17. [36]

    • Strongly Disagree • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree 188 • Strongly Agree

    I feel confident about my knowledge and understanding of Middle Eastern (South West Asian and North African) history, traditions, values, family systems, and artistic expressions. • Strongly Disagree • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree 188 • Stron...

  18. [37]

    • Strongly Disagree • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree • Strongly Agree

    I feel confident about my knowledge and understanding of women’s history, traditions, values, family systems, and artistic expressions. • Strongly Disagree • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree • Strongly Agree

  19. [38]

    • Strongly Disagree • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree • Strongly Agree

    I feel confident about my knowledge and understanding of gay/lesbian/bisexual/transgender history, traditions, values, family systems, and artistic expressions. • Strongly Disagree • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree • Strongly Agree

  20. [39]

    • Strongly Disagree • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree • Strongly Agree 189

    I feel confident about my knowledge and understanding of Jewish history, traditions, values, family systems, and artistic expressions. • Strongly Disagree • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree • Strongly Agree 189

  21. [40]

    • Strongly Disagree • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree • Strongly Agree

    I am aware of ways in which institutional oppression and the misuse of power constrain the human and legal rights of individuals and groups within American society. • Strongly Disagree • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree • Strongly Agree

  22. [41]

    • Strongly Disagree • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree • Strongly Agree

    I feel confident about my knowledge and understanding of Native American history, traditions, values, family systems, and artistic expressions. • Strongly Disagree • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree • Strongly Agree

  23. [42]

    • Strongly Disagree • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree 190 • Strongly Agree

    I have the knowledge to critique and apply culturally competent and social justice approaches to influence assessment, planning, access to resources, intervention, and research. • Strongly Disagree • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree 190 • Strongly Agree

  24. [43]

    • Strongly Disagree • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree • Strongly Agree A.2 Interviews A.2.1 Interview 1

    I feel confident about my knowledge and understanding of Asian and Asian American history, traditions, values, family systems, and artistic expressions. • Strongly Disagree • Somewhat Disagree • Neither agree nor disagree • Somewhat Agree • Strongly Agree A.2 Interviews A.2.1 ...

  25. [44]

    Tell me about a recent project you worked on where you built an AI system that others would potentially use. a. When did you know you’ve reached the end of the design process for your system? i. When the current assignment is complete. b. When did you know you’ve reached the e...

  26. [45]

    What are the metrics you use for evaluating the performance of your system? 191

  27. [46]

    If they say no, ASK - Can you think of any other interaction that might be similar to the interaction you developed? b

    What real-world interactions can you think of that resemble your user interaction? a. If they say no, ASK - Can you think of any other interaction that might be similar to the interaction you developed? b. What are the identities of the people in these interactions? c. Explain...

  28. [47]

    How would that be different with a person who is funny? b

    What behavioral or communicational changes could you make to the bot that would impact the user’s interaction with the bot? a. How would that be different with a person who is funny? b. How would that impact the user’s interactions with folks who are and aren’t funny/don’t use...

  29. [48]

    What are the best practices in bot design you have learned in your work? b

    What conclusions would researchers who are building off of your work reach or assume about human interactions, behavior, or AI? What about end-users? What conclusions might they reach? 192 a. What are the best practices in bot design you have learned in your work? b. How do en...

  30. [49]

    Who are the people most likely to use your chatbot? b

    In your opinion, should we design chatbots, robots, or other AI agents to meet the needs of various community groups? a. Who are the people most likely to use your chatbot? b. What would happen if someone from [specific group] used it? c. What would happen if you adjusted the ...

  31. [50]

    How would you describe your current professional position?

  32. [51]

    In your opinion, what experiences from your life greatly impacted or led you to your interest in your current career?

  33. [52]

    What made you stay or continue to pursue your current career?

  34. [53]

    193 Intersectionality Questions

    Tell me about a time when your personal background overlapped with your current position or your work. 193 Intersectionality Questions

  35. [54]

    Has an ism ever impacted you in your life? b

    Has an ism (e.g., racism, sexism, ableism, homophobia) ever impacted you during your computer science journey? a. Has an ism ever impacted you in your life? b. If no, ask: At work, do you interact with a number of people who have a similar racial, ethnic, or gender background ...

  36. [55]

    Has it ever impacted your work or research? If so, how?

    In your opinion, what is intersectionality? a. Has it ever impacted your work or research? If so, how?

  37. [56]

    Tell me about a time when you used or experienced intersectionality in your work/research. a. If non-tech, let them provide it. If research-related, ask: Do you have any examples related to your current position?

  38. [57]

    Tell me about a time when you used or experienced a push for inclusivity in your work/research

  39. [58]

    Tell me a time when intersectionality has impacted you at work

  40. [59]

    If so, what did the training look like? Can you describe it? b

    Have you had any training in working with or building tools that support inclusivity? a. If so, what did the training look like? Can you describe it? b. Did it help you?

  41. [60]

    If so, what did the support look like? Can you describe it? b

    Have you had any support in working with or building tools that support inclusivity? a. If so, what did the support look like? Can you describe it? b. Did it help you? Situational Awareness 194

  42. [61]

    How would you approach a conversation with them? a

    Suppose your new colleague shows up wearing headphones to a research conference and continues wearing them even while presenting their work. How would you approach a conversation with them? a. What if your coworker mentions they’re wearing it because the room is too loud?

  43. [62]

    Ramadan is a Muslim holiday where Muslims refrain from eating and drinking from sunrise to sunset

    Imagine that your company celebrates its anniversary during Ramadan. Ramadan is a Muslim holiday where Muslims refrain from eating and drinking from sunrise to sunset. This year is their 50th anniversary, and the plan is to have a week full of free buffet-style lunch and brunc...

  44. [63]

    How would you address this situation?

    Suppose your research assistant doodles during meetings, and your supervisor approaches you about the lack of professionalism exhibited by your RA. How would you address this situation?

  45. [64]

    deficient

    Max and Alex are working on a research project and are expected to provide feedback to each other to ensure the project stays on track. [Share a conversation transcript.] a. Why is Alex responding this way? b. Why is Max responding this way? c. How might you respond as their s...

  46. [65]

    ‘I need therapy’)

    Using medical terminology in a more personal manner (e.g. ‘I need therapy’)

  47. [66]

    Discussions + suggestions from neurodivergent people (community-generated discus- sions)

  48. [67]

    General discussions of the medical processes (unrelated to neurodivergence)

  49. [68]

    I am autistic

    Example: “I am autistic”, “As an autistic person, I think . . . ” Implicitly Ableist (Round 1-3) Critical disability theory (CDT) is a framework centering disability which challenges the ableist assumptions present in our society. Using this theory, we define “implicitly ablei...

  50. [69]

    Making assumptions about a disabled person’s abilities that would not be made about an able-bodied person

  51. [70]

    Inspiration porn

    “Inspiration porn”→looking at disabled people as “inspiration” for able-bodied people

  52. [71]

    Using condescending language such as saying a disabled person “suffers” with their disability

  53. [72]

    innocent

    Infantilizing disabled people by portraying them as unable to make their own decisions, or “innocent” and “pure”

  54. [73]

    She is confined to a wheelchair

    Example: “She is confined to a wheelchair”, “it must be awful living with bipolar disorder” 197 Implicitly Anti-Autistic (Round 3) For the purposes of this study, we are applying a critical disability framework and classifying any sentences that describe autism using medical t...

  55. [1097]

    Social" in

    doi:10.1109/ICEARS56391.2023.10084943 Shilpa Aggarwal and Beth Angus. 2015. Misdiagnosis versus missed diagnosis: diagnosing autism spectrum disorder in adolescents.Australasian Psychiatry23, 2 (2015), 120–123. Basma Alharbi, Hind Alamro, Manal Alshehri, Zuhair Khayyat, Manal ...

  56. [2011]

    race hygiene

    Classroom-based assistive technology: Collective use of interactive visual schedules by students with autism.. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems. ACM, Vancouver, BC, Canada, 1–10. Catherine J. Crompton, Martha Sharp, Harriet Axbey, Su...

  57. [2017]

    Weaponized autism

    Hate Me, Hate Me Not: Hate Speech Detection on Facebook. InItalian Conference on Cybersecurity. https://api.semanticscholar.org/CorpusID:8293149 Somin Wadhwa, Vivek Khetan, Silvio Amir, and Byron Wallace. 2023. RedHOT: A Corpus of Annotated Medical Questions, Experiences, and ...

  58. [2019]

    doi:10.1186/s13229-019-0267-5 Mirko Gelsomini, Giulia Leonardi, Marzia Degiorgi, Franca Garzotto, Simone Penati, Jacopo Silvestri, Noëlie Ramuzat, and Francesco Clasadonte

    The role of gender in the perception of autism symptom severity and future behavioral development.Molecular Autism10, Article 16 (2019). doi:10.1186/s13229-019-0267-5 Mirko Gelsomini, Giulia Leonardi, Marzia Degiorgi, Franca Garzotto, Simone Penati, Jacopo Silvestri, Noëlie Ra...

  59. [2020]

    doi:10.1007/s40732-019-00363-4 Luke Beardon

    The Effect of Educational Messages on Implicit and Explicit Attitudes towards Individ- uals on the Autism Spectrum versus Normally Developing Individuals.The Psychological Record70, 1 (2020), 123–145. doi:10.1007/s40732-019-00363-4 Luke Beardon. 2017.Autism and Asperger syndro...

  60. [2021]

    Xiaoyu from the Stars

    Agreeing to disagree: Annotating offensive language datasets with annotators’ disagree- ment.arXiv preprint arXiv:2109.13563(2021). Calvin A. Liang, Sean A. Munson, and Julie A. Kientz. 2021. Embracing Four Tensions in Human-Computer Interaction Research with Marginalized Peop...

  61. [2022]

    Vulnerable, Victimized, and Objectified

    SciTweets-A Dataset and Annotation Framework for Detecting Scientific Online Discourse. InProceedings of the 31st ACM International Conference on Information & Knowledge Management. 3988–3992. Martin S. Hagger, Nikos L. D. Chatzisarantis, and Stuart J. H. Biddle. 2002. A meta-...

  62. [2023]

    InProceedings of the 2023 Conference on Empirical Methods in Natural 160 Language Processing

    mRedditSum: A Multimodal Abstractive Summarization Dataset of Reddit Threads with Images. InProceedings of the 2023 Conference on Empirical Methods in Natural 160 Language Processing. Association for Computational Linguistics, Singapore, 4117–4132. doi:10.18653/v1/2023.emnlp-m...

  63. [2024]

    InProceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24)

    Are Robots Ready to Deliver Autism Inclusion?: A Critical Review. InProceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24). ACM, Honolulu, HI, USA, Article 487. doi:10.1145/3613904.3642681 163 Ben Robins, Kerstin Dautenhahn, and Janek Dubowski. 2006....

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.