{"id":"72b5d365-b5f0-4d2b-b7e9-57409274583c","arxiv_id":"2507.21077","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Across five studies, the dissertation documents exclusion of autistic perspectives in human-robot interaction research and releases AUTALIC, a Reddit-derived benchmark for anti-autistic hate speech detection.","lead":"This dissertation argues that AI systems built to mimic human communication encode neurotypical norms and marginalize autistic people. It contributes interview, annotation, and benchmark studies, including a new dataset for detecting anti-autistic language, AUTALIC.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline 90% exclusion statistic is load-bearing for the abstract's central claim, but it rests on a hand-coded corpus with no reported inter-coder reliability and an abstract that generalizes from HRI studies to all human-like AI agents.","rationale":"Chapter 2 is the dissertation's quantitative anchor; the abstract and introduction repeat the 90% figure as if it measured exclusion of autistic perspectives across human-like AI. The coding process is qualitative and hypothesis-driven, but the output is used as a descriptive statistic. The corpus is assembled from Google Scholar relevance ranking and venue lists; no PRISMA-style flow or sensitivity analysis is provided. Because the categories are author-generated and the coding was done by three authors with discussion to consensus, the absence of inter-coder reliability makes the point estimate unverifiable. This is not a disagreement with the critical-autism-studies consensus; it is a measurement concern. The concrete test (independent re-coding of a sample) would settle it. If re-coding confirms the rate, the concern is resolved; if not, the conditional verdict stands with the 90% claim removed or qualified. I therefore keep the reader's CONDITIONAL verdict unchanged.","tokens_in":48564,"tokens_out":5701,"duration_ms":62804,"concrete_test":"Release the full Chapter 2 codebook and have two independent coders, blind to the study hypotheses, re-apply the coding scheme (Tables 2.2-2.3, especially the 'autistic input in design process' variable) to a random sample of at least 30 papers from the posted Figshare corpus. Report Cohen's kappa per code and recompute the 'nearly 90%' rate with a 95% confidence interval, excluding surveys/editorials where 'design process' is undefined. If kappa < 0.6 or the confidence interval includes rates below 80%, the abstract's 90% claim should be revised and the thesis should rest on the qualitative rather than the quantitative framing.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central quantified claim ('90% of human-like AI agents exclude autistic perspectives') is the most load-bearing evidence for the thesis that AI systems built to mimic humanness reproduce anti-autistic ableism. It inherits two fragilities from Chapter 2. First, §2.2.2 builds the 142-paper corpus from venue lists plus the first 50 Google Scholar results per venue; this is a relevance-ranked convenience sample, not a systematic corpus with a documented recall/precision audit. Second, §2.2.4 reports that three authors generated the thematic codes (medical model, pathologizing, essentialism, power imbalance) but gives no codebook excerpt for the 'design input' variable, no inter-coder reliability statistic, and no denominator rule for survey/editorial papers where 'design process' may not apply. §2.3.3 then converts the 'no design input' code into 'nearly 90% of papers did not include the input of autistic people in the design process,' and the abstract rephrases this as excluding autistic perspectives. If re-coding shifts even 10-15 of 142 papers, the headline number falls below 80%, and the claim 'human-like AI agents' generalizes from HRI papers to all human-like agents without a sampling frame. The qualitative themes may survive, but the quantitative headline does not.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The dissertation, arXiv:2507.21077, argues that current AI systems built to mimic human communication reproduce anti-autistic ableism. It defines \"Neuro-Inclusive AI\" as an approach that de-centers neuronormative benchmarks, then presents five studies: a formative pilot of an autism-inclusive communication course for IT workers (Chapter 1); a critical review of 142 human-robot interaction papers quantifying the exclusion of autistic perspectives (Chapter 2); interviews with 16 creators of human-like AI about ethics, accessibility, and neurodiversity (Chapter 3); a participatory annotation study with six annotators comparing labeling schemes for anti-autistic language (Chapter 4); and AUTALIC, a benchmark dataset of Reddit posts annotated for anti-autistic ableist language, used to evaluate four LLMs (Chapter 5). The abstract claims \"90% of human-like AI agents exclude autistic perspectives,\" a headline that generalizes beyond the HRI corpus in Chapter 2, and the dissertation also claims that LLMs frequently misclassify autistic community speech while failing to identify ableist speech. The work is framed as a participatory, community-oriented contribution with practical resources.","tokens_in":48797,"tokens_out":3607,"duration_ms":44556,"significance":"If the claims hold, the dissertation makes a valuable intervention: it names a concrete mechanism by which human-like AI can encode neuronormative assumptions, provides AUTALIC as a reusable resource for fine-tuning and evaluating models on anti-autistic language in context, and centers the perspectives of autistic and neurodivergent annotators and researchers throughout. The qualitative themes—pathologization, essentialism, and power imbalance in HRI research—are internally consistent and align with critical disability studies and prior work. The author also consistently acknowledges limitations: small samples, snowball recruitment, formative assessment, and US-centric perspectives. However, the most quantified headline, the \"90%\" exclusion statistic, rests on a hand-coded convenience corpus with no reported inter-coder reliability, and the abstract extends this from HRI papers to all human-like AI agents. That overreach is load-bearing for the abstract's central claim and must be fixed before the quantitative conclusions can be relied upon.","major_comments":[{"comment":"The abstract's claim that \"90% of human-like AI agents exclude autistic perspectives\" is not supported by the evidence in Chapter 2. What Chapter 2 actually reports is that, in a hand-assembled corpus of 142 HRI papers from selected venues between 2016 and 2022, \"nearly 90% of the papers did not include the input of autistic people in the design process.\" The leap from a venue-specific, keyword-searched HRI corpus to all human-like AI agents (chatbots, voice assistants, embodied agents beyond robots) has no sampling frame. The recommendation should be to soften the abstract to state exactly what was measured, e.g., \"in our corpus of HRI studies,\" or to provide additional evidence covering other classes of human-like agents.","section":"Abstract; §2.2.2; §2.3.3; Figure 2.11"},{"comment":"The headline 90% exclusion statistic inherits two unaddressed methodological fragilities. First, the corpus is built from venue lists plus the first 50 Google Scholar results per venue, which is a relevance-ranked convenience sample; no recall/precision audit, exclusion log, or inter-coder reliability statistic is reported for the manual coding of the \"design input\" variable. Second, the codebook for what counts as \"input in the design process\" is not provided, and no denominator rule is given for survey papers, editorials, or papers where no design process is described. If even 10–15 of the 142 papers were recoded, the headline would fall below 80%. The qualitative themes may be robust, but the quantitative headline is not; the manuscript should either report reliability and a clear coding protocol or demote this statistic from headline to context.","section":"§2.2.2; §2.2.4; §2.3.3"},{"comment":"The claim that LLMs \"frequently misclassify autistic community speech\" and \"fail to identify ableist speech due to reliance on simplistic keyword-based methods\" needs a stronger quantitative basis. As presented, Table 5.3 reports F1 scores across prompts and in-context learning examples, and Figure 5.4 reports mean Cohen's Kappa values, but there is no baseline comparison (random classifier, keyword-matching classifier, or majority-class classifier), no confidence intervals or variance estimates across prompt repetitions, and no statistical test comparing human-LLM agreement against human-human agreement. Without these, the conclusion that LLM performance is deficient rather than merely moderate is not fully established. The AUTALIC resource remains useful, but the evaluation section should be explicit about the threshold for \"frequent misclassification.\"","section":"§5.4.2; Table 5.3; Figure 5.4"},{"comment":"The claim that a binary annotation scheme \"sufficiently captures the nuances\" of labeling anti-autistic language is based on a study with six annotators and a co-design session. The manuscript acknowledges the small sample, but the generalizability of this conclusion to other annotator pools, other platforms, and other social media contexts is limited. Since this claim is used to justify the binary labeling in AUTALIC, a brief statement tying the binary scheme's sufficiency to the annotators' own reported preferences, rather than to a general claim about all annotation tasks, would make the inference more precise.","section":"§4.3–§4.5"}],"minor_comments":[{"comment":"There are multiple typographical errors, including \"Pyschology\" (Figure 2.5 caption), \"contexualize\" (§2.3.3), \"immitation\" (§3.2.1), \"neuornormative\" and \"neuronomativity\" (§3.1 and §3.2), and \"AUTALICdataset\" (Table 5.1). These should be corrected in a careful copyedit.","section":"Throughout"},{"comment":"Two tables are both labeled \"Table 1\" (the key-terms glossary appears twice with identical numbering). The duplicated numbering should be fixed, and the glossary table should appear once or be renumbered consistently.","section":"Table 1 and Table 1 (page 4 vs. page 9)"},{"comment":"The sentence \"There is a notable correlation between thehumanness\" of robots and a power imbalance\" contains a formatting error (missing space and an unmatched quote mark). Also, \"correlation\" here appears to mean an observed association in categorical codes, not a statistical correlation; the wording should be clarified.","section":"Figure 2.2"},{"comment":"The claim that the majority of referenced works were published before Critical Autism Studies introduced inclusive theories is based on an annualized comparison, but no base-rate comparison is given for the overall publication volume in those fields. Without such a baseline, the figure may overstate the shift. A brief caveat would help.","section":"§2.2.3; Figure 2.6"}],"recommendation":"major_revision","confidential_remarks":"The dissertation has several genuine strengths: it is candid about its own limitations, it provides a practical dataset resource with responsible-use guidelines, and it centers autistic and neurodivergent perspectives in both the research team and the annotation process. The main issue is the gap between the abstract's confident 90% generalization and the fragile corpus behind it. I would advise the editor to require the author to either restrict the headline to the HRI corpus with appropriate hedging or add the missing reliability and sampling evidence. The LLM evaluation chapter would also benefit from explicit baselines and variance reporting. These are fixable within the manuscript's scope, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing to know up front: this dissertation has a solid core—the AUTALIC benchmark, the creator interview study, and the HRI critical review—but the headline statistic that '90% of human-like AI agents exclude autistic perspectives' is not supported by the methods as reported. That claim comes from a hand-coded corpus of 142 HRI papers, and the coding process in §2.2.4 reports no inter-coder reliability. The corpus itself is a convenience sample from Google Scholar's relevance-ranked results and selected venues, not a systematic review with a documented recall/precision audit. Shift 10-15 codes and the number drops below 80%. The abstract then inflates 'HRI studies' to 'human-like AI agents,' which is a jump with no sampling frame.\n\nThat said, there is real value here. The author is honest: positionality statements, explicit limitations on small samples and snowball recruitment, and a consistent neurodiversity framework. The qualitative themes in Chapter 2—pathologizing, essentialism, power imbalance—are likely robust even if the percentages are fragile. The creator interviews (16 professionals) give a rare look at how AI practitioners think about accessibility and ethics, and the finding that many defer responsibility is important. The annotation co-design study with binary labels is a useful empirical result for hate-speech annotation. And AUTALIC is a genuine new resource: Reddit sentences labeled in context for anti-autistic ableism, with human ground truth and LLM evaluation.\n\nThe soft spots beyond the 90% claim: AUTALIC is not shipped—no data, no code, no annotation platform link—so the benchmark can't be used or verified yet. The LLM evaluation results are described but the exact prompts and model versions are only partially specified. And the dissertation's own limitations sections are sometimes two sentences after a strong claim; they need to be reflected in the abstract.\n\nBottom line: this is a serious piece of work that deserves referee time, but it needs major revision before publication: tighten the quantitative claims, report inter-coder reliability (or reframe as qualitative), ship the dataset with responsible-use guidelines, and make the abstract match the evidence. If those are done, it becomes a citeable contribution.","headline":"A worthwhile dissertation with a real contribution in AUTALIC, but the abstract's 90% claim outruns the evidence in Chapter 2.","tokens_in":49328,"tokens_out":2452,"would_cite":false,"duration_ms":28541,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The dissertation argues that AI systems built to mimic human communication systematically reproduce anti-autistic ableism, and that the path forward is to de-center neuronormative 'humanness' as the benchmark for machine intelligence.","keywords":["neuro-inclusive AI","anti-autistic ableism","human-robot interaction","LLM bias","hate speech detection","participatory design","neurodiversity","AUTALIC"],"falsifier":"An independent replication of the Chapter 2 corpus selection with a preregistered search protocol, full-text screening, multiple blinded coders, and inter-coder reliability statistics would settle whether the 90% exclusion figure holds; a substantially lower rate, or the discovery that many excluded studies reported participatory design outside the searched text, would undercut the quantitative centerpiece. For the AUTALIC claim, re-annotating a random sample with a larger autistic community panel and comparing LLM agreement would test whether the misclassification results generalize.","tokens_in":48331,"feed_emoji":"🤖","tokens_out":7980,"duration_ms":69745,"temperature":0.7,"pith_summary":"The dissertation argues that AI systems built to mimic human communication systematically reproduce anti-autistic ableism: the medical model treats autism as a deficit of neurotypical skills, autistic people are excluded from designing the technologies aimed at them, and large language models both censor autistic speech and fail to catch ableist speech. The author offers a definition of Neuro-Inclusive AI as an explicit alternative to benchmarks that equate intelligence with mimicking humanness. Across five studies—an education intervention, a 142-paper critical review, interviews with 16 AI creators, an annotation study, and a new benchmark—the dissertation makes the case that the exclusion is structural, not incidental, and provides tools to correct it. A reader should care because if the claim is right, the default evaluation criteria for human-like AI need to change, and content-moderation systems built on current models would systematically harm autistic users.","feed_headline":"AI built to mimic humans reproduces anti-autistic bias","feed_subtitle":"A new benchmark, AUTALIC, lets developers detect and fine-tune away ableist language in context.","key_machinery":"The machinery is threefold: (1) the concept of Neuro-Inclusive AI, defined as design that de-centers neuronormative benchmarks; (2) the AUTALIC benchmark, a dataset of Reddit posts annotated for anti-autistic ableist language in context, with a binary labeling scheme developed through co-design with annotators; and (3) the critical-analytic categories of pathologizing, essentialism, and power imbalance used to code the human-robot interaction corpus. The Turing Test functions as the negative benchmark: the dissertation treats 'mimicking humanness' as the assumption to be dismantled.","core_discovery":"On its own terms, the dissertation establishes the following: current AI systems designed to imitate human communication are built on a neuronormative standard of humanness, and this reproduces anti-autistic ableism. In support, it reports a critical review of 142 human-robot interaction studies, finding that nearly 90% did not include autistic input in design and that a large majority applied the medical model; interviews with 16 creators of human-like agents, most of whom did not treat accessibility or neurodiversity as their responsibility; and an annotation study in which a binary ableist/not-ableist label captured annotator nuances. It then introduces AUTALIC, a dataset of Reddit sentences labeled in context for anti-autistic ableist language, and shows that four open-source LLMs frequently misclassify autistic community speech and miss ableist speech. The positive claim is the definition of Neuro-Inclusive AI: systems that de-center neuronormative benchmarks and do not take 'mimicking humanness' as the goal.","pith_inferences":["A broader implication not stated in the dissertation: the same critique should apply to other neurodivergent and disabled populations, so any benchmark that defines 'human-like' behavior needs a representativeness check before deployment.","The AUTALIC findings imply keyword-based toxicity filters will be brittle in general; a concrete extension is stress-testing moderation pipelines on context-heavy examples from other marginalized groups.","The binary-label result suggests that finer-grained annotation schemes are not automatically better; a testable hypothesis is that co-designed binary schemes improve agreement in other hate-speech labeling tasks.","The dissertation's use of the contact hypothesis points to an untested opportunity: if human-like agents displayed neurodiverse communication styles, they might reduce real-world dehumanization of autistic people."],"forward_implications":["If the dissertation is right, any AI system that treats 'human-like communication' as its goal should specify whose communication style is being modeled; otherwise it will inherit neuronormative assumptions.","Human-robot interaction research on autism should be reoriented from diagnosis and treatment toward support and participation, with autistic people as co-designers rather than subjects.","AUTALIC gives a concrete tool for fine-tuning or evaluating LLMs on anti-autistic ableist language; current popular models do poorly on it, so content-moderation deployments should not rely on them without such tuning.","A simplified binary annotation scheme can replace more complex labeling schemes for this task, making it cheaper and more consistent to build datasets for anti-autistic language.","Because many makers of human-like AI do not see ethics and accessibility as their responsibility, education, funding priorities, and organizational standards need to embed neuro-inclusion rather than treating it as an add-on."],"supporting_citations":[{"why":"Defines machine intelligence as the ability to mimic human communication, the benchmark the dissertation argues is neuronormative and must be de-centered.","marker":"(Turing, 2009)"},{"why":"Supplies the neurodiversity framing that contrasts the medical/deficit model with a difference-based understanding of autism, used throughout the corpus coding.","marker":"(Kapp et al., 2013)"},{"why":"Exemplifies the foundational dehumanizing theory-of-mind work whose deficit view the dissertation traces into robotics and AI design.","marker":"(Baron-Cohen, 1997)"},{"why":"Proposes the double empathy problem, the theoretical basis for shifting communication responsibility away from autistic people alone.","marker":"(Milton, 2012)"},{"why":"Argues that robots placed in mentor roles move autistic users toward 'humanness', a key power-imbalance mechanism quantified in Chapter 2.","marker":"(Williams, 2021b)"},{"why":"Provides evidence that autistic children prefer non-anthropomorphic robots, used to show that HRI design contradicts known user preferences.","marker":"(Ricks and Colton, 2010)"},{"why":"Supplies the critical literature-review methodology and the emphasis on centering disabled end-users that Chapter 2 applies to the 142-paper corpus.","marker":"(Spiel et al., 2022)"},{"why":"Documents identity-first language preference among autistic adults, informing the dissertation's language choices and annotation design.","marker":"(Taboas et al., 2023)"}],"fun_headline_variants":["New AUTALIC benchmark detects ableist AI language","90% of human-like AI exclude autistic perspectives","Stop mimicking humans: define AI with autistic input","AUTALIC: a tool to fine-tune away ableist language"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The dissertation's headline statistic—that nearly 90% of human-like AI agents exclude autistic perspectives—rests on a manually assembled corpus of 142 human-robot interaction papers, selected from venue lists and the first 50 Google Scholar results per venue, with thematic codes generated by the authors and no reported inter-coder reliability.","fun_headline_variants_meta":{"raw":{"variants":["New AUTALIC benchmark detects ableist AI language","90% of human-like AI exclude autistic perspectives","Stop mimicking humans: define AI with autistic input","AUTALIC: a tool to fine-tune away ableist language"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000962,"raw_usage":{"total_tokens":4097,"prompt_tokens":946,"completion_tokens":3151,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":562,"completion_tokens_details":{"reasoning_tokens":3085}},"tokens_in":562,"tokens_out":3151,"duration_ms":23604,"temperature":1.0,"reasoning_tokens":3085,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:12:40.028139+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An independent replication of the Chapter 2 corpus selection with a preregistered search protocol, full-text screening, multiple blinded coders, and inter-coder reliability statistics would settle whether the 90% exclusion figure holds; a substantially lower rate, or the discovery that many excluded studies reported participatory design outside the searched text, would undercut the quantitative centerpiece. For the AUTALIC claim, re-annotating a random sample with a larger autistic community panel and comparing LLM agreement would test whether the misclassification results generalize.","supporting_citations":[],"review_version":1}