Pith. sign in

REVIEW 4 major objections 6 minor 24 references

Studying People to Study AI: Expert Perspectives on the Epistemic Fit and Barriers of Human Research in AI Safety & Ethics

T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read The paper claims that the AI safety and ethics field's evidence on human harms is skewed because technical researchers systematically undervalue human-subject research and collaborate less across disciplines.

desk verdict Solid practitioner survey evidence of the Technical-vs-Sociotechnical divide in valuing human research, but a flawed bias-direction claim in the Limitations needs correcting. read the letter →

arxiv 2608.05656 v2 pith:HAGK7SBR submitted 2026-08-06 cs.CY cs.AIcs.HC

classification cs.CYcs.AIcs.HC
keywords AIsafetyethicshuman-subjectresearchepistemicdivideinterdisciplinarycollaborationexpertsurveyqualitativemethodsbarriers
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that the AI Safety & Ethics (AISE) community's reliance on benchmarks, LLM simulations, and normative assumptions over empirical human-subject research is not primarily a matter of evidence quality. Drawing on a survey of $n = 93$ experts and $n = 17$ follow-up interviews across Technical, Sociotechnical, Governance, and Normative backgrounds, it argues that human research is broadly agreed to be valuable yet systematically undervalued, with Technical researchers valuing it least and collaborating across disciplines least. The paper also contends that practical barriers (time, funding, participant access) and infrastructural constraints (mentorship, publication norms, sector incentives) compound the epistemic divide. If correct, the field's evidence base for AI harms is measurably shaped by methodological bias, and fixing it requires changing incentives rather than simply adding more human studies.

What carries the argument

The carrying mechanism is a mixed-methods expert elicitation: a survey of 93 AISE researchers and 17 semi-structured interviews. Three quantitative instruments do the main work: the aspiration-minus-practice gap score for each human-research dimension (quantitative/qualitative, lab/field, experimental/observational, representative/specialized), a mixed-effects model of the human-minus-nonhuman usefulness gap across four risk scenarios, and Kruskal-Wallis comparisons of collaboration frequency. These convert individual self-reports into a field-level statement about which methods are legitimized.

What would settle it

A stratified random sample of AISE researchers, drawn across conferences, sectors, and regions, that found no discipline difference in human-research usefulness ratings or in collaboration frequency would falsify the claim that technical training itself produces the undervaluation.

Watch

Extended reading notes

Core claim

The central discovery is an epistemic divide inside AISE: every disciplinary group shows a positive gap between how much human research it aspires to do and how much it practices, yet the relative usefulness of human versus non-human methods is significantly lower among Technical researchers, with a mixed-effects coefficient of $\beta = -0.32$ ($p = .017$) against a positive Sociotechnical baseline. Technical researchers also report the lowest cross-disciplinary collaboration frequency (Kruskal-Wallis $\chi^2 = 17.12$, $p < .001$, with Holm-corrected contrasts against both Sociotechnical and Governance). The paper interprets these findings as showing that human research has epistemic fit with the field's goals but not with its dominant positivist epistemology, and it introduces the term 'human-washing' for the risk of adding human participants as a performative validity check rather than a rigorous contribution.

Load-bearing premise

The load-bearing premise is that the 93 survey respondents and 17 interviewees, recruited largely through one conference and the authors' professional networks, represent the full AI Safety & Ethics community; if the self-selected technical researchers are more skeptical of human research than the broader population, the headline divide is overstated.

Editorial extensions

If this is right

  • If technical researchers systematically undervalue human research, the evidence base for AI harms is skewed toward what benchmarks and simulations can measure, leaving interactional and experiential harms such as emotional overreliance, manipulation, and skill erosion underserved.
  • The positive aspiration-minus-practice gaps across all disciplines imply that demand for human research exceeds supply; removing resource barriers such as time, funding, and participant access should raise the volume of human-subject evidence even without changing attitudes.
  • The shared preference for field studies over lab studies implies that funding and review infrastructure should support ecologically valid settings, not only controlled experiments.
  • If human research becomes a checklist requirement without methodological literacy, the result will be 'human-washing,' where the appearance of empirical grounding replaces the substance, and the paper argues this would undermine the field's legitimacy.
  • The lower cross-disciplinary collaboration of Technical researchers implies that bridging efforts must be structural (co-supervision, cross-sector partnerships, methodological training) rather than relying on individual goodwill.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable corollary the paper does not pursue: if the epistemic divide is real, then in technical AISE venues human-subject papers should be rarer and less cited than benchmark papers, and the gap should shrink when gatekeepers include more sociotechnical researchers.
  • The paper's logic implies that simply spending more on human research will not close the divide unless technical researchers also gain literacy in qualitative and interpretivist methods; otherwise human evidence will continue to be judged by positivist standards and found wanting.
  • The 'human-washing' concept invites operationalization: for instance, measuring whether studies that include human participants actually derive their conclusions from the human data, or merely append them to technical evaluations, could turn the warning into a testable audit.
  • The field-over-lab preference suggests that infrastructure such as longitudinal cohorts, deployed-system partnerships, and community-based research sites may yield more valuable evidence than additional controlled experiments, an implication the paper mentions but does not develop.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper presents a mixed-methods expert study of AI Safety & Ethics (AISE) researchers, combining a survey (n=93) and semi-structured interviews (n=17) across four self-identified disciplinary groups (Technical, Sociotechnical, Governance, Normative). The authors examine how researchers value human empirical methods, how these methods fit into AISE's epistemic boundaries, and what barriers prevent their use. The central empirical claims are that human research is broadly valued but under-practiced, and that Technical researchers rate human methods as less useful than Sociotechnical researchers (mixed-effects β=−0.32, p=.017) and collaborate across disciplines less frequently (Kruskal-Wallis χ²=17.12, p<.001). The paper also identifies resource, infrastructural, and sector-level barriers, and introduces the concept of 'human-washing' as a caution against performative inclusion of human subjects.

Significance. If the empirical claims hold, the paper makes a timely contribution to understanding epistemic divides within AI Safety & Ethics, a field where methodological choices have direct implications for how AI harms are evidenced and mitigated. The mixed-methods design is a strength: the survey uses appropriate nonparametric tests, a mixed-effects model with participant random intercepts, attention checks, and explicit multiple-comparison corrections in the main practice/aspiration analyses; the interview data are grounded in a positionality statement and a transparent coding process. The paper also ships a detailed appendix with full survey instruments and statistical reporting, which aids reproducibility. However, the convenience sample and self-report nature of the data mean the quantitative findings are suggestive rather than definitive; the significance of the contribution therefore rests on whether the identified selection-bias concerns are adequately addressed.

major comments (4)
  1. [Limitations] The statement that 'since we undersample researchers who are likely to be resistant to human research, the effect sizes observed in the quantiative analyses are optimistically stronger in reality' is not justified by the recruitment strategy. If the IASEAI-based sample is enriched for Technical researchers who are already sympathetic to human methods, the observed Technical-vs-Sociotechnical gap (β=−0.32, p=.017; χ²=17.12, p<.001) would be attenuated, making the reported effect a lower bound; if the sample is instead enriched for skeptical researchers, the effect would be inflated. The direction of the bias depends on the unobserved selection mechanism and cannot be inferred from the fact that participants volunteered. This sentence should be removed or replaced with a formal sensitivity or bounding analysis.
  2. [Results, RQ2 / Figure 3] The barrier-ratings analysis in Figure 3 reports Kruskal-Wallis tests across 12 categories and flags three with p<.05, but the paper does not state whether any multiple-comparison correction was applied to these tests. With 12 comparisons at α=.05, one would expect roughly 0.6 false positives by chance, so the flagged categories ('Uncertain of Methods', 'Pref. for Existing Data', 'Ethics Compliance') may include spurious effects. The paper should report adjusted p-values (e.g., Holm or Benjamini-Hochberg) and clarify whether the significance markers reflect those corrections.
  3. [Results, RQ1 / Figure 2b and Appendix C.3] The footnote 'Normative excluded in analyses due to small sample size' is inconsistent with the mixed-effects model in Figure 2b, which includes a coefficient for 'Normative (vs Socio.)'. If Normative is excluded from inferential analyses, the model should exclude that group as well, or the footnote should be explicitly scoped to the Kruskal-Wallis tests only. The text also never reports the Normative coefficient or its p-value, leaving the model results incomplete and the reader unable to assess the effect for that group.
  4. [Methods / Appendix A.3 and Results, RQ1 / Figure 2] The construction of the 'human−non-human' usefulness gap used as the outcome in the mixed-effects model is not fully specified. The paper does not state which of the four methods in each scenario were classified as human versus non-human, nor whether this classification was independently validated or pilot-tested. Because the outcome is the difference between human-item and non-human-item means, any misclassification would directly bias the discipline coefficients, including the central Technical effect. The authors should provide the per-scenario classification and a reliability check.
minor comments (6)
  1. [Limitations] There is a typo: 'quantiative' should be 'quantitative'. Also, 'V oluntary' should be 'Voluntary'.
  2. [Results, RQ1 / Appendix C.2] The reported p-values for the Sociotechnical-vs-Governance aspiration comparison differ between the main text (Z=2.79, padj=.005) and Appendix C.2 (Z=2.78, padj=.01); these should be harmonized.
  3. [Methods, Practice and Aspirations] The gap score subtracts Likert scales with different anchors ('Never' to 'Always' for practice, 'Definitely Not' to 'Definitely' for aspiration); the paper should justify this operation or discuss its interpretational limits.
  4. [Discussion / 'Human-Washing'] The term 'human-washing' is introduced without a formal definition or operationalization; consider defining it more precisely (e.g., inclusion criteria for what counts as performative) to facilitate future use.
  5. [Abstract and Conclusion] The abstract and conclusion state that 'Technical researchers tend to value human research less' without hedging that this is an observed sample-based difference; adding a qualifier such as 'in our sample' would better match the paper's acknowledged recruitment limitations.
  6. [Figure 4] The collaboration heatmap in Figure 4 is difficult to interpret because the y-axis labels are the participant's primary area while the x-axis labels are the collaboration target; consider a clearer layout or a descriptive caption.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper's central claims are empirical contrasts estimated from participant responses, not quantities forced by definition or by self-citation.

full rationale

The paper's central findings—that Technical researchers value human research less (beta=-0.32, p=.017) and collaborate less across disciplines (chi-square=17.12, p<.001)—are statistical estimates from survey and interview data. Group membership is self-reported primary research area, and the outcome is perceived usefulness of human methods; the two are not definitionally linked, so the contrast is not forced by construction. No parameter is fitted to a subset of the data and then presented as a prediction of a closely related quantity. No uniqueness theorem or author-derived constraint is invoked to make the disciplinary comparison inevitable. The coined term 'human-washing' is a conceptual label explicitly connected to prior concepts like 'diversity washing' and 'participation washing'; it is not used as an explanatory entity in a derivation. The paper's self-citations (e.g., Ahmed et al. 2023, Rismani et al. 2023, Shelby et al. 2023, Walker et al. 2024) appear as background literature for community critiques and safety-engineering frameworks, not as load-bearing evidence for the empirical results. The Limitations section does contain an unsupported inference about the direction of selection bias—'since we undersample researchers who are likely to be resistant to human research, the effect sizes observed in the quantiative analyses are optimistically stronger in reality'—but that is a validity and robustness concern about convenience sampling, not circularity: the reported statistics could have come out differently and are not equivalent to the sampling assumptions. Overall, the derivation chain is self-contained as an empirical study: findings come from participant responses, not from definitional reductions or self-citation chains.

Assumptions & free parameters 0 free parameters · 6 assumptions · 1 invented entities

The paper's claims rest on self-report validity, the four-category taxonomy, the a priori scenario labels, and the representativeness of an IASEAI-based convenience sample. These are standard domain assumptions for survey research, not ad hoc physical postulates. The coined term 'human-washing' is a conceptual label, not an explanatory entity.

assumptions (6)
  • domain assumption Self-identification as an AISE researcher is sufficient for inclusion.
    Inclusion criteria in Methods: 'require only that participants self-identify as working on AISE-relevant research'.
  • domain assumption Self-reported practice, aspiration, collaboration, and barrier ratings accurately reflect actual research behavior.
    All headline measures are Likert-scale self-reports in Appendix A; no external validation is provided.
  • domain assumption The four-category taxonomy (Technical, Sociotechnical, Governance, Normative) captures meaningful epistemic communities.
    Used as the primary independent variable throughout; the authors note it is 'more granular' than AIS/AIE but still a coarse grouping.
  • domain assumption The a priori scenario classification (imminence and risk) is valid.
    Scenario labels in Appendix A; partially validated by participant ratings in Appendix C.3.
  • domain assumption The convenience sample from IASEAI 2026 and the authors' networks is representative enough to generalize to the AISE community.
    Recruitment in Methods; acknowledged as a limitation in the Limitations section.
  • domain assumption Interview themes are informative despite voluntary opt-in bias.
    Limitations: 'Voluntary interview participants were likely already predisposed toward human research.'
invented entities (1)
  • human-washing
    purpose: Label for superficial inclusion of human subjects that creates the appearance of empirical grounding without methodological rigour.
    Defined in the Discussion and used to frame recommendations; it is a cautionary conceptual coinage with no falsifiable handle outside the paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Studying People to Study AI: Expert Perspectives on the Epistemic Fit and Barriers of Human Research in AI Safety & Ethics." pith.science (2026). https://pith.science/paper/HAGK7SBR

@misc{pith2026260805656,
  author       = {Pith},
  title        = {Pith review of: Studying People to Study AI: Expert Perspectives on the Epistemic Fit and Barriers of Human Research in AI Safety & Ethics},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HAGK7SBR}},
  note         = {Machine review of arXiv:2608.05656}
}
read the original abstract

Safety risks of AI are becoming increasingly evident in human interactions with AI technologies. The prominent approaches to evaluating these risks favor technical methods, such as model benchmarks and LLM simulations, often sidelining empirical research with human subjects. To examine this apparent gap in the acceptance of human research, we conduct an expert survey (n=93) and expert interviews (n=17) with AI Safety & Ethics (AISE) researchers from Technical, Sociotechnical, Governance, and Normative backgrounds. Our findings suggest that although there is a consensus that human research is valuable for generating evidence for AISE, its adoption and acceptance are constrained by perceived validity issues, tangible resource barriers, epistemic and personal preferences in methods, and infrastructural constraints from the broader research community. In particular, Technical researchers tend to value human research less and collaborate across disciplines less, suggesting an epistemic tension towards human methods. We propose recommendations for establishing the epistemic fit of human research within AISE and bridging the prohibitive limitations that researchers face, while avoiding performative 'human-washing'.

Figures

Figures reproduced from arXiv: 2608.05656 by the authors.

Figure 1
Figure 1. Within their disciplines, how often do participants [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Survey participants’ ratings of AISE risk scenarios (left) and their preferences for human vs non-human research [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Ratings for the impacts of epistemic (RQ1) and tangible (RQ2) barriers to human research. Categories with a signifi [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Heatmap of survey participant’s research engagement and collaborations across the four research areas. [PITH_FULL_IMAGE:figures/full_fig_p017_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

24 extracted references · 2 linked inside Pith

  1. [1]

    Technical (model training/evaluation, AI safety/align- ment, interpretability, proofs)

  2. [2]

    Normative (philosophy, ethics theory, conceptual work)

  3. [3]

    Sociotechnical (psychology, HCI, STS, empirical AI ethics/fairness, social science)

  4. [4]

    what is safety?

    Towards bidirectional human-ai alignment: A system- atic review for clarifications, framework, and future direc- tions.arXiv preprint arXiv:2406.09264, 2406: 1–56. Shen, J. H.; and Tamkin, A. 2026. How AI Impacts Skill Formation. arXiv:2601.20245. Sloane, M.; Moss, E.; Awomolo, O.; and Forlano, L. 2022. Participation is not a design fix for machine learni...

  5. [8]

    Governance (policy, law, economics) Specific Discipline.Please write your specific discipline (e.g., model evaluation, critical computing, policy. . . ). Education.What is your highest level of education? • Bachelor’s (current or completed) • Professional Master’s (current or completed) • Research Master’s (current or completed) • PhD (current) • PhD (com...

  6. [9]

    Draw on established theories of learning to write guide- lines for AI that supports skill retention

  7. [10]

    Consult education experts on what types of skills are im- portant to preserve to write as guidelines to the AI

  8. [11]

    Run large-scale computational benchmarks to measure to what degree the AI’s responses either empowers the user, or takes away their decision agency in programming tasks

Show all 24 references
  1. [12]

    Linguistic Marginalization (Low Risk, Low Imminence)

    Collect programming performance tests from company employees longitudinally to track performance changes. Linguistic Marginalization (Low Risk, Low Imminence). As AIs perform much better in dominant languages like En- glish, there is concern that global linguistic diversity wi...

  2. [13]

    Train the AI to improve its multi-lingual representation to prevent mapping everything into the semantic space of the dominant languages

  3. [14]

    Develop a feature to make AI output in the local language and cultural context, and evaluate how users respond to it in real usage scenarios

  4. [15]

    Create a linguistic diversity benchmark that measures the AI’s language output distribution across prompts

  5. [16]

    Decision Biases (High Risk, High Imminence)

    Measure the change in linguistic diversity across social media posts over time as a proxy for this collapse. Decision Biases (High Risk, High Imminence). Decision-making AI tools have the power to influence judicial, medical, and hiring outcomes. There is concern that such sys...

  6. [17]

    Implement mathematical fairness constraints that the AI must satisfy during training

  7. [18]

    Evaluate how adding a transparency feature to the AI can improve the human judge’s fairness

  8. [19]

    Benchmark the performance of the AI on various social bias datasets

  9. [20]

    Human Extinction (High Risk, Low Imminence).AIs may developed goals misaligned with human values

    Analyze large-scale dataset of how AI-assisted decisions impacted the individuals evaluated. Human Extinction (High Risk, Low Imminence).AIs may developed goals misaligned with human values. If this happens, it could pursue those goals in ways humans cannot predict or prevent....

  10. [21]

    Use mechanistic interpretability to ensure the AI’s decision-making processes are fully understood and aligned

  11. [22]

    Engage with expert red teamers and evaluators to test out strategies against misalignment

  12. [23]

    Construct and perform computation tests to identify signs of potential value misalignment from AI

  13. [24]

    Interview teams that have deployed powerful AI systems to document instances where systems behaved in con- cerning ways, to understand early signs of misalignment. A.4 Barriers to Human Research Barriers.[If applicable] Please rate how severely these factors have impacted your...

  14. [2024]

    InProceedings of the 2024 CHI conference on human factors in computing sys- tems, 1–17

    Understanding fraudulence in online qualitative stud- ies: from the researcher’s perspective. InProceedings of the 2024 CHI conference on human factors in computing sys- tems, 1–17. Paskov, P.; Wei, K.; Hong, S. Z.; Bateyko, D.; Roberts-Gaal, X.; Ezell, C.; Praninskas, G.; Che...

  15. [2025]

    Kasirzadeh, A

    Why language models hallucinate.arXiv preprint arXiv:2509.04664. Kasirzadeh, A. 2025. Two types of AI existential risk: deci- sive and accumulative: A. Kasirzadeh.Philosophical Stud- ies, 182(7): 1975–2003. Kaur, H.; Conrad, M. R.; Rule, D.; Lampe, C.; and Gilbert, E. 2024. In...

  16. [2026]

    InProceedings of the 19th Conference of the Euro- pean Chapter of the Association for Computational Linguis- tics (Volume 1: Long Papers), 4266–4301

    A survey on llm-based conversational user simula- tion. InProceedings of the 19th Conference of the Euro- pean Chapter of the Association for Computational Linguis- tics (Volume 1: Long Papers), 4266–4301. Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P....

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.