{"id":"5cf21219-c07e-42bd-9170-6f72993805bd","arxiv_id":"2608.05656","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"AI Safety and Ethics researchers broadly value human-subject research but practice it less than they aspire to, with technical researchers undervaluing it and collaborating least across disciplines.","lead":"This paper surveys 93 and interviews 17 AI safety and ethics researchers, finding that experts agree human-subject research is valuable but underused, and that technical researchers value it least and collaborate least across disciplines. It maps the barriers (funding, time, participants, mentorship) and warns against performative 'human-washing' rather than genuine integration.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Convenience self-selected sample leaves the Technical-vs-Sociotechnical gap unidentified under plausible selection on affinity for human research; the Limitations' bias-direction claim is unsubstantiated and possibly backwards.","rationale":"The reader's verdict is CONDITIONAL, with sample representativeness as the weakest assumption. My stress-test concurs: representativeness is the load-bearing assumption for the central claim about Technical researchers. The paper's own Limitations flags non-systematic recruitment and volunteer predisposition, but the attempted corrective statement about 'optimistically stronger in reality' is not derived and is ambiguous about whether the reported gap is under- or over-estimated. The proposed sensitivity analysis would settle the direction and magnitude of selection bias. I am not raising an ad hominem or a consensus disagreement: the epistemic-tension hypothesis is plausible, and the mixed-methods design is a reasonable way to probe it. However, the absence of data and code means the reported statistics cannot be independently checked, and the collaboration claim's displayed heatmap (Figure 4b) is not obviously consistent with the reported Technical-vs-Governance contrast; this is a second reason to require the raw data rather than to accept the numbers on faith. Conditional on the sensitivity analysis and data release, the verdict stays CONDITIONAL; if the analysis shows the gap is robust and the data match the text, the paper could be accepted with minor revisions. If the gap disappears under plausible selection, the central claim should be substantially weakened. Because the reader already made the verdict CONDITIONAL, I recommend UNCHANGED.","tokens_in":24352,"tokens_out":15899,"duration_ms":150202,"concrete_test":"Run a formal selection-on-the-outcome sensitivity analysis. Assume response propensity is monotonically increasing in each respondent's latent affinity for human research within each discipline; for selection odds ratios of, say, 1.5, 3, and 5, re-weight or impute the missing lower tail of the affinity distributions (or approximate using the observed early- versus late-responder contrast as an empirical proxy) and re-estimate the mixed-effects Technical coefficient and the collaboration Kruskal-Wallis test. If the Technical coefficient remains negative and significant under all plausible selection strengths, the divide is robust; if it attenuates to non-significance or changes sign, the central claim is not identified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that Technical AISE researchers value human research less (beta=-0.32, p=.017) and collaborate less (chi-square=17.12, p<.001)—is only as strong as the sample from which it is estimated. Recruitment was through IASEAI 2026, the authors' professional networks, and opt-in social media and Scholar outreach (Methods, Recruitment; Limitations). This is a convenience sample with self-selection on a trait (affinity for human research) that is also the outcome. The paper's Limitations concedes that recruitment was 'targetted to IASEAI attendees and the researchers' existing professional networks rather than a systematic sampling of the entire AISE community' and that interview volunteers 'were likely already predisposed toward human research.' The very next sentence, however, asserts without derivation that 'since we undersample researchers who are likely to be resistant to human research, the effect sizes observed in the quantiative analyses are optimistically stronger in reality.' This does not follow: under monotone selection on affinity, the bias in the Technical-Sociotechnical contrast depends on the variance and tail shape of each group's affinity distribution and can attenuate, inflate, or even reverse the gap. As written, the paper does not establish whether the reported divide is an upper bound, a lower bound, or an artifact. If sampled Technical researchers are the human-research-friendly tail of their discipline, the true undervaluation is larger; if selection is stronger among Sociotechnical respondents, the true gap is smaller. The interview-based mechanism (epistemic bias) is also drawn from volunteers already engaged with human research, so it cannot independently certify the quantitative result.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a mixed-methods expert study of AI Safety & Ethics (AISE) researchers, combining a survey (n=93) and semi-structured interviews (n=17) across four self-identified disciplinary groups (Technical, Sociotechnical, Governance, Normative). The authors examine how researchers value human empirical methods, how these methods fit into AISE's epistemic boundaries, and what barriers prevent their use. The central empirical claims are that human research is broadly valued but under-practiced, and that Technical researchers rate human methods as less useful than Sociotechnical researchers (mixed-effects β=−0.32, p=.017) and collaborate across disciplines less frequently (Kruskal-Wallis χ²=17.12, p<.001). The paper also identifies resource, infrastructural, and sector-level barriers, and introduces the concept of 'human-washing' as a caution against performative inclusion of human subjects.","tokens_in":24644,"tokens_out":9393,"duration_ms":82101,"significance":"If the empirical claims hold, the paper makes a timely contribution to understanding epistemic divides within AI Safety & Ethics, a field where methodological choices have direct implications for how AI harms are evidenced and mitigated. The mixed-methods design is a strength: the survey uses appropriate nonparametric tests, a mixed-effects model with participant random intercepts, attention checks, and explicit multiple-comparison corrections in the main practice/aspiration analyses; the interview data are grounded in a positionality statement and a transparent coding process. The paper also ships a detailed appendix with full survey instruments and statistical reporting, which aids reproducibility. However, the convenience sample and self-report nature of the data mean the quantitative findings are suggestive rather than definitive; the significance of the contribution therefore rests on whether the identified selection-bias concerns are adequately addressed.","major_comments":[{"comment":"The statement that 'since we undersample researchers who are likely to be resistant to human research, the effect sizes observed in the quantiative analyses are optimistically stronger in reality' is not justified by the recruitment strategy. If the IASEAI-based sample is enriched for Technical researchers who are already sympathetic to human methods, the observed Technical-vs-Sociotechnical gap (β=−0.32, p=.017; χ²=17.12, p<.001) would be attenuated, making the reported effect a lower bound; if the sample is instead enriched for skeptical researchers, the effect would be inflated. The direction of the bias depends on the unobserved selection mechanism and cannot be inferred from the fact that participants volunteered. This sentence should be removed or replaced with a formal sensitivity or bounding analysis.","section":"Limitations"},{"comment":"The barrier-ratings analysis in Figure 3 reports Kruskal-Wallis tests across 12 categories and flags three with p<.05, but the paper does not state whether any multiple-comparison correction was applied to these tests. With 12 comparisons at α=.05, one would expect roughly 0.6 false positives by chance, so the flagged categories ('Uncertain of Methods', 'Pref. for Existing Data', 'Ethics Compliance') may include spurious effects. The paper should report adjusted p-values (e.g., Holm or Benjamini-Hochberg) and clarify whether the significance markers reflect those corrections.","section":"Results, RQ2 / Figure 3"},{"comment":"The footnote 'Normative excluded in analyses due to small sample size' is inconsistent with the mixed-effects model in Figure 2b, which includes a coefficient for 'Normative (vs Socio.)'. If Normative is excluded from inferential analyses, the model should exclude that group as well, or the footnote should be explicitly scoped to the Kruskal-Wallis tests only. The text also never reports the Normative coefficient or its p-value, leaving the model results incomplete and the reader unable to assess the effect for that group.","section":"Results, RQ1 / Figure 2b and Appendix C.3"},{"comment":"The construction of the 'human−non-human' usefulness gap used as the outcome in the mixed-effects model is not fully specified. The paper does not state which of the four methods in each scenario were classified as human versus non-human, nor whether this classification was independently validated or pilot-tested. Because the outcome is the difference between human-item and non-human-item means, any misclassification would directly bias the discipline coefficients, including the central Technical effect. The authors should provide the per-scenario classification and a reliability check.","section":"Methods / Appendix A.3 and Results, RQ1 / Figure 2"}],"minor_comments":[{"comment":"There is a typo: 'quantiative' should be 'quantitative'. Also, 'V oluntary' should be 'Voluntary'.","section":"Limitations"},{"comment":"The reported p-values for the Sociotechnical-vs-Governance aspiration comparison differ between the main text (Z=2.79, padj=.005) and Appendix C.2 (Z=2.78, padj=.01); these should be harmonized.","section":"Results, RQ1 / Appendix C.2"},{"comment":"The gap score subtracts Likert scales with different anchors ('Never' to 'Always' for practice, 'Definitely Not' to 'Definitely' for aspiration); the paper should justify this operation or discuss its interpretational limits.","section":"Methods, Practice and Aspirations"},{"comment":"The term 'human-washing' is introduced without a formal definition or operationalization; consider defining it more precisely (e.g., inclusion criteria for what counts as performative) to facilitate future use.","section":"Discussion / 'Human-Washing'"},{"comment":"The abstract and conclusion state that 'Technical researchers tend to value human research less' without hedging that this is an observed sample-based difference; adding a qualifier such as 'in our sample' would better match the paper's acknowledged recruitment limitations.","section":"Abstract and Conclusion"},{"comment":"The collaboration heatmap in Figure 4 is difficult to interpret because the y-axis labels are the participant's primary area while the x-axis labels are the collaboration target; consider a clearer layout or a descriptive caption.","section":"Figure 4"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of the venue and addresses a timely topic. The main quantitative results are statistically appropriate, but the unsubstantiated bias-direction claim in the Limitations and the incomplete reporting of the barrier and scenario analyses require correction before the central claims can be considered fully supported. The convenience-sample generalizability concern is real but not, by itself, disqualifying if the authors fix the internal inconsistencies and moderate their language."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is the first direct practitioner-level data showing that Technical AI Safety researchers value human-subjects research less than Sociotechnical and Governance researchers do, and collaborate less across disciplines. That's a genuine empirical contribution, not another literature review. It also coins a useful cautionary term: 'human-washing.'\n\nThe new piece is the data itself: a survey (n=93) and interviews (n=17) with active AISE researchers, using scenario-based ratings to capture field-level beliefs about human versus non-human methods. The main quantitative result holds together: in a mixed-effects model with participant random intercepts, Technical (vs Sociotechnical) shows a significantly smaller preference for human methods (β=-0.32, p=.017), and collaboration frequency differs by discipline (K-W χ²=17.12, p<.001), with Technical lower than both Sociotechnical and Governance. The methods are generally sound: nonparametric tests with Holm corrections, Dunn post-hoc tests, attention checks, a positionality statement, and a detailed appendix. The authors are also transparent about the WEIRD skew and the volunteer nature of the interview sample.\n\nThe main soft spot is in the Limitations section, not the core analysis. After conceding that recruitment was through IASEAI 2026 and professional networks, the paper asserts that 'since we undersample researchers who are likely to be resistant to human research, the effect sizes observed... are optimistically stronger in reality.' That does not follow. Under self-selection on affinity for human research, the bias in the Technical-vs-Sociotechnical contrast depends on how selection differs between groups. If the sampled Technical researchers are the human-friendly tail of their discipline, the true gap is larger; but if the Sociotechnical sample is also selection-biased upward, the gap could be smaller or even reversed. The paper simply doesn't know the direction. That sentence should be dropped or replaced with proper uncertainty language.\n\nTwo smaller issues: no data or code is posted, so the exact numbers can't be verified; and the barriers questionnaire has a duplicated item, plus a typo ('quantiative') in the Limitations. The Normative group (n=10) is sensibly excluded from inferential statistics, but the abstract might overstate the breadth by including them in the headline.\n\nWho this serves: researchers and funders working on AI safety/ethics methods and infrastructure, and anyone studying epistemic divides in interdisciplinary fields. It offers concrete levers—funding, mentorship, venue incentives—that make it a useful paper for policy-adjacent readers.\n\nRecommendation: send it to peer review. The convenience sample and the bias-direction overreach are real, but fixable. The core finding is a valuable, appropriately hedged empirical result that deserves referee time. Ask for data release and the Limitations fix.","headline":"Solid practitioner survey evidence of the Technical-vs-Sociotechnical divide in valuing human research, but a flawed bias-direction claim in the Limitations needs correcting.","tokens_in":25173,"tokens_out":4945,"would_cite":true,"duration_ms":43120,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that the AI safety and ethics field's evidence on human harms is skewed because technical researchers systematically undervalue human-subject research and collaborate less across disciplines.","keywords":["AI safety","AI ethics","human-subject research","epistemic divide","interdisciplinary collaboration","expert survey","qualitative methods","research barriers"],"falsifier":"A stratified random sample of AISE researchers, drawn across conferences, sectors, and regions, that found no discipline difference in human-research usefulness ratings or in collaboration frequency would falsify the claim that technical training itself produces the undervaluation.","tokens_in":24217,"feed_emoji":"🤖","tokens_out":8816,"duration_ms":68454,"temperature":0.7,"pith_summary":"The paper tries to establish that the AI Safety & Ethics (AISE) community's reliance on benchmarks, LLM simulations, and normative assumptions over empirical human-subject research is not primarily a matter of evidence quality. Drawing on a survey of $n = 93$ experts and $n = 17$ follow-up interviews across Technical, Sociotechnical, Governance, and Normative backgrounds, it argues that human research is broadly agreed to be valuable yet systematically undervalued, with Technical researchers valuing it least and collaborating across disciplines least. The paper also contends that practical barriers (time, funding, participant access) and infrastructural constraints (mentorship, publication norms, sector incentives) compound the epistemic divide. If correct, the field's evidence base for AI harms is measurably shaped by methodological bias, and fixing it requires changing incentives rather than simply adding more human studies.","feed_headline":"Technical AI safety researchers value human studies least","feed_subtitle":"Survey of 93 experts finds technical AI safety researchers undervalue human research and collaborate across disciplines least.","key_machinery":"The carrying mechanism is a mixed-methods expert elicitation: a survey of 93 AISE researchers and 17 semi-structured interviews. Three quantitative instruments do the main work: the aspiration-minus-practice gap score for each human-research dimension (quantitative/qualitative, lab/field, experimental/observational, representative/specialized), a mixed-effects model of the human-minus-nonhuman usefulness gap across four risk scenarios, and Kruskal-Wallis comparisons of collaboration frequency. These convert individual self-reports into a field-level statement about which methods are legitimized.","core_discovery":"The central discovery is an epistemic divide inside AISE: every disciplinary group shows a positive gap between how much human research it aspires to do and how much it practices, yet the relative usefulness of human versus non-human methods is significantly lower among Technical researchers, with a mixed-effects coefficient of $\\beta = -0.32$ ($p = .017$) against a positive Sociotechnical baseline. Technical researchers also report the lowest cross-disciplinary collaboration frequency (Kruskal-Wallis $\\chi^2 = 17.12$, $p < .001$, with Holm-corrected contrasts against both Sociotechnical and Governance). The paper interprets these findings as showing that human research has epistemic fit with the field's goals but not with its dominant positivist epistemology, and it introduces the term 'human-washing' for the risk of adding human participants as a performative validity check rather than a rigorous contribution.","pith_inferences":["A testable corollary the paper does not pursue: if the epistemic divide is real, then in technical AISE venues human-subject papers should be rarer and less cited than benchmark papers, and the gap should shrink when gatekeepers include more sociotechnical researchers.","The paper's logic implies that simply spending more on human research will not close the divide unless technical researchers also gain literacy in qualitative and interpretivist methods; otherwise human evidence will continue to be judged by positivist standards and found wanting.","The 'human-washing' concept invites operationalization: for instance, measuring whether studies that include human participants actually derive their conclusions from the human data, or merely append them to technical evaluations, could turn the warning into a testable audit.","The field-over-lab preference suggests that infrastructure such as longitudinal cohorts, deployed-system partnerships, and community-based research sites may yield more valuable evidence than additional controlled experiments, an implication the paper mentions but does not develop."],"forward_implications":["If technical researchers systematically undervalue human research, the evidence base for AI harms is skewed toward what benchmarks and simulations can measure, leaving interactional and experiential harms such as emotional overreliance, manipulation, and skill erosion underserved.","The positive aspiration-minus-practice gaps across all disciplines imply that demand for human research exceeds supply; removing resource barriers such as time, funding, and participant access should raise the volume of human-subject evidence even without changing attitudes.","The shared preference for field studies over lab studies implies that funding and review infrastructure should support ecologically valid settings, not only controlled experiments.","If human research becomes a checklist requirement without methodological literacy, the result will be 'human-washing,' where the appearance of empirical grounding replaces the substance, and the paper argues this would undermine the field's legitimacy.","The lower cross-disciplinary collaboration of Technical researchers implies that bridging efforts must be structural (co-supervision, cross-sector partnerships, methodological training) rather than relying on individual goodwill."],"supporting_citations":[{"why":"Documents that AISE research with human subjects is rare, establishing the evidence-gap premise.","marker":"Weidinger et al. 2023"},{"why":"Supplies the construct-validity critique of LLM benchmarks that motivates the need for human evidence.","marker":"Bean et al. 2026"},{"why":"Exemplifies the benchmark-based evaluation paradigm the paper contrasts with human research.","marker":"Reuel et al. 2024b"},{"why":"Maps the AI safety/ethics community divide that frames the disciplinary analysis.","marker":"Gyevnár and Kasirzadeh 2026"},{"why":"Prior literature-level attempt to unify AI safety and ethics that the paper extends to researchers' own attitudes.","marker":"Roytburg and Miller 2025"},{"why":"Provides the template for analyzing resource and infrastructure barriers in cross-disciplinary research teams.","marker":"Agapie, Haldar, and Poblete 2022"},{"why":"Supplies the epistemological pluralism framework used to argue methods should be complementary.","marker":"Miller et al. 2008"},{"why":"Represents the RCT/uplift paradigm, the dominant quantitative human method whose ecological-validity limits the paper discusses.","marker":"Paskov et al. 2026"}],"fun_headline_variants":["Tech AI safety researchers undervalue human studies","Survey: technical AI safety researchers devalue human research","AI safety's technical experts rank human research lowest","AI safety's epistemic divide: tech experts undervalue human studies"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 93 survey respondents and 17 interviewees, recruited largely through one conference and the authors' professional networks, represent the full AI Safety & Ethics community; if the self-selected technical researchers are more skeptical of human research than the broader population, the headline divide is overstated.","fun_headline_variants_meta":{"raw":{"variants":["Tech AI safety researchers undervalue human studies","Survey: technical AI safety researchers devalue human research","AI safety's technical experts rank human research lowest","AI safety's epistemic divide: tech experts undervalue human studies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00079,"raw_usage":{"total_tokens":3461,"prompt_tokens":905,"completion_tokens":2556,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":521,"completion_tokens_details":{"reasoning_tokens":2501}},"tokens_in":521,"tokens_out":2556,"duration_ms":17561,"temperature":1.0,"reasoning_tokens":2501,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T04:13:01.160348+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A stratified random sample of AISE researchers, drawn across conferences, sectors, and regions, that found no discipline difference in human-research usefulness ratings or in collaboration frequency would falsify the claim that technical training itself produces the undervaluation.","supporting_citations":[],"review_version":2}