{"id":"5738b904-1d86-4add-9b91-9b7cfb56eab6","arxiv_id":"2506.04482","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Practitioners trying to measure representational harms in LLM-based systems often cannot use public measurement instruments, either because the instruments lack validity, specificity, interpretability, or actionability, or because practical and institutional barriers block their uptake.","lead":"This study interviewed 12 practitioners who evaluate LLM-based systems for representational harms. It finds that publicly available datasets, benchmarks, and tools often go unused because they are misaligned with practitioner needs or are blocked by practical and institutional barriers.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's 'often' overstates what the sampling design can support: a self-selected, snowball-recruited sample of 12 practitioners cannot establish prevalence, and the paper's own limitations section concedes this.","rationale":"I agree with the reader that the paper is transparent, methodologically detailed, and makes a genuine contribution in distinguishing instrument usefulness from instrument uptake. The reader's weakest assumption identifies two concerns: possible protocol-induced alignment from the Q20 prompts in Appendix B, and the non-representativeness of the 12-person snowball sample. Of these, the sample representativeness issue is the more load-bearing because it directly undercuts the prevalence quantifier in the central claim as stated in the abstract ('often unable'). The paper's own Limitations section explicitly says the sample is likely skewed toward practitioners who faced challenges and that the findings cannot answer prevalence questions. That is a direct internal tension with the abstract's wording. The Q20 prompting concern is real but secondary: the interviews began with open-ended questions (Q6, Q18–Q19, Q22, Q27–Q28), and the two-type taxonomy emerged from thematic analysis of full transcripts, so the prompting alone would not fully manufacture the central distinction. A scoped rewrite—replacing 'often' with 'in our interview sample' or 'among the practitioners we interviewed'—would resolve the main concern while preserving the paper's contribution. The reader's verdict of CONDITIONAL is therefore appropriate, and my stress-test does not move it. No concerns about author integrity, data fabrication, or fundamental unworkability of the qualitative method were identified.","tokens_in":20133,"tokens_out":4618,"duration_ms":62038,"concrete_test":"Administer a short pre-registered survey to a broader and independently recruited sample of practitioners who evaluate LLM-based systems, for example by sampling from industry job postings, professional associations, or multiple company lists rather than the authors' networks and snowball referrals. Ask whether respondents have tried to use publicly available instruments for measuring representational harms and, if so, which challenges prevented use. Pre-register a threshold for 'often' (e.g., more than 50% reporting inability despite desire). If the rate falls below that threshold, the abstract's prevalence claim is unsupported and should be scoped to the interview sample; if the rate is at or above threshold, the prevalence wording gains independent support.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, as stated in the abstract and conclusion, is a prevalence claim: practitioners are 'often' unable to use publicly available instruments for measuring representational harms. The supporting evidence is 12 semi-structured interviews recruited through professional networks, social media, cold emails, and snowball sampling, with a low response rate (1 of 73 cold-emailed practitioners participated). The authors themselves state in the Limitations section that the participant pool is 'likely skewed toward practitioners who faced challenges' and that the findings 'do not enable us to answer questions about the prevalence of the challenges.' A sample that is admitted to be non-representative and enriched for practitioners with negative experiences cannot support a population-level quantifier like 'often.' It can establish that such challenges exist and were experienced by the interviewed practitioners, and it can support the two-type taxonomy as an inductively derived characterization. But the headline claim, as worded, requires the load-bearing assumption that the sample is representative enough to estimate prevalence, and that assumption is explicitly disavowed in the paper's own limitations. This is an internal inconsistency between the central claim and the stated scope of the method, not merely an outside objection. The paper's substantive contribution—the usefulness-versus-uptake distinction—may survive a scoped rewrite, but the prevalence wording should be removed or explicitly limited to the interview sample.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a semi-structured interview study with 12 practitioners who evaluate LLM-based systems, focusing on their experiences using publicly available instruments for measuring representational harms. The authors find that practitioners in their sample faced two broad types of challenges: instruments are sometimes not useful because they lack validity, specificity, interpretability, or actionability for the practitioner's context, and even useful instruments are sometimes not used because of practical and institutional barriers such as security requirements, data licensing, and organizational culture. Drawing on measurement theory and pragmatic measurement, the paper offers recommendations for instrument designers, including systematic concept systematization, providing interpretability resources, and designing extensible instruments. The paper includes a detailed interview guide, a systematic literature review of desiderata, a positionality statement, and an explicit limitations section.","tokens_in":20350,"tokens_out":3199,"duration_ms":34661,"significance":"If the findings are interpreted as a qualitative characterization of challenges rather than a prevalence estimate, the paper makes a useful contribution to the responsible AI and NLP evaluation literature. Its distinction between instruments that are not useful and instruments that are useful but not used is a valuable framing that connects practitioner experience to measurement theory. The paper also gives concrete, actionable recommendations grounded in established measurement frameworks, and it is unusually transparent about recruitment difficulties, coding procedures, and study limitations. The interview guide and PRISMA-based literature review are strengths that will facilitate replication and extension. However, the headline claim in the abstract and conclusion goes beyond what the sampling design can support, and the interview protocol's prompting of the seven a priori desiderata weakens the claim that those desiderata emerged from practitioners' own accounts.","major_comments":[{"comment":"The abstract, introduction, and conclusion claim that practitioners are 'often' unable to use publicly available instruments for measuring representational harms. This is a prevalence claim, but the Limitations section explicitly states that the participant pool is 'likely skewed toward practitioners who faced challenges' and that the findings 'do not enable us to answer questions about the prevalence of the challenges.' A self-selected, snowball-recruited sample of 12 practitioners cannot support the quantifier 'often.' I recommend rewording the central claim to an existence claim about the challenges practitioners can face (for example, 'practitioners in our sample reported being unable to use...' or 'practitioners can face challenges that leave them unable to use...'), and removing or qualifying 'often' in the abstract, §1, and §6.","section":"Abstract; §1; §6; Limitations"},{"comment":"The claim in §4 that 'the desiderata identified in Table 3 aligned closely with considerations participants described' and that 'participants did not report considering any additional desiderata' is partly an artifact of the interview protocol. Appendix B Q20 explicitly asks participants to confirm or deny each of the seven a priori desiderata, so the observed alignment is partly by construction. The paper should either analyze unprompted mentions separately from prompted confirmations, or substantially soften the claim to say that participants affirmed these desiderata when asked. As written, the inductive–deductive coding description (§3) does not fully address the circularity introduced by the scaffolding.","section":"§3 Interview scaffolding; §4 Desiderata alignment; Appendix B Q20"}],"minor_comments":[{"comment":"Table 2 uses participant IDs P01–P12 while the text (e.g., §3 and §4) refers to P1–P12; please unify the notation.","section":"Table 2; §3"},{"comment":"The sentence beginning 'Finally, we note that multiple participants lived in countries where English is not the primary language...' introduces a topic that is not strictly a challenge 'exacerbated by, but extend[ing] beyond, representational harms'; consider moving it to the Limitations or presenting it as a separate observation.","section":"§4.3"},{"comment":"The annotation strategy states that the desiderata list was developed before mapping passages and that no additional desiderata emerged; this is a limitation of the systematic review that should be acknowledged in the main text alongside the interview scaffolding limitation.","section":"Appendix A.1.2"},{"comment":"The row labeled 'Other' contains the phrase 'Matched guise probing'; for consistency with the other entries, consider using a full sentence or consistent formatting.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The reviewer's concern about the 'often' claim is well founded and should be addressed in revision. The paper's substantive qualitative contribution is sound and the transparency about limitations is commendable, but the central claim needs to be scoped to avoid an internal inconsistency with the stated methodological limitations."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: read this for the usefulness/uptake distinction, which is a real contribution. But the abstract's \"often\" is not supported by a sample of 12 snowball-recruited practitioners, and the paper itself concedes as much in the limitations. That internal inconsistency needs fixing before this can be cited as evidence about prevalence.\n\nWhat's new and good: prior HCI work on fairness toolkits focused on allocative harms, and prior NLP work assessed instruments analytically. This is the first interview study I know of that asks practitioners specifically about measuring representational harms in LLM-based systems. The two-type taxonomy—instruments that are not useful versus useful instruments that are not used—is a clean and practically meaningful division. It separates design failures from organizational and institutional barriers like data licensing, security, and lack of time or incentive. That distinction is worth having. The methods are also transparent: the interview guide is in the appendix, the coding process is described, the authors give a positionality statement, and they are honest about the low response rate and the likely skew of the participant pool. The recommendations grounded in measurement theory and pragmatic measurement are concrete, especially the points about systematization, interpretability, and extensibility.\n\nWhere it gets soft: the stress-test concern lands. The abstract and conclusion say practitioners are \"often\" unable to use these instruments, which is a prevalence claim. The evidence is 12 self-selected interviews, with a 1-in-73 response rate to cold emails, and the authors themselves write that the pool is \"likely skewed toward practitioners who faced challenges\" and that the findings \"do not enable us to answer questions about the prevalence of the challenges.\" The word \"often\" should be cut or explicitly scoped to the interview sample. There is also some circularity: the interview protocol prompts participants to confirm each of the seven a priori desiderata (Appendix B, Q20), so the observed alignment between those desiderata and practitioners' concerns is partly constructed. The authors do invite open-ended discussion, and participants did raise unprompted barriers, so the taxonomy is not purely an artifact. But the paper should acknowledge this more directly and discuss how much the prompts may have shaped the reported alignment. These are fixable problems, not fatal ones.\n\nWho is this for: anyone designing or studying evaluation instruments for social harms in NLP, and practitioners who want vocabulary for why these tools fail in real settings. It deserves a serious referee—a good reviewer will push for a scoped abstract and a fuller treatment of the prompt artifact, and the paper will be stronger for it.","headline":"A genuinely useful qualitative study of why practitioners don't use public representational-harm instruments, but the abstract overstates prevalence in a way the authors' own limitations section disavows.","tokens_in":20882,"tokens_out":2050,"would_cite":true,"duration_ms":20868,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Public instruments for measuring LLM representational harms often go unused by practitioners.","keywords":["representational harms","LLM evaluation","measurement instruments","practitioner needs","qualitative interviews","measurement theory","pragmatic measurement","responsible AI"],"falsifier":"A random-sample survey of practitioners with a high response rate, or an audit of actual instrument usage logs across organizations, that found most practitioners routinely use public instruments as-is and rate them valid, specific, and actionable, would contradict the paper's central claim.","tokens_in":19943,"feed_emoji":"📏","tokens_out":4517,"duration_ms":36162,"temperature":0.7,"pith_summary":"This paper reports that practitioners tasked with evaluating large language model systems for representational harms are often unable to use the publicly available measurement instruments the NLP community has produced, even when they want to. Through semi-structured interviews with 12 practitioners, the authors find two distinct sources of failure: instruments are not useful when they lack validity, specificity, interpretability, or actionability, and instruments go unused when security, licensing, or organizational barriers block uptake. The authors argue that this gap is not simply a matter of better metrics, and they draw on measurement theory and pragmatic measurement to recommend how instrument designers, practitioners, and organizations could close it. If the finding holds, it would imply that a large body of public evaluation research is currently failing its intended users.","feed_headline":"Practitioners can't use public tools for measuring LLM harms","feed_subtitle":"Interviews with 12 evaluators show instruments fail on validity and specificity, or hit workplace barriers.","key_machinery":"The study is carried by semi-structured interviews with 12 practitioners, scaffolded by a set of seven instrument desiderata: validity, reliability, specificity, extensibility, scalability, interpretability, and actionability. The interview guide prompts participants to discuss each desideratum and to describe any additional challenges. The analysis then maps reported challenges onto the \"not useful\" versus \"not used\" distinction, and the authors use measurement theory and pragmatic measurement as the interpretive lens for their recommendations, particularly the ideas that instruments should document their systematized concept and that extensible, modular instruments can be adapted by practitioners.","core_discovery":"The paper's central claim is that the public measurement instruments for representational harms are frequently unusable in practice, and that the reasons fall into two classes. In one class, instruments are not useful: practitioners reported that concepts are often undefined or detached from theory, that datasets contain mislabeled examples, that benchmarks may be contaminated by training data, that instruments are not specific to a system's context, and that outputs cannot be interpreted or acted upon. In the other class, instruments are not used even when potentially useful, because organizations impose security, data-licensing, time, or incentive constraints. The authors report that every participant discussed challenges that prevented use, that validity and specificity were the primary considerations, and that practitioners often responded by building their own bespoke instruments.","pith_inferences":["A natural extension of the paper's logic is that public evaluation artifacts in NLP more broadly may need to be assessed for uptake, not just for technical performance.","The paper's interviews were not designed to estimate prevalence, so a large survey of practitioners would be a reasonable next step to test how widespread the reported barriers are.","Because the authors note that participants saw their difficulties as extending to other abstract or contested concepts, the \"not useful, not used\" split may generalize well beyond representational harms.","Practitioners measuring harms in low-resource languages likely face even steeper shortages of usable instruments, a gap the paper only touches on."],"forward_implications":["Instrument designers should document the systematized concept behind each instrument, separating definition from operationalization, so practitioners can judge what is being measured.","Public instruments should ship with interpretability aids, such as distributions of scores on known datasets and guidance on what a given score means.","Designing instruments to be open, modular, and extensible can help practitioners adapt them to local contexts while preserving validity and reliability.","Organizations and regulators can remove practical and institutional barriers, such as security constraints and missing incentives, rather than leaving uptake to designers alone.","If instruments remain unused, practitioners will continue to build bespoke instruments, making measurements harder to compare across organizations."],"supporting_citations":[{"why":"Supplies the measurement-theory framework of systematization, operationalization, application, and interrogation that grounds the recommendations.","marker":"Adcock and Collier, 2001"},{"why":"Supplies the pragmatic measurement perspective that instruments must be both useful and used, which organizes the uptake analysis.","marker":"Glasgow and Riley, 2013"},{"why":"Provides prior evidence that practitioner needs differ from research assumptions, motivating the interview approach.","marker":"Holstein et al., 2019"},{"why":"Provides the definition of representational harms and the critique that bias constructs in NLP are often poorly conceptualized.","marker":"Blodgett et al., 2020"},{"why":"Supplies the measurement-theoretic framing of validity and reliability for fairness metrics.","marker":"Jacobs and Wallach, 2021"},{"why":"Prior work on actionability of bias metrics that the paper builds on for its recommendations.","marker":"Delobelle et al., 2024"},{"why":"Documents recruitment challenges for studying AI industry practitioners, supporting the paper's low response-rate limitations.","marker":"Scheuerman, 2024"}],"fun_headline_variants":["Public LLM harm meters fail real-world evaluators","Two reasons practitioners ditch LLM harm tools","LLM harm measures often unusable, interviews show","Practitioners forced to build own LLM harm metrics","Public tools for LLM harms hit validity and walls"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 12 recruited practitioners, most reached through the authors' networks and snowball sampling, are telling the truth about their measurement practices and are not just echoing the challenges the interview guide put in front of them.","fun_headline_variants_meta":{"raw":{"variants":["Public LLM harm meters fail real-world evaluators","Two reasons practitioners ditch LLM harm tools","LLM harm measures often unusable, interviews show","Practitioners forced to build own LLM harm metrics","Public tools for LLM harms hit validity and walls"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000166,"raw_usage":{"total_tokens":1204,"prompt_tokens":847,"completion_tokens":357,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":463,"completion_tokens_details":{"reasoning_tokens":283}},"tokens_in":463,"tokens_out":357,"duration_ms":11705,"temperature":1.0,"reasoning_tokens":283,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:40:58.482886+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A random-sample survey of practitioners with a high response rate, or an audit of actual instrument usage logs across organizations, that found most practitioners routinely use public instruments as-is and rate them valid, specific, and actionable, would contradict the paper's central claim.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior work on actionability of bias metrics that the paper builds on for its recommendations."}],"review_version":1}