Pith. sign in

REVIEW 3 major objections 6 minor 13 references

Is the ACL Responsible NLP Checklist a Box-Ticking Exercise? A Large-Scale Analysis of EMNLP 2025

T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read The paper argues that the EMNLP 2025 Responsible NLP Checklist is mostly box-ticking: 44.9% of 'no' justifications are brief or empty, 15.5% of main-track checklists contain parent-child contradictions, and 53% of authors dismiss…

desk verdict First large-scale look at EMNLP 2025 checklist justifications with a useful data release, but the headline 44.9% failure rate overreaches its word-count proxy and the abstract contradicts the body on the contradiction rate. read the letter →

arxiv 2608.09280 v1 pith:CIWL2MVI submitted 2026-08-10 cs.CL

classification cs.CL
keywords responsibleNLPchecklistanalysisEMNLP2025ethicsreproducibilitysocietalimpactsurfacecompliancebox-ticking
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

EMNLP 2025 made the ACL Responsible NLP Checklist public for every accepted paper, and this study asks whether a required form can actually push researchers toward transparent, ethical, and socially aware practice. Analyzing 73,922 responses from 3,214 main- and findings-track checklists, the paper reports three patterns: 44.9% of 'no' justifications are brief or empty, 15.5% of main-track checklists contain logical contradictions between a parent question and its child questions, and 53% of authors answer 'no' to the mandatory risks question A2. The authors read these patterns as surface compliance, with ethics and societal-impact items answered minimally and effectively isolated from the paper's body. If the analysis is right, the checklist is not serving its stated purpose, and the paper's proposed fixes—disabling child questions when a parent is 'no', banning empty justifications, and adding reviewer scrutiny to risk answers—offer a concrete path for reform.

What carries the argument

The central object is the ARR Responsible NLP Research Checklist: 23 questions grouped into five parent items (A mandatory, B artifacts, C experiments, D human annotators, E AI assistance), each with child subquestions answered YES, NO, or N/A. The analysis exploits three structural features: the parent-child gating rule (if a parent is NO or N/A, every child must also be NO or N/A), which defines logical contradictions; the requirement that NO subquestion answers carry a written justification, which the authors grade by word count (adequate >10, brief 1–10, empty 0); and the requirement that YES answers cite paper sections, from which the authors build a section-reference linking dataset and a concentration score (total section references divided by unique sections cited) to detect vague or repetitive pointing. The pipeline that extracts these from 3,214 PDFs uses a document parser plus a constrained LLM pass, with a reported 99.9% item-level accuracy on a 71-paper validation sample and 90.9% section-text retrieval.

What would settle it

Manually annotate a random sample of, say, 200 'brief' NO justifications from the released corpus and ask whether each actually explains the missing practice; if most short answers are judged adequate, the headline failure rate collapses, though the contradiction and risk-dismissal findings would remain.

Watch

Extended reading notes

Core claim

The central claim is that the EMNLP 2025 Responsible NLP Checklist is being treated as a bureaucratic form rather than a reflective instrument. The evidence is quantitative: across 41,607 main-track responses, the paper finds that questions about limitations and risks (A1, A2) and human-annotator ethics (D) receive far fewer 'yes' answers than reproducibility questions (B, C); 44.9% of 'no' justifications fall into 'brief' (1–10 words) or empty categories; and 281 main-track papers (15.5%) contain at least one case where a parent question is answered 'no' but a child question is answered 'yes', which the checklist's own gating rules make logically impossible. The paper also finds that 53% of authors say their work has no potential risks, and that the same patterns appear in the findings track, with 18.1% of checklists containing contradictions. The authors interpret this not as individual carelessness but as a design problem: the form gives no feedback, does not enforce its own parent-child logic, and does not require justifications to meet any standard.

Load-bearing premise

The 44.9% figure depends on treating any justification of 10 or fewer words as inadequate, an assumption the paper itself flags could mislabel a perfectly good short answer.

Editorial extensions

If this is right

  • If the checklist is made self-consistent by disabling child questions when a parent is answered 'no', the detected 15.5% contradiction rate should drop to near zero and authors would be forced to reconsider their parent answer.
  • Enforcing a minimum justification length would push the 44.9% failure rate down, though it would not by itself guarantee meaningful engagement.
  • Giving A2 (risks) explicit reviewer scrutiny would likely raise the 53% 'no risk' rate only if reviewers can request revisions; the paper shows the current form offers no such check.
  • The similarity between main and findings tracks suggests the checklist does not discriminate paper quality, so acceptance-track comparisons that rely on checklist compliance would need another measure.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A conservative reading of the data is that the 44.9% figure is an upper bound on bad faith: word count cannot distinguish a terse but honest explanation from a dismissive one, and the paper's own limitation acknowledges this.
  • The high N/A rates on data-ethics and human-annotator items (B4, D) could mean researchers genuinely avoid private or human data, or it could mean N/A is used as a low-effort escape hatch; the paper does not separate these, and a follow-up could compare N/A rates with paper content.
  • The section-reference dataset opens a natural next test: automatically check whether the cited section actually answers the checklist question, which would expose 'lazy section pointer' compliance more directly than word count.
  • Cross-venue replication (for example, at other ACL conferences with public checklists) would clarify whether these patterns are specific to EMNLP 2025 or characteristic of the checklist format itself.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper presents a large-scale analysis of the EMNLP 2025 Responsible NLP Research Checklist. The authors release the CheckBox corpus of 3,214 main- and findings-track checklists (73,922 responses) plus a linking of checklist references to paper sections, and they use it to report three headline findings: 44.9% of NO justifications are brief or empty, 15.5% of main-track checklists contain parent-child logical contradictions, and 53% of authors answer NO to the mandatory risk question A2. They interpret these numbers as evidence that the checklist is often filled out as a box-ticking exercise, and they propose reforms such as enforcing a minimum word count and disabling child questions when a parent is NO.

Significance. The dataset is a potentially valuable resource: it is, to my knowledge, the first public corpus of ACL checklist responses with justifications and section links, and the paper addresses a question of direct policy relevance to ACL venues. The pipeline description is transparent, the validation sample reports 99.9% item-level accuracy, and the authors explicitly acknowledge limitations. However, the validity of the headline 44.9% figure rests on an unvalidated word-count proxy that the Limitations section concedes can misclassify adequate justifications, and the abstract and body disagree on the contradiction rate (6% vs 15.5%). Because these two numbers are the quantitative pillars of the central 'box-ticking' claim, the conclusion is currently stronger than the evidence supports. The contribution is therefore significant and publishable in principle, but only after the measurement validity and reporting consistency issues are addressed.

major comments (3)
  1. [§5.3, Table 3, Limitations] The headline '44.9% of NO justifications are poor or bad-faith' depends entirely on the unvalidated cutoff in §5.3: justifications are 'adequate (>10 words), brief (1-10), empty (0)', and 'Failure=Empty+Brief'. The Limitations explicitly concede that 'a ten word justification could be perfectly adequate but our software would classify it as inadequate', yet the Abstract and Conclusion still use 'poor or bad-faith' without hedging. Since the two required components of a NO justification in §3.1 are content-based (why the practice was omitted, and the missing information), the authors should either manually annotate a sample of brief justifications and report the fraction that actually satisfy §3.1, or replace the adequacy claim with a purely descriptive 'brief or empty' claim. Without this validation, the box-ticking conclusion is not supported. I also note that 213 of the 2,074 justifications are empty (about 10.3%), so even a stricter empty-only measure would still show a problem, but it would not justify the 'nearly half' framing.
  2. [Abstract and §5.4/Table 4] The Abstract reports that '6% of all checklists contained logical contradictions between parent and child responses', while §5.4 and Table 4 report 281 of 1,809 main-track papers (15.5%), and the Findings comparison reports 18.1%. These numbers are mutually inconsistent, and the paper nowhere explains the discrepancy or defines the denominator separately for the abstract figure. The authors must correct the abstract or the body and state precisely whether each rate is per paper, per checklist, or per question item.
  3. [§5.1, §6, Abstract] The claim that '53% of authors dismiss potential risks or social impacts of their work' over-interprets a NO answer to A2. A2 asks whether the authors 'discussed any potential risks of your work', so a NO response means only that the paper/checklist does not report a risk discussion; it does not establish that the authors dismissed or ignored risks. Phrases such as 'violation of the ACM Code of Ethics' and 'surface compliance' are therefore stronger than the data support. The authors should rephrase these statements to say that a majority of papers do not report risk discussion, and separate that observation from a substantive judgment about whether risks exist.
minor comments (6)
  1. [§4.2] The text says 'all 3,124 papers and checklists are included', but Figure 1 and the pipeline description state 3,214 (1,809 main + 1,405 findings); the number 3,124 appears to be a typo.
  2. [Abstract and throughout] 'Main and Finding tracks' should be 'Main and Findings tracks' for consistency with the body, which uses 'Findings track'.
  3. [Abstract] The phrase 'risks of appliances' in the recommendations should presumably read 'risks of applications'.
  4. [§5.2, Figure 3] The 'problematic' threshold of Score >4 for citation concentration is introduced without a justification or sensitivity analysis; since the text draws conclusions about problematic referencing from this cutoff, a brief rationale or robustness check would strengthen the claim.
  5. [§5.5, Figures 9-10] Comparative statements such as 'significantly less adequate' are not accompanied by confidence intervals or significance tests; adding them would make the findings-track comparison more rigorous.
  6. [§3.1] The conditional rules for when justifications are required are formatted as a bullet list that is hard to parse; a small truth-table or cleaner notation would improve readability.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the study is an empirical measurement of externally released checklist data, with any concerns limited to measurement validity rather than derivation.

full rationale

The paper's derivation chain is empirical, not definitional. It uses a documented extraction pipeline (Marker plus constrained LLM prompting) with 99.9% item-level agreement on a manually annotated sample of 71 papers, then reports descriptive statistics over 73,922 extracted checklist responses. The headline figures—44.9% brief/empty NO justifications, 53% NO on A2, and parent-child inconsistency rates of 15.5% (main) and 18.1% (findings)—are direct counts of the external checklist data under the checklist's own response rules. The word-count classification in Section 5.3 is a transparent operationalization (adequate >10 words; Failure = Empty + Brief), and the Limitations section explicitly concedes that a ten-word justification could be adequate; this is a measurement-validity caveat, not a circular step, because the threshold is not fitted from the data and the paper does not use the conclusion to justify the metric. No load-bearing self-citations, imported uniqueness theorems, or ansatz-by-citation moves appear; the related-work discussion grounds the checklist in prior external work. The LLM extraction schema is a potential source of systematic bias, but the reported manual validation and the lack of any reduction of the claims to the schema's configuration mean this remains a speculation, not demonstrated circularity.

Assumptions & free parameters 2 free parameters · 3 assumptions · 0 invented entities

The central claims rest on two hand-set analytical thresholds (word count and concentration score) and on three domain assumptions about the extraction pipeline, sample representativeness, and the validity of word count as a quality proxy. No new entities are introduced.

free parameters (2)
  • NO justification adequacy threshold = 10 words
    Justifications with 10 or fewer words are classified as 'brief' and grouped with empty responses as failures; this threshold directly determines the 44.9% headline figure.
  • Concentration score problematic threshold = >4
    Checklists citing more than 4 sections per distinct section are labeled 'problematic'; this threshold shapes the claim that authors over-cite sections.
assumptions (3)
  • domain assumption The LLM-based extraction pipeline preserves parent-child answer relationships with sufficient accuracy for the contradiction analysis.
    The manual validation covered only 71 of 3,214 papers, and the authors note potential issues with checkbox symbol misreads and cut-off justifications.
  • domain assumption The EMNLP 2025 main and findings track checklists are representative of ACL community practice.
    The study excludes workshop, industry, and system demonstration tracks because checklists were unavailable, so the conclusions are limited to two tracks.
  • domain assumption Word count is a valid proxy for justification quality.
    The authors acknowledge in Limitations that a ten-word justification can be perfectly adequate, yet the failure-rate metric relies on this proxy.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Is the ACL Responsible NLP Checklist a Box-Ticking Exercise? A Large-Scale Analysis of EMNLP 2025." pith.science (2026). https://pith.science/paper/CIWL2MVI

@misc{pith2026260809280,
  author       = {Pith},
  title        = {Pith review of: Is the ACL Responsible NLP Checklist a Box-Ticking Exercise? A Large-Scale Analysis of EMNLP 2025},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CIWL2MVI}},
  note         = {Machine review of arXiv:2608.09280}
}
abstract

Responsible NLP practice includes a) transparency, b) ethics, and c) societal impacts. The Responsible NLP Checklist aims to push these goals, and promote responsible practice. Recently, ACL released the EMNLP 2025 Checklists to aid transparency on the current research practice, which we focus on. We curate and release the first two datasets of: a) all the checklist responses and justifications from the EMNLP 2025 Main and Finding tracks; b) checklist reference linking to paper sections. We also provide the first analysis of recent EMNLP Checklists, by examining $73,922$ responses and justifications to them. For the Main track, we find that authors isolate ethics questions of the Checklist from the paper's bulk, mimicking the trend of ethics being an afterthought. We then examine \texttt{NO} responses. We find $44.9\%$ of justifications are poor or bad-faith, being brief or empty. Then, we find significant issues with the checklist design and effort of authors, namely that $6\%$ of all checklists contained logical contradictions between parent and child responses. We also find evidence of surface compliance for responsible ethics, with $53\%$ authors dismissing potential risks or social impacts of their work, for which there should be none. We compare this to the Findings track, noticing a similar trend in both tracks. Lastly, we discuss the implications of the checklist design and provide recommendations for future checklist iterations. Including: a) enforcing a minimum word count, b) enforcing more scrutiny on the risks of appliances.

Figures

Figures reproduced from arXiv: 2608.09280 by the authors.

Figure 1
Figure 1. The ARR Responsible NLP Research Check￾list has questions relating to ethics, social impact, and re￾producibility. With EMNLP 2025 releasing said check￾lists publicly (for transparency), we made a pipeline to analyse the practices. We provide the first analysis with our own pipeline which created the Checkbox Corpus of 3,214 papers and checklists. In CheckBox, we crucially found 2 3 are not addressed properly. We pr… view at source ↗
Figure 3
Figure 3. Citation concentration scores distribution. [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 2
Figure 2. Checklist subquestion answer distribution [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Co-citation matrix. Each cell records how [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Justification quality for S=YES where the P = YES, grouped by empty, brief (<10 words) and adequate (>10 words). or empty words (see [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 7
Figure 7. Figure 7: Citation concentration overlay for both tracks. [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 9
Figure 9. Figure 9: Justification failure rates (empty + brief) by [PITH_FULL_IMAGE:figures/full_fig_p013_9.png]
Figure 10
Figure 10. Figure 10: Percentage-point differences (Main minus [PITH_FULL_IMAGE:figures/full_fig_p013_10.png]
Figure 11
Figure 11. Figure 11: Parent-child inconsistency rates by section, [PITH_FULL_IMAGE:figures/full_fig_p014_11.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

13 extracted references · 9 canonical work pages

  1. [7]

    Ethical con- cern identification in NLP: A corpus of ACL Anthol- ogy ethics statements. InProceedings of the 2025 Conference of the Nations of the Americas Chap- ter of the Association for Computational Linguistics: Human Language Technologies (V olume 1: Long Pa- pers), pages 11618–11635, Albuquerque, New Mex- ico. Association for Computational Linguisti...

  2. [8]

    Association for Computa- tional Linguistics

    Reproducibility in NLP: What have we learned from the checklist? InFindings of the Association for Computational Linguistics: ACL 2023, pages 12789– 12811, Toronto, Canada. Association for Computa- tional Linguistics. Vik Paruchuri

  3. [9]

    Anna Rogers, Timothy Baldwin, and Kobi Leins

    Improving reproducibility in machine learning research(a report from the neurips 2019 reproducibil- ity program).Journal of Machine Learning Research, 22(164):1–20. Anna Rogers, Timothy Baldwin, and Kobi Leins

  4. [10]

    InFindings of the Association for Computational Linguistics: EMNLP 2021, pages 4821–4833, Punta Cana, Do- minican Republic

    ‘just what do you think you’re doing, dave?’ a check- list for responsible data use in NLP. InFindings of the Association for Computational Linguistics: EMNLP 2021, pages 4821–4833, Punta Cana, Do- minican Republic. Association for Computational Linguistics. Deven Santosh Shah, H. Andrew Schwartz, and Dirk Hovy

  5. [11]

    InFindings of the Association for Computational Linguistics: EMNLP 2025, pages 25738–25760, Suzhou, China

    Evolving stances on reproducibility: A longitudinal study of NLP and ML researchers’ views and experience of reproducibility. InFindings of the Association for Computational Linguistics: EMNLP 2025, pages 25738–25760, Suzhou, China. Associa- tion for Computational Linguistics. David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman...

  6. [13]

    Talukder Hasnat Zadid, Shahidul Morsalin Jahin, Aninda Roy Dhruba, Sharia Arfin Tanim, Ab- dullah Muhammad Hamja, Md

    Mineru: An open-source solution for precise document content extraction.Preprint, arXiv:2409.18839. Talukder Hasnat Zadid, Shahidul Morsalin Jahin, Aninda Roy Dhruba, Sharia Arfin Tanim, Ab- dullah Muhammad Hamja, Md. Rakib Hasan, Md Saef Ullah Miah, Md. Imamul Islam, and Ahmed Al Mansur

  7. [2017]

    CoRR, abs/1707.00061

    Racial disparity in natural language processing: A case study of social media african-american english. CoRR, abs/1707.00061. Joy Buolamwini and Timnit Gebru

  8. [2019]

    Show your work: Improved reporting of experimental results. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Lan- guage Processing (EMNLP-IJCNLP), pages 2185– 2194, Hong Kong, China. Association for Computa- tional Linguistics. Michael Galarnyk, Rutwik Routu, Vi...

Show all 13 references
  1. [2020]

    InProceedings of the 2020 Con- ference on Empirical Methods in Natural Language Processing (EMNLP), pages 7534–7550, Online

    Fact or fiction: Verifying scientific claims. InProceedings of the 2020 Con- ference on Empirical Methods in Natural Language Processing (EMNLP), pages 7534–7550, Online. As- sociation for Computational Linguistics. Bin Wang, Chao Xu, Xiaomeng Zhao, Linke Ouyang, Fan Wu, Zhiyu...

  2. [2021]

    Lukas Blecher, Guillem Cucurull, Thomas Scialom, and Robert Stojnic

    Introducing the neurips 2021 paper checklist. Lukas Blecher, Guillem Cucurull, Thomas Scialom, and Robert Stojnic

  3. [2023]

    Su Lin Blodgett and Brendan O’Connor

    Nougat: Neural optical understanding for academic documents.Preprint, arXiv:2308.13418. Su Lin Blodgett and Brendan O’Connor

  4. [2024]

    Alina Beygelzimer, Yann Dauphin, Percy Liang, and Jennifer Vaughan

    Docling technical report.Preprint, arXiv:2408.09869. Alina Beygelzimer, Yann Dauphin, Percy Liang, and Jennifer Vaughan

  5. [2025]

    InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 667–677, Suzhou, China

    ConfReady: A RAG based as- sistant and dataset for conference checklist responses. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 667–677, Suzhou, China. As- sociation for Computational Linguistics. Antoni...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.