{"id":"539ab5f5-2cf2-410a-a109-3ee842224c60","arxiv_id":"2608.09280","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Analysis of EMNLP 2025 Responsible NLP Checklists shows 44.9% of NO justifications are brief or empty and 15.5% of main-track checklists contain parent-child logical contradictions.","lead":"A study collected 73,922 answers to the ACL Responsible NLP Checklist from EMNLP 2025 papers and analyzed how carefully authors filled them out. It finds many justifications are missing or brief, and some checklists contain contradictory answers, suggesting the form is often treated as a formality.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 44.9% 'poor or bad-faith' rate rests entirely on a 10-word cutoff that the authors concede can misclassify adequate justifications; manual annotation is needed before that headline can stand.","rationale":"The reader's weakest_assumption correctly identifies the word-count threshold as the load-bearing assumption, and I agree. The paper has independent strengths: it releases the dataset, reports a manual validation of the extraction pipeline (99.9% item-level accuracy on 2,911 items), and the parent-child contradiction analysis is objectively checkable from the released data. The abstract/body discrepancy in the contradiction rate (6% vs 15.5%) is a real credibility issue, but it is secondary because the body provides a detailed table and a clear counting method; the 15.5% figure can be verified directly. The most consequential claim, however, is the 44.9% failure rate and the characterization of those responses as 'poor or bad-faith.' That claim is only as strong as the assumption that 10 or fewer words equals inadequate justification, an assumption the authors themselves concede in Limitations can produce false positives. Because the paper explicitly calls for future manual annotation by checklist researchers, the proposed test is feasible and would settle whether the headline overstates the problem. The verdict should remain CONDITIONAL: the analysis is valuable and largely reproducible, but the strongest headline needs either manual validation or more cautious wording before full acceptance.","tokens_in":13722,"tokens_out":5877,"duration_ms":63034,"concrete_test":"Manually annotate a stratified random sample of the 934 main-track NO justifications currently labeled brief (721) or empty (213), with two annotators per item, blind to the word-count label. Use the checklist's own rule in §3.1: a justification is adequate if it gives (i) why the practice was not included and (ii) the missing information, or if the omission is genuinely inapplicable to the specific subquestion. Compute the proportion of word-count-inadequate items judged adequate. Also rerun the headline rate with thresholds of 5 and 15 words and with a semantic non-empty check. If the human-adequate share among brief items exceeds about 20%, the 44.9% figure and 'bad-faith' language require revision; if it is near zero, the claim survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5.3 defines adequate NO justifications as longer than 10 words and brief as 1–10 words, then equates brief+empty with 'poor or bad-faith' in the abstract and conclusion. This is the load-bearing step: the paper's headline number and its box-ticking narrative depend on treating word count as a valid proxy for whether a justification satisfies the two required components from §3.1 (why the practice was omitted, and the missing information). The Limitations section explicitly concedes that a perfectly adequate 10-word justification would be classified as inadequate, but the body and conclusion still use 'poor or bad-faith' and 'bare minimum' language without hedging. Since 44.9% is the first quantitative pillar of the central claim, and since the 53% 'dismissing risks' finding is similarly an interpretation of NO answers rather than a content analysis, the strength of the whole argument hinges on this proxy. If a nontrivial fraction of the 934 brief or empty justifications are actually responsive to their specific subquestions, the headline failure rate is overstated and the 'box-ticking' conclusion is weakened. This is a correctness risk in the measurement, not merely a framing choice.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a large-scale analysis of the EMNLP 2025 Responsible NLP Research Checklist. The authors release the CheckBox corpus of 3,214 main- and findings-track checklists (73,922 responses) plus a linking of checklist references to paper sections, and they use it to report three headline findings: 44.9% of NO justifications are brief or empty, 15.5% of main-track checklists contain parent-child logical contradictions, and 53% of authors answer NO to the mandatory risk question A2. They interpret these numbers as evidence that the checklist is often filled out as a box-ticking exercise, and they propose reforms such as enforcing a minimum word count and disabling child questions when a parent is NO.","tokens_in":13925,"tokens_out":6169,"duration_ms":58838,"significance":"The dataset is a potentially valuable resource: it is, to my knowledge, the first public corpus of ACL checklist responses with justifications and section links, and the paper addresses a question of direct policy relevance to ACL venues. The pipeline description is transparent, the validation sample reports 99.9% item-level accuracy, and the authors explicitly acknowledge limitations. However, the validity of the headline 44.9% figure rests on an unvalidated word-count proxy that the Limitations section concedes can misclassify adequate justifications, and the abstract and body disagree on the contradiction rate (6% vs 15.5%). Because these two numbers are the quantitative pillars of the central 'box-ticking' claim, the conclusion is currently stronger than the evidence supports. The contribution is therefore significant and publishable in principle, but only after the measurement validity and reporting consistency issues are addressed.","major_comments":[{"comment":"The headline '44.9% of NO justifications are poor or bad-faith' depends entirely on the unvalidated cutoff in §5.3: justifications are 'adequate (>10 words), brief (1-10), empty (0)', and 'Failure=Empty+Brief'. The Limitations explicitly concede that 'a ten word justification could be perfectly adequate but our software would classify it as inadequate', yet the Abstract and Conclusion still use 'poor or bad-faith' without hedging. Since the two required components of a NO justification in §3.1 are content-based (why the practice was omitted, and the missing information), the authors should either manually annotate a sample of brief justifications and report the fraction that actually satisfy §3.1, or replace the adequacy claim with a purely descriptive 'brief or empty' claim. Without this validation, the box-ticking conclusion is not supported. I also note that 213 of the 2,074 justifications are empty (about 10.3%), so even a stricter empty-only measure would still show a problem, but it would not justify the 'nearly half' framing.","section":"§5.3, Table 3, Limitations"},{"comment":"The Abstract reports that '6% of all checklists contained logical contradictions between parent and child responses', while §5.4 and Table 4 report 281 of 1,809 main-track papers (15.5%), and the Findings comparison reports 18.1%. These numbers are mutually inconsistent, and the paper nowhere explains the discrepancy or defines the denominator separately for the abstract figure. The authors must correct the abstract or the body and state precisely whether each rate is per paper, per checklist, or per question item.","section":"Abstract and §5.4/Table 4"},{"comment":"The claim that '53% of authors dismiss potential risks or social impacts of their work' over-interprets a NO answer to A2. A2 asks whether the authors 'discussed any potential risks of your work', so a NO response means only that the paper/checklist does not report a risk discussion; it does not establish that the authors dismissed or ignored risks. Phrases such as 'violation of the ACM Code of Ethics' and 'surface compliance' are therefore stronger than the data support. The authors should rephrase these statements to say that a majority of papers do not report risk discussion, and separate that observation from a substantive judgment about whether risks exist.","section":"§5.1, §6, Abstract"}],"minor_comments":[{"comment":"The text says 'all 3,124 papers and checklists are included', but Figure 1 and the pipeline description state 3,214 (1,809 main + 1,405 findings); the number 3,124 appears to be a typo.","section":"§4.2"},{"comment":"'Main and Finding tracks' should be 'Main and Findings tracks' for consistency with the body, which uses 'Findings track'.","section":"Abstract and throughout"},{"comment":"The phrase 'risks of appliances' in the recommendations should presumably read 'risks of applications'.","section":"Abstract"},{"comment":"The 'problematic' threshold of Score >4 for citation concentration is introduced without a justification or sensitivity analysis; since the text draws conclusions about problematic referencing from this cutoff, a brief rationale or robustness check would strengthen the claim.","section":"§5.2, Figure 3"},{"comment":"Comparative statements such as 'significantly less adequate' are not accompanied by confidence intervals or significance tests; adding them would make the findings-track comparison more rigorous.","section":"§5.5, Figures 9-10"},{"comment":"The conditional rules for when justifications are required are formatted as a bullet list that is hard to parse; a small truth-table or cleaner notation would improve readability.","section":"§3.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript makes a useful empirical contribution and the dataset release is a genuine asset, but the two headline numbers have validity or reporting problems: the 44.9% figure rests on an admitted word-count proxy, and the abstract's 6% contradicts the body's 15.5% for the same logical-contradiction measure. Both issues affect the central 'box-ticking' thesis, so they should be resolved before publication. The paper would be suitable for the journal once these are fixed, ideally by adding a manual annotation study for NO justifications or by substantially softening the causal and motivational language."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper gives us the first large-scale look at actual checklist responses and justifications from EMNLP 2025, and the authors released the dataset. That alone is worth something. The section-linking pipeline is new, and the descriptive analysis is straightforward but sensible. The finding that ethics questions get more N/A and less justification than reproducibility questions is likely robust and matches what many of us suspected. The parent-child contradiction analysis is also a nice idea; 15.5% of main-track checklists containing logical contradictions is a concrete, actionable signal for the ACL to fix the form.\n\nThe soft spots are real but not fatal. The headline 44.9% \"poor or bad-faith\" rate rests on a word-count threshold, and the Limitations section explicitly concedes that a ten-word justification can be perfectly adequate. That concession should have tempered the abstract and conclusion, which still use \"poor or bad-faith\" and \"bare minimum\" language without hedging. Also, the abstract reports 6% contradictions while Section 5.4 reports 15.5%; that is a clear internal inconsistency that should have been caught before submission. The 53% \"dismissing risks\" claim is similarly an interpretation of NO answers rather than a content analysis; some NO answers are legitimate. These issues mean the strong language in the abstract is ahead of the evidence, but they do not undermine the core descriptive findings.\n\nThe paper deserves peer review. The dataset release and first analysis of justification quality are genuine contributions, and the central observation that authors often fill out checklists with minimal effort is credible. A serious referee could help the authors either validate the word-count proxy with manual annotation or soften the claims, and fix the contradiction-rate discrepancy. I would cite the dataset in my own work, and I would bring this to a reading group for the discussion it would provoke.","headline":"First large-scale look at EMNLP 2025 checklist justifications with a useful data release, but the headline 44.9% failure rate overreaches its word-count proxy and the abstract contradicts the body on the contradiction rate.","tokens_in":14441,"tokens_out":1272,"would_cite":true,"duration_ms":12719,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that the EMNLP 2025 Responsible NLP Checklist is mostly box-ticking: 44.9% of 'no' justifications are brief or empty, 15.5% of main-track checklists contain parent-child contradictions, and 53% of authors dismiss…","keywords":["responsible NLP","checklist analysis","EMNLP 2025","ethics","reproducibility","societal impact","surface compliance","box-ticking"],"falsifier":"Manually annotate a random sample of, say, 200 'brief' NO justifications from the released corpus and ask whether each actually explains the missing practice; if most short answers are judged adequate, the headline failure rate collapses, though the contradiction and risk-dismissal findings would remain.","tokens_in":13504,"feed_emoji":"☑️","tokens_out":9122,"duration_ms":74095,"temperature":0.7,"pith_summary":"EMNLP 2025 made the ACL Responsible NLP Checklist public for every accepted paper, and this study asks whether a required form can actually push researchers toward transparent, ethical, and socially aware practice. Analyzing 73,922 responses from 3,214 main- and findings-track checklists, the paper reports three patterns: 44.9% of 'no' justifications are brief or empty, 15.5% of main-track checklists contain logical contradictions between a parent question and its child questions, and 53% of authors answer 'no' to the mandatory risks question A2. The authors read these patterns as surface compliance, with ethics and societal-impact items answered minimally and effectively isolated from the paper's body. If the analysis is right, the checklist is not serving its stated purpose, and the paper's proposed fixes—disabling child questions when a parent is 'no', banning empty justifications, and adding reviewer scrutiny to risk answers—offer a concrete path for reform.","feed_headline":"44.9% of checklist 'no' answers are brief or empty","feed_subtitle":"Audit of 73,922 EMNLP 2025 checklist answers finds 53% dismiss risks and 15.5% have contradictions","key_machinery":"The central object is the ARR Responsible NLP Research Checklist: 23 questions grouped into five parent items (A mandatory, B artifacts, C experiments, D human annotators, E AI assistance), each with child subquestions answered YES, NO, or N/A. The analysis exploits three structural features: the parent-child gating rule (if a parent is NO or N/A, every child must also be NO or N/A), which defines logical contradictions; the requirement that NO subquestion answers carry a written justification, which the authors grade by word count (adequate >10, brief 1–10, empty 0); and the requirement that YES answers cite paper sections, from which the authors build a section-reference linking dataset and a concentration score (total section references divided by unique sections cited) to detect vague or repetitive pointing. The pipeline that extracts these from 3,214 PDFs uses a document parser plus a constrained LLM pass, with a reported 99.9% item-level accuracy on a 71-paper validation sample and 90.9% section-text retrieval.","core_discovery":"The central claim is that the EMNLP 2025 Responsible NLP Checklist is being treated as a bureaucratic form rather than a reflective instrument. The evidence is quantitative: across 41,607 main-track responses, the paper finds that questions about limitations and risks (A1, A2) and human-annotator ethics (D) receive far fewer 'yes' answers than reproducibility questions (B, C); 44.9% of 'no' justifications fall into 'brief' (1–10 words) or empty categories; and 281 main-track papers (15.5%) contain at least one case where a parent question is answered 'no' but a child question is answered 'yes', which the checklist's own gating rules make logically impossible. The paper also finds that 53% of authors say their work has no potential risks, and that the same patterns appear in the findings track, with 18.1% of checklists containing contradictions. The authors interpret this not as individual carelessness but as a design problem: the form gives no feedback, does not enforce its own parent-child logic, and does not require justifications to meet any standard.","pith_inferences":["A conservative reading of the data is that the 44.9% figure is an upper bound on bad faith: word count cannot distinguish a terse but honest explanation from a dismissive one, and the paper's own limitation acknowledges this.","The high N/A rates on data-ethics and human-annotator items (B4, D) could mean researchers genuinely avoid private or human data, or it could mean N/A is used as a low-effort escape hatch; the paper does not separate these, and a follow-up could compare N/A rates with paper content.","The section-reference dataset opens a natural next test: automatically check whether the cited section actually answers the checklist question, which would expose 'lazy section pointer' compliance more directly than word count.","Cross-venue replication (for example, at other ACL conferences with public checklists) would clarify whether these patterns are specific to EMNLP 2025 or characteristic of the checklist format itself."],"forward_implications":["If the checklist is made self-consistent by disabling child questions when a parent is answered 'no', the detected 15.5% contradiction rate should drop to near zero and authors would be forced to reconsider their parent answer.","Enforcing a minimum justification length would push the 44.9% failure rate down, though it would not by itself guarantee meaningful engagement.","Giving A2 (risks) explicit reviewer scrutiny would likely raise the 53% 'no risk' rate only if reviewers can request revisions; the paper shows the current form offers no such check.","The similarity between main and findings tracks suggests the checklist does not discriminate paper quality, so acceptance-track comparisons that rely on checklist compliance would need another measure."],"supporting_citations":[{"why":"Defines the 23-question checklist and its parent-child gating rules that the analysis uses to detect contradictions.","marker":"(ACL Rolling Review, 2024)"},{"why":"The NeurIPS 2021 checklist that the ARR checklist is based on, establishing the risk and broader-impact questions.","marker":"(Beygelzimer et al., 2021)"},{"why":"Introduced the reproducibility checklist format whose reporting questions feed the experiment (C) section.","marker":"(Dodge et al., 2019)"},{"why":"The NeurIPS 2019 reproducibility program whose checklist design underpins the ACL checklist's transparency goals.","marker":"(Pineau et al., 2021)"},{"why":"Responsible Data Use Checklist that shaped the artifacts/data-ethics (B) items, including the B4 anonymisation question.","marker":"(Rogers et al., 2021)"},{"why":"Prior quantitative checklist analysis showing gamification and acceptance-rate links; the baseline this study extends to justifications.","marker":"(Magnusson et al., 2023)"},{"why":"ConfReady dataset of ACL 2023 checklist responses; this paper contrasts by analyzing justification content rather than token statistics.","marker":"(Galarnyk et al., 2025)"},{"why":"Supports the claim that NLP systems commonly carry dual-use or data-ethics risks, underlining why A2 and B4 compliance matters.","marker":"(Leins et al., 2020)"}],"fun_headline_variants":["44.9% of EMNLP 'no' justifications are brief or empty","15.5% of EMNLP 2025 checklists have logical contradictions","53% of EMNLP authors dismiss risks in their work","Audit of 73,922 checklist answers reveals box-ticking","EMNLP 2025 responsible checklist: more form than function"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 44.9% figure depends on treating any justification of 10 or fewer words as inadequate, an assumption the paper itself flags could mislabel a perfectly good short answer.","fun_headline_variants_meta":{"raw":{"variants":["44.9% of EMNLP 'no' justifications are brief or empty","15.5% of EMNLP 2025 checklists have logical contradictions","53% of EMNLP authors dismiss risks in their work","Audit of 73,922 checklist answers reveals box-ticking","EMNLP 2025 responsible checklist: more form than function"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000297,"raw_usage":{"total_tokens":1790,"prompt_tokens":1084,"completion_tokens":706,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":700,"completion_tokens_details":{"reasoning_tokens":609}},"tokens_in":700,"tokens_out":706,"duration_ms":6478,"temperature":1.0,"reasoning_tokens":609,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T20:13:08.760940+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Manually annotate a random sample of, say, 200 'brief' NO justifications from the released corpus and ask whether each actually explains the missing practice; if most short answers are judged adequate, the headline failure rate collapses, though the contradiction and risk-dismissal findings would remain.","supporting_citations":[],"review_version":1}