{"id":"471a9a03-7aa2-43c6-a0af-1873a44a1667","arxiv_id":"2606.07936","paper_version":2,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic analysis of 284 manually reviewed papers plus 1.8k+ others from 2023-2025 reveals under-reporting of human evaluation study design details, creating ambiguity in what was measured and how.","lead":"The paper conducts a large-scale review of human evaluation protocols in recent NLP papers on long-form text generation and finds widespread under-reporting of key details like rater instructions and demographics. This matters because clearer reporting standards could make human evaluations more reliable and comparable across studies.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"The authors' 20 criteria may omit or over-weight aspects critical to reproducibility, making the 'widespread under-reporting of important aspects' claim dependent on an unvalidated definition of importance.","rationale":"The reader's weakest_assumption directly identifies the load-bearing point for the strongest_claim. The methodology (manual review of 284 papers plus shared artifacts) is otherwise transparent and low-risk for correctness. No internal inconsistency or stronger methodological flaw (e.g., sampling bias or statistical error) is evident from the provided description that would override the criteria-completeness issue.","tokens_in":1681,"tokens_out":431,"duration_ms":19703,"concrete_test":"Run a short expert elicitation: present the paper's 20 criteria to 15-20 NLP researchers who have published human evaluations of long-form generation; ask each to (a) rate each criterion's necessity on a 1-5 scale and (b) list up to 5 additional criteria they consider essential for reproducibility. Compute overlap and average necessity scores; if >25% of expert additions receive mean necessity ≥4 and are absent from the original list, the completeness assumption requires qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central finding rests on applying a fixed set of 20 reportable criteria (defined by the authors) to 284 manually reviewed + 1.8k LLM-analyzed papers. The claim that missing details create 'ambiguity about what was measured and how, who contributed judgments, and how judgments should be interpreted' is only as strong as the premise that these 20 items are the right, complete, and non-redundant set. No external validation (expert survey, comparison to existing reporting guidelines such as those from ACL or prior meta-studies on evaluation protocols) is described in the abstract or reader's summary. If the list systematically excludes well-reported but still consequential details (e.g., exact payment structures, rater training materials, or statistical power calculations) or includes low-impact items, both the prevalence numbers and the downstream recommendations become sensitive to author framing rather than community consensus.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper conducts a large-scale observational analysis of human evaluation protocols for long-form text generation in *CL conference papers (2023–2025). It performs a full manual review of 284 papers plus LLM-assisted analysis of an additional 1.8k+ papers against a fixed set of 20 reportable criteria for reproducibility. The central claim is that widespread under-reporting of study-design details creates ambiguity about what was measured, who provided judgments, and how results should be interpreted; the authors provide recommendations and release code plus an annotated dataset.","tokens_in":1878,"tokens_out":586,"duration_ms":16779,"significance":"If the central observational claim holds, the work is significant because human evaluation remains a primary method for assessing long-form generation quality, and documented under-reporting directly affects reproducibility and interpretability in the field. The explicit release of analysis code and the annotated dataset is a clear strength that supports verification and follow-up studies. The findings could usefully inform community guidelines, provided the 20 criteria are shown to be well-aligned with existing reporting standards.","major_comments":[{"comment":"Section defining the 20 criteria: the manuscript presents these criteria as capturing 'important aspects' necessary for reproducibility and interpretability, yet provides no external validation (e.g., expert survey, comparison against ACL or prior meta-study reporting guidelines, or inter-rater agreement on criterion importance). Because the prevalence statistics and the downstream claim of 'widespread under-reporting of important aspects' rest directly on this author-defined set, the absence of such validation makes the quantitative conclusions sensitive to the particular framing chosen.","section":"Section defining the 20 criteria"},{"comment":"LLM-assisted analysis section (1.8k+ papers): the extension from the 284 manually reviewed papers to the larger corpus is load-bearing for the 'widespread' claim, but the manuscript does not report prompt details, few-shot examples, or measured agreement/error rates between the LLM outputs and the manual annotations. Without these, systematic biases in the automated labeling could materially affect the reported under-reporting rates.","section":"LLM-assisted analysis section"}],"minor_comments":[{"comment":"Table 1 (or equivalent summary table of criteria): the mapping from each criterion to the specific ambiguity it addresses (measurement, contributors, or interpretation) could be made more explicit to help readers trace how missing items produce the claimed ambiguities.","section":"Table 1"},{"comment":"The GitHub link is provided, but the README should include a clear description of how the 284 manual annotations were performed (annotator background, resolution process) to strengthen reproducibility of the core dataset.","section":"Data and code release"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback. We address each major comment below and outline planned revisions to improve the manuscript.","responses":[{"response":"We agree that additional justification for the criteria would strengthen the work. The 20 criteria were synthesized from recurring elements in prior NLP literature on human evaluation reproducibility and reporting standards. In revision we will add a new subsection explicitly mapping each criterion to relevant ACL guidelines and earlier meta-studies, together with a brief rationale for inclusion. While we did not conduct a new expert survey, this explicit alignment will reduce sensitivity to the chosen framing and make the prevalence claims more robust.","revision_made":"partial","referee_comment":"[Section defining the 20 criteria] Section defining the 20 criteria: the manuscript presents these criteria as capturing 'important aspects' necessary for reproducibility and interpretability, yet provides no external validation (e.g., expert survey, comparison against ACL or prior meta-study reporting guidelines, or inter-rater agreement on criterion importance). Because the prevalence statistics and the downstream claim of 'widespread under-reporting of important aspects' rest directly on this author-defined set, the absence of such validation makes the quantitative conclusions sensitive to the particular framing chosen."},{"response":"We concur that full transparency on the LLM-assisted labeling is required. The original submission omitted these details for brevity. The revised manuscript will include the complete prompts, few-shot examples, and a dedicated error-analysis subsection reporting agreement rates (and disagreement categories) between the LLM and the manual annotations on a held-out validation set. This addition will allow readers to evaluate potential biases directly.","revision_made":"yes","referee_comment":"[LLM-assisted analysis section] LLM-assisted analysis section (1.8k+ papers): the extension from the 284 manually reviewed papers to the larger corpus is load-bearing for the 'widespread' claim, but the manuscript does not report prompt details, few-shot examples, or measured agreement/error rates between the LLM outputs and the manual annotations. Without these, systematic biases in the automated labeling could materially affect the reported under-reporting rates."}],"tokens_in":1433,"tokens_out":453,"duration_ms":18965,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The core takeaway is that this work measures reporting gaps across hundreds of recent *CL papers on long-form text generation and finds many details missing. They manually checked 284 papers and used LLM assistance on another 1.8k, applying a fixed list of 20 criteria they defined for reproducibility. That scale is new; earlier papers on evaluation practices were smaller or less focused on long-form tasks.\n\nWhat works is the release of the annotated dataset and code. That lets others check or extend the analysis. The observational claim lines up with what many of us see in reviews: papers often skip who the raters were, how instructions were written, or how scores should be read. The recommendations at the end are practical and tied directly to the gaps they counted.\n\nThe soft spot is the criteria themselves. The authors chose the 20 items without showing external validation against existing guidelines or a survey of what practitioners actually need for reproducibility. If some items are low-impact or if key details like payment structures or power calculations are left out, the prevalence numbers and the call for better norms become sensitive to that choice. The LLM-assisted part on the larger set also needs the full paper to confirm error rates were low enough not to move the main findings.\n\nThis is for anyone running or reviewing human evaluations in generation work. It supplies concrete data on current norms rather than just complaints. The paper deserves peer review because the artifacts are there and the question is real, even if the criteria list will need discussion in revision.","headline":"The paper gives the first big numbers on under-reporting in human eval protocols for long-form generation, but the 20 criteria are unvalidated so the strength of the 'widespread' claim is still open.","tokens_in":2390,"tokens_out":391,"would_cite":true,"duration_ms":8491,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Human evaluation protocols for long-form text generation in recent NLP conference papers are often incompletely reported, creating ambiguity about what was measured and by whom.","keywords":["human evaluation","reproducibility","long-form text generation","NLP conferences","reporting practices","evaluation protocols","under-reporting","study design"],"falsifier":"A re-analysis of the same papers that applies a different or expanded list of criteria and finds high rates of complete reporting on the missing items.","tokens_in":2590,"feed_emoji":"📊","tokens_out":593,"duration_ms":13276,"temperature":0.7,"pith_summary":"The paper examines reporting practices for human evaluations in papers on long-form text generation from *CL conferences between 2023 and 2025. It applies a defined set of 20 criteria for reproducibility to a manual review of 284 papers plus LLM-assisted analysis of more than 1800 additional papers. The analysis shows frequent omission of details on study design, participant information, judgment processes, and interpretation guidelines. This pattern leaves readers uncertain about the reliability of the evaluations and how to compare results across papers. The authors propose concrete steps to improve documentation in future work.","feed_headline":"NLP papers omit key human eval protocol details","feed_subtitle":"Review of 2023-2025 *CL papers finds missing design and participant information creates ambiguity in results.","key_machinery":"A set of 20 reportable criteria related to reproducibility of human evaluation studies, applied to check what details papers include about design, participants, and judgment processes.","core_discovery":"A systematic review of human evaluation protocols in *CL publications reveals widespread under-reporting of important aspects of study design, who contributed judgments, and how judgments should be interpreted, which produces ambiguity about what was actually measured and how the results should be understood.","pith_inferences":["The same under-reporting pattern likely appears in evaluations of short-form or other generation tasks outside the long-form focus.","Conferences could reduce ambiguity by adding checklist items based on the 20 criteria during submission.","Greater transparency might shift research incentives toward more careful study design rather than just reporting results."],"forward_implications":["Adopting the 20 criteria would make it easier to interpret and compare human evaluation results across different papers.","Papers would need to document participant recruitment, training, and agreement measures more consistently.","Ambiguity in current evaluations would decrease if journals and conferences required explicit reporting on these points.","Future comparisons of generation systems could rest on clearer evidence of evaluation quality."],"fun_headline_variants":["*CL papers underreport human eval protocols","*CL review shows missing judgment details","Human eval designs often lack transparency","Underreporting creates eval result ambiguity"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The authors' chosen set of 20 criteria is enough to determine whether a human evaluation study is reproducible and interpretable.","fun_headline_variants_meta":{"raw":{"variants":["*CL papers underreport human eval protocols","*CL review shows missing judgment details","Human eval designs often lack transparency","Underreporting creates eval result ambiguity"]},"model":"grok-4.3","cost_usd":0.006087,"raw_usage":{"total_tokens":2768,"prompt_tokens":613,"num_sources_used":0,"completion_tokens":46,"cost_in_usd_ticks":60865500,"prompt_tokens_details":{"text_tokens":613,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2109,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":613,"tokens_out":46,"duration_ms":11707,"temperature":1.0,"reasoning_tokens":2109,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-27T20:18:32.801057+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A re-analysis of the same papers that applies a different or expanded list of criteria and finds high rates of complete reporting on the missing items.","supporting_citations":[],"review_version":1}