{"id":"d57f0006-341a-41b9-a2e7-7c94ebc62080","arxiv_id":"2505.17043","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"QRA++ scores the degree of reproducibility of NLP results continuously at system, quality-criterion, and study levels, and grounds expectations in experiment-property similarity.","lead":"This paper introduces QRA++, a framework that turns 'was the result reproduced?' into continuous reproducibility scores for NLP experiments, using statistics like coefficient of variation, correlation, and agreement measures. It also connects how similar two experiment setups are to how reproducible the outcomes should be, and illustrates the framework on three sets of reproduction studies.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 6, the paper's key evidence that reproducibility tracks experiment similarity, compares subsets matched on random seed—but seed is not in the Section 4 property set, so the example undercuts the framework's own grounding.","rationale":"The reader's weakest assumption was the minimality and objective grounding of the Section 4 property set. My concern is a concrete, internal instantiation of that assumption: the paper's own key demonstration, Table 6, uses random seed as a matching criterion even though seed is absent from the property set, and it does not control for other Section 4 properties that Section 7.3 acknowledges differ. This means the central claim that QRA++ 'grounds expectations about degree of reproducibility in degree of similarity between experiments' is not actually demonstrated by the paper's main evidence, and the property set is shown to be incomplete by the paper's own choice of variables. This is more load-bearing than the statistical issues (missing CIs, significance tests) because even perfect statistics would not fix a mismatch between the framework's similarity concept and the variables used to operationalize it. I still do not think the framework should be rejected: the metrology-inspired measures and reporting template are coherent and potentially useful, and the conditional path—adding seed or acknowledging the property set's incompleteness, providing uncertainty estimates, and softening the 'clear evidence' wording—remains appropriate. The reader's CONDITIONAL verdict therefore stands unchanged.","tokens_in":16490,"tokens_out":5901,"duration_ms":63216,"concrete_test":"Reconstruct from the eight cited REPROLANG papers a full Section 4 property vector for each experiment, including test dataset, metric implementation, and run-time environment but excluding random seed. Then compute mean CV*, W, and P for experiment pairs grouped by number of differing Section 4 properties (0, 1, 2, 3+), and repeat the grouping after adding seed as an extra variable. If the monotonic worsening with property differences seen in Table 6 disappears or reverses when seed is excluded, the claimed dependence on the framework's property set is not supported. Additionally, run a paired permutation test over the 11 systems comparing each system's CV* in subset (a) versus subset (b) of Table 6; if the p-value exceeds 0.05, the 'clear evidence' is not statistically established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that the Section 4 experiment properties form a minimal set such that matching properties makes matching outcomes expected. Section 4 itself concedes there is 'not a distinction for which there is an objective basis.' The problem is sharper than that: the paper's only direct evidence for property-based expectations, Table 6, compares a subset of four REPROLANG experiments that 'share the same test dataset and random seed' against a subset with 'test samples and seeds that differ.' Random seed is not among the properties in Table 1, nor in the HEDS 3.0 questions in Appendix A. If seed is a reproducibility-relevant condition, the property set is incomplete—two experiments can match on every listed property yet differ in seed, so equal property values do not imply equal expected outcomes, contradicting the framework's grounding. If seed is not a QRA++ property, then Table 6 is not an application of the framework and cannot support the abstract's 'clear evidence' claim. The comparison is further confounded because Section 7.3 says metric implementation and run-time environment also differ between subsets of the eight experiments, so even the 'same dataset and seed' subset is not homogeneous with respect to the Section 4 set. No confidence intervals or significance tests are reported for Table 6; with n=4 experiments per subset, differences such as P=0.706 versus P=0.558 are within plausible sampling variability. The framework may still be useful, but its headline empirical demonstration is not internally consistent with its own similarity mechanism.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents QRA++, a framework for quantitative reproducibility assessment in NLP, built on metrology concepts. It defines a set of experiment properties (measurement conditions) in Table 1, four types of experimental results (Type I–IV), corresponding reproducibility measures (CV*, Pearson/Spearman/Kendall correlations, Fleiss's κ/Krippendorff's α, and a new P measure for ordinal findings), and three assessment levels (system, quality criterion, study). The framework is illustrated on three example sets of comparable experiments from ReproNLP 2024 and REPROLANG 2020. The paper's central claim, stated in the abstract, is that applying QRA++ reveals 'clear evidence' that degree of reproducibility depends on similarity of experiment properties, system type, and evaluation method.","tokens_in":16624,"tokens_out":6777,"duration_ms":63088,"significance":"The framework is a coherent synthesis of standard statistical measures with a metrological grounding, and the explicit specification of experiment properties is a useful step toward comparable reproducibility assessments. The paper provides careful definitions, including the small-sample-corrected CV* and the P measure. However, the empirical support for the central claim is currently weak: the key comparison in Table 6 is confounded and uses a variable (random seed) not in the property set, the property set is admitted to lack an objective basis, and no inferential statistics are provided. If the property set can be validated and the examples restricted to the framework's own variables, the approach would be a valuable contribution; as it stands, the abstract overstates the evidence.","major_comments":[{"comment":"The comparison in Table 6 groups experiments by 'same test dataset and random seed' versus 'different datasets and seeds,' but random seed is not among the experiment properties in Table 1 nor in the HEDS 3.0 questions in Appendix A. Therefore, Table 6 does not instantiate the QRA++ property-based expectation mechanism: either the property set is incomplete (in which case the grounding in Section 4 is not minimal), or the comparison is made on a variable outside the framework (in which case it cannot support the abstract's claim of 'clear evidence' about property similarity). Furthermore, Section 7.3 states that metric implementation and run-time environment also differ between the two subsets, so the 'same dataset and seed' subset is not homogeneous with respect to the Section 4 properties; the observed differences in Table 6 are confounded. No confidence intervals or significance tests are reported; with only four experiments per subset, differences such as P=0.706 versus P=0.558 are within plausible sampling variability.","section":"§7.3, Table 6"},{"comment":"The paper explicitly concedes that the minimal set of experiment properties is 'not a distinction for which there is an objective basis' and calls it 'a way of controlling the strictness of the reproducibility assessment.' This is a load-bearing limitation for the framework's central claim (abstract item iii) that QRA++ 'grounds expectations about degree of reproducibility in degree of similarity between experiments.' If the property set is arbitrary, those expectations are not externally valid, and the framework can only describe outcome differences, not explain them. The authors should either validate the property set on a larger body of reproduction studies or substantially weaken this claim in the abstract and conclusion.","section":"§4"},{"comment":"The abstract and conclusion claim that applying QRA++ reveals 'clear evidence' about factors affecting reproducibility. The evidence base is limited to three illustrative example sets, with no hypothesis tests, confidence intervals, or correction for multiple comparisons. For instance, Table 6 compares four experiments per group; with n=4, the differences in mean CV* (8.7 vs 14.9) and P (0.706 vs 0.558) are not shown to be statistically reliable. The paper should reframe these as illustrative findings rather than 'clear evidence,' or provide appropriate statistical support.","section":"Abstract and §9"}],"minor_comments":[{"comment":"The text says 'Table 3 reports QRA++ results based on two comparable experiments,' but the table referred to is Table 4; it also says 'the experiment properties ... were exactly the same in all three experiments' when there are only two experiments.","section":"§7.2"},{"comment":"There is a typo, 'beccause,' and the text says 'Cohen's κ for n = 2, and Krippendorff's α for n>2,' but Table 2 appears to list both κ and α for both cases; please clarify which measure applies when.","section":"§6.3"},{"comment":"The note under the table says 'P = proportion of differences between pairs of systems that have the same sign (see Section 1),' but the measure is defined in Section 6.4, not Section 1.","section":"Table 2"},{"comment":"The thresholds for 'good reproducibility' (below about 12 for human evaluations, below 1 for metric-based evaluations) are stated without a citation or derivation; please provide a source or describe how these values were determined.","section":"§6.1"},{"comment":"The formula for P uses the sign function; please clarify how zero differences are treated (e.g., if one experiment yields exactly equal scores for a pair of systems and another does not).","section":"Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript has a serious internal inconsistency in Table 6, and I recommend that the editor require a revision that either adds random seed to the property set with justification or re-analyzes the examples using only the declared properties. The paper's contribution as a descriptive framework is otherwise sound, but the abstract's 'clear evidence' claim should be moderated."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear [Colleague],\n\nQRA++ is a genuine attempt to give NLP a common yardstick for reproducibility, and the core framework is worth engaging with. The four result types, the three assessment levels, and the reporting template are a sensible synthesis; the small-sample-corrected CV* and the new P measure for pairwise rank agreement are reasonable choices. If you work on reproduction studies or shared tasks like ReproNLP, this gives you a concrete protocol to apply. Credit where due: the paper is clearly written, the metrology framing is used seriously, and the author explicitly flags the absence of an objective basis for the property set (Section 4). That honesty is a point in its favor.\n\nThe soft spots are mostly in the empirical demonstration, and one is load-bearing. The abstract claims 'clear evidence' that reproducibility depends on experiment similarity, but the examples are three convenience sets with no confidence intervals, no significance tests, and no raw score pairs provided. That alone would justify a conditional verdict. The sharper problem is Table 6. The paper splits the eight REPROLANG experiments into groups sharing 'same test dataset and random seed' versus different datasets and seeds, yet random seed is not among the Section 4 experiment properties, and the text in 7.3 says metric implementation and run-time environment also differ between subsets. So the comparison does not isolate the framework's own similarity variables. If seed is a variable that matters, the property set is incomplete; if it is not, Table 6 is not an application of QRA++. Either way, it cannot be the basis for the abstract's claim.\n\nThere are also small internal inconsistencies: Section 7.2 refers to 'three experiments' where there are two, and Table 4 is mislabelled as Table 3 in the text; Section 6.3 says Cohen's kappa for n=2 while Table 2 lists Fleiss's kappa. These are fixable but should be cleaned up.\n\nWho should read this: anyone designing or adjudicating reproduction studies in NLP. The framework itself deserves referee time; the empirical claims need revision before it is published as is. I would send it to review, with a request that the author either make the Table 6 demonstration consistent with the property set or present it as a separate observation, and that the 'clear evidence' language be toned down to 'suggestive'. The framework is a net positive for the field.\n\nBest,\n[You]","headline":"A useful synthesis of reproducibility assessment for NLP, but the empirical demonstration is overclaimed and the Table 6 example is not consistent with the framework's own property set.","tokens_in":17262,"tokens_out":3258,"would_cite":true,"duration_ms":28554,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"QRA++ measures reproducibility as a continuous degree, comparable across NLP studies, grounded in how similar experiments are.","keywords":["reproducibility","natural language processing","evaluation","metrology","measurement precision","coefficient of variation","experiment properties","degree of reproducibility"],"falsifier":"Compile many pairs of studies that share all QRA++ experiment properties, including the same system outputs and metric implementation, and compare their CV* and P values with pairs that differ in properties; if matched pairs show no better reproducibility than unmatched ones, the claimed grounding of reproducibility in experiment similarity fails.","tokens_in":16160,"feed_emoji":"📏","tokens_out":7551,"duration_ms":68974,"temperature":0.7,"pith_summary":"The paper argues that reproducibility in NLP should be assessed as a matter of degree rather than a binary success/failure, and presents QRA++, a framework that produces continuous-valued reproducibility scores at three levels of granularity: individual systems, quality criteria, and whole studies. The framework grounds expectations about reproducibility in the degree of similarity between experiments, using an explicit set of experiment properties as measurement conditions. Because the measures are unitless and defined uniformly across studies, QRA++ makes reproducibility assessments directly comparable across different papers and shared tasks. Applied to three sets of comparable experiments, QRA++ shows that reproducibility is better when experiments share the same test data and seeds, and that system type and evaluation method also drive differences. If adopted, the framework would let the field move from narrative judgments to quantitative, diagnostic reproducibility reporting.","feed_headline":"NLP reproducibility gets a continuous grade, not a yes/no","feed_subtitle":"QRA++ scores system, criterion, and study level from experiment-property similarity, so studies become comparable.","key_machinery":"The load-bearing object is the equation that identifies reproducibility with the precision of a set of measured values: $R(M_1,\\dots,M_n) := \\mathrm{Precision}(v_1,\\dots,v_n)$, where each measurement $M_i : (m,O,t_i,C_i) \\mapsto v_i$ carries measurand, object, time, and measurement conditions. In NLP terms, objects are systems, measurands are evaluation measures, and conditions are experiment properties. QRA++ supplies a curated checklist of such properties (task, input and output type, test dataset, metric implementation, run-time environment, and several human-evaluation specifics) and treats it as controlling the strictness of the assessment: equal properties make any outcome difference a possible reproducibility failure, while differing properties make differences expected. On top of this, the framework pairs each of its four result types with scale-free measures of agreement or precision, so the same numbers can be read across studies.","core_discovery":"The paper's central claim is that reproducibility in NLP should be assessed quantitatively, as a degree, by treating evaluation as a measurement and reproducibility as measurement precision. Formally, for measurements $M_i$ that map a measurand $m$, object $O$, time $t_i$, and conditions $C_i$ to a quantity value $v_i$, reproducibility is $R(M_1,\\dots,M_n) := \\mathrm{Precision}(v_1,\\dots,v_n)$, with repeatability as the special case where all $C_i$ are identical. QRA++ instantiates this by fixing a set of experiment properties as the conditions, by distinguishing four types of results (single scores, sets of scores, categorical labels, and findings about system differences), and by assigning each a family of degree-of-reproducibility measures: $\\mathrm{CV}^{*}$ for single scores, Pearson and Spearman correlation plus Kendall's $\\tau$ and $W$ for score sets, Fleiss's $\\kappa$ and Krippendorff's $\\alpha$ for labels, and $P$, the proportion of identical pairwise system ranks, for findings. Three example applications show higher degrees of reproducibility when experiment properties are the same, and clear differences by system type and evaluation method.","pith_inferences":["A natural next step is to apply QRA++ across many existing reproduction studies to build a public corpus of quantitative reproducibility scores; the paper itself notes that such a body is currently lacking.","The same measurement-precision machinery could transfer to other fields that report evaluation metrics over shared tasks, such as information retrieval or vision-language evaluation, since the four result types are generic.","One testable extension is using QRA++ to monitor reproducibility over time within a fixed research community, checking whether improved reporting standards actually raise reproducibility scores.","Because the framework separates conditions (experiment properties) from outcomes, it could be combined with controlled perturbation studies that vary one property at a time, such as evaluator expertise or rating scale, to estimate its causal effect on reproducibility."],"forward_implications":["Reproducibility assessments from different studies become numerically comparable, because the chosen measures are unitless and defined uniformly.","A single reproduction attempt no longer decides success or failure; assessments over two or more comparable experiments yield a more stable degree of reproducibility.","Researchers can diagnose which experiment properties are associated with worse reproducibility (e.g., differing test datasets and seeds) and which system types are more fragile, as shown in the eight-experiment example.","The framework reveals that even when absolute scores reproduce only moderately, system rankings and findings can reproduce perfectly, so reproducibility conclusions depend on result type and level of aggregation.","Adding or removing properties from the experiment-property list changes the strictness of the assessment, giving researchers a dial for how demanding a reproducibility check should be."],"supporting_citations":[{"why":"Supplies the metrological framing, the coefficient-of-variation measure, and the treatment of repeatability as identical measurement conditions.","marker":"(Belz, 2022)"},{"why":"Provides the definitions of repeatability and reproducibility as measurement precision under specified conditions, which ground the whole framework.","marker":"(JCGM, 2012)"},{"why":"Provides the evaluation data sheet whose questions are the source of the experiment properties used as measurement conditions.","marker":"(Belz and Thomson, 2024b)"},{"why":"Reports the ReproNLP shared task example experiments and the observation that two re-runs of the same original can give very different reproducibility results.","marker":"(Belz and Thomson, 2024a)"},{"why":"Provides the REPROLANG shared task context and the eight-experiment set used in the third illustrative application.","marker":"(Branco et al., 2020)"},{"why":"Documents 513 pairs of scores from system comparisons, motivating the need for a quantitative and comparable assessment method.","marker":"(Belz et al., 2021)"},{"why":"Gives the small-sample correction used in the CV* measure.","marker":"(Sokal and Rohlf, 1971)"},{"why":"Gives the unbiased standard deviation correction used in the CV* measure.","marker":"(Rao, 1973)"}],"fun_headline_variants":["NLP reproducibility: from yes/no to a precise degree","Reproducibility as precision: QRA++ quantifies NLP studies","Continuous reproducibility scores for NLP experiments","QRA++: how similar experiments reproduce, measured"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework assumes that the chosen list of experiment properties is sufficient to determine when evaluation outcomes should match; the paper itself admits this list has no objective basis and serves as a way to set how strict the reproducibility assessment is.","fun_headline_variants_meta":{"raw":{"variants":["NLP reproducibility: from yes/no to a precise degree","Reproducibility as precision: QRA++ quantifies NLP studies","Continuous reproducibility scores for NLP experiments","QRA++: how similar experiments reproduce, measured"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000231,"raw_usage":{"total_tokens":1498,"prompt_tokens":969,"completion_tokens":529,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":585,"completion_tokens_details":{"reasoning_tokens":465}},"tokens_in":585,"tokens_out":529,"duration_ms":5056,"temperature":1.0,"reasoning_tokens":465,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:52:13.259456+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compile many pairs of studies that share all QRA++ experiment properties, including the same system outputs and metric implementation, and compare their CV* and P values with pairs that differ in properties; if matched pairs show no better reproducibility than unmatched ones, the claimed grounding of reproducibility in experiment similarity fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the definitions of repeatability and reproducibility as measurement precision under specified conditions, which ground the whole framework."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the REPROLANG shared task context and the eight-experiment set used in the third illustrative application."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents 513 pairs of scores from system comparisons, motivating the need for a quantitative and comparable assessment method."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the unbiased standard deviation correction used in the CV* measure."}],"review_version":1}