{"id":"2e7c442f-8a7a-4db6-8c0a-a6ce8823d199","arxiv_id":"2607.19231","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A preregistered replication shows that the monotonicity–label-agreement boundary found in ChaosNLI fails to generalize to unselected SNLI/MNLI dev sets, with all effects small and slightly positive.","lead":"A preregistered replication finds that a previously reported link between monotonicity and label disagreement in NLI disappears—and slightly reverses—when tested on unselected SNLI and MNLI development data. The result suggests that the earlier boundary was a consequence of ChaosNLI's disagreement-based selection, not a general property of human label variation.","discovery_kind":"replication","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Outcome-scale confound threatens the central claim: the unselected-population test uses a five-label ordinal that cannot see the entropy-level boundary, so 'selection shapes the boundary' is not established.","rationale":"The reader's CONDITIONAL verdict is well calibrated, and their flagged tagger non-differential-error assumption is a genuine measurement concern. However, the more load-bearing issue is the non-commensurable outcome scale: the earlier delta that defines the 'boundary' was measured on 100-label entropy in the 3/5-selected ChaosNLI subset, while the unselected-population test uses a five-label ordinal that has zero variance on that subset. A negative entropy effect can coexist with a positive ordinal effect in the same population, so the observed reversal does not identify selection as the cause. This means the title and Section 6 conclusion require either a commensurable-outcome measurement in the unselected population or explicitly weakened causal language. The paper's own Limitations correctly state the non-commensurability, but the conclusion is not adjusted accordingly. Therefore the reader's CONDITIONAL verdict stands, with the condition now including the need for a commensurable-outcome replication or softened causal claims.","tokens_in":11028,"tokens_out":9303,"duration_ms":111650,"concrete_test":"Annotate a random sample of unselected SNLI/MNLI dev items (about 500 per corpus, stratified by monotonicity tag) with 100 labels per item under the ChaosNLI protocol, then compute non-upward vs upward Cliff's delta on entropy. If delta is negative and comparable to -0.284, the boundary is a population-level entropy property and the paper's selection-conditional conclusion fails; if delta is near zero or positive, the conclusion is supported on a commensurable outcome.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline inference—that the ChaosNLI monotonicity boundary is conditional on low-agreement selection—rests on comparing the earlier study's Cliff's delta on 100-label entropy inside the 3/5-selected ChaosNLI subset with deltas on a five-label ordinal in unselected SNLI/MNLI dev. These are not merely different resolutions: the ordinal is constant by construction on the ChaosNLI overlap (Appendix B), so the entropy signal that produced delta=-0.284 lives entirely within a single ordinal level. The current analysis never measures the unselected-population analogue of the outcome on which the boundary was defined. The defense that coarsening attenuates toward zero (Section 4) is insufficient: the ordinal is not a coarsening of entropy but a different functional of the label distribution. A latent distribution in which non-upward items are slightly more likely to reach 5/5 (positive ordinal delta) yet have higher entropy conditional on non-unanimity (negative entropy delta) would reproduce both observations without any selection effect. Figure 2's level-wise shares cannot rule this out. Thus the data support only 'no monotonicity effect is detectable on the five-label ordinal in the unselected dev,' not 'the entropy boundary is conditional on selection.' The paper's Limitations explicitly state the non-commensurability, but the title and Section 6 conclusion assert the stronger causal claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a preregistered replication attempt of a previously observed association between non-upward monotonicity operators and lower label agreement in NLI. The earlier study (Choi, 2026) measured this on ChaosNLI, a disagreement-selected resource restricted to items with a 3/5 majority label, using 100-label entropy as the outcome (Cliff's delta -0.284). The present study applies the same frozen tagger to the unselected SNLI and MNLI development sets and uses a four-level ordinal agreement outcome built from the five original labels. Across seven contrasts, Cliff's delta is positive (non-upward items agree slightly more), all effects are below the preregistered SESOI of 0.10, and the only significant confirmatory contrast has the opposite sign to the registration. The paper concludes that the earlier negative boundary is plausibly a structure conditional on low-agreement selection rather than a population-level property, and recommends that HLV structure claims state their selection conditional. The manuscript includes a thorough reproducibility appendix, a registration-versus-result concordance, a misclassification simulation grid, and a manual audit.","tokens_in":11269,"tokens_out":4537,"duration_ms":64226,"significance":"If the conclusion is taken at face value, the paper is a useful cautionary result for the perspectivist NLI literature: a pronounced disagreement-related boundary measured inside a disagreement-selected resource can be absent—and even slightly reversed—in the unselected population on a coarser outcome. The preregistration, frozen predictor, SESOI, sensitivity analyses, simulation grid, and manual audit are genuine methodological strengths, and the reported effects are consistently below the SESOI across all robustness checks. The paper also honestly discloses its own limitations, including outcome-scale non-commensurability and single-author audit. However, the central interpretive claim outruns the evidence because the replication used a different outcome functional than the original study, and the measurement-validity analysis does not rule out agreement-correlated tagger error.","major_comments":[{"comment":"The central claim that the earlier monotonicity–agreement boundary is conditional on low-agreement selection is not directly supported by the reported analysis. The original δ = -0.284 was computed on 100-label entropy inside ChaosNLI, which by construction contains only items at the 3/5 ordinal level; Appendix B shows the four-level ordinal has zero variance on the entire ChaosNLI overlap. The entropy signal that produced the original boundary therefore lives entirely within a single ordinal level. A positive ordinal delta in unselected dev (Table 1) is compatible with a latent distribution in which non-upward items are slightly more likely to reach 5/5 unanimous agreement yet have higher entropy conditional on non-unanimity. The Section 4 argument that coarsening attenuates toward zero is not sufficient because the ordinal is not a coarsening of 100-label entropy; it is a different fun","section":"Section 6 and Appendix B"},{"comment":"The misclassification simulation in Appendix C applies random label flips at MED-anchored or fixed rates. This cannot detect tagger errors that are systematically correlated with the agreement ordinal, because the flip probabilities are assumed to be independent of the item's agreement level. If the tagger misclassifies non-upward items more often in low-agreement items—plausible if complex syntax drives both disagreement and tagger failure—then the observed positive reversal could be manufactured rather than a property of the unselected population. The Tier 3 audit reports overall agreement (0.875 four-class, κ=0.607) but does not stratify the 25 disagreements by the item's agreement level. The claim that measurement error 'shrinks the effects rather than manufacturing them' is true only under non-differential misclassification, which is not established. Please test or explicitly constr","section":"Section 5, Tier 2 and Tier 3"}],"minor_comments":[{"comment":"The gray reference row for the earlier entropy-scale estimate is plotted on the same axis as the ordinal-scale deltas even though the paper argues the scales are not numerically commensurable. A separate panel or clear visual separation would avoid implying interval comparability.","section":"Figure 1"},{"comment":"The text alternates between 'four-level ordinal' (Abstract, Section 3) and 'five-label ordinal' (Appendix B, Figure 2). Consistent terminology—e.g., 'four-level agreement ordinal derived from five labels'—would help.","section":"Appendix B and Figure 2 caption"},{"comment":"The statement that 'every quantity this paper reports is printed in the paper itself' is in slight tension with Section 5, where the layer counts are explicitly described as approximate and not regenerated by any script. This is not a substantive problem, but the wording could be softened.","section":"Section 7"},{"comment":"The term 'boundary' is used both for the entropy-scale effect inside ChaosNLI and for the ordinal-scale comparison in unselected dev. Defining the two senses early would reduce the risk of conflating the outcome scales.","section":"Title and Introduction"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper does a lot right. The preregistration, frozen tagger, SESOI, sensitivity analyses, and the Tier 2 simulation grid are careful and honestly reported. The empirical fact it nails down is real: on the four-level agreement ordinal in unselected SNLI/MNLI dev sets, non-upward monotonicity does not predict lower agreement; if anything, the association is faintly positive. That is a useful negative result for the HLV crowd, and the manual audit with the codebook strengthens confidence in the predictor measurement.\n\nBut the central claim in the title and Discussion goes beyond what the design can support. The original ChaosNLI effect was on 100-label entropy; this replication measures a five-label ordinal. Those are not the same construct, and the ordinal is constant by construction on the entire ChaosNLI overlap — so the entropy signal that produced the original delta lives entirely within a single ordinal level. The paper's defense that coarsening attenuates rather than reverses is not fully persuasive, because the ordinal is not a coarsening of entropy; it is a different functional of the label distribution. The stress-test scenario is the right worry: it is entirely possible that non-upward items are slightly more likely to reach 5/5 (positive ordinal delta) while also having higher entropy conditional on non-unanimity (negative entropy delta). That would reproduce both the original and the present results with no selection effect at all. The level-wise shares in Figure 2 cannot rule this out, and the paper's own Table 5 shows substantial entropy variation within the 3/5 level.\n\nThe paper does flag the non-commensurability in the Limitations section, and it restricts the comparison to 'sign and magnitude class.' But even that comparison is compromised, because the outcomes are different functionals. The honest conclusion is: no monotonicity effect is detectable on the five-label ordinal in unselected populations. That is not the same as showing that the entropy boundary is selection-dependent.\n\nA second, minor soft spot: the Tier 2 simulation assumes random misclassification, and the audit does not test whether tagger errors correlate with agreement level. That concern is secondary, but worth a sentence in revision.\n\nThis paper deserves serious peer review — the methodological transparency alone merits it — but the authors should be pushed to either soften the causal language or, better, add a commensurable analysis (e.g., entropy on subsampled labels or a direct test of the within-level association). As written, I would not cite it as evidence that selection shapes the boundary; I would cite it as a cautionary example of outcome-scale mismatch in replication.","headline":"A transparent, well-executed preregistered replication whose headline claim overreaches: the ordinal outcome cannot test the entropy-based boundary, so 'selection shapes the boundary' is not actually established.","tokens_in":11789,"tokens_out":2217,"would_cite":true,"duration_ms":27093,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A preregistered replication shows the monotonicity–label agreement boundary in NLI is conditional on disagreement-selected resources, not a population property.","keywords":["human label variation","natural language inference","monotonicity","label agreement","preregistered replication","selection bias","ChaosNLI","Cliff's delta"],"falsifier":"Tag the monotonicity property on a large random sample of unselected dev items with high-quality human annotation, and compute tagger error rates separately within each agreement level. If error rates differ across levels (e.g., higher misclassification among 5/5 items) such that agreement-corrected δ becomes negative and approaches −0.284, the reversal is a measurement artifact; if errors are uniform and corrected δ remains positive, the selection-conditional conclusion stands.","tokens_in":10848,"feed_emoji":"📉","tokens_out":5017,"duration_ms":50617,"temperature":0.7,"pith_summary":"Earlier work found that NLI hypotheses containing non-upward monotonicity operators (downward-entailing or non-monotone triggers) show lower annotator agreement, with a Cliff's delta of −0.284, in ChaosNLI — but ChaosNLI only re-annotates items whose five original labels split 3-to-2. This paper preregistered a replication of that boundary in the unselected SNLI and MNLI development sets, using the same tagger and a four-level agreement ordinal from the original five labels. The registered prediction fails: all seven contrasts return positive Cliff's deltas (non-upward items agree slightly more), the only significant contrast has the opposite sign, and every effect is below the preregistered smallest effect size of interest of 0.10. The paper argues the earlier negative boundary is a structure conditional on low-agreement selection, and that HLV structure claims built on selected re-annotation resources should state their selection conditional explicitly.","feed_headline":"Monotonicity disagreement gap vanishes outside selected NLI data","feed_subtitle":"In unselected SNLI/MNLI dev sets, non-upward hypotheses agree slightly more, not less; the negative result was selection-dependent.","key_machinery":"The load-bearing machinery is a three-part measurement chain: a rule-based monotonicity operator tagger (v0.3, frozen before analysis) that labels each hypothesis upward, downward, non-monotone, or mixed; a four-level agreement ordinal (no majority < 3/5 < 4/5 < 5/5) derived from the five original validation labels; and tie-corrected Cliff's delta with a preregistered smallest effect size of interest at |δ| = 0.10. The tagger's reliability is defended by a MED-anchored misclassification simulation and a codebook-based manual audit. A deliberate design point is that the ChaosNLI bridge correlation cannot be computed — the ordinal has zero variance on the overlap — so comparisons are limited t","core_discovery":"The central discovery is the measured selection dependence itself. Inside ChaosNLI's 3-of-5-majority stratum, non-upward monotonicity operators separated items by residual disagreement; in the unselected SNLI dev (9,986 items), MNLI matched dev (10,000), and MNLI mismatched dev (9,946), the association reverses: every contrast gives a positive Cliff's delta, largest +0.059, none reaching the SESOI. The natural bridge analysis is undefined by arithmetic because the agreement ordinal is constant on the ChaosNLI overlap, so the paper compares sign and magnitude class only. Robustness checks — simulated tagger misclassification on a 36-cell grid (maximum 97.5th percentile δ 0.064) and a 200-item","pith_inferences":["An implication the paper leaves implicit: any linguistic predictor found to 'predict disagreement' inside a selected resource should be re-tested on the unselected population before being treated as a general property; selection may not just attenuate but flip the sign of such boundaries.","A testable extension: re-annotating a large random unselected sample with many labels per item would show whether a monotonicity effect appears at finer granularity, or whether the boundary is genuinely absent outside low-agreement strata.","Because the ordinal is constant on the ChaosNLI overlap, any future work wanting to compare selected and unselected populations must either build selection into the design or use a measure that varies inside strata, such as entropy with many labels.","One could extend the same preregistered design to other re-annotation resources with different selection rules (e.g., selecting by entropy or by model uncertainty) to map how selection shape changes apparent linguistic structure."],"forward_implications":["HLV structure claims estimated on disagreement-selected re-annotation resources must state their selection conditional explicitly; the ChaosNLI 3-of-5 rule is strong enough to zero out the variance of the four-level agreement ordinal on the overlap.","The monotonicity–agreement boundary is not a population-level property of SNLI/MNLI; non-upward hypotheses do not agree less in unselected dev sets.","Coarsening an outcome cannot reverse an association's sign, so the positive reversal is not an artifact of the coarser four-level ordinal.","The contested semantics of operators like 'only' and 'many' means part of the tagger–human gap is irreducible and itself a form of human label variation.","The registered prediction failure and below-SESOI effects suggest item-level discriminability of monotonicity for agreement is essentially absent in these populations (pseudo-R² near zero, AUC near chance)."],"fun_headline_variants":["Monotonicity NLI bias vanishes outside selected data","Replication fails: NLI agreement gap reverses in unselected sets","No monotonicity boundary in unselected NLI populations","Selection explains NLI disagreement gap, not monotonicity","ChaosNLI monotonicity effect not found in dev sets"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The replication's verdict rests on the assumption that the frozen monotonicity tagger's errors are not systematically correlated with agreement level; if high- or low-agreement items are mis-tagged more often, the observed positive reversal could be manufactured rather than real.","fun_headline_variants_meta":{"raw":{"variants":["Monotonicity NLI bias vanishes outside selected data","Replication fails: NLI agreement gap reverses in unselected sets","No monotonicity boundary in unselected NLI populations","Selection explains NLI disagreement gap, not monotonicity","ChaosNLI monotonicity effect not found in dev sets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000159,"raw_usage":{"total_tokens":1103,"prompt_tokens":822,"completion_tokens":281,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":566,"completion_tokens_details":{"reasoning_tokens":210}},"tokens_in":566,"tokens_out":281,"duration_ms":3580,"temperature":1.0,"reasoning_tokens":210,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T13:00:37.462519+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Tag the monotonicity property on a large random sample of unselected dev items with high-quality human annotation, and compute tagger error rates separately within each agreement level. If error rates differ across levels (e.g., higher misclassification among 5/5 items) such that agreement-corrected δ becomes negative and approaches −0.284, the reversal is a measurement artifact; if errors are uniform and corrected δ remains positive, the selection-conditional conclusion stands.","supporting_citations":[],"review_version":1}