REVIEW 2 major objections 4 minor 13 references
A preregistered replication shows the monotonicity–label agreement boundary in NLI is conditional on disagreement-selected resources, not a population property.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
2026-08-01 13:00 UTC pith:V3LJT6IZ
load-bearing objection A transparent, well-executed preregistered replication whose headline claim overreaches: the ordinal outcome cannot test the entropy-based boundary, so 'selection shapes the boundary' is not actually established. the 2 major comments →
Selection Shapes the Boundary: A Preregistered Replication of Monotonicity and Label Agreement in Unselected NLI Populations
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central discovery is the measured selection dependence itself. Inside ChaosNLI's 3-of-5-majority stratum, non-upward monotonicity operators separated items by residual disagreement; in the unselected SNLI dev (9,986 items), MNLI matched dev (10,000), and MNLI mismatched dev (9,946), the association reverses: every contrast gives a positive Cliff's delta, largest +0.059, none reaching the SESOI. The natural bridge analysis is undefined by arithmetic because the agreement ordinal is constant on the ChaosNLI overlap, so the paper compares sign and magnitude class only. Robustness checks — simulated tagger misclassification on a 36-cell grid (maximum 97.5th percentile δ 0.064) and a 200-item
What carries the argument
The load-bearing machinery is a three-part measurement chain: a rule-based monotonicity operator tagger (v0.3, frozen before analysis) that labels each hypothesis upward, downward, non-monotone, or mixed; a four-level agreement ordinal (no majority < 3/5 < 4/5 < 5/5) derived from the five original validation labels; and tie-corrected Cliff's delta with a preregistered smallest effect size of interest at |δ| = 0.10. The tagger's reliability is defended by a MED-anchored misclassification simulation and a codebook-based manual audit. A deliberate design point is that the ChaosNLI bridge correlation cannot be computed — the ordinal has zero variance on the overlap — so comparisons are limited t
Load-bearing premise
The replication's verdict rests on the assumption that the frozen monotonicity tagger's errors are not systematically correlated with agreement level; if high- or low-agreement items are mis-tagged more often, the observed positive reversal could be manufactured rather than real.
What would settle it
Tag the monotonicity property on a large random sample of unselected dev items with high-quality human annotation, and compute tagger error rates separately within each agreement level. If error rates differ across levels (e.g., higher misclassification among 5/5 items) such that agreement-corrected δ becomes negative and approaches −0.284, the reversal is a measurement artifact; if errors are uniform and corrected δ remains positive, the selection-conditional conclusion stands.
If this is right
- HLV structure claims estimated on disagreement-selected re-annotation resources must state their selection conditional explicitly; the ChaosNLI 3-of-5 rule is strong enough to zero out the variance of the four-level agreement ordinal on the overlap.
- The monotonicity–agreement boundary is not a population-level property of SNLI/MNLI; non-upward hypotheses do not agree less in unselected dev sets.
- Coarsening an outcome cannot reverse an association's sign, so the positive reversal is not an artifact of the coarser four-level ordinal.
- The contested semantics of operators like 'only' and 'many' means part of the tagger–human gap is irreducible and itself a form of human label variation.
- The registered prediction failure and below-SESOI effects suggest item-level discriminability of monotonicity for agreement is essentially absent in these populations (pseudo-R² near zero, AUC near chance).
Where Pith is reading between the lines
- An implication the paper leaves implicit: any linguistic predictor found to 'predict disagreement' inside a selected resource should be re-tested on the unselected population before being treated as a general property; selection may not just attenuate but flip the sign of such boundaries.
- A testable extension: re-annotating a large random unselected sample with many labels per item would show whether a monotonicity effect appears at finer granularity, or whether the boundary is genuinely absent outside low-agreement strata.
- Because the ordinal is constant on the ChaosNLI overlap, any future work wanting to compare selected and unselected populations must either build selection into the design or use a measure that varies inside strata, such as entropy with many labels.
- One could extend the same preregistered design to other re-annotation resources with different selection rules (e.g., selecting by entropy or by model uncertainty) to map how selection shape changes apparent linguistic structure.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a preregistered replication attempt of a previously observed association between non-upward monotonicity operators and lower label agreement in NLI. The earlier study (Choi, 2026) measured this on ChaosNLI, a disagreement-selected resource restricted to items with a 3/5 majority label, using 100-label entropy as the outcome (Cliff's delta -0.284). The present study applies the same frozen tagger to the unselected SNLI and MNLI development sets and uses a four-level ordinal agreement outcome built from the five original labels. Across seven contrasts, Cliff's delta is positive (non-upward items agree slightly more), all effects are below the preregistered SESOI of 0.10, and the only significant confirmatory contrast has the opposite sign to the registration. The paper concludes that the earlier negative boundary is plausibly a structure conditional on low-agreement selection rather than a population-level property, and recommends that HLV structure claims state their selection conditional. The manuscript includes a thorough reproducibility appendix, a registration-versus-result concordance, a misclassification simulation grid, and a manual audit.
Significance. If the conclusion is taken at face value, the paper is a useful cautionary result for the perspectivist NLI literature: a pronounced disagreement-related boundary measured inside a disagreement-selected resource can be absent—and even slightly reversed—in the unselected population on a coarser outcome. The preregistration, frozen predictor, SESOI, sensitivity analyses, simulation grid, and manual audit are genuine methodological strengths, and the reported effects are consistently below the SESOI across all robustness checks. The paper also honestly discloses its own limitations, including outcome-scale non-commensurability and single-author audit. However, the central interpretive claim outruns the evidence because the replication used a different outcome functional than the original study, and the measurement-validity analysis does not rule out agreement-correlated tagger error.
major comments (2)
- [Section 6 and Appendix B] The central claim that the earlier monotonicity–agreement boundary is conditional on low-agreement selection is not directly supported by the reported analysis. The original δ = -0.284 was computed on 100-label entropy inside ChaosNLI, which by construction contains only items at the 3/5 ordinal level; Appendix B shows the four-level ordinal has zero variance on the entire ChaosNLI overlap. The entropy signal that produced the original boundary therefore lives entirely within a single ordinal level. A positive ordinal delta in unselected dev (Table 1) is compatible with a latent distribution in which non-upward items are slightly more likely to reach 5/5 unanimous agreement yet have higher entropy conditional on non-unanimity. The Section 4 argument that coarsening attenuates toward zero is not sufficient because the ordinal is not a coarsening of 100-label entropy; it is a different fun
- [Section 5, Tier 2 and Tier 3] The misclassification simulation in Appendix C applies random label flips at MED-anchored or fixed rates. This cannot detect tagger errors that are systematically correlated with the agreement ordinal, because the flip probabilities are assumed to be independent of the item's agreement level. If the tagger misclassifies non-upward items more often in low-agreement items—plausible if complex syntax drives both disagreement and tagger failure—then the observed positive reversal could be manufactured rather than a property of the unselected population. The Tier 3 audit reports overall agreement (0.875 four-class, κ=0.607) but does not stratify the 25 disagreements by the item's agreement level. The claim that measurement error 'shrinks the effects rather than manufacturing them' is true only under non-differential misclassification, which is not established. Please test or explicitly constr
minor comments (4)
- [Figure 1] The gray reference row for the earlier entropy-scale estimate is plotted on the same axis as the ordinal-scale deltas even though the paper argues the scales are not numerically commensurable. A separate panel or clear visual separation would avoid implying interval comparability.
- [Appendix B and Figure 2 caption] The text alternates between 'four-level ordinal' (Abstract, Section 3) and 'five-label ordinal' (Appendix B, Figure 2). Consistent terminology—e.g., 'four-level agreement ordinal derived from five labels'—would help.
- [Section 7] The statement that 'every quantity this paper reports is printed in the paper itself' is in slight tension with Section 5, where the layer counts are explicitly described as approximate and not regenerated by any script. This is not a substantive problem, but the wording could be softened.
- [Title and Introduction] The term 'boundary' is used both for the entropy-scale effect inside ChaosNLI and for the ordinal-scale comparison in unselected dev. Defining the two senses early would reduce the risk of conflating the outcome scales.
Circularity Check
No significant circularity: the replication's predictor, outcome, and audit are independent of the conclusion.
full rationale
The paper's derivation chain is self-contained and does not reduce to its inputs. The monotonicity predictor is a rule-based tagger (v0.3) frozen before analysis under preregistration, not fitted to the agreement outcome; the audit codebook was explicitly written from semantic definitions rather than the tagger's rule inventory, and the authors deliberately refrained from patching the tagger after the audit to avoid circular validation. The outcome is a four-level agreement ordinal built from the original SNLI/MNLI labels, independent of the monotonicity tagger. The earlier Choi (2026) result is cited as the empirical target being replicated, not as a load-bearing proof of the current paper's claims; the current study tests it on new unselected data and finds the opposite sign, so the self-citation is the object of falsification rather than a circular justification. The acknowledged non-commensurability between the 100-label entropy outcome and the four-level ordinal is a validity limitation on the strength of the comparative conclusion, not a definitional equivalence or a fitted-input prediction. No parameter is fitted to data and then renamed as a prediction; the Tier 2 simulation uses MED-anchored external error rates, and the Tier 3 audit is a blind manual check. Therefore no circular step is present.
Axiom & Free-Parameter Ledger
free parameters (1)
- SESOI bound =
0.10 (absolute Cliff's delta)
axioms (4)
- standard math Cliff's delta, Mann-Whitney U, and percentile bootstrap CIs are valid for ordinal comparison
- domain assumption SNLI and MNLI dev five-label sets are the unselected populations from which ChaosNLI items were drawn, and the 3/5-majority rule is the operative selection rule
- domain assumption The monotonicity tagger v0.3, frozen before analysis, gives a valid classification of hypotheses into upward vs non-upward, with misclassification non-differential with respect to the agreement ordinal; MED-anchored error rates are a reasonable proxy for dev-domain error rates
- domain assumption The four-level agreement ordinal and the ChaosNLI 100-label entropy measure the same latent agreement construct, so sign comparison is meaningful
read the original abstract
Prior work on human label variation (HLV) in natural language inference (NLI) has often relied on re-annotation resources that select items by disagreement level. An earlier study (arXiv:2607.15870) found that hypotheses containing non-upward monotonicity operators showed lower label agreement in ChaosNLI (Cliff's delta = -0.284), which is restricted to items whose majority label carries exactly three of five votes. We preregistered a replication of this boundary in the unselected populations that ChaosNLI was drawn from: the SNLI and MultiNLI development sets, using the same operator tagger and a four-level ordinal agreement outcome. The registered prediction fails. All seven contrasts return a positive Cliff's delta (non-upward items agree slightly more, not less), the only significant confirmatory contrast has the opposite sign to the registration, and every effect is far below our smallest effect size of interest (0.10). Robustness checks support the measurement: simulated tagger misclassification shrinks the effects rather than manufacturing them, and a manual re-tagging audit reaches four-class agreement of 0.875 on a fresh 200-item sample. We conclude that the earlier negative boundary is plausibly a structure conditional on low-agreement selection rather than a population-level property, and that HLV structure claims built on selected re-annotation resources should state their selection conditional explicitly.
Figures
Reference graph
Works this paper leans on
-
[1]
Bowman, Gabor Angeli, Christopher Potts, and Christopher D
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. https://doi.org/10.18653/v1/D15-1075 A large annotated corpus for learning natural language inference . In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 632--642, Lisbon, Portugal. Association for Computational Linguistics
-
[2]
Rollin Brant. 1990. https://doi.org/10.2307/2532457 Assessing proportionality in the proportional odds model for ordinal logistic regression . Biometrics, 46(4):1171--1178
-
[3]
Haram Choi. 2026. https://arxiv.org/abs/2607.15870 How much human label variation does formal semantic structure explain?: Group-level effects and item-level ceilings in NLI . Preprint, arXiv:2607.15870. ArXiv preprint, v1, submitted 2026-07-17
Pith/arXiv arXiv 2026
-
[4]
Pritam Kadasi and Mayank Singh. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.96 Unveiling the multi-annotation process: Examining the influence of annotation quantity and instance difficulty on model performance . In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 1371--1388, Singapore. Association for Computational L...
-
[5]
Bill MacCartney. 2009. https://nlp.stanford.edu/ wcmac/papers/nli-diss.pdf Natural Language Inference . Ph.D. thesis, Stanford University
2009
-
[6]
Yixin Nie, Xiang Zhou, and Mohit Bansal. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.734 What can we learn from collective human opinions on natural language inference data? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9131--9143, Online. Association for Computational Linguistics
-
[7]
Barbara H. Partee. 1989. Many quantifiers. In Proceedings of the 5th Eastern States Conference on Linguistics, pages 383--402, Columbus, OH. Ohio State University. Reprinted in Compositionality in Formal Semantics: Selected Papers by Barbara H. Partee, pages 241--258, Blackwell, Oxford, 2004
1989
-
[8]
Ellie Pavlick and Tom Kwiatkowski. 2019. https://doi.org/10.1162/tacl_a_00293 Inherent disagreements in human textual inferences . Transactions of the Association for Computational Linguistics, 7:677--694
-
[9]
Barbara Plank. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.731 The ``problem'' of human label variation: On ground truth in data, modeling and evaluation . In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 10671--10682, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics
-
[10]
Jessica Rett. 2018. https://doi.org/10.1111/lnc3.12269 The semantics of many, much, few, and little . Language and Linguistics Compass, 12(1):e12269
-
[11]
Kai von Fintel. 1999. https://doi.org/10.1093/jos/16.2.97 NPI licensing, S trawson entailment, and context dependency . Journal of Semantics, 16(2):97--148
-
[12]
Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2018. https://doi.org/10.18653/v1/N18-1101 A broad-coverage challenge corpus for sentence understanding through inference . In Proceedings of the 2018 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers) , pages 1...
-
[13]
Hitomi Yanaka, Koji Mineshima, Daisuke Bekki, Kentaro Inui, Satoshi Sekine, Lasha Abzianidze, and Johan Bos. 2019. https://doi.org/10.18653/v1/W19-4804 Can neural networks understand monotonicity reasoning? In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 31--40, Florence, Italy. Association fo...
This paper was first reviewed by deepseek-v4-flash on August 1, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.