Pith. sign in

REVIEW 2 major objections 4 minor 13 references

A preregistered replication shows the monotonicity–label agreement boundary in NLI is conditional on disagreement-selected resources, not a population property.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review

2026-08-01 13:00 UTC pith:V3LJT6IZ

load-bearing objection A transparent, well-executed preregistered replication whose headline claim overreaches: the ordinal outcome cannot test the entropy-based boundary, so 'selection shapes the boundary' is not actually established. the 2 major comments →

arxiv 2607.19231 v1 pith:V3LJT6IZ submitted 2026-07-21 cs.CL

Selection Shapes the Boundary: A Preregistered Replication of Monotonicity and Label Agreement in Unselected NLI Populations

classification cs.CL
keywords human label variationnatural language inferencemonotonicitylabel agreementpreregistered replicationselection biasChaosNLICliff's delta
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Earlier work found that NLI hypotheses containing non-upward monotonicity operators (downward-entailing or non-monotone triggers) show lower annotator agreement, with a Cliff's delta of −0.284, in ChaosNLI — but ChaosNLI only re-annotates items whose five original labels split 3-to-2. This paper preregistered a replication of that boundary in the unselected SNLI and MNLI development sets, using the same tagger and a four-level agreement ordinal from the original five labels. The registered prediction fails: all seven contrasts return positive Cliff's deltas (non-upward items agree slightly more), the only significant contrast has the opposite sign, and every effect is below the preregistered smallest effect size of interest of 0.10. The paper argues the earlier negative boundary is a structure conditional on low-agreement selection, and that HLV structure claims built on selected re-annotation resources should state their selection conditional explicitly.

Core claim

The central discovery is the measured selection dependence itself. Inside ChaosNLI's 3-of-5-majority stratum, non-upward monotonicity operators separated items by residual disagreement; in the unselected SNLI dev (9,986 items), MNLI matched dev (10,000), and MNLI mismatched dev (9,946), the association reverses: every contrast gives a positive Cliff's delta, largest +0.059, none reaching the SESOI. The natural bridge analysis is undefined by arithmetic because the agreement ordinal is constant on the ChaosNLI overlap, so the paper compares sign and magnitude class only. Robustness checks — simulated tagger misclassification on a 36-cell grid (maximum 97.5th percentile δ 0.064) and a 200-item

What carries the argument

The load-bearing machinery is a three-part measurement chain: a rule-based monotonicity operator tagger (v0.3, frozen before analysis) that labels each hypothesis upward, downward, non-monotone, or mixed; a four-level agreement ordinal (no majority < 3/5 < 4/5 < 5/5) derived from the five original validation labels; and tie-corrected Cliff's delta with a preregistered smallest effect size of interest at |δ| = 0.10. The tagger's reliability is defended by a MED-anchored misclassification simulation and a codebook-based manual audit. A deliberate design point is that the ChaosNLI bridge correlation cannot be computed — the ordinal has zero variance on the overlap — so comparisons are limited t

Load-bearing premise

The replication's verdict rests on the assumption that the frozen monotonicity tagger's errors are not systematically correlated with agreement level; if high- or low-agreement items are mis-tagged more often, the observed positive reversal could be manufactured rather than real.

What would settle it

Tag the monotonicity property on a large random sample of unselected dev items with high-quality human annotation, and compute tagger error rates separately within each agreement level. If error rates differ across levels (e.g., higher misclassification among 5/5 items) such that agreement-corrected δ becomes negative and approaches −0.284, the reversal is a measurement artifact; if errors are uniform and corrected δ remains positive, the selection-conditional conclusion stands.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • HLV structure claims estimated on disagreement-selected re-annotation resources must state their selection conditional explicitly; the ChaosNLI 3-of-5 rule is strong enough to zero out the variance of the four-level agreement ordinal on the overlap.
  • The monotonicity–agreement boundary is not a population-level property of SNLI/MNLI; non-upward hypotheses do not agree less in unselected dev sets.
  • Coarsening an outcome cannot reverse an association's sign, so the positive reversal is not an artifact of the coarser four-level ordinal.
  • The contested semantics of operators like 'only' and 'many' means part of the tagger–human gap is irreducible and itself a form of human label variation.
  • The registered prediction failure and below-SESOI effects suggest item-level discriminability of monotonicity for agreement is essentially absent in these populations (pseudo-R² near zero, AUC near chance).

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • An implication the paper leaves implicit: any linguistic predictor found to 'predict disagreement' inside a selected resource should be re-tested on the unselected population before being treated as a general property; selection may not just attenuate but flip the sign of such boundaries.
  • A testable extension: re-annotating a large random unselected sample with many labels per item would show whether a monotonicity effect appears at finer granularity, or whether the boundary is genuinely absent outside low-agreement strata.
  • Because the ordinal is constant on the ChaosNLI overlap, any future work wanting to compare selected and unselected populations must either build selection into the design or use a measure that varies inside strata, such as entropy with many labels.
  • One could extend the same preregistered design to other re-annotation resources with different selection rules (e.g., selecting by entropy or by model uncertainty) to map how selection shape changes apparent linguistic structure.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper reports a preregistered replication attempt of a previously observed association between non-upward monotonicity operators and lower label agreement in NLI. The earlier study (Choi, 2026) measured this on ChaosNLI, a disagreement-selected resource restricted to items with a 3/5 majority label, using 100-label entropy as the outcome (Cliff's delta -0.284). The present study applies the same frozen tagger to the unselected SNLI and MNLI development sets and uses a four-level ordinal agreement outcome built from the five original labels. Across seven contrasts, Cliff's delta is positive (non-upward items agree slightly more), all effects are below the preregistered SESOI of 0.10, and the only significant confirmatory contrast has the opposite sign to the registration. The paper concludes that the earlier negative boundary is plausibly a structure conditional on low-agreement selection rather than a population-level property, and recommends that HLV structure claims state their selection conditional. The manuscript includes a thorough reproducibility appendix, a registration-versus-result concordance, a misclassification simulation grid, and a manual audit.

Significance. If the conclusion is taken at face value, the paper is a useful cautionary result for the perspectivist NLI literature: a pronounced disagreement-related boundary measured inside a disagreement-selected resource can be absent—and even slightly reversed—in the unselected population on a coarser outcome. The preregistration, frozen predictor, SESOI, sensitivity analyses, simulation grid, and manual audit are genuine methodological strengths, and the reported effects are consistently below the SESOI across all robustness checks. The paper also honestly discloses its own limitations, including outcome-scale non-commensurability and single-author audit. However, the central interpretive claim outruns the evidence because the replication used a different outcome functional than the original study, and the measurement-validity analysis does not rule out agreement-correlated tagger error.

major comments (2)
  1. [Section 6 and Appendix B] The central claim that the earlier monotonicity–agreement boundary is conditional on low-agreement selection is not directly supported by the reported analysis. The original δ = -0.284 was computed on 100-label entropy inside ChaosNLI, which by construction contains only items at the 3/5 ordinal level; Appendix B shows the four-level ordinal has zero variance on the entire ChaosNLI overlap. The entropy signal that produced the original boundary therefore lives entirely within a single ordinal level. A positive ordinal delta in unselected dev (Table 1) is compatible with a latent distribution in which non-upward items are slightly more likely to reach 5/5 unanimous agreement yet have higher entropy conditional on non-unanimity. The Section 4 argument that coarsening attenuates toward zero is not sufficient because the ordinal is not a coarsening of 100-label entropy; it is a different fun
  2. [Section 5, Tier 2 and Tier 3] The misclassification simulation in Appendix C applies random label flips at MED-anchored or fixed rates. This cannot detect tagger errors that are systematically correlated with the agreement ordinal, because the flip probabilities are assumed to be independent of the item's agreement level. If the tagger misclassifies non-upward items more often in low-agreement items—plausible if complex syntax drives both disagreement and tagger failure—then the observed positive reversal could be manufactured rather than a property of the unselected population. The Tier 3 audit reports overall agreement (0.875 four-class, κ=0.607) but does not stratify the 25 disagreements by the item's agreement level. The claim that measurement error 'shrinks the effects rather than manufacturing them' is true only under non-differential misclassification, which is not established. Please test or explicitly constr
minor comments (4)
  1. [Figure 1] The gray reference row for the earlier entropy-scale estimate is plotted on the same axis as the ordinal-scale deltas even though the paper argues the scales are not numerically commensurable. A separate panel or clear visual separation would avoid implying interval comparability.
  2. [Appendix B and Figure 2 caption] The text alternates between 'four-level ordinal' (Abstract, Section 3) and 'five-label ordinal' (Appendix B, Figure 2). Consistent terminology—e.g., 'four-level agreement ordinal derived from five labels'—would help.
  3. [Section 7] The statement that 'every quantity this paper reports is printed in the paper itself' is in slight tension with Section 5, where the layer counts are explicitly described as approximate and not regenerated by any script. This is not a substantive problem, but the wording could be softened.
  4. [Title and Introduction] The term 'boundary' is used both for the entropy-scale effect inside ChaosNLI and for the ordinal-scale comparison in unselected dev. Defining the two senses early would reduce the risk of conflating the outcome scales.

Circularity Check

0 steps flagged

No significant circularity: the replication's predictor, outcome, and audit are independent of the conclusion.

full rationale

The paper's derivation chain is self-contained and does not reduce to its inputs. The monotonicity predictor is a rule-based tagger (v0.3) frozen before analysis under preregistration, not fitted to the agreement outcome; the audit codebook was explicitly written from semantic definitions rather than the tagger's rule inventory, and the authors deliberately refrained from patching the tagger after the audit to avoid circular validation. The outcome is a four-level agreement ordinal built from the original SNLI/MNLI labels, independent of the monotonicity tagger. The earlier Choi (2026) result is cited as the empirical target being replicated, not as a load-bearing proof of the current paper's claims; the current study tests it on new unselected data and finds the opposite sign, so the self-citation is the object of falsification rather than a circular justification. The acknowledged non-commensurability between the 100-label entropy outcome and the four-level ordinal is a validity limitation on the strength of the comparative conclusion, not a definitional equivalence or a fitted-input prediction. No parameter is fitted to data and then renamed as a prediction; the Tier 2 simulation uses MED-anchored external error rates, and the Tier 3 audit is a blind manual check. Therefore no circular step is present.

Axiom & Free-Parameter Ledger

1 free parameters · 4 axioms · 0 invented entities

The paper introduces no new entities and fits no parameters to data beyond the preregistered SESOI threshold. Its core free parameters and axioms are measurement assumptions: the tagger's validity/non-differential errors and the comparability of the two agreement operationalizations. The statistical machinery is standard.

free parameters (1)
  • SESOI bound = 0.10 (absolute Cliff's delta)
    Preregistered smallest effect size of interest; the paper's conclusion that all effects are below SESOI depends on this hand-chosen threshold. A smaller SESOI (e.g., 0.05) would change that particular conclusion, though the sign reversal would remain.
axioms (4)
  • standard math Cliff's delta, Mann-Whitney U, and percentile bootstrap CIs are valid for ordinal comparison
    Section 3 statistics; standard nonparametric inference applied to the four-level ordinal outcome.
  • domain assumption SNLI and MNLI dev five-label sets are the unselected populations from which ChaosNLI items were drawn, and the 3/5-majority rule is the operative selection rule
    Appendix B overlap arithmetic confirms every ChaosNLI row maps to a 3/5 (or 3/4) item in the dev sets, establishing the resource as a selected subset.
  • domain assumption The monotonicity tagger v0.3, frozen before analysis, gives a valid classification of hypotheses into upward vs non-upward, with misclassification non-differential with respect to the agreement ordinal; MED-anchored error rates are a reasonable proxy for dev-domain error rates
    Section 3 (predictor freeze) and Section 5 (Tier 2 simulation and Tier 3 audit). This is the weakest premise: the simulation assumes random flips and cannot model agreement-correlated systematic tagger errors.
  • domain assumption The four-level agreement ordinal and the ChaosNLI 100-label entropy measure the same latent agreement construct, so sign comparison is meaningful
    Sections 3 and 6; the paper explicitly notes the scales are non-commensurable and restricts comparison to sign and magnitude class.

pith-pipeline@v1.3.0-alltime-deepseek · 10605 in / 15264 out tokens · 154913 ms · 2026-08-01T13:00:37.462519+00:00 · methodology

0 comments
read the original abstract

Prior work on human label variation (HLV) in natural language inference (NLI) has often relied on re-annotation resources that select items by disagreement level. An earlier study (arXiv:2607.15870) found that hypotheses containing non-upward monotonicity operators showed lower label agreement in ChaosNLI (Cliff's delta = -0.284), which is restricted to items whose majority label carries exactly three of five votes. We preregistered a replication of this boundary in the unselected populations that ChaosNLI was drawn from: the SNLI and MultiNLI development sets, using the same operator tagger and a four-level ordinal agreement outcome. The registered prediction fails. All seven contrasts return a positive Cliff's delta (non-upward items agree slightly more, not less), the only significant confirmatory contrast has the opposite sign to the registration, and every effect is far below our smallest effect size of interest (0.10). Robustness checks support the measurement: simulated tagger misclassification shrinks the effects rather than manufacturing them, and a manual re-tagging audit reaches four-class agreement of 0.875 on a fresh 200-item sample. We conclude that the earlier negative boundary is plausibly a structure conditional on low-agreement selection rather than a population-level property, and that HLV structure claims built on selected re-annotation resources should state their selection conditional explicitly.

Figures

Figures reproduced from arXiv: 2607.19231 by Haram Choi.

Figure 2
Figure 2. Figure 2: Share of non-upward hypotheses at each agreement level, with Wilson 95% confidence inter￾vals, for SNLI dev and MNLI matched dev. The share does not decrease as agreement rises. Generated by analysis/fig_bridge_shares.py. annotation count from one to one hundred labels per item on ChaosNLI make that dependence ex￾plicit (Kadasi and Singh, 2023). Complexity confounders do not explain the re￾versal. As an au… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

13 extracted references · 5 canonical work pages

  1. [1]

    Bowman, Gabor Angeli, Christopher Potts, and Christopher D

    Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. https://doi.org/10.18653/v1/D15-1075 A large annotated corpus for learning natural language inference . In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 632--642, Lisbon, Portugal. Association for Computational Linguistics

  2. [2]

    Rollin Brant. 1990. https://doi.org/10.2307/2532457 Assessing proportionality in the proportional odds model for ordinal logistic regression . Biometrics, 46(4):1171--1178

  3. [3]

    Haram Choi. 2026. https://arxiv.org/abs/2607.15870 How much human label variation does formal semantic structure explain?: Group-level effects and item-level ceilings in NLI . Preprint, arXiv:2607.15870. ArXiv preprint, v1, submitted 2026-07-17

  4. [4]

    Pritam Kadasi and Mayank Singh. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.96 Unveiling the multi-annotation process: Examining the influence of annotation quantity and instance difficulty on model performance . In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 1371--1388, Singapore. Association for Computational L...

  5. [5]

    Bill MacCartney. 2009. https://nlp.stanford.edu/ wcmac/papers/nli-diss.pdf Natural Language Inference . Ph.D. thesis, Stanford University

  6. [6]

    Yixin Nie, Xiang Zhou, and Mohit Bansal. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.734 What can we learn from collective human opinions on natural language inference data? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9131--9143, Online. Association for Computational Linguistics

  7. [7]

    Barbara H. Partee. 1989. Many quantifiers. In Proceedings of the 5th Eastern States Conference on Linguistics, pages 383--402, Columbus, OH. Ohio State University. Reprinted in Compositionality in Formal Semantics: Selected Papers by Barbara H. Partee, pages 241--258, Blackwell, Oxford, 2004

  8. [8]

    Ellie Pavlick and Tom Kwiatkowski. 2019. https://doi.org/10.1162/tacl_a_00293 Inherent disagreements in human textual inferences . Transactions of the Association for Computational Linguistics, 7:677--694

  9. [9]

    Barbara Plank. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.731 The ``problem'' of human label variation: On ground truth in data, modeling and evaluation . In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 10671--10682, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics

  10. [10]

    Jessica Rett. 2018. https://doi.org/10.1111/lnc3.12269 The semantics of many, much, few, and little . Language and Linguistics Compass, 12(1):e12269

  11. [11]

    Kai von Fintel. 1999. https://doi.org/10.1093/jos/16.2.97 NPI licensing, S trawson entailment, and context dependency . Journal of Semantics, 16(2):97--148

  12. [12]

    Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2018. https://doi.org/10.18653/v1/N18-1101 A broad-coverage challenge corpus for sentence understanding through inference . In Proceedings of the 2018 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers) , pages 1...

  13. [13]

    Hitomi Yanaka, Koji Mineshima, Daisuke Bekki, Kentaro Inui, Satoshi Sekine, Lasha Abzianidze, and Johan Bos. 2019. https://doi.org/10.18653/v1/W19-4804 Can neural networks understand monotonicity reasoning? In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 31--40, Florence, Italy. Association fo...

This paper was first reviewed by deepseek-v4-flash on August 1, 2026.