Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

Formal semantic structure explains only a small slice of human label variation in NLI, yet it reliably separates high- from low-disagreement items.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 22:04 UTC pith:QWZVFU6X

load-bearing objection A transparent, preregistered measurement that gives formal semantics a small but real group-level role in NLI disagreement; the tagger's in-domain validity is the main soft spot. the 3 major comments →

arxiv 2607.15870 v1 pith:QWZVFU6X submitted 2026-07-17 cs.CL

How Much Human Label Variation Does Formal Semantic Structure Explain?: Group-Level Effects and Item-Level Ceilings in NLI

classification cs.CL
keywords human label variationnatural language inferencemonotonicityformal semanticsannotation disagreementChaosNLIlabel entropypreregistered analysis
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper asks a direct, quantitative question: how much of human label variation in natural language inference is explained by formal semantic structure? Using the 3,113 low-agreement SNLI/MNLI items of ChaosNLI, the author tags premises and hypotheses for negation, quantifiers, and monotonicity and compares label entropy across formal classes. Three bounds emerge. First, hypotheses that are not purely upward monotone have reliably higher annotator disagreement (Cliff's delta = -0.284), and this group-level gap survives controls for operator presence and sentence length. Second, at the item level the same formal profiles explain only 3.3 to 3.6 percent of the variance in entropy and reach an AUC of 0.606, so they cannot pick out high-disagreement items. Third, the composition of disagreement—shares of annotation error and of explanation types—does not detectably differ across the boundary. The paper concludes that in this sample formal semantic structure shifts how much annotators disagree by a small amount, and not detectably what they disagree about.

Core claim

On the paper's own terms, the central discovery is a calibrated bound: formal semantic structure shifts how much annotators disagree by a small amount and not detectably what they disagree about. A rule-based monotonicity tagger splits ChaosNLI items into purely upward versus other hypotheses; the non-upward group has reliably higher label entropy (Cliff's delta = -0.284), defended against operator-presence and length controls. But the same profiles explain only 3.3 to 3.6 percent of entropy variance (AUC 0.606), and three preregistered contrasts on error shares and explanation types are null. The result is a group-level boundary with an item-level ceiling.

What carries the argument

The load-bearing instrument is a shallow, rule-based operator and monotonicity tagger over dependency parses, which assigns each sentence a four-way summary (upward, downward, non-monotone, mixed) from a closed inventory of negation, quantifier, NPI, and monotonicity triggers. What carries the argument is the binary boundary derived from it: purely upward hypotheses (no downward or non-monotone trigger) versus all others, measured against label entropy and majority margin with Cliff's delta and Holm-corrected tests. The item-level ceiling is established by the same profiles in variance-explained and AUC form.

Load-bearing premise

All analyses depend on the tagger's sentence-level monotonicity summary, which agrees with the MED benchmark at only 0.807 and is unmeasured for accuracy inside the ChaosNLI domain; if that summary misrepresents the true formal structure of the items—especially in conditional and comparative environments, which the tagger skips—the group-level boundary could be an artifact of tagging blind spots rather than of formal semantics.

What would settle it

Recompute the entropy comparison using a compositional monotonicity reasoner or human expert annotation of monotonicity environments on the same ChaosNLI items, including conditionals and comparatives; if the non-upward group no longer shows higher entropy, or the variance explained rises above a few percent, the paper's ceiling and boundary claims would need revision. Alternatively, run the same preregistered protocol on an unrestricted sample of SNLI/MNLI items to see whether the group-level boundary persists outside the low-agreement selection.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Formal semantic structure belongs in the inventory of disagreement sources, but with a small, measured weight rather than a presumed large one.
  • Item-level disagreement prediction cannot rely on formal features alone; other item properties and annotator-side variables carry most of the signal.
  • The null composition contrasts suggest formal structure does not change the mix of annotation error, logical-structure conflict, or pragmatic inference behind disagreement.
  • The group-level boundary may attenuate or shift on unrestricted NLI samples, outside the low-agreement selection that defines ChaosNLI-S/M.
  • The registered audit-trail design offers a template for preregistered variance-decomposition studies of label variation.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Because the sample is restricted to low-agreement items, the 3.3 to 3.6 percent ceiling might underestimate formal structure's explanatory power on the full range of NLI items; a test on unselected items is a natural extension.
  • A compositional monotonicity tagger, covering conditionals and comparatives (which the current tagger skips and which co-occur with higher entropy), could raise the item-level ceiling; the paper's own coverage counts suggest this is testable.
  • The group-level effect might interact with genre or annotator demographics; the paper's subset dummy dominates regressions, hinting that formal structure's role could be genre-dependent.
  • The invariance of explanation-type composition across the boundary suggests that formal structure may change the difficulty of an item without changing the reasons annotators give; this distinction merits a dedicated study.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper measures, on the 3,113 SNLI/MNLI-matched items of ChaosNLI, how much formal semantic structure—operationalized by a rule-based operator and monotonicity tagger—explains human label variation. Using preregistered analyses, it reports three bounds: (1) a group-level boundary: hypotheses whose sentence-level monotonicity summary is not purely upward have higher label entropy (Cliff's δ = −0.284 for upward vs. rest) and lower majority margin, with rank-based tests claimed to survive controls for operator presence and length/complexity, though a beta-regression sensitivity check weakens the length defense; (2) an item-level ceiling: formal profiles explain 3.3–3.6% of entropy variance with a median-split AUC of 0.606; and (3) composition invariance: three preregistered contrasts on VariErr error shares and LiTEx explanation categories are null. All claims are explicitly conditioned on the low-agreement selection of ChaosNLI-S/M. The paper also discloses its preregistration audit trail, one corrected interpretation rule, and a tagged blind spot for conditional/comparative environments.

Significance. Should the findings hold, the paper is a useful calibration for perspectivist NLI: formal semantics contributes a small, reliable shift in the amount of disagreement but does not identify high-disagreement items or change the composition of disagreement. The strengths are real: all confirmatory analyses were preregistered in a version-controlled log; negative results and the failed sensitivity check are reported in full; the post-hoc correction of the R-A interpretation rule is disclosed with timing; and the reproduction repository with pinning scripts is described. This level of process transparency is rare and valuable. The principal risk is construct validity of the tagger's sentence-level monotonicity summary, which is the input to every analysis and is only indirectly validated. The paper's own limitation section and C-R4 make that risk concrete; additional stratification or validation would make the contribution solid.

major comments (3)
  1. [§3.2, §7 C-R4, Limitations 4] The central attribution of the group-level boundary to formal structure is not yet established because every analysis consumes the tagger's sentence-level monotonicity summary, whose agreement with MED is only 0.807 and whose accuracy inside ChaosNLI is unmeasured. C-R4 documents that 231/3,113 items contain 'if' or 'than'; these items are over-represented among the non-upward classes (13.9% of downward, 10.4% of non-monotone, 14.3% of mixed, vs. 6.2% of upward) and have higher median entropy (1.166 vs. 0.971 bits). Because the headline boundary contrasts upward against all other hypotheses, the observed δ = −0.284 could be inflated by the tagger's blind spots rather than by monotonicity per se. The manuscript states that it does not know in which direction full coverage would move the effect; that unknown is directly load-bearing for the main claim. Please report the boundary within the
  2. [§4.4 R-C and §7 C-R2] The abstract and §8 claim that the boundary survives preregistered reductions against length and complexity. But the preregistered beta-regression sensitivity check—the bounded-outcome model—fails the registered criterion: the marked-hypothesis coefficient loses significance in both M0 (p = 0.064) and M1 (p = 0.167). The manuscript says the defense should rest on 'rank-based results,' yet R-C itself is an OLS regression, and no rank-based test in §4 actually controls for length or parse depth. Thus the length/complexity reduction is only partially rebutted. Either add a permutation or rank-based test that controls these covariates, or explicitly downgrade the claim and report that the length defense is sensitive to the distributional model. The disclosed HC3-vs-model-based standard-error asymmetry is relevant but does not by itself satisfy the registered criterion.
  3. [§4.3] The binary boundary and its δ = −0.284 are described as the 'headline number' and are used in the abstract and conclusion as if they were a confirmatory result. The paper itself discloses that this binary contrast was defined after seeing the four-class result and 'is not a new hypothesis test.' The confirmatory evidence for the group-level effect consists of the four-class Kruskal-Wallis test and the two powered upward-vs-marked contrasts, not the binary test. Please present the binary contrast strictly as a descriptive effect size and remove confirmatory wording attached to δ = −0.284 in the abstract and conclusion.
minor comments (4)
  1. [§3.3, Table 5] The paper correctly notes a two-item discrepancy between the published MED table and the released file, and uses the released file. Please add a footnote or table caption so readers do not mistake the discrepancy for a typo in the reported validation statistics.
  2. [§2.2] The reference to Choi (2026) as the source of the dependency-based classification pattern is surprising in a formal-semantics paper, since that work concerns rhetorical invocation in institutional art writing. Please clarify what specific tagger machinery is adapted from that paper, or remove the citation if it is not load-bearing.
  3. [§3.3] For both MED agreement metrics (0.883 edit-site and 0.807 sentence-level), reporting chance-adjusted agreement (e.g., Cohen's kappa) would help interpret the sentence-level figure, which is the operative one for all downstream analyses.
  4. [§7.2] The disclosure of the corrected R-A interpretation rule is unusually transparent. Since readers may want to verify the timing, please include the original preregistered wording verbatim in an appendix or supplementary file, rather than paraphrasing it in the main text.

Circularity Check

0 steps flagged

No circularity: the predictor is externally validated, the outcome is independently recomputed, and the headline association is an empirical contrast, not a derivation from its own inputs.

full rationale

The paper's derivation chain does not reduce to its inputs. The monotonicity predictor (tagger v0.3) is a rule-based inventory validated against the external MED benchmark (0.883 edit-site agreement, 0.807 sentence-level agreement); it is not fitted to ChaosNLI label entropy or to the entropy contrasts of Section 4. The outcome variables are computed from released ChaosNLI 100-annotation label distributions and checked against published distributions (Section 3.4), and the group-level boundary is a preregistered Mann-Whitney/Kruskal-Wallis comparison of observed entropy across tagger classes. There is no equation in which the predicted quantity is defined in terms of the fit: the binary boundary is a hypothesis-side monotonicity contrast, and entropy is measured independently. Section 4.4's reductions (operator presence, length/complexity) are additional controls, not renaming of the outcome. The only self-citation, Choi (2026), supplies a classification pattern for discourse connectives and is not used to establish entropy differences; the tagger's validity rests on MED agreement. Limitations 4 and C-R4 flag measurement validity and coverage blind spots (e.g., conditionals/comparatives) that co-occur with higher entropy; these are genuine empirical threats to attribution, but they are threats of construct validity, not circularity, because they do not make the target result true by construction. The paper's negative results (3.3-3.6% variance explained, AUC 0.606, null composition contrasts) are also inconsistent with a circular design that would be expected to force the headline. No load-bearing self-citation chain, fitted-parameter-as-prediction, or definitional equivalence was found.

Axiom & Free-Parameter Ledger

0 free parameters · 3 axioms · 0 invented entities

The paper introduces no new free parameters or postulated entities. Its load-bearing assumptions are about the validity and transferability of the tagger's sentence-level monotonicity summary, and the representativeness of the cross-dataset overlap. These are domain assumptions rather than circular definitions.

axioms (3)
  • domain assumption The rule-based tagger's sentence-level monotonicity summary is a valid proxy for the formal semantic structure of NLI items.
    The paper validates the tagger against MED but explicitly notes (§3.3, Limitations 4) that agreement is 0.807 at the sentence level and unmeasured in the ChaosNLI domain. All downstream analyses consume this summary.
  • domain assumption The MED benchmark's monotonicity classes transfer to the ChaosNLI items, despite differences in genre and item construction.
    The tagger is validated on MED and then applied to SNLI/MNLI items; the paper acknowledges the external validation is only indirect (Limitations 4).
  • domain assumption The three-way overlap of ChaosNLI, VariErr, and LiTEx is representative enough for composition-invariance tests.
    The 498-item overlap is a non-random subset (all MNLI-m, all VariErr-selected), and the paper notes non-independence and reuse of this sample (§6 caveats, Limitations 6).

pith-pipeline@v1.3.0-alltime-deepseek · 167 in / 5305 out tokens · 99741 ms · 2026-08-01T22:04:14.348817+00:00 · methodology

0 comments
read the original abstract

Human label variation in natural language inference is increasingly treated as signal rather than noise, but how much of it formal semantic structure explains has not been measured directly. We measure it on the 3,113 SNLI and MNLI items of ChaosNLI, using a rule-based operator and monotonicity tagger validated against MED (0.883 agreement at the edit site, 0.807 on the sentence-level summary our analyses consume), three preregistered analysis blocks, and full reporting of negative results. Three bounds emerge. First, a group-level boundary: hypotheses that are not purely upward monotone show reliably higher label entropy (Cliff's delta = -0.284), and rank-based tests defend the effect against operator-presence and length reductions, though a bounded-outcome sensitivity check weakens the regression form of the length defense. Second, an item-level ceiling: the same formal profiles explain only 3.3 to 3.6 percent of entropy variance and reach a median-split AUC of 0.606, too weak to identify high-disagreement items. Third, composition invariance: across the boundary, three high-powered preregistered contrasts on validated error shares and explanation-type shares (VariErr, LiTEx) all return null results. In this sample, formal semantic structure shifts how much annotators disagree by a small amount and does not detectably change what they disagree about. ChaosNLI-S/M consists of items selected for low original agreement, and every claim is conditioned on that scope. All analyses were preregistered in a version-controlled research log, whose audit trail, including one corrected interpretation rule, the paper discloses.

Figures

Figures reproduced from arXiv: 2607.15870 by Haram Choi (University of Bremen).

Figure 1
Figure 1. Figure 1: Distributions of label entropy across the four [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Selection Shapes the Boundary: A Preregistered Replication of Monotonicity and Label Agreement in Unselected NLI Populations

    cs.CL 2026-07 conditional novelty 6.0

    A preregistered replication shows that the monotonicity–label-agreement boundary found in ChaosNLI fails to generalize to unselected SNLI/MNLI dev sets, with all effects small and slightly positive.

Reference graph

Works this paper leans on

22 extracted references · 5 canonical work pages · cited by 1 Pith paper

  1. [1]

    Idris Abdulmumin, Mokgadi Penelope Matloga, Tadesse Destaw Belay, Botshelo Kondowe, Letlhogonolo Mohleleng, Hareaipha Nkopo Letsoalo, Shamsuddeen Hassan Muhammad, and Vukosi Marivate. 2026. https://arxiv.org/abs/2605.27239 Temporal simultaneity predicts annotation quality in sentiment corpora . Preprint, arXiv:2605.27239

  2. [2]

    Yoav Benjamini and Yosef Hochberg. 1995. https://doi.org/10.1111/j.2517-6161.1995.tb02031.x Controlling the false discovery rate: A practical and powerful approach to multiple testing . Journal of the Royal Statistical Society: Series B (Methodological), 57(1):289--300

  3. [3]

    Haram Choi. 2026. https://doi.org/10.31235/osf.io/4au72_v2 Rhetorical invocation: A four-layer computational analysis of discourse vocabulary in institutional art writing . SocArXiv preprint, version 2, posted 2026-03-24

  4. [4]

    Norman Cliff. 1993. https://doi.org/10.1037/0033-2909.114.3.494 Dominance statistics: Ordinal analyses to answer ordinal questions . Psychological Bulletin, 114(3):494--509

  5. [5]

    S. Holm. 1979. https://www.jstor.org/stable/4615733 A simple sequentially rejective multiple test procedure . Scandinavian Journal of Statistics, 6(2):65--70

  6. [6]

    Pingjun Hong, Beiduo Chen, Siyao Peng, Marie-Catherine de Marneffe, and Barbara Plank. 2025. https://doi.org/10.18653/v1/2025.emnlp-main.1728 L i TE x: A linguistic taxonomy of explanations for understanding within-label variation in natural language inference . In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pag...

  7. [7]

    Pingjun Hong, Beiduo Chen, Siyao Peng, Marie-Catherine de Marneffe, Benjamin Roth, and Barbara Plank. 2026. https://doi.org/10.18653/v1/2026.findings-acl.1342 Agree, disagree, explain: Decomposing human label variation in NLI through the lens of explanations . In Findings of the A ssociation for C omputational L inguistics: ACL 2026 , pages 26922--26934, ...

  8. [8]

    Chathuri Jayaweera and Bonnie J. Dorr. 2025. https://arxiv.org/abs/2507.15114 From disagreement to understanding: The case for ambiguity detection in NLI . Preprint, arXiv:2507.15114

  9. [9]

    Nan-Jiang Jiang and Marie-Catherine de Marneffe. 2022. https://doi.org/10.1162/tacl_a_00523 Investigating reasons for disagreement in natural language inference . Transactions of the Association for Computational Linguistics, 10:1357--1374

  10. [10]

    Nan-Jiang Jiang, Chenhao Tan, and Marie-Catherine de Marneffe. 2023. https://arxiv.org/abs/2304.12443 Understanding and predicting human label variation in natural language inference through explanation . Preprint, arXiv:2304.12443

  11. [11]

    Maximilian Maurer, Maximilian Linde, and Gabriella Lapesa. 2026. https://arxiv.org/abs/2605.06318 Who and what? using linguistic features and annotator characteristics to analyze annotation variation . Preprint, arXiv:2605.06318

  12. [12]

    Yixin Nie, Xiang Zhou, and Mohit Bansal. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.734 What can we learn from collective human opinions on natural language inference data? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9131--9143, Online. Association for Computational Linguistics

  13. [13]

    Ellie Pavlick and Tom Kwiatkowski. 2019. https://doi.org/10.1162/tacl_a_00293 Inherent disagreements in human textual inferences . Transactions of the Association for Computational Linguistics, 7:677--694

  14. [14]

    Barbara Plank. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.731 The ``problem'' of human label variation: On ground truth in data, modeling and evaluation . In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 10671--10682, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics

  15. [15]

    Romano, J

    J. Romano, J. D. Kromrey, J. Coraggio, and J. Skowronek. 2006. Appropriate statistics for ordinal level data: Should we really be using t-test and C ohen's d for evaluating group differences on the NSSE and other surveys? In Annual meeting of the Florida Association of Institutional Research. Bibliographic entry adopted from the effsize package documentat...

  16. [16]

    Michael Smithson and Jay Verkuilen. 2006. https://doi.org/10.1037/1082-989X.11.1.54 A better lemon squeezer? maximum-likelihood regression with beta-distributed dependent variables . Psychological Methods, 11(1):54--71

  17. [17]

    Torchiano

    M. Torchiano. 2020. https://CRAN.R-project.org/package=effsize effsize: Efficient Effect Size Computation . R package version 0.8.1

  18. [18]

    Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, and Massimo Poesio

    Alexandra N. Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, and Massimo Poesio. 2021. https://doi.org/10.1613/jair.1.12752 Learning from disagreement: A survey . Journal of Artificial Intelligence Research, 72:1385--1470

  19. [19]

    Leon Weber-Genzel, Siyao Peng, Marie-Catherine De Marneffe, and Barbara Plank. 2024. https://doi.org/10.18653/v1/2024.acl-long.123 V ari E rr NLI : Separating annotation error from human label variation . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2256--2269, Bangkok, Thailand....

  20. [20]

    Hitomi Yanaka, Koji Mineshima, Daisuke Bekki, Kentaro Inui, Satoshi Sekine, Lasha Abzianidze, and Johan Bos. 2019 a . https://doi.org/10.18653/v1/W19-4804 Can neural networks understand monotonicity reasoning? In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 31--40, Florence, Italy. Association...

  21. [21]

    Hitomi Yanaka, Koji Mineshima, Daisuke Bekki, Kentaro Inui, Satoshi Sekine, Lasha Abzianidze, and Johan Bos. 2019 b . https://doi.org/10.18653/v1/S19-1027 HELP : A dataset for identifying shortcomings of neural models in monotonicity reasoning . In Proceedings of the Eighth Joint Conference on Lexical and Computational Semantics (* SEM 2019) , pages 250--...

  22. [22]

    Leixin Zhang and C a g r C \"o ltekin. 2026. https://arxiv.org/abs/2605.01168 Quantifying and predicting disagreement in graded human ratings . Preprint, arXiv:2605.01168