REVIEW 3 major objections 4 minor 1 cited by
Formal semantic structure explains only a small slice of human label variation in NLI, yet it reliably separates high- from low-disagreement items.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 22:04 UTC pith:QWZVFU6X
load-bearing objection A transparent, preregistered measurement that gives formal semantics a small but real group-level role in NLI disagreement; the tagger's in-domain validity is the main soft spot. the 3 major comments →
How Much Human Label Variation Does Formal Semantic Structure Explain?: Group-Level Effects and Item-Level Ceilings in NLI
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On the paper's own terms, the central discovery is a calibrated bound: formal semantic structure shifts how much annotators disagree by a small amount and not detectably what they disagree about. A rule-based monotonicity tagger splits ChaosNLI items into purely upward versus other hypotheses; the non-upward group has reliably higher label entropy (Cliff's delta = -0.284), defended against operator-presence and length controls. But the same profiles explain only 3.3 to 3.6 percent of entropy variance (AUC 0.606), and three preregistered contrasts on error shares and explanation types are null. The result is a group-level boundary with an item-level ceiling.
What carries the argument
The load-bearing instrument is a shallow, rule-based operator and monotonicity tagger over dependency parses, which assigns each sentence a four-way summary (upward, downward, non-monotone, mixed) from a closed inventory of negation, quantifier, NPI, and monotonicity triggers. What carries the argument is the binary boundary derived from it: purely upward hypotheses (no downward or non-monotone trigger) versus all others, measured against label entropy and majority margin with Cliff's delta and Holm-corrected tests. The item-level ceiling is established by the same profiles in variance-explained and AUC form.
Load-bearing premise
All analyses depend on the tagger's sentence-level monotonicity summary, which agrees with the MED benchmark at only 0.807 and is unmeasured for accuracy inside the ChaosNLI domain; if that summary misrepresents the true formal structure of the items—especially in conditional and comparative environments, which the tagger skips—the group-level boundary could be an artifact of tagging blind spots rather than of formal semantics.
What would settle it
Recompute the entropy comparison using a compositional monotonicity reasoner or human expert annotation of monotonicity environments on the same ChaosNLI items, including conditionals and comparatives; if the non-upward group no longer shows higher entropy, or the variance explained rises above a few percent, the paper's ceiling and boundary claims would need revision. Alternatively, run the same preregistered protocol on an unrestricted sample of SNLI/MNLI items to see whether the group-level boundary persists outside the low-agreement selection.
If this is right
- Formal semantic structure belongs in the inventory of disagreement sources, but with a small, measured weight rather than a presumed large one.
- Item-level disagreement prediction cannot rely on formal features alone; other item properties and annotator-side variables carry most of the signal.
- The null composition contrasts suggest formal structure does not change the mix of annotation error, logical-structure conflict, or pragmatic inference behind disagreement.
- The group-level boundary may attenuate or shift on unrestricted NLI samples, outside the low-agreement selection that defines ChaosNLI-S/M.
- The registered audit-trail design offers a template for preregistered variance-decomposition studies of label variation.
Where Pith is reading between the lines
- Because the sample is restricted to low-agreement items, the 3.3 to 3.6 percent ceiling might underestimate formal structure's explanatory power on the full range of NLI items; a test on unselected items is a natural extension.
- A compositional monotonicity tagger, covering conditionals and comparatives (which the current tagger skips and which co-occur with higher entropy), could raise the item-level ceiling; the paper's own coverage counts suggest this is testable.
- The group-level effect might interact with genre or annotator demographics; the paper's subset dummy dominates regressions, hinting that formal structure's role could be genre-dependent.
- The invariance of explanation-type composition across the boundary suggests that formal structure may change the difficulty of an item without changing the reasons annotators give; this distinction merits a dedicated study.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper measures, on the 3,113 SNLI/MNLI-matched items of ChaosNLI, how much formal semantic structure—operationalized by a rule-based operator and monotonicity tagger—explains human label variation. Using preregistered analyses, it reports three bounds: (1) a group-level boundary: hypotheses whose sentence-level monotonicity summary is not purely upward have higher label entropy (Cliff's δ = −0.284 for upward vs. rest) and lower majority margin, with rank-based tests claimed to survive controls for operator presence and length/complexity, though a beta-regression sensitivity check weakens the length defense; (2) an item-level ceiling: formal profiles explain 3.3–3.6% of entropy variance with a median-split AUC of 0.606; and (3) composition invariance: three preregistered contrasts on VariErr error shares and LiTEx explanation categories are null. All claims are explicitly conditioned on the low-agreement selection of ChaosNLI-S/M. The paper also discloses its preregistration audit trail, one corrected interpretation rule, and a tagged blind spot for conditional/comparative environments.
Significance. Should the findings hold, the paper is a useful calibration for perspectivist NLI: formal semantics contributes a small, reliable shift in the amount of disagreement but does not identify high-disagreement items or change the composition of disagreement. The strengths are real: all confirmatory analyses were preregistered in a version-controlled log; negative results and the failed sensitivity check are reported in full; the post-hoc correction of the R-A interpretation rule is disclosed with timing; and the reproduction repository with pinning scripts is described. This level of process transparency is rare and valuable. The principal risk is construct validity of the tagger's sentence-level monotonicity summary, which is the input to every analysis and is only indirectly validated. The paper's own limitation section and C-R4 make that risk concrete; additional stratification or validation would make the contribution solid.
major comments (3)
- [§3.2, §7 C-R4, Limitations 4] The central attribution of the group-level boundary to formal structure is not yet established because every analysis consumes the tagger's sentence-level monotonicity summary, whose agreement with MED is only 0.807 and whose accuracy inside ChaosNLI is unmeasured. C-R4 documents that 231/3,113 items contain 'if' or 'than'; these items are over-represented among the non-upward classes (13.9% of downward, 10.4% of non-monotone, 14.3% of mixed, vs. 6.2% of upward) and have higher median entropy (1.166 vs. 0.971 bits). Because the headline boundary contrasts upward against all other hypotheses, the observed δ = −0.284 could be inflated by the tagger's blind spots rather than by monotonicity per se. The manuscript states that it does not know in which direction full coverage would move the effect; that unknown is directly load-bearing for the main claim. Please report the boundary within the
- [§4.4 R-C and §7 C-R2] The abstract and §8 claim that the boundary survives preregistered reductions against length and complexity. But the preregistered beta-regression sensitivity check—the bounded-outcome model—fails the registered criterion: the marked-hypothesis coefficient loses significance in both M0 (p = 0.064) and M1 (p = 0.167). The manuscript says the defense should rest on 'rank-based results,' yet R-C itself is an OLS regression, and no rank-based test in §4 actually controls for length or parse depth. Thus the length/complexity reduction is only partially rebutted. Either add a permutation or rank-based test that controls these covariates, or explicitly downgrade the claim and report that the length defense is sensitive to the distributional model. The disclosed HC3-vs-model-based standard-error asymmetry is relevant but does not by itself satisfy the registered criterion.
- [§4.3] The binary boundary and its δ = −0.284 are described as the 'headline number' and are used in the abstract and conclusion as if they were a confirmatory result. The paper itself discloses that this binary contrast was defined after seeing the four-class result and 'is not a new hypothesis test.' The confirmatory evidence for the group-level effect consists of the four-class Kruskal-Wallis test and the two powered upward-vs-marked contrasts, not the binary test. Please present the binary contrast strictly as a descriptive effect size and remove confirmatory wording attached to δ = −0.284 in the abstract and conclusion.
minor comments (4)
- [§3.3, Table 5] The paper correctly notes a two-item discrepancy between the published MED table and the released file, and uses the released file. Please add a footnote or table caption so readers do not mistake the discrepancy for a typo in the reported validation statistics.
- [§2.2] The reference to Choi (2026) as the source of the dependency-based classification pattern is surprising in a formal-semantics paper, since that work concerns rhetorical invocation in institutional art writing. Please clarify what specific tagger machinery is adapted from that paper, or remove the citation if it is not load-bearing.
- [§3.3] For both MED agreement metrics (0.883 edit-site and 0.807 sentence-level), reporting chance-adjusted agreement (e.g., Cohen's kappa) would help interpret the sentence-level figure, which is the operative one for all downstream analyses.
- [§7.2] The disclosure of the corrected R-A interpretation rule is unusually transparent. Since readers may want to verify the timing, please include the original preregistered wording verbatim in an appendix or supplementary file, rather than paraphrasing it in the main text.
Circularity Check
No circularity: the predictor is externally validated, the outcome is independently recomputed, and the headline association is an empirical contrast, not a derivation from its own inputs.
full rationale
The paper's derivation chain does not reduce to its inputs. The monotonicity predictor (tagger v0.3) is a rule-based inventory validated against the external MED benchmark (0.883 edit-site agreement, 0.807 sentence-level agreement); it is not fitted to ChaosNLI label entropy or to the entropy contrasts of Section 4. The outcome variables are computed from released ChaosNLI 100-annotation label distributions and checked against published distributions (Section 3.4), and the group-level boundary is a preregistered Mann-Whitney/Kruskal-Wallis comparison of observed entropy across tagger classes. There is no equation in which the predicted quantity is defined in terms of the fit: the binary boundary is a hypothesis-side monotonicity contrast, and entropy is measured independently. Section 4.4's reductions (operator presence, length/complexity) are additional controls, not renaming of the outcome. The only self-citation, Choi (2026), supplies a classification pattern for discourse connectives and is not used to establish entropy differences; the tagger's validity rests on MED agreement. Limitations 4 and C-R4 flag measurement validity and coverage blind spots (e.g., conditionals/comparatives) that co-occur with higher entropy; these are genuine empirical threats to attribution, but they are threats of construct validity, not circularity, because they do not make the target result true by construction. The paper's negative results (3.3-3.6% variance explained, AUC 0.606, null composition contrasts) are also inconsistent with a circular design that would be expected to force the headline. No load-bearing self-citation chain, fitted-parameter-as-prediction, or definitional equivalence was found.
Axiom & Free-Parameter Ledger
axioms (3)
- domain assumption The rule-based tagger's sentence-level monotonicity summary is a valid proxy for the formal semantic structure of NLI items.
- domain assumption The MED benchmark's monotonicity classes transfer to the ChaosNLI items, despite differences in genre and item construction.
- domain assumption The three-way overlap of ChaosNLI, VariErr, and LiTEx is representative enough for composition-invariance tests.
read the original abstract
Human label variation in natural language inference is increasingly treated as signal rather than noise, but how much of it formal semantic structure explains has not been measured directly. We measure it on the 3,113 SNLI and MNLI items of ChaosNLI, using a rule-based operator and monotonicity tagger validated against MED (0.883 agreement at the edit site, 0.807 on the sentence-level summary our analyses consume), three preregistered analysis blocks, and full reporting of negative results. Three bounds emerge. First, a group-level boundary: hypotheses that are not purely upward monotone show reliably higher label entropy (Cliff's delta = -0.284), and rank-based tests defend the effect against operator-presence and length reductions, though a bounded-outcome sensitivity check weakens the regression form of the length defense. Second, an item-level ceiling: the same formal profiles explain only 3.3 to 3.6 percent of entropy variance and reach a median-split AUC of 0.606, too weak to identify high-disagreement items. Third, composition invariance: across the boundary, three high-powered preregistered contrasts on validated error shares and explanation-type shares (VariErr, LiTEx) all return null results. In this sample, formal semantic structure shifts how much annotators disagree by a small amount and does not detectably change what they disagree about. ChaosNLI-S/M consists of items selected for low original agreement, and every claim is conditioned on that scope. All analyses were preregistered in a version-controlled research log, whose audit trail, including one corrected interpretation rule, the paper discloses.
Figures
Forward citations
Cited by 1 Pith paper
-
Selection Shapes the Boundary: A Preregistered Replication of Monotonicity and Label Agreement in Unselected NLI Populations
A preregistered replication shows that the monotonicity–label-agreement boundary found in ChaosNLI fails to generalize to unselected SNLI/MNLI dev sets, with all effects small and slightly positive.
Reference graph
Works this paper leans on
-
[1]
Idris Abdulmumin, Mokgadi Penelope Matloga, Tadesse Destaw Belay, Botshelo Kondowe, Letlhogonolo Mohleleng, Hareaipha Nkopo Letsoalo, Shamsuddeen Hassan Muhammad, and Vukosi Marivate. 2026. https://arxiv.org/abs/2605.27239 Temporal simultaneity predicts annotation quality in sentiment corpora . Preprint, arXiv:2605.27239
Pith/arXiv arXiv 2026
-
[2]
Yoav Benjamini and Yosef Hochberg. 1995. https://doi.org/10.1111/j.2517-6161.1995.tb02031.x Controlling the false discovery rate: A practical and powerful approach to multiple testing . Journal of the Royal Statistical Society: Series B (Methodological), 57(1):289--300
arXiv 1995
-
[3]
Haram Choi. 2026. https://doi.org/10.31235/osf.io/4au72_v2 Rhetorical invocation: A four-layer computational analysis of discourse vocabulary in institutional art writing . SocArXiv preprint, version 2, posted 2026-03-24
-
[4]
Norman Cliff. 1993. https://doi.org/10.1037/0033-2909.114.3.494 Dominance statistics: Ordinal analyses to answer ordinal questions . Psychological Bulletin, 114(3):494--509
-
[5]
S. Holm. 1979. https://www.jstor.org/stable/4615733 A simple sequentially rejective multiple test procedure . Scandinavian Journal of Statistics, 6(2):65--70
arXiv 1979
-
[6]
Pingjun Hong, Beiduo Chen, Siyao Peng, Marie-Catherine de Marneffe, and Barbara Plank. 2025. https://doi.org/10.18653/v1/2025.emnlp-main.1728 L i TE x: A linguistic taxonomy of explanations for understanding within-label variation in natural language inference . In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pag...
-
[7]
Pingjun Hong, Beiduo Chen, Siyao Peng, Marie-Catherine de Marneffe, Benjamin Roth, and Barbara Plank. 2026. https://doi.org/10.18653/v1/2026.findings-acl.1342 Agree, disagree, explain: Decomposing human label variation in NLI through the lens of explanations . In Findings of the A ssociation for C omputational L inguistics: ACL 2026 , pages 26922--26934, ...
-
[8]
Chathuri Jayaweera and Bonnie J. Dorr. 2025. https://arxiv.org/abs/2507.15114 From disagreement to understanding: The case for ambiguity detection in NLI . Preprint, arXiv:2507.15114
Pith/arXiv arXiv 2025
-
[9]
Nan-Jiang Jiang and Marie-Catherine de Marneffe. 2022. https://doi.org/10.1162/tacl_a_00523 Investigating reasons for disagreement in natural language inference . Transactions of the Association for Computational Linguistics, 10:1357--1374
-
[10]
Nan-Jiang Jiang, Chenhao Tan, and Marie-Catherine de Marneffe. 2023. https://arxiv.org/abs/2304.12443 Understanding and predicting human label variation in natural language inference through explanation . Preprint, arXiv:2304.12443
Pith/arXiv arXiv 2023
-
[11]
Maximilian Maurer, Maximilian Linde, and Gabriella Lapesa. 2026. https://arxiv.org/abs/2605.06318 Who and what? using linguistic features and annotator characteristics to analyze annotation variation . Preprint, arXiv:2605.06318
Pith/arXiv arXiv 2026
-
[12]
Yixin Nie, Xiang Zhou, and Mohit Bansal. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.734 What can we learn from collective human opinions on natural language inference data? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9131--9143, Online. Association for Computational Linguistics
-
[13]
Ellie Pavlick and Tom Kwiatkowski. 2019. https://doi.org/10.1162/tacl_a_00293 Inherent disagreements in human textual inferences . Transactions of the Association for Computational Linguistics, 7:677--694
-
[14]
Barbara Plank. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.731 The ``problem'' of human label variation: On ground truth in data, modeling and evaluation . In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 10671--10682, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics
-
[15]
Romano, J
J. Romano, J. D. Kromrey, J. Coraggio, and J. Skowronek. 2006. Appropriate statistics for ordinal level data: Should we really be using t-test and C ohen's d for evaluating group differences on the NSSE and other surveys? In Annual meeting of the Florida Association of Institutional Research. Bibliographic entry adopted from the effsize package documentat...
2006
-
[16]
Michael Smithson and Jay Verkuilen. 2006. https://doi.org/10.1037/1082-989X.11.1.54 A better lemon squeezer? maximum-likelihood regression with beta-distributed dependent variables . Psychological Methods, 11(1):54--71
-
[17]
Torchiano
M. Torchiano. 2020. https://CRAN.R-project.org/package=effsize effsize: Efficient Effect Size Computation . R package version 0.8.1
2020
-
[18]
Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, and Massimo Poesio
Alexandra N. Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, and Massimo Poesio. 2021. https://doi.org/10.1613/jair.1.12752 Learning from disagreement: A survey . Journal of Artificial Intelligence Research, 72:1385--1470
-
[19]
Leon Weber-Genzel, Siyao Peng, Marie-Catherine De Marneffe, and Barbara Plank. 2024. https://doi.org/10.18653/v1/2024.acl-long.123 V ari E rr NLI : Separating annotation error from human label variation . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2256--2269, Bangkok, Thailand....
-
[20]
Hitomi Yanaka, Koji Mineshima, Daisuke Bekki, Kentaro Inui, Satoshi Sekine, Lasha Abzianidze, and Johan Bos. 2019 a . https://doi.org/10.18653/v1/W19-4804 Can neural networks understand monotonicity reasoning? In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 31--40, Florence, Italy. Association...
-
[21]
Hitomi Yanaka, Koji Mineshima, Daisuke Bekki, Kentaro Inui, Satoshi Sekine, Lasha Abzianidze, and Johan Bos. 2019 b . https://doi.org/10.18653/v1/S19-1027 HELP : A dataset for identifying shortcomings of neural models in monotonicity reasoning . In Proceedings of the Eighth Joint Conference on Lexical and Computational Semantics (* SEM 2019) , pages 250--...
-
[22]
Leixin Zhang and C a g r C \"o ltekin. 2026. https://arxiv.org/abs/2605.01168 Quantifying and predicting disagreement in graded human ratings . Preprint, arXiv:2605.01168
Pith/arXiv arXiv 2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.