{"id":"d58ddce9-fb32-429a-bd06-c78d80c1fe16","arxiv_id":"2607.15870","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Formal semantic structure explains only 3.3–3.6% of label-entropy variance in ChaosNLI-S/M and does not detectably change what annotators disagree about.","lead":"This paper measures how much formal semantic structure (monotonicity classes and operator tags) explains disagreement among annotators in natural language inference items. It finds a small, reliable group-level effect but almost no item-level predictive power or change in the kind of disagreement.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The tagger's sentence-level monotonicity summary is validated only against MED, not on ChaosNLI; its blind spots for conditionals/comparatives co-occur with the higher-entropy classes, so the headline boundary's attribution to formal structure is not yet established.","rationale":"The reader's weakest assumption is also the most load-bearing concern: the central claim attributes label-entropy differences to formal semantic structure, but the operative measurement instrument is a lexical tagger whose sentence-level summary is validated only externally (0.807 against MED) and whose in-domain accuracy on ChaosNLI is unmeasured. The paper's own C-R4 provides a concrete mechanism for concern: 7.4% of items contain untagged conditional/comparative environments, these items are more frequent among the marked classes, and they have substantially higher entropy. If this co-occurrence drives the binary boundary, Contribution 1 collapses to a tagger-artifact claim. The concern does not justify rejection: the paper is unusually transparent, the within-MNLI-m effect is still present (δ = −0.110), and the direction of the bias is not known—the tagger's blind spots could also attenuate the true effect. The proposed gold-label check would settle the question. Until it is run, the current CONDITIONAL verdict is appropriate; I would not move to ACCEPT or REJECT on the existing evidence.","tokens_in":12251,"tokens_out":9626,"duration_ms":111252,"concrete_test":"Hand-label the monotonicity profile for a stratified random sample of ~200 ChaosNLI hypotheses, oversampling the 231 items containing 'if' or 'than', following the MED class definitions. Compute the tagger's sentence-level agreement on this sample, then recompute the Section 4 upward-versus-rest entropy contrast using only gold labels. If agreement is far below 0.807, or if the gold-labeled Cliff's δ is no longer in the small/significant range (or conversely becomes materially larger), the headline effect is an artifact of tagger blind spots and the item-level ceiling is not a valid bound on formal structure.","verdict_should_be":"UNCHANGED","load_bearing_attack":"All analyses in Sections 4–6 consume the tagger's sentence-level monotonicity summary, whose agreement with MED is only 0.807 and whose accuracy on ChaosNLI is unmeasured (Limitation 4). The paper's own C-R4 shows the tagger is blind to conditional/comparative environments: 231/3,113 items (7.4%) contain 'if' or 'than'. These items are over-represented among the non-upward classes (13.9% of downward, 10.4% of non-monotone, 14.3% of mixed hypotheses, versus 6.2% of upward) and have median entropy 1.166 bits versus 0.971 for the rest. Because the headline binary boundary compares upward versus all other hypotheses, this co-occurrence means the observed δ = −0.284 could be inflated by tagger-blind formal structure rather than by monotonicity per se. The paper correctly says it does not know in which direction full coverage would move the effect; that unknown is precisely the central attribution. The 0.883 edit-site agreement does not resolve this, since the analyses do not use edit-site labels.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper measures, on the 3,113 SNLI/MNLI-matched items of ChaosNLI, how much formal semantic structure—operationalized by a rule-based operator and monotonicity tagger—explains human label variation. Using preregistered analyses, it reports three bounds: (1) a group-level boundary: hypotheses whose sentence-level monotonicity summary is not purely upward have higher label entropy (Cliff's δ = −0.284 for upward vs. rest) and lower majority margin, with rank-based tests claimed to survive controls for operator presence and length/complexity, though a beta-regression sensitivity check weakens the length defense; (2) an item-level ceiling: formal profiles explain 3.3–3.6% of entropy variance with a median-split AUC of 0.606; and (3) composition invariance: three preregistered contrasts on VariErr error shares and LiTEx explanation categories are null. All claims are explicitly conditioned on the low-agreement selection of ChaosNLI-S/M. The paper also discloses its preregistration audit trail, one corrected interpretation rule, and a tagged blind spot for conditional/comparative environments.","tokens_in":12417,"tokens_out":6482,"duration_ms":62768,"significance":"Should the findings hold, the paper is a useful calibration for perspectivist NLI: formal semantics contributes a small, reliable shift in the amount of disagreement but does not identify high-disagreement items or change the composition of disagreement. The strengths are real: all confirmatory analyses were preregistered in a version-controlled log; negative results and the failed sensitivity check are reported in full; the post-hoc correction of the R-A interpretation rule is disclosed with timing; and the reproduction repository with pinning scripts is described. This level of process transparency is rare and valuable. The principal risk is construct validity of the tagger's sentence-level monotonicity summary, which is the input to every analysis and is only indirectly validated. The paper's own limitation section and C-R4 make that risk concrete; additional stratification or validation would make the contribution solid.","major_comments":[{"comment":"The central attribution of the group-level boundary to formal structure is not yet established because every analysis consumes the tagger's sentence-level monotonicity summary, whose agreement with MED is only 0.807 and whose accuracy inside ChaosNLI is unmeasured. C-R4 documents that 231/3,113 items contain 'if' or 'than'; these items are over-represented among the non-upward classes (13.9% of downward, 10.4% of non-monotone, 14.3% of mixed, vs. 6.2% of upward) and have higher median entropy (1.166 vs. 0.971 bits). Because the headline boundary contrasts upward against all other hypotheses, the observed δ = −0.284 could be inflated by the tagger's blind spots rather than by monotonicity per se. The manuscript states that it does not know in which direction full coverage would move the effect; that unknown is directly load-bearing for the main claim. Please report the boundary within the","section":"§3.2, §7 C-R4, Limitations 4"},{"comment":"The abstract and §8 claim that the boundary survives preregistered reductions against length and complexity. But the preregistered beta-regression sensitivity check—the bounded-outcome model—fails the registered criterion: the marked-hypothesis coefficient loses significance in both M0 (p = 0.064) and M1 (p = 0.167). The manuscript says the defense should rest on 'rank-based results,' yet R-C itself is an OLS regression, and no rank-based test in §4 actually controls for length or parse depth. Thus the length/complexity reduction is only partially rebutted. Either add a permutation or rank-based test that controls these covariates, or explicitly downgrade the claim and report that the length defense is sensitive to the distributional model. The disclosed HC3-vs-model-based standard-error asymmetry is relevant but does not by itself satisfy the registered criterion.","section":"§4.4 R-C and §7 C-R2"},{"comment":"The binary boundary and its δ = −0.284 are described as the 'headline number' and are used in the abstract and conclusion as if they were a confirmatory result. The paper itself discloses that this binary contrast was defined after seeing the four-class result and 'is not a new hypothesis test.' The confirmatory evidence for the group-level effect consists of the four-class Kruskal-Wallis test and the two powered upward-vs-marked contrasts, not the binary test. Please present the binary contrast strictly as a descriptive effect size and remove confirmatory wording attached to δ = −0.284 in the abstract and conclusion.","section":"§4.3"}],"minor_comments":[{"comment":"The paper correctly notes a two-item discrepancy between the published MED table and the released file, and uses the released file. Please add a footnote or table caption so readers do not mistake the discrepancy for a typo in the reported validation statistics.","section":"§3.3, Table 5"},{"comment":"The reference to Choi (2026) as the source of the dependency-based classification pattern is surprising in a formal-semantics paper, since that work concerns rhetorical invocation in institutional art writing. Please clarify what specific tagger machinery is adapted from that paper, or remove the citation if it is not load-bearing.","section":"§2.2"},{"comment":"For both MED agreement metrics (0.883 edit-site and 0.807 sentence-level), reporting chance-adjusted agreement (e.g., Cohen's kappa) would help interpret the sentence-level figure, which is the operative one for all downstream analyses.","section":"§3.3"},{"comment":"The disclosure of the corrected R-A interpretation rule is unusually transparent. Since readers may want to verify the timing, please include the original preregistered wording verbatim in an appendix or supplementary file, rather than paraphrasing it in the main text.","section":"§7.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's disclosure practices are exemplary, and the negative results are reported with unusual honesty. The central attribution, however, rests on a tagger whose sentence-level output is the only measurement instrument and whose coverage gaps are correlated with the outcome. The requested sensitivity analyses should be feasible within the paper's scope and would materially strengthen the contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Zoe—quick take: this is worth your time. The paper does something nobody has done: it quantifies how much formal semantic structure explains NLI label variation in variance-explained and AUC terms, and it does it with a public audit trail, preregistered analyses, and full disclosure of a corrected interpretation rule. The headline result is honest: non-upward hypotheses have reliably higher entropy (Cliff's delta -0.284), but formal structure explains only ~3.5% of entropy variance, and three high-powered contrasts on disagreement composition come back null. The low ceiling is exactly the kind of negative result the perspectivist subfield needs.\n\nThe strengths are real. The tagger is validated against MED (0.883 at edit site, 0.807 sentence-level), and the paper is explicit that all analyses consume the sentence-level summary. The power-derived reporting rule, the Holm corrections, the five robustness checks, and the accident-and-recovery audit trail all signal careful empirical practice. The R-A interpretation rule correction is disclosed honestly; even though it was noticed after the null result, the paper tells you exactly what happened.\n\nNow the soft spots, in order. First, the tagger's sentence-level validity on ChaosNLI itself is unmeasured. The stress-test concern lands: 7.4% of items contain 'if' or 'than', those are over-represented among non-upward classes, and they have higher median entropy. So the boundary could be inflated by tagger blindness to conditionals/comparatives. The paper acknowledges this and says direction is unknown; that unknown is exactly the attribution problem. Second, the binary boundary was defined after inspecting the four-class result. The paper says this openly—it's not a new test, just the effect size of record—but it does mean the headline number is post-hoc. Third, the beta-regression sensitivity check fails, so the regression-based length defense is weaker than the rank-based one. The paper admits this too.\n\nNone of this kills the paper. The group-level effect is rank-based, preregistered, and survives operator-presence and length controls in that form. But the central claim is conditional on tagger coverage, and the paper's own C-R4 shows that coverage is incomplete in the very places where disagreement is higher. A referee should push for in-domain tagger validation or at least a manual error audit on a ChaosNLI subsample.\n\nVerdict: send to peer review. The field needs more of this—modest, falsifiable, transparent. I'd cite it, and I'd put it on the reading group list when the revised version lands.","headline":"A transparent, preregistered measurement that gives formal semantics a small but real group-level role in NLI disagreement; the tagger's in-domain validity is the main soft spot.","tokens_in":12977,"tokens_out":1992,"would_cite":true,"duration_ms":20139,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Formal semantic structure explains only a small slice of human label variation in NLI, yet it reliably separates high- from low-disagreement items.","keywords":["human label variation","natural language inference","monotonicity","formal semantics","annotation disagreement","ChaosNLI","label entropy","preregistered analysis"],"falsifier":"Recompute the entropy comparison using a compositional monotonicity reasoner or human expert annotation of monotonicity environments on the same ChaosNLI items, including conditionals and comparatives; if the non-upward group no longer shows higher entropy, or the variance explained rises above a few percent, the paper's ceiling and boundary claims would need revision. Alternatively, run the same preregistered protocol on an unrestricted sample of SNLI/MNLI items to see whether the group-level boundary persists outside the low-agreement selection.","tokens_in":12035,"feed_emoji":"📊","tokens_out":4472,"duration_ms":36256,"temperature":0.7,"pith_summary":"The paper asks a direct, quantitative question: how much of human label variation in natural language inference is explained by formal semantic structure? Using the 3,113 low-agreement SNLI/MNLI items of ChaosNLI, the author tags premises and hypotheses for negation, quantifiers, and monotonicity and compares label entropy across formal classes. Three bounds emerge. First, hypotheses that are not purely upward monotone have reliably higher annotator disagreement (Cliff's delta = -0.284), and this group-level gap survives controls for operator presence and sentence length. Second, at the item level the same formal profiles explain only 3.3 to 3.6 percent of the variance in entropy and reach an AUC of 0.606, so they cannot pick out high-disagreement items. Third, the composition of disagreement—shares of annotation error and of explanation types—does not detectably differ across the boundary. The paper concludes that in this sample formal semantic structure shifts how much annotators disagree by a small amount, and not detectably what they disagree about.","feed_headline":"Formal semantics explains only 3.6% of label variance in NLI","feed_subtitle":"A preregistered study of 3,113 ChaosNLI items finds monotonicity shifts how much annotators disagree, but not what they disagree about.","key_machinery":"The load-bearing instrument is a shallow, rule-based operator and monotonicity tagger over dependency parses, which assigns each sentence a four-way summary (upward, downward, non-monotone, mixed) from a closed inventory of negation, quantifier, NPI, and monotonicity triggers. What carries the argument is the binary boundary derived from it: purely upward hypotheses (no downward or non-monotone trigger) versus all others, measured against label entropy and majority margin with Cliff's delta and Holm-corrected tests. The item-level ceiling is established by the same profiles in variance-explained and AUC form.","core_discovery":"On the paper's own terms, the central discovery is a calibrated bound: formal semantic structure shifts how much annotators disagree by a small amount and not detectably what they disagree about. A rule-based monotonicity tagger splits ChaosNLI items into purely upward versus other hypotheses; the non-upward group has reliably higher label entropy (Cliff's delta = -0.284), defended against operator-presence and length controls. But the same profiles explain only 3.3 to 3.6 percent of entropy variance (AUC 0.606), and three preregistered contrasts on error shares and explanation types are null. The result is a group-level boundary with an item-level ceiling.","pith_inferences":["Because the sample is restricted to low-agreement items, the 3.3 to 3.6 percent ceiling might underestimate formal structure's explanatory power on the full range of NLI items; a test on unselected items is a natural extension.","A compositional monotonicity tagger, covering conditionals and comparatives (which the current tagger skips and which co-occur with higher entropy), could raise the item-level ceiling; the paper's own coverage counts suggest this is testable.","The group-level effect might interact with genre or annotator demographics; the paper's subset dummy dominates regressions, hinting that formal structure's role could be genre-dependent.","The invariance of explanation-type composition across the boundary suggests that formal structure may change the difficulty of an item without changing the reasons annotators give; this distinction merits a dedicated study."],"forward_implications":["Formal semantic structure belongs in the inventory of disagreement sources, but with a small, measured weight rather than a presumed large one.","Item-level disagreement prediction cannot rely on formal features alone; other item properties and annotator-side variables carry most of the signal.","The null composition contrasts suggest formal structure does not change the mix of annotation error, logical-structure conflict, or pragmatic inference behind disagreement.","The group-level boundary may attenuate or shift on unrestricted NLI samples, outside the low-agreement selection that defines ChaosNLI-S/M.","The registered audit-trail design offers a template for preregistered variance-decomposition studies of label variation."],"fun_headline_variants":["Semantics explains only 3.6% of NLI label variance","Monotonicity shifts NLI disagreement amount, not type","NLI label variance: semantics has small effect","Semantics barely explains annotator disagreement","3.6% variance: semantics' small role in NLI"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"All analyses depend on the tagger's sentence-level monotonicity summary, which agrees with the MED benchmark at only 0.807 and is unmeasured for accuracy inside the ChaosNLI domain; if that summary misrepresents the true formal structure of the items—especially in conditional and comparative environments, which the tagger skips—the group-level boundary could be an artifact of tagging blind spots rather than of formal semantics.","fun_headline_variants_meta":{"raw":{"variants":["Semantics explains only 3.6% of NLI label variance","Monotonicity shifts NLI disagreement amount, not type","NLI label variance: semantics has small effect","Semantics barely explains annotator disagreement","3.6% variance: semantics' small role in NLI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000153,"raw_usage":{"total_tokens":1099,"prompt_tokens":856,"completion_tokens":243,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":600,"completion_tokens_details":{"reasoning_tokens":162}},"tokens_in":600,"tokens_out":243,"duration_ms":2985,"temperature":1.0,"reasoning_tokens":162,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T22:04:14.348817+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the entropy comparison using a compositional monotonicity reasoner or human expert annotation of monotonicity environments on the same ChaosNLI items, including conditionals and comparatives; if the non-upward group no longer shows higher entropy, or the variance explained rises above a few percent, the paper's ceiling and boundary claims would need revision. Alternatively, run the same preregistered protocol on an unrestricted sample of SNLI/MNLI items to see whether the group-level boundary persists outside the low-agreement selection.","supporting_citations":[],"review_version":1}