{"id":"003c10a2-b7dd-4131-84fe-ad0d22374332","arxiv_id":"1908.05253","paper_version":3,"verdict":"REJECT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":6,"one_line_summary":"A large-scale neg-raising judgment dataset and a latent-factor model suggest both verbs and complement structures matter, but the structural component is not statistically reliable.","lead":"The authors collected a new large-scale dataset of neg-raising judgments for 925 English clause-embedding verbs and fit a latent-factor model to see whether the inference comes from the verb or from the structure of its complement. The dataset is a usable resource, but the paper's headline conclusion that sentence structure independently drives some neg-raising inferences is not supported by its own model comparisons.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's own model comparison does not support the structural-source claim: the pure lexical model has the lowest held-out loss and the top three models are not reliably different.","rationale":"The paper contributes a large, carefully constructed dataset and a model that extends White and Rawlins' s-selection framework, and the validation experiments show that the measurement method tracks canonical neg-raising verbs. However, the theoretical conclusion is exactly the part that fails. Section 5 states that the three best held-out models are, in order, pure lexical, pure structural, and joint, and that pairwise differences among them are not reliable. The joint model is selected because it has the best fit among the three, but with held-out losses that are not reliably different, the data cannot distinguish lexical from structural sources. The latent structural factor is also uninterpreted: P(pi_tf) in Table 3 is fit to the neg-raising judgments themselves, so it could absorb any frame-level bias, including a per-frame intercept, rather than a named grammatical property. This double problem, non-significant model comparison plus an uninterpreted latent factor, makes the abstract's claim load-bearing and unsupported. The reader's weakest_assumption targeted the interpretability issue; my primary concern is the statistical indecisiveness, so agreement is partial. A conservative model-selection reanalysis would settle whether the joint model is genuinely preferred, and if not, the appropriate finding is that the data support only a lexical-source model, with the structural-source claim unsubstantiated.","tokens_in":16027,"tokens_out":3930,"duration_ms":36728,"concrete_test":"Re-run the five-fold cross-validation and apply the one-standard-error rule to select among |I|,|T| in {0,...,4}: if the chosen model is |I|=1,|T|=0 (or any model without structural properties), the claim that subordinate-clause structure independently contributes to neg-raising is not supported. Report per-fold paired differences in held-out weighted KL loss between |I|=1,|T|=0 and |I|=1,|T|=1; the joint model must be consistently better across folds to license the abstract's conclusion.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim, that some neg-raising inferences are attributable to subordinate clause structure, is not supported by the paper's own statistics. Section 5 reports that the three best held-out models are |I|=1,|T|=0 (pure lexical), |I|=0,|T|=1 (pure structural), and |I|=1,|T|=1 (joint), and that 'none of these models' performance is reliably different from the others.' The pure lexical model has the lowest held-out loss according to the order listed. The paper nevertheless selects |I|=|T|=1 and concludes that structural properties contribute. Because the held-out loss differences are within noise and the pure lexical model is numerically better, the data do not license the binary attribution made in the abstract. Additionally, the structural factor P(pi_tf) in Table 3 is a per-frame fitted latent parameter, not tied to finiteness, overt subjecthood, or eventivity independently of the neg-raising judgments. Thus the abstract's claim fails on statistical grounds before interpretability is even considered.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper collects a large-scale dataset of neg-raising judgments for 925 English clause-embedding verbs across six syntactic frames, two matrix tenses, and two matrix subjects. It extends White and Rawlins' boolean matrix factorization model of s-selection with latent lexical properties and latent structural properties, and fits variants with different numbers of lexical and structural properties using held-out weighted KL loss. On the basis of the fitted model with one lexical and one structural property, the paper claims that some neg-raising inferences are attributable to properties of particular predicates while others are attributable to subordinate clause structure (abstract, §5, §7).","tokens_in":16280,"tokens_out":3857,"duration_ms":38729,"significance":"If the central claim were supported, the paper would provide a novel large-scale empirical constraint on theories of neg-raising, and the dataset of 7,936 verb-tense-frame-subject items would be a useful resource. The validation experiments in Appendix C and the care taken with stimulus construction are strengths. However, the two load-bearing premises — that the model comparison favors a joint lexical/structural model, and that the fitted latent 'structural property' corresponds to an identifiable property of subordinate clause structure — are not established by the paper's own analyses. Since the abstract's main finding depends on both premises, the contribution as stated is not supported.","major_comments":[{"comment":"The model comparison reported in §5 does not support the claim that subordinate clause structure contributes to neg-raising. The paper lists the three best held-out models (Figure 2, starred) 'in order': |I|=1, |T|=0 (pure lexical), |I|=0, |T|=1 (pure structural), and |I|=1, |T|=1 (joint), and states that none is reliably different from the others. The subsequent sentence, 'Among these three, the model with the best fit to the dataset has |I|=1 and |T|=1,' is contradicted by that ordering; and even if the joint model were numerically best, the reported 95% confidence intervals for pairwise differences imply the difference is within noise. The abstract's attribution of some inferences to subordinate clause structure is therefore not licensed by the held-out loss comparison.","section":"§5, Results"},{"comment":"The structural property is a latent factor fitted to the neg-raising judgments themselves: P(πtf) is a free parameter in Eq. (21), optimized against the same responses it is later used to explain. Table 3 therefore reports fitted frame-level values of this latent factor, not evidence that any independently identifiable syntactic property (finiteness, overt subject, eventivity) drives neg-raising. With |I|=|T|=1, the model has only one structural factor, and the six nonzero P(πtf) values in Table 3 are, in effect, a per-frame latent intercept. Nothing outside the model confirms the mapping from this factor to 'subordinate clause structure' asserted in the abstract, so the central claim rests on an interpretation that the modeling framework does not provide.","section":"§4, Eq. (21); §6, Table 3"},{"comment":"The analysis of the selected model does not provide independent evidence for two distinct sources of neg-raising. Table 2 reports P(φijk) and P(ωtjk) values that are all near 1 and nearly identical; Footnote 9 explains that the optimizer sets these to the square root of the largest expected neg-raising value. Footnote 10 further states that P(ψvi) and P(λvt) are effectively a single parameter p. As a result, the decomposition into 'lexical' and 'structural' sources is not identified: the same predictions could be obtained by a single per-verb propensity combined with a per-frame effect, which is exactly what the pure-lexical boundary model provides. The paper does not supply an identifiability or model-comparison test that would distinguish these alternatives beyond the reported, non-significant held-out loss differences.","section":"§6, Table 2 and Footnotes 9–10"}],"minor_comments":[{"comment":"The description of the |T|=0 boundary model, which fixes Λ=1|V| and Π=1|F|, is formally unclear if the tensor dimensions are |V|×0 and 0×|F|; the boundary model should be defined directly, e.g., by omitting the λ and π factors from Eq. (21).","section":"§4, boundary models"},{"comment":"The fold-construction constraint that every verb-frame combination appears at least once in the training folds makes the held-out loss closer to interpolation over frames than to prediction of unseen verb-frame combinations; this should be stated when interpreting the model comparison.","section":"§5, Method"},{"comment":"The y-axis label reads 'Held-out weighted KL divergence (nats)' while the text and caption describe the sum of the weighted KL loss across all five folds; the figure should make explicit whether one fold or the summed value is plotted.","section":"Figure 2"},{"comment":"The reference 'Pollack 1976' may be a misspelling of 'Pollock'; please verify the cited author's name.","section":"Reference list"}],"recommendation":"reject","confidential_remarks":"The paper's own model-selection results, as reported, contradict the selection of |I|=|T|=1, and the latent 'structural property' is not tied to any independently measured syntactic feature. The abstract's central claim therefore appears unsupported. I would be willing to reconsider a substantially revised analysis that resolves the model-selection ambiguity and demonstrates that the structural factor is identifiable and interpretable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear X,\n\nThe short version: the dataset is the real contribution; the theoretical conclusion is not. I largely agree with the reader's take, though I'd put a bit more weight on the empirical resource.\n\nWhat's genuinely good: they collect neg-raising judgments for 925 clause-embedding verbs across six frames, two tenses, two subjects, with acceptability weights. That's a lot of careful work. The validation experiments against literature-known neg-raisers (think, believe, want) and non-neg-raisers (know, say, try) come out clean. The data is released. Future work on the lexical semantics of clause-embedding verbs will use this.\n\nThe model extension of White & Rawlins is straightforward: add a per-verb lexical property and a per-frame structural property, both latent, and let them jointly license neg-raising. That's a reasonable thing to try. The cross-validation setup is sound as far as I can tell.\n\nThe soft spot is exactly where the reader points. The paper's own Section 5 says the three best models—pure lexical, pure structural, and joint—are not reliably different in held-out loss, and the pure lexical model is numerically the best. Then the abstract says some inferences are attributable to properties of predicates and others to subordinate clause structure. That simply doesn't follow. The joint model's point estimate being slightly better is not evidence for a structural source when the difference is within noise. The paper should either report the comparison honestly and say 'we cannot distinguish these accounts,' or argue for the joint model on some other grounds.\n\nThe second problem is interpretability. The 'structural property' is a latent factor fitted to the same neg-raising judgments. The per-frame values in Table 3 are exactly what you'd get from a frame-level random intercept. The paper tries to link it to finiteness, overt subject, eventivity, but nothing constrains the latent factor to correspond to any named syntactic feature. Even the model's own footnote 9 shows the lexical and structural parameters collapse into a single effective parameter, so the two 'sources' are not separately identifiable in the fitted model. That makes the 'grammatical source' claim doubly unsupported.\n\nSo: the empirical resource deserves to be published, and the model is a reasonable baseline. But the current framing oversells what the statistics show. If the authors reframe the contribution as a dataset plus a modeling exercise, and present the model comparison as 'no evidence for a structural source beyond lexical knowledge'—or at least as inconclusive—the paper is publishable. As is, I would not accept the abstract's claim.\n\nI'd still send it to peer review—an editor should not desk-reject a paper with this much careful empirical work and a clear, fixable statistical overreach. With major revisions, it could become a solid contribution.","headline":"A valuable new dataset and a model that is honestly reported, but the structural-source conclusion is not supported by the paper's own model comparison.","tokens_in":16825,"tokens_out":2943,"would_cite":false,"duration_ms":28892,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Neg-raising inferences come from both the verb and the clause it embeds, a new model of 925 English verbs suggests.","keywords":["neg-raising","clause-embedding verbs","semantic selection","boolean matrix factorization","lexical semantics","subordinate clause structure","probabilistic model","entailment judgments"],"falsifier":"A permutation test that shuffles verbs across frames while preserving each verb's estimated lexical property should leave the frame-level structural parameters $\\pi_{tf}$ unchanged if those parameters reflect real structure; if the $\\pi_{tf}$ values collapse toward equality under such shuffling, the structural effect is an artifact of which verbs happen to appear in each frame.","tokens_in":15770,"feed_emoji":"💬","tokens_out":9912,"duration_ms":82836,"temperature":0.7,"pith_summary":"This paper investigates neg-raising, the phenomenon where negation on a verb like 'think' is interpreted inside its embedded clause: 'Jo doesn't think Bo left' implies 'Jo thinks Bo didn't leave'. The authors collect a large dataset of neg-raising judgments for 925 English clause-embedding verbs across six syntactic frames, two tenses, and two subject persons, and fit a statistical model that jointly learns which properties of verbs and which properties of subordinate clauses license the inference. The best-fitting model has exactly one lexical property and one structural property, and it attributes some neg-raising inferences to the predicate itself and others to the structure of the clause it embeds. This is a new empirical constraint on theories of neg-raising, which have traditionally tied the inference to individual predicates and not to complement structure.","feed_headline":"Neg-raising springs from verbs and clause structure alike","feed_subtitle":"A 7,936-item judgment study of 925 English verbs shows complement structure shapes the inference.","key_machinery":"The key object is a probabilistic boolean matrix factorization model that extends a prior s-selection model from the literature. It factorizes a tensor of neg-raising judgments $N$ into a verb–lexical-property matrix $\\Psi$, a lexical-property–structure–subject–tense tensor $\\Phi$, a verb–structure-type matrix $\\Lambda$, a structure-type–frame matrix $\\Pi$, and a structure-type–subject–tense tensor $\\Omega$, via the approximation $n_{vfjk} \\approx \\bigvee_{t,i} \\lambda_{vt} \\wedge \\psi_{vi} \\wedge \\varphi_{ijk} \\wedge \\pi_{tf} \\wedge \\omega_{tjk}$. The number of lexical properties $|I|$ and structural properties $|T|$ are hyperparameters; cross-validation selects $|I|=|T|=1$ as the best trade-off between fit and generalization.","core_discovery":"The central discovery is that neg-raising is not purely a product of lexical knowledge. In a five-fold cross-validation, the model that generalizes best posits one lexical property and one structural property; models with only lexical properties or only structural properties perform comparably, but the joint model is the best fit. The fitted parameters show that the lexical property and the structural property each contribute near-deterministically to neg-raising across subjects and tenses, while variability across verbs is captured by the probability that a verb has the lexical property and selects the structural property. Differences across syntactic frames are captured by the probability that the structural property maps onto each frame. The authors interpret this as evidence that properties of the subordinate clause—such as finiteness, presence of an overt subject, or eventivity—influence whether neg-raising is triggered, a role that prior proposals do not assign to complement structure.","pith_inferences":["If the structural property is a real syntactic feature, neg-raising diagnostics such as negative-polarity-item licensing and Horn-clause behavior should differ between finite and non-finite complements of the same verb; this is a testable consequence the paper does not draw.","Because the single lexical and structural properties are almost perfectly correlated, the model may be reducible to a one-dimensional 'neg-raising propensity' per verb modulated by frame-specific weights; a simpler model should be tested before accepting two separate sources.","The authors' treatment of subject and tense variation as noise suggests that first-person present-tense effects such as 'I don't know' are pragmatic speaker-commitment effects rather than lexical semantic facts, which could be tested with controlled decontextualized prompts."],"forward_implications":["Theories of neg-raising must explain why subordinate clause structure modulates the inference, not just which predicates trigger it.","The same verb can differ in neg-raising behavior across complement types; the model predicts this is systematic, not noise.","Subject and tense variability in neg-raising is captured as predicate-specific idiosyncrasy, supporting pragmatic rather than purely syntactic accounts of that variability.","The dataset and model provide a reusable tool for measuring neg-raising at scale and for jointly modeling related lexically triggered inferences such as veridicality."],"supporting_citations":[{"why":"Supplies the s-selection model that is extended and the MegaAcceptability dataset used to construct the neg-raising items.","marker":"White and Rawlins (2016)"},{"why":"Provides the likelihood-of-inference measurement method, adapted here for neg-raising judgments.","marker":"White et al. (2018)"},{"why":"Another source of the veridicality measurement and clause-selection theory that the neg-raising measure is analogous to.","marker":"White and Rawlins (2018)"},{"why":"Supplies the excluded-middle inference account and the list of neg-raising predicates used to validate the measurement.","marker":"Gajewski (2007)"},{"why":"Classic characterization of neg-raising as tied to particular predicates, the view the new structural result challenges.","marker":"Horn (1978)"}],"fun_headline_variants":["Neg-raising: verb meaning and clause shape both matter","Complement structure drives neg-raising inferences","Neg-raising: not just lexical, structure also plays a role","Study: clause structure influences neg-raising alongside verbs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's central claim depends on interpreting the model's single 'structural property' as a real property of subordinate clause structure, such as finiteness or eventivity, rather than as a free latent intercept that merely absorbs frame-level differences in the data.","fun_headline_variants_meta":{"raw":{"variants":["Neg-raising: verb meaning and clause shape both matter","Complement structure drives neg-raising inferences","Neg-raising: not just lexical, structure also plays a role","Study: clause structure influences neg-raising alongside verbs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000269,"raw_usage":{"total_tokens":1538,"prompt_tokens":777,"completion_tokens":761,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":393,"completion_tokens_details":{"reasoning_tokens":701}},"tokens_in":393,"tokens_out":761,"duration_ms":7894,"temperature":1.0,"reasoning_tokens":701,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:19:52.362360+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A permutation test that shuffles verbs across frames while preserving each verb's estimated lexical property should leave the frame-level structural parameters $\\pi_{tf}$ unchanged if those parameters reflect real structure; if the $\\pi_{tf}$ values collapse toward equality under such shuffling, the structural effect is an artifact of which verbs happen to appear in each frame.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the s-selection model that is extended and the MegaAcceptability dataset used to construct the neg-raising items."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Classic characterization of neg-raising as tied to particular predicates, the view the new structural result challenges."}],"review_version":1}