{"id":"9ba13af3-3ffe-460d-984e-89741a7f5669","arxiv_id":"1909.02597","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"BERT's apparent knowledge of English NPI licensing varies across five evaluation methods, from near-perfect on gradient minimal pairs to inconsistent on absolute judgments and scope probing.","lead":"This paper compares five common ways of testing what language models know about grammar, using negative polarity words like 'any' as the test case. It finds that BERT seems to know NPI rules under some tests but not others, so the choice of test changes the conclusion.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Dataset label validity is the load-bearing assumption: weak MTurk validation, especially for Conditionals and Simple Questions, could make the measured method differences reflect template label noise rather than BERT's NPI knowledge.","rationale":"The paper is a comparative evaluation of five methods, and its central claim is that BERT possesses systematic but feature-unequal knowledge of English NPI licensing. For that claim to be true, the artificial labels must be a valid proxy for English acceptability, and the evaluation must not be gameable by template statistics. I focused on label validity because all five methods and all training configurations consume the same 136,000 labels; if those labels are noisy in specific environments, then the paper's environment-level comparisons (e.g., Simple Questions' low MCC under All-but-1) and the method-level contrasts (absolute vs. gradient) may be artifacts. The paper's own validation is too thin to rule this out: 500 sentences, 82.8% overall agreement, and clearly weak agreement in Conditionals and Simple Questions. The reader's weakest assumption flagged exactly this issue, so I agree with that assessment. The paper deserves credit for releasing data and code and for including a Wilcox et al. held-out set, but that set is used for only a subset of evaluations and BERT's scores there are not high enough to independently establish general NPI knowledge. Therefore, I do not think the paper should be rejected; the methodology and finding are plausible and worth conditional acceptance. A check that recomputes results on higher-confidence labels or correlates environment performance with human agreement would settle whether the central claim survives.","tokens_in":23116,"tokens_out":9270,"duration_ms":105470,"concrete_test":"Recruit trained linguists (or a substantially larger MTurk panel) to judge a stratified sample of at least 2,000 generated sentences, oversampling Conditionals, Simple Questions, and 'any'-sentences; then recompute each of the five methods on the subset with high inter-annotator agreement. Alternatively, compute the Spearman rank correlation across the nine environments between the per-environment human agreement and BERT's per-environment MCC and gradient-preference scores under the All-but-1 and CoLA settings; a high correlation (rho > 0.7) would indicate that BERT's apparent environment-level differences are explained by label reliability rather than by unequal NPI-licensing knowledge.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that BERT has systematic, feature-dependent knowledge of NPI licensing presupposes that the 136,000 generated sentences actually instantiate English NPI acceptability. The paper's own validation in Section 4 checks only 500 sentences, with 82.8% overall agreement and much lower agreement in environments where BERT's performance is weakest. In Table 4, Conditionals show only 50% of intended-acceptable sentences rated acceptable and 37.5% of intended-unacceptable sentences rated acceptable; Simple Questions show 33.3% of intended-unacceptable sentences rated acceptable. Because the same generated labels are used both for fine-tuning and for all five evaluation methods, high agreement among methods on the template distribution would not establish knowledge of English NPI licensing if the labels are noisy or if BERT learns generation-template correlations rather than licensing constraints. The paper's one naturalistic check, the Wilcox et al. (2019) handcrafted set, is used only for acceptability classification under some fine-tuning settings, and BERT's MCC there is modest at best (e.g., 0.55 under CoLA fine-tuning and lower under NPI fine-tuning in Figure 1), so it does not independently rescue the claim. Without confidence intervals or seed variation, the environment-level differences (e.g., MCC 0.58 for Simple Questions under All-but-1) could also be unstable. The validity of the synthetic labels is therefore the load-bearing assumption for the conclusion.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper compares five methods for evaluating grammatical knowledge in sentence encoders, using English NPI licensing as a case study. The authors introduce a generated dataset of 136,000 sentences labeled for licensor presence, NPI presence, and scope, and evaluate BERT-large (with several fine-tuning configurations) and a GloVe bag-of-words baseline via Boolean acceptability classification, absolute and gradient minimal-pair tests, an unsupervised cloze test, and feature probing. The results show that BERT performs near ceiling on gradient minimal pairs but markedly worse on absolute minimal pairs and on scope detection in probing, with systematic variation across licensing environments. The paper concludes that different evaluation methods reveal different aspects of a model's grammatical knowledge and that BERT's knowledge of NPI licensing is real but unequal across features and environments.","tokens_in":23380,"tokens_out":5908,"duration_ms":63665,"significance":"If the empirical claims hold, the paper makes a valuable methodological contribution: it provides one of the first direct comparisons of five evaluation paradigms on a single linguistic phenomenon, and it ships the generated data, generation code, and a simple baseline. The unsupervised cloze control and the handcrafted Wilcox et al. (2019) test set are useful external checks. The main risk is that the central conclusion depends on the validity of the automatically generated acceptability labels, which the paper's own validation only partially supports, and on single-run point estimates without uncertainty quantification. With the validity and stability concerns addressed, the contribution would be solid and likely influential for future work on evaluating linguistic knowledge in pretrained encoders.","major_comments":[{"comment":"The MTurk validation covers only 500 of 136,000 generated sentences, and the per-environment separation is weak exactly in the environments where BERT's performance is lowest: for Conditionals, 37.5% of sentences labeled unacceptable were rated acceptable and only 50.0% of intended-acceptable sentences were rated acceptable; for Simple Questions the corresponding numbers are 33.3% and 63.0%. Because these same labels are used for fine-tuning and evaluation in all five methods, the environment-level differences in the results (e.g., acceptability MCC 0.58 for SMP-Q under All-but-1 NPI in Figure 1) could partly reflect label noise rather than BERT's NPI knowledge. The authors should report per-environment inter-rater agreement and the exact Wilcoxon test procedure, and ideally verify the main environment-level claims on a validated subset or under alternative binarization thresholds.","section":"§4, Table 4"},{"comment":"All reported scores are single point estimates without confidence intervals or multiple runs. The paper's key claims are about relative differences—for example, near-ceiling gradient minimal-pair accuracy versus much lower absolute minimal-pair accuracy, and low scope-detection MCC in probing. It is unclear whether the environment-level differences (e.g., the drop in Simple Questions) are stable or within run-to-run noise. The authors should provide variance estimates via repeated training with different seeds or bootstrap resampling over test items.","section":"§6, Fig. 2 and Fig. 4"},{"comment":"The conclusion that 'BERT has systematic knowledge of NPI licensing' is only partially supported by the external checks. The unsupervised cloze test is a useful control, but it applies only to length-matched minimal pairs with a single token difference, and the only naturalistic test, the Wilcox et al. (2019) handcrafted set, yields modest MCC values for BERT (0.55 under CoLA fine-tuning and lower under NPI fine-tuning, Figure 1). The paper should either temper the conclusion to acknowledge that the evidence is primarily on the template-generated distribution, or add a stronger validation of the labels in the environments that drive the main claims.","section":"§4, Data validation; §6, Acceptability Judgments"}],"minor_comments":[{"comment":"There is a typo in the paragraph on Marvin and Linzen: 'sentencew' should be 'sentences'.","section":"§2, Related Work"},{"comment":"The example 'Mary has n’teaten any cookies' is missing a space; it should read 'Mary hasn't eaten any cookies.'","section":"§1, example (1)"},{"comment":"The definition of 'Diff' is not consistent with the displayed numbers: for Adverb, the difference between 61.67 and 8.33 is 53.34, not 70.00. The authors should correct the table or clarify how 'Diff' is computed.","section":"§4, Table 4 caption"},{"comment":"The paper should state explicitly whether the reported agreement percentages (81.3%, 85.2%, 82.8%) are based on per-sentence majority votes or on individual participant ratings, and should describe the sampling procedure for the 500 validated sentences in more detail.","section":"§4, Data validation"},{"comment":"The cloze test uses only minimal pairs that are length-matched and differ in exactly one token; this restriction should be stated in the main text when interpreting the cloze results, since it limits the comparison with the other minimal-pair conditions.","section":"§6, Cloze Test"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a solid and useful comparative study, but the central empirical claim rests on the validity of the generated labels and on single-run results. I do not see grounds for rejection; the paper is publishable after the data-validation analysis is strengthened, uncertainty is quantified, and the conclusions are calibrated to the strength of the evidence. The low per-environment agreement in the two environments where BERT also performs worst is particularly important to address."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The main thing to know: this paper demonstrates that which evaluation method you choose changes what you conclude about BERT's grammatical knowledge. That is a real result, not a slogan. Gradient minimal pairs make BERT look nearly perfect on NPI licensing, while absolute judgments and probing show clear gaps, especially for scope. The authors compare five methods on one phenomenon, with multiple training/evaluation splits and several model variants, and they release the data and code. The unsupervised cloze test is a nice addition because it doesn't depend on the generated labels at all, and it broadly agrees with the supervised results.\n\nThe soft spots are real but mostly affect the quantitative details. The synthetic dataset is the load-bearing part: 136k sentences, but only 500 were validated by MTurk, with 82.8% agreement overall. In Conditionals and Simple Questions, a third to half of the intended-unacceptable sentences were rated acceptable, which is noisy enough to blur environment-level differences. The paper lacks confidence intervals and repeated runs, so point estimates like the 0.58 MCC for Simple Questions might not be stable. There is also partial circularity: the labels are generated from the same template features used to probe the model. That said, the cloze results and the modest external check on the Wilcox et al. handcrafted set provide some independent support, though not enough to rescue precise quantitative claims.\n\nThe central qualitative conclusion—that no single method gives a complete verdict and that BERT's knowledge is systematic but feature-dependent—holds up. The paper is honest about its limitations, including the free-choice 'any' over-acceptance, and it does not overclaim. For someone designing evaluation benchmarks or studying what pretrained models know about grammar, this is a genuinely useful study. I would cite it, and I would send it to peer review: it deserves serious refereeing, with the main request being more careful validation of the synthetic labels and error bars on the headline numbers.","headline":"Useful multi-method comparison shows evaluation method shapes conclusions about BERT's NPI knowledge, despite weak label validation and missing error bars.","tokens_in":23940,"tokens_out":1724,"would_cite":true,"duration_ms":21571,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"BERT has systematic but uneven knowledge of NPI licensing, and the evaluation method changes what knowledge appears.","keywords":["BERT","negative polarity items","acceptability judgments","minimal pairs","probing classifiers","masked language modeling","grammatical knowledge","evaluation methodology"],"falsifier":"Replace the template labels with a fully human-validated set in which every sentence is rated by multiple native speakers, then rerun the five methods: if the discrepancies between gradient and absolute results disappear, those discrepancies were artifacts of label noise; if they persist, they reflect genuine properties of BERT.","tokens_in":22928,"feed_emoji":"🧠","tokens_out":5008,"duration_ms":48335,"temperature":0.7,"pith_summary":"This paper asks whether BERT, a sentence representation model, knows the grammar of negative polarity items (NPIs)—words like 'any' and 'ever' that are only acceptable in negative or similar licensing environments. The authors build a generated dataset of 136,000 sentences that independently manipulate three licensing features (presence of a licensor, presence of an NPI, and whether the NPI falls inside the licensor's scope) across nine licensing environments, and then evaluate BERT with five different methods: Boolean acceptability classification, absolute minimal pairs, gradient minimal pairs, cloze tests, and feature probing. A sympathetic reader would care because the paper argues that the answer to 'does BERT know this grammar?' is not a single yes or no: BERT shows near-perfect sensitivity on gradient comparisons, but absolute judgments and probing reveal that its knowledge is uneven across features and environments. The conclusion is methodological: a single evaluation task can misrepresent a model's grammatical knowledge, so multiple complementary methods are needed.","feed_headline":"Five tests show BERT's grammar knowledge is uneven","feed_subtitle":"The same model looks near-perfect on NPI licensing under one test and flawed under another. Method choice changes the verdict.","key_machinery":"The load-bearing object is a generated dataset of 136,000 English sentences organized into nine NPI licensing environments (adverbs, conditionals, determiner negation, sentential negation, only, quantifiers, questions, simple questions, superlatives), each following a 2×2×2 paradigm over three Boolean features: licensor presence, NPI presence, and scope. Each paradigm yields minimal pairs isolating one feature at a time, with a reduced 2×2 paradigm for simple questions. Around this dataset, the paper constructs five evaluation methods: a Boolean acceptability classifier, an absolute minimal-pair test, a gradient minimal-pair test, an unsupervised cloze test using BERT's masked language modeling head, and probing classifiers that predict the three metadata features from frozen representations. The dataset is what makes the method comparison possible, and the divergence among the five methods is the evidence for the paper's main conclusion.","core_discovery":"The central claim is that BERT has systematic knowledge of all the features needed to judge NPI sentences—licensor presence, NPI presence, and scope—but that this knowledge is not equal across features and does not behave like the Boolean acceptability contrast it supports. On gradient minimal pairs, where the model only has to rank the acceptable sentence above the unacceptable one, BERT is near perfect across almost all environments; the bag-of-words baseline also reaches ceiling on detecting licensor and NPI presence, apparently through word co-occurrence. Under the stricter absolute minimal-pair measure and in probing classifiers, BERT's performance drops, and scope detection—knowing whether the NPI is inside the licensor's syntactic scope—is the weakest ability. Additional fine-tuning on NLI or CCG data does not improve BERT's NPI knowledge, contrary to what prior work on structurally supervised LSTMs would suggest.","pith_inferences":["If method-dependence generalizes beyond NPIs, then 'a model knows grammar X' is best understood as a relation between the model and a family of probes, not a fixed internal fact (this is my editorial inference, not the paper's claim).","A natural extension is to run the same five-method battery on other non-local dependencies, such as reflexive licensing or filler-gap dependencies, to test whether the observed asymmetry between gradient and absolute measures is a general property of transformer language models.","The data validation results suggest a testable correction: building a fully human-validated subset of the generated sentences and rerunning all five methods would show whether the method discrepancies shrink when label noise is removed."],"forward_implications":["A near-perfect score on gradient minimal pairs is not enough to conclude that a model has categorical grammatical knowledge, because absolute judgments reveal gaps.","The same model can look near-perfect under one method and substantially weaker under another, so claims about grammatical knowledge should report the evaluation method.","BERT's scope detection is the weakest of the three NPI features; tasks that isolate long-range structural dependencies will expose limitations that lexical cues hide.","Co-occurrence statistics can explain part of the apparent NPI knowledge: the bag-of-words baseline detects licensor and NPI presence at ceiling while failing on scope.","Intermediate fine-tuning on natural language inference or CCG data does not transfer to NPI licensing for BERT, so structural supervision gains found for LSTMs do not automatically carry over."],"supporting_citations":[{"why":"Provides BERT, the pretrained model whose grammatical knowledge is under evaluation.","marker":"Devlin et al. (2018)"},{"why":"Supplies the minimal-pair evaluation method and the earlier finding that LSTM language models do not systematically prefer licensed NPIs.","marker":"Marvin and Linzen (2018)"},{"why":"Supplies the handcrafted NPI test set, the structural-supervision comparison, and the baseline that BERT is tested against.","marker":"Wilcox et al. (2019)"},{"why":"Supplies CoLA, the acceptability-judgment corpus used to train the supervised classifiers.","marker":"Warstadt et al. (2018)"},{"why":"Inspires the automatically generated diagnostic data and the probing approach for testing compositional knowledge.","marker":"Ettinger et al. (2016, 2018)"},{"why":"Provides the GloVe embeddings used to build the bag-of-words baseline.","marker":"Pennington et al. (2014)"},{"why":"Introduces probability-based minimal-pair evaluation of grammatical knowledge in language models.","marker":"Linzen et al. (2016)"},{"why":"Supplies the probing methodology used to read grammatical features from frozen representations.","marker":"Tenney et al. (2019)"}],"fun_headline_variants":["BERT's grammar IQ flips depending on the test","Method choice flips BERT's grammar verdict","BERT's scope detection lags in NPI tests","Five methods yield different BERT grammar scores","BERT's NPI mastery depends on the test used"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The template-generated labels correctly instantiate real English NPI acceptability; if the labels are noisy or confounded, the measured differences between methods could reflect label artifacts rather than BERT's grammatical knowledge.","fun_headline_variants_meta":{"raw":{"variants":["BERT's grammar IQ flips depending on the test","Method choice flips BERT's grammar verdict","BERT's scope detection lags in NPI tests","Five methods yield different BERT grammar scores","BERT's NPI mastery depends on the test used"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000691,"raw_usage":{"total_tokens":3100,"prompt_tokens":888,"completion_tokens":2212,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":504,"completion_tokens_details":{"reasoning_tokens":2153}},"tokens_in":504,"tokens_out":2212,"duration_ms":15761,"temperature":1.0,"reasoning_tokens":2153,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:45:12.376709+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Replace the template labels with a fully human-validated set in which every sentence is rated by multiple native speakers, then rerun the five methods: if the discrepancies between gradient and absolute results disappear, those discrepancies were artifacts of label noise; if they persist, they reflect genuine properties of BERT.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the minimal-pair evaluation method and the earlier finding that LSTM language models do not systematically prefer licensed NPIs."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the handcrafted NPI test set, the structural-supervision comparison, and the baseline that BERT is tested against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Inspires the automatically generated diagnostic data and the probing approach for testing compositional knowledge."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the GloVe embeddings used to build the bag-of-words baseline."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces probability-based minimal-pair evaluation of grammatical knowledge in language models."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the probing methodology used to read grammatical features from frozen representations."}],"review_version":1}