{"id":"a0284451-deca-4b01-944a-ef1826722ea9","arxiv_id":"2412.11318","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A new dataset and surprisal-based metric estimate that roughly 20% of naturally occurring generic sentences are weak generalisations and that generics are more context-sensitive than explicit quantifiers.","lead":"This paper introduces a dataset of naturally occurring generic sentences in context and a surprisal-based metric to infer which quantifier (all, most, some, or none) best fits each sentence. Using language models, it finds that about one in five generics expresses a weak generalisation and that generics respond to context more than explicit quantifiers do.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The p-acceptability metric is not calibrated against human quantifier judgments; the §5.1 original-label validation can be satisfied by surface-form familiarity, so the 20% weak-generic estimate and the context-sensitivity asymmetry are not yet supported.","rationale":"I read the paper as making a valuable empirical contribution: a new naturally occurring corpus, a new surprisal-based metric, and several interpretable findings about generics. The strongest claims, however, are only as good as the p-acceptability metric, and the paper does not establish that the metric tracks human quantifier semantics. The reader's weakest assumption identifies exactly this gap. I would sharpen it: the only validation against original labels is especially weak for explicit quantifiers, because reinserting the original quantifier restores the exact observed surface string and can therefore win by frequency or memorization rather than semantics. This does not invalidate the generic-specific claims, since generics have no original quantifier to recover, but it means the metric's calibration for implicit quantification is unknown. The context-sensitivity result is partly protected by the random-context control in Appendix F, which is a point in the paper's favor, but without human ratings the asymmetry could still be a generic-surface-form artifact rather than a semantic effect. The proposed human study would directly test the load-bearing assumption: if agreement with human quantifier-fit judgments is high, the central claims gain real support; if agreement is at chance, the 20% weak-generic estimate and the context-sensitivity conclusion should be revised. For these reasons I do not move the reader's verdict; the paper should remain conditional pending this calibration.","tokens_in":16504,"tokens_out":7461,"duration_ms":73122,"concrete_test":"Run a human-calibration study on a stratified sample of 150 CONGEN generics (50 with p-acceptable all, 50 with most, 50 with some under Mistral-7B) plus 50 explicit-quantifier controls. For each item, show the original context and the sentence, and have at least 5 native speakers choose which of all/most/some/generic best captures the intended generalisation, or rate the naturalness of each quantized variant on a 1-5 scale. Compare the metric's choice to the modal human choice using percentage agreement and Cohen's kappa. If agreement is not significantly above chance (e.g., majority agreement below 50% or kappa not above 0.2), the §5.2 weak-generic estimate and §5.3 context-sensitivity asymmetry are unsupported. As a secondary check, compare validation accuracy on DOLMA versus Reddit-2024 subsets to test whether explicit-quantifier recovery is driven by memorized surface strings.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claims rest on Definition 4.1's p-acceptability, but the metric is never calibrated against human quantifier judgments. The §5.1 validation is weaker than it appears: for sentences originally labeled all/most/some, the candidate that deletes and reinserts the original quantifier recreates the exact surface string (with context) that the model has likely seen or that matches corpus distribution, so recovering the original label can reflect surface-form familiarity rather than quantificational semantics. For generic sentences there is no original quantifier to recover, so the 20% weak-generic estimate (§5.2) and the context-sensitivity asymmetry (§5.3) rest on an unverified assumption that lower property-token surprisal tracks human perceived quantificational strength. The fact that 'some' is the semantically weakest candidate also makes it a default choice if the model is uncertain, potentially inflating the weak-generic count. Section 8 acknowledges classifier and annotator biases, but not this calibration gap.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces CONGEN, a new dataset of 2,873 naturally occurring bare-plural generic and explicitly quantified sentences (all/most/some) with linguistic context, and proposes p-acceptability, a surprisal-based metric that selects the quantifier (including the null generic) that minimizes the surprisal of the property tokens after the verb. Using three Mistral models, the paper reports three main findings: (i) p-acceptability recovers intuitive quantifier patterns on CONGEN and GenericsKB-BP, (ii) generics are more context-sensitive than determiner quantifiers because increasing left context improves p-acceptability accuracy more for generic than for quantified sentences, and (iii) about 20% of naturally occurring generics are weak generalisations, operationally defined as those whose p-acceptable quantifier is 'some'. The paper also applies the metric to stereotype sentences, finding that negative stereotypes are predominantly quantified as 'all' and that the 'people who are' paraphrase reduces this tendency.","tokens_in":16745,"tokens_out":4085,"duration_ms":36916,"significance":"The paper addresses a real gap: existing generic datasets are synthetic or lack context, and the proposed CONGEN dataset is a potentially valuable resource for studying generics in naturalistic settings. The p-acceptability metric is clearly defined and the paper includes useful controls (whole-sequence versus property-token surprisal in Appendix C, random-context control in Appendix F). If the metric were shown to track human quantificational judgments, the reported findings on weak generics and context-sensitivity would be of substantial interest to both theoretical semantics and NLP. However, the current manuscript does not provide that validation, and the core claims therefore rest on an unverified assumption. The paper is honest about several limitations in Section 8, but it does not flag the absence of human calibration, which is the most load-bearing gap.","major_comments":[{"comment":"The p-acceptability metric is never calibrated against human quantifier judgments, and the validation in §5.1 is vulnerable to a surface-form familiarity confound. For each sentence, the variation that reinserts the original quantifier (or preserves the generic's null quantifier) reproduces the original surface string, which is likely the most probable string in the model regardless of quantificational semantics. This is especially acute for generic sentences, where the 'original label' is the absence of a quantifier, i.e., the original surface form itself. Consequently, the fact that p-acceptability recovers original labels does not establish that lower property-token surprisal corresponds to the semantically most appropriate quantifier. A human judgment baseline (e.g., paraphrase acceptance or quantifier-strength ratings) on a sample of sentences is needed to support the metric's validity, and the downstream claims in §§5.2–5.4 currently stand or fall with this missing calibration.","section":"§4.2, Definition 4.1; §5.1, Figure 1"},{"comment":"The operationalization of weak generics as 'generics whose p-acceptable quantifier is some' is an unvalidated axiom. The paper provides no independent evidence that the sentences so classified are actually perceived by speakers as expressing weak generalisations, nor that 'some' is the appropriate paraphrase rather than, say, 'many' or 'often'. Because 'some' is the semantically weakest candidate, it may also act as a default when the model has low confidence across all candidates, artificially inflating the weak-generic count. The abstract's claim that 'about 20% of naturally occurring generics ... express weak generalisations' therefore requires either a human annotation study on a subset of the classified sentences or a demonstration that the p-acceptability choice is not driven by uncertainty. Without such evidence, the 20% figure is a model-internal statistic rather than an empirical claim about language use.","section":"§5.2, Figure 2, text after Eq. (1)"},{"comment":"The context-sensitivity claim is measured as improvement in accuracy relative to original labels, not as a comparison with human context-sensitivity judgments. For generic sentences, the 'correct' answer is the original surface form, and adding more context may simply increase the model's probability for the exact original sentence, producing a spurious 'context effect' that reflects familiarity rather than semantic sensitivity. The random-context control in Appendix F rules out some generic benefits of context, but it does not remove the surface-form confound because the accuracy metric still uses the original label as ground truth. To support the conclusion that 'generics are more context-sensitive than determiner quantifiers', the paper needs either a human baseline for how context changes perceived quantificational strength, or a metric whose validity has been established independently of original-label recovery. In addition, the arbitrary choice of 4-token context chunks (Table I.6) is not motivated; a sensitivity analysis over chunk sizes would help establish robustness.","section":"§5.3, Figures 3 and 4"},{"comment":"The stereotype experiment interprets p-acceptability results as evidence about human bias in generic language (e.g., 'negative stereotypes are overwhelmingly implicitly quantified as universals'). This interpretation inherits the same calibration gap: without knowing whether p-acceptability tracks human quantificational strength, the model's 'all' preference for negative stereotypes could reflect corpus statistics or the metric's design rather than the psychological strikingness effect posited by Cimpian et al. (2010). Section 8 acknowledges classifier and annotator biases and the synthetic nature of the stereotype stimuli, but it does not acknowledge that the central metric has not been validated against human judgments. A focused validation study on stereotype sentences—e.g., collecting human ratings of how quantifier-like each stereotype generic feels—would substantially strengthen this section.","section":"§5.4 (stereotypes) and §8 (Limitations)"}],"minor_comments":[{"comment":"The text contains a typo: 'indentifies' should be 'identifies' in the sentence defining the purpose of p-acceptability.","section":"§4.2"},{"comment":"The word 'p-acceptablility' appears instead of 'p-acceptability' in the experimental setup paragraph.","section":"§5.3"},{"comment":"The example in Table I.6 is a sentence originally labeled 'All' ('All wolf spiders are sensitive to vibrations in the ground'), but the table caption labels it as a generic sample; please clarify the mapping between the example and the original-quantifier categories.","section":"Table I.6 caption and §5.3"},{"comment":"The appendix contains the typo 'p-aceptable' in the description of the random-context control.","section":"Appendix F"},{"comment":"The paper notes that the first author annotates most of the CONGEN data, but it does not report inter-annotator agreement or a second pass by other annotators; adding such information would improve the dataset's credibility.","section":"Section 8, Limitations"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses an interesting and timely question, and the CONGEN dataset is a potentially valuable contribution. However, the central empirical claims (20% weak generics and the context-sensitivity asymmetry) are not yet supported because the p-acceptability metric lacks human validation. The original-label recovery in §5.1 is too weak to establish the metric's validity, especially for generic sentences where the label is the surface form itself. I recommend that the authors add a human judgment experiment or an alternative validation (e.g., known weak generics such as 'mosquitoes carry malaria' versus strong generics such as 'bees reproduce') before the findings are presented with the current level of confidence. The paper is otherwise well-structured and should be suitable for publication after this gap is addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the genuinely good part: CONGEN is a real contribution. It is, as far as I know, the first collection of naturally occurring bare-plural generics with their context, and the dataset construction (classifier plus human annotation) is documented carefully. The p-acceptability metric is also clever—restricting surprisal to property tokens and showing that whole-sequence surprisal fails is a useful methodological data point for anyone working on quantifier semantics in LMs. The experiments are run on three Mistral models with a control for random context, which gives the results a solid base.\n\nThe weak spot is validation. The metric is never checked against human quantifier judgments. The section 5.1 validation recovers original labels on the same datasets used for the main claims. For quantified sentences, that recovery can be driven by surface-form familiarity: \"all tigers have stripes\" is a more natural string than \"some tigers have stripes\", so lower surprisal might reflect corpus frequency, not quantificational semantics. For generic sentences, there is no original quantifier to recover, so the 20% weak-generic estimate depends entirely on the assumption that lower property-token surprisal for \"some\" corresponds to the speaker meaning a weak generalization. Since \"some\" is the weakest candidate, it is also the default under uncertainty, which could inflate the count. Section 8 acknowledges classifier and annotator biases, but not this calibration gap. The context-sensitivity result is interesting, and the random-context control rules out a simple baseline, but the interpretation that generics are context-sensitive in a way quantifiers are not goes beyond what the data show—it shows the model's label confidence changes with context, not that the semantic content changes.\n\nThe stereotype section is more exploratory, and the invented-demonyn materials are synthetic, but it does raise a real question about instruction tuning.\n\nAll of this is fixable. A human-rating study on a sample of CONGEN, plus uncertainty estimates for the 20% figure, would move this from conditional to solid. I think the paper deserves a serious referee—the dataset alone is worth publishing, and the metric is worth scrutiny. I would also tell the authors to engage directly with the calibration objection in the revision.\n\nWho is this for? Semantics folks (philosophy and linguistics) and NLP people interested in LM behavior. I'd bring it to a reading group, but mainly to discuss what counts as validation for a metric like p-acceptability.","headline":"A genuinely new corpus and a clever metric, but the headline numbers depend on a validation gap: p-acceptability is never checked against human quantifier judgments.","tokens_in":17226,"tokens_out":4065,"would_cite":true,"duration_ms":31939,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Generic sentences hide quantifier strength, and surprisal finds it","keywords":["generic sentences","implicit quantification","surprisal","p-acceptability","weak generics","context-sensitivity","stereotypes","language models"],"falsifier":"A direct check would be to ask native speakers to choose among all, most, some, or no quantifier for a sample of CONGEN generics, with and without context, and measure agreement with p-acceptability; chance-level agreement would falsify the metric's semantic validity, and the context-sensitivity claim with it.","tokens_in":16346,"feed_emoji":"🧩","tokens_out":4721,"duration_ms":39508,"temperature":0.7,"pith_summary":"The paper argues that the implicit quantificational strength of generic sentences can be recovered from language model surprisal. It introduces CONGEN, a dataset of naturally occurring generic and quantified sentences with contexts, and p-acceptability, a metric that picks the explicit quantifier (all, most, some, or none) that makes the property tokens least surprising. Using this metric, it finds that generics respond to preceding context much more than explicit quantifiers do, and that roughly one in five naturally occurring generics express weak generalisations — properties true of only a minority of the kind. It also shows that negative stereotypes pattern as universal quantification, positive ones as most, and that a 'people who are' paraphrase shifts negative stereotypes toward some. If correct, these results give philosophers and linguists a corpus-scale way to test theories of generics against actual usage.","feed_headline":"LM surprisal reveals the hidden quantifier in generic sentences","feed_subtitle":"About 20% of generic sentences are weak generalisations, and context matters for them more than for explicit quantifiers.","key_machinery":"The central object is p-acceptability, a metric defined on a sentence with a bare plural subject and a verb-plus-property predicate. For each candidate quantifier q in {all, most, some, none}, the metric prepends q to the sentence, computes the per-token surprisal of only the property tokens under a language model, and selects the q with the lowest surprisal. This turns implicit quantificational strength into a measurable, comparable quantity, and it is the device that generates all three main claims: the prevalence of weak generics, the context-sensitivity asymmetry, and the stereotype quantification profile. The metric is deliberately restricted to property tokens because whole-sentence surprisal is shown to be insensitive to quantifiers, replicating earlier failures.","core_discovery":"On its own terms, the central discovery is that the surprisal of property tokens (the words after the verb) in a language model is sensitive to explicit quantifiers, and that the quantifier which minimizes this surprisal tracks how speakers use generics. Validated against two datasets, this p-acceptability criterion recovers the expected ordering all/most stronger than some, reproduces known effects such as generic overgeneralisation, and assigns explicit quantifiers to generics in a way that matches semantic intuitions. Applied to naturally occurring generics, it estimates that roughly 18–23% of generic sentences are weak generics, and it shows that adding preceding context improves quantifier recovery for generics by about 20 percentage points within the first sentence, while explicit quantifiers gain little from context. The paper further claims that negative stereotypes are perceived as universals (all), positive ones as most, and that paraphrasing a stereotype as 'people who are X' shifts its implicit quantification toward some for real stereotypes, though the effect weakens for invented social kinds.","pith_inferences":["An extension the paper does not pursue: p-acceptability could track how the implicit quantification of a single generic shifts across a discourse, not just aggregate context windows; if the chosen quantifier changes with topic or preceding content, that would give contextualism a dynamic, per-occurrence test.","The 20% weak-generic figure suggests that philosophical theories calibrated on striking examples (dangerous predators, stereotyped groups) may underestimate the role of mundane, non-striking weak generics; revisiting those theories with this base rate in mind is a natural next step.","A testable cross-linguistic extension would adapt p-acceptability to languages without bare plurals, using definite plurals or singular generics; if the context-sensitivity asymmetry persists across such adaptations, it would indicate a general cognitive pattern rather than an artefact of English morphology.","The observed shift for 'people who are' paraphrases is compatible with the view that generics express primitive, kind-level generalisations: the paraphrase disrupts the kind-level reading, and the metric's sensitivity to that disruption is itself evidence that p-acceptability tracks the intended semantic distinction."],"forward_implications":["The p-acceptability metric gives a continuous, corpus-scale measure of implicit quantificational strength, allowing theories of generics to be tested on naturally occurring language rather than only on synthetic examples.","If generics are indeed more context-sensitive than determiner quantifiers, contextualist accounts of generics gain empirical support, and models of generic semantics must incorporate preceding discourse as a variable.","A stable estimate of weak generics at roughly one in five provides a base rate for philosophical debates about striking generics, suggesting that non-striking minority generalisations are a common rather than exceptional phenomenon.","The stereotype results imply that language models encode the human bias of interpreting negative group statements as universal claims, and that surface paraphrases can measurably shift this perceived quantification, pointing to a concrete intervention target for bias mitigation.","Because the metric works without prompting or fine-tuning, it can be applied to any autoregressive language model, making quantification-sensitive analysis available for models that lack instruction tuning."],"supporting_citations":[{"why":"Provides GenericsKB, one of the two datasets on which p-acceptability is validated and the large-scale source for the implicit-quantification estimates.","marker":"(Bhakthavatsalam et al., 2020)"},{"why":"Earlier work that failed to find quantifier sensitivity in whole-sequence surprisal; the paper replicates that failure and contrasts its property-token metric.","marker":"(Collacciani et al., 2024)"},{"why":"Supplies the generics-as-defaults theory, the stable-content view that the context-sensitivity experiment is designed to test against.","marker":"(Leslie, 2008)"},{"why":"Supplies the contextualist theory of generics, the rival account that predicts the context-sensitivity the paper reports.","marker":"(Sterken, 2015a)"},{"why":"Introduces the notion of weak generics used to define and interpret the roughly 20% of generics implicitly quantified as some.","marker":"(Almotahari, 2022)"},{"why":"Documents generic overgeneralisation in language models, the comparison point for the high prevalence of 'all' as an implicit quantifier.","marker":"(Allaway et al., 2024)"},{"why":"Provides psychological evidence that striking properties lead to overestimation of prevalence, the hypothesis tested in the stereotype experiment.","marker":"(Cimpian et al., 2010)"},{"why":"Supplies the Social Bias Frames dataset from which real negative stereotypes are extracted for the stereotype quantification analysis.","marker":"(Sap et al., 2020)"},{"why":"Defines the three autoregressive language models of increasing size used across all experiments.","marker":"(Jiang et al., 2023)"}],"fun_headline_variants":["Surprisal exposes the implicit quantifier in generics","Generics lean on context more than all or some do","Around 20% of generics are weak generalisations, LMs show","Context matters more for generics than for explicit quantifiers","LM surprisal links generics to quantifiers and stereotype biases"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the quantifier which makes a language model least surprised by the property words is the quantifier a speaker would judge as the best fit for that generic sentence, and this is tested only against the original labels of two datasets, not against fresh human judgments.","fun_headline_variants_meta":{"raw":{"variants":["Surprisal exposes the implicit quantifier in generics","Generics lean on context more than all or some do","Around 20% of generics are weak generalisations, LMs show","Context matters more for generics than for explicit quantifiers","LM surprisal links generics to quantifiers and stereotype biases"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001108,"raw_usage":{"total_tokens":4588,"prompt_tokens":886,"completion_tokens":3702,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":502,"completion_tokens_details":{"reasoning_tokens":3614}},"tokens_in":502,"tokens_out":3702,"duration_ms":23837,"temperature":1.0,"reasoning_tokens":3614,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T15:03:09.590026+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct check would be to ask native speakers to choose among all, most, some, or no quantifier for a sample of CONGEN generics, with and without context, and measure agreement with p-acceptability; chance-level agreement would falsify the metric's semantic validity, and the context-sensitivity claim with it.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Earlier work that failed to find quantifier sensitivity in whole-sequence surprisal; the paper replicates that failure and contrasts its property-token metric."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the generics-as-defaults theory, the stable-content view that the context-sensitivity experiment is designed to test against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the notion of weak generics used to define and interpret the roughly 20% of generics implicitly quantified as some."},{"cited_title":"Hwang, Kathleen McKeown, and Sarah-Jane Leslie","cited_arxiv_id":null,"evidence_quote":"Documents generic overgeneralisation in language models, the comparison point for the high prevalence of 'all' as an implicit quantifier."}],"review_version":1}