{"id":"97f5ea6f-0244-464e-8e8e-6064ed29af1e","arxiv_id":"2509.10860","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Across seven LLMs, surface-scope readings dominate \"every...a\" sentences while \"a...every\" sentences favor inverse scope, and only BERT-family models reproduce the English-Chinese contrast seen in humans.","lead":"Seven large language models were tested on how they interpret sentences with two quantifiers, like \"Every child climbed a tree\", in English and Chinese. Most models prefer the simpler surface reading for some sentence types, but only some show the human-like difference between the two languages.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract and conclusion overgeneralize the surface-scope preference: for the EU construction (Table 3), no model significantly prefers surface scope and most significant contrasts favor inverse scope.","rationale":"The reader's weakest_assumption field identified the Experiment 2 item mismatch as the key vulnerability, but their rationale also notes that the headline claim is contradicted by Table 3. My stress-test focuses on that internal contradiction because it threatens the paper's central claim independently of Experiment 2. The EU results are not a consensus disagreement; they come from the authors' own statistical comparisons. This strengthens rather than replaces the reader's conditional verdict: the paper needs a major revision that either restricts the surface-scope claim to UE constructions or reframes the EU inverse-scope preference as the headline finding. No ad hominem is intended; the data collection and descriptive analyses appear sincere. The Section 6 limitation is acknowledged, but it cannot rescue the unqualified conclusion because the abstract and conclusion draw on Experiment 1. I would keep the verdict at CONDITIONAL: the underlying item-level data would be useful if made available and if the claims are reworded, but the current presentation is not acceptable as-is.","tokens_in":12569,"tokens_out":8298,"duration_ms":72367,"concrete_test":"Reconstruct or rerun a mixed-effects model on all LLM surprisal data with fixed effects for Structure (UE/EU), Interpretation (SS/IS), and Language, plus their interactions, with LLM and item random intercepts, using the released or reconstructed item-level surprisals. Test whether the main effect of Interpretation favors SS and whether the Structure-by-Interpretation interaction is significant. If the simple effect of Interpretation within EU favors IS in the majority of models, the abstract's unqualified 'most LLMs prefer SS' claim is unsupported. A quicker check from Tables 2 and 3 is to count significant cells by structure: UE has 10 significant SS cells, while EU has 7 significant IS cells and zero significant SS cells.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing concern is internal to Experiment 1, not the Section 6 stimulus mismatch. The abstract and conclusion state that 'most LLMs prefer the surface scope interpretations, aligning with human tendencies.' This is only observed for UE sentences (Table 2). For EU sentences, Section 2.4 reports that 'all LLMs predominantly preferred IS readings,' and Table 3 shows that every statistically significant model-by-language comparison favors inverse scope (e.g., DistilGPT2, GPT-2En, LlamaEn, LlamaCh in English; BERT-large and LlamaEn in Chinese); no EU cell shows a significant surface-scope preference. The Discussion itself calls the EU results 'unexpected' and admits humans favored SS for both structures. Since EU is half of the materials, an unqualified surface-scope claim is contradicted by the paper's own data. The Section 6 limitation about non-identical human/LLM items concerns Experiment 2 only and does not fix this Experiment 1 mismatch. At minimum, every occurrence of 'most LLMs prefer SS' must be restricted to UE constructions, and the EU inverse-scope preference should be reported as a central negative result.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies how seven LLMs interpret doubly quantified sentences in English and Chinese, using surprisal-based probabilities to infer preference for surface scope (SS) versus inverse scope (IS) in two structures: UE (e.g., \"Every child climbed a tree\") and EU (e.g., \"A child climbed every tree\"). Experiment 1 reports that for UE, most models prefer SS over IS, with some language-dependent variation; for EU, models predominantly prefer IS, a pattern the authors call unexpected. Experiment 2 compares LLM surprisal distributions with human truth-value judgment data from Fang (2023) using a Jensen-Shannon-divergence-based Human Similarity score, finding that GPT and LLaMA models are more human-like than BERT models and that L2-learner-like patterns emerge. The abstract and conclusion claim that most LLMs prefer surface scope and align with human tendencies, and that only some models differentiate English from Chinese in IS availability.","tokens_in":12767,"tokens_out":3617,"duration_ms":33025,"significance":"If the surface-scope claim were restricted to UE constructions, the study would contribute a useful cross-linguistic, cross-architecture dataset for LLM quantifier-scope behavior, and the direct computation of surprisals from pretrained weights without any fitting to the target results is a methodological strength. The inclusion of both UE and EU constructions and Chinese materials goes beyond prior LLM scope work, and the EU inverse-scope preference is a falsifiable negative result. However, the overgeneralized headline claim and the non-identical stimuli in Experiment 2 mean that the human-alignment conclusions, as currently stated, are not supported by the evidence presented.","major_comments":[{"comment":"The claim that \"most LLMs prefer the surface scope interpretations\" is contradicted by the paper's own Experiment 1 results for EU sentences. Section 2.4 states that for EU structures \"all LLMs predominantly preferred IS readings,\" and Table 3 shows that every statistically significant SS-vs-IS contrast favors inverse scope (e.g., DistilGPT2, GPT-2En, LlamaEn, and LlamaCh in English; BERT-large and LlamaEn in Chinese), with no EU cell showing a significant surface-scope preference. The authors themselves describe the EU results as \"unexpected\" (§4.1). Since EU constitutes half the materials, all occurrences of the surface-scope preference claim must be restricted to UE constructions, and the EU inverse-scope preference should be reported as a central negative result rather than as an aside.","section":"Abstract, §5 Conclusion, §4.1 Discussion"},{"comment":"The Human Similarity comparison in Experiment 2 relies on non-identical stimuli for humans and LLMs: human ratings came from the original story contexts in Fang (2023), while LLM surprisals were computed on expanded and enriched contexts (Section 2.1). As the authors acknowledge in Section 6, the items \"were not identical,\" and this discrepancy threatens the construct validity of HS scores, because surprisal differences may reflect context length or detail rather than scope interpretation. The HS analysis also lacks a baseline (e.g., human-human agreement or a chance-level reference) and Figure 4 shows no error bars or per-item dispersion. Without identical items, a baseline, and uncertainty estimates, the claim that HS scores show LLMs \"approximate human language use\" is not fully supported; the authors should either rerun the comparison on identical items or substantially temper these alignment claims.","section":"§3.2, §3.3, §6 Limitations"},{"comment":"The only significant Language effect in the per-LLM analyses is BERT-large in the UE condition, reported as b=1.1, p=.0499. Because seven models and two structures are tested, this borderline p-value, uncorrected for multiple comparisons, is too weak to support the Discussion's emphasis on BERT models exhibiting cross-linguistic contrasts. The authors should report multiplicity-adjusted p-values or, at minimum, all model-specific p-values, and in the Discussion distinguish strong effects (e.g., the EU patterns for LlamaCh) from this borderline finding.","section":"§2.4 Results"}],"minor_comments":[{"comment":"There is a typo: \"ANOV A tests\" should read \"ANOVA tests.\"","section":"§3.3"},{"comment":"\"Experimental 2\" should be \"Experiment 2.\"","section":"Figure 3 caption"},{"comment":"The inline text contains \"allps\" without a space; it should read \"all ps.\"","section":"§2.4"},{"comment":"The sentence \"As shown in Tables 1 and 2\" appears to refer to the results tables; it should refer to Tables 2 and 3, since Table 1 is a gloss of the Chinese sentence.","section":"§4.2"},{"comment":"The labels UE and EU are not defined explicitly; \"30 existential quantifier (UE)\" is confusing because UE denotes the universal-existential structure \"Every ... a ...\" and EU denotes the existential-universal structure \"A ... every ...\". Please define the labels at first use.","section":"§2.1"},{"comment":"The exploratory Deepseek-R1 analysis lacks details about model version, sampling parameters, and prompt robustness; since it is not central to the paper's claims, it would fit better in a clearly marked exploratory subsection or in the appendix.","section":"§4.3"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's what to know: this is a legitimate extension of LLM scope probing to Chinese and to EU structures, but the abstract and conclusion say 'most LLMs prefer SS' when the EU results in Table 3 show the opposite. That mismatch is the first thing to fix.\n\nWhat's new: Kamath et al. (2024) only tested English 'every...a' sentences. This paper adds existential-first (EU) constructions, Chinese materials, L2 human comparison groups, seven models, and a Deepseek-R1 exploratory probe. That is a real step forward, and the surprisal-based method is sound. The BERT-vs-GPT/LLaMA contrast—BERT tracks English/Chinese IS availability, the autoregressive models do not—is the most interesting finding. The HS scores add a descriptive layer, though they're only as good as the comparison.\n\nThe load-bearing soft spot: the headline claim is only true for UE. For EU, no model significantly prefers SS; every significant contrast favors IS (DistilGPT2, GPT-2En, LlamaEn, LlamaCh in English; BERT-large, LlamaEn in Chinese). The Discussion calls EU 'unexpected' and admits humans favored SS for both structures. So you can't say 'most LLMs prefer the surface scope interpretations, aligning with human tendencies' without restricting it to UE. This is not a minor wording issue; it changes the takeaway. The EU inverse-scope preference should be reported as a central negative result.\n\nExperiment 2 is weaker. The LLM items are enriched expansions of the human items, not identical. The authors flag this in Limitations but then say the HS metric does not strictly require identical items. That's defensible for a descriptive metric, but it undercuts any specific claim about how closely LLMs match human judgments. There are also no error bars or baselines on the HS scores, and the p=.0499 Language effect for BERT-large on UE is thin. The human data come from the first author's thesis; that's not a problem by itself, but it's a single source.\n\nThe citation pattern is fine, and there's no circularity—surprisals are computed directly from pretrained weights. This paper is for psycholinguists and LLM interpretability researchers. It deserves a serious referee. My recommendation: send to peer review, but require the authors to restrict every SS-preference claim to UE, report the EU result prominently, and either share code/data or explain why not. With those revisions, it's a solid empirical contribution.","headline":"Useful cross-linguistic LLM scope benchmark, but the abstract's surface-scope preference claim is contradicted by the paper's own EU results.","tokens_in":13315,"tokens_out":2986,"would_cite":true,"duration_ms":22804,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Most LLMs prefer surface-scope readings, like humans, but only a subset reproduce the English–Chinese contrast in inverse-scope availability.","keywords":["quantifier scope","large language models","cross-linguistic variation","human similarity score","surprisal","truth-value judgment task","English and Chinese","semantic transparency"],"falsifier":"Run the same seven LLMs on the exact item set given to human participants and compare surprisal-ranked preferences with human ratings item by item. If the models no longer show inverse scope more likely in English than in Chinese, or if a direct 7-point acceptability probe of Chinese inverse-scope items returns ratings equal to surface-scope ratings, the claim that a subset of LLMs captures human-like cross-linguistic scope contrast fails.","tokens_in":12358,"feed_emoji":"🧠","tokens_out":8888,"duration_ms":70421,"temperature":0.7,"pith_summary":"This paper asks whether large language models interpret doubly quantified sentences the way humans do, in both English and Chinese. Its central claim is that most LLMs prefer the surface-scope reading, the interpretation humans also favor, while only a subset of models reproduce the human-like contrast where inverse scope is more available in English than in Chinese. The authors read this as partial evidence that LLMs can approximate human interpretive preferences, with model architecture, scale, and especially the language of pre-training data affecting how close the approximation is. The reason to care is that quantifier scope is a case where a sentence's meaning depends on syntax, semantics, and context, so it can reveal whether a model represents meaning rather than only surface word statistics.","feed_headline":"Most LLMs mirror human quantifier-scope preferences; gaps persist","feed_subtitle":"The language of pre-training data and model architecture shape how closely LLMs approximate human scope judgments.","key_machinery":"The load-bearing machinery is surprisal-based interpretation scoring: each model assigns a conditional probability to a target sentence given a story context, and the reading whose context yields lower surprisal is treated as preferred. Masked models are scored by pseudo-log-likelihood and autoregressive models by chain-rule token probabilities. The second component is the Human Similarity score, a Jensen-Shannon divergence between an LLM's response distribution and human per-item acceptability ratings, with lower divergence meaning greater human-likeness. This turns the psycholinguistic truth-value judgment task into a probability comparison and makes cross-model and cross-language differences measurable.","core_discovery":"On the paper's own terms, the discovery is that seven LLMs, scored by the conditional probability that a sentence is accepted in a context favoring one reading, mostly prefer surface scope for universal-existential sentences such as “Every child climbed a tree,” matching humans. For the reverse construction “A child climbed every tree,” most LLMs prefer the inverse reading, which is the opposite of the human pattern. Only the two BERT models show inverse scope being more likely in English than in Chinese for both constructions, mirroring the human cross-linguistic contrast. Human Similarity scores, computed with Jensen-Shannon divergence, show the GPT-family models closest to human judgments, LLaMA intermediate, and BERT lowest, especially against Chinese speakers; training-language effects appear but are not consistent, with the Chinese-trained GPT-2 failing to replicate the Chinese-trained LLaMA's behavior.","pith_inferences":["If the LLM items were made identical to the human items, the reported Human Similarity ordering could change; the paper's own limitation note says the LLM items were expanded, so the ranking is not final.","The universal LLM preference for inverse scope in EU sentences could mean LLMs exploit the entailment from 'a child climbed every tree' differently than humans do; a test using contexts that block that entailment would show whether the preference disappears.","A direct 7-point acceptability probe of inverse-scope Chinese items, along the lines of the exploratory Deepseek test, would settle whether the human-like English–Chinese difference is a grammatical contrast or only a relative surprisal effect.","The 'LLMs behave like L2 learners' pattern suggests cross-linguistic transfer in training data; training the same architecture on carefully balanced English and Chinese corpora could test whether the English-like bias disappears."],"forward_implications":["Quantifier scope becomes a reusable diagnostic: a model's preference between surface and inverse scope can be read off surprisal without prompting.","If this result generalizes, autoregressive models (GPT, LLaMA) are better global proxies for human judgments, while masked models (BERT) are better at capturing one specific cross-linguistic contrast.","The language of pre-training data does not guarantee human-like cross-linguistic behavior: the Chinese-trained GPT-2 did not reproduce the Chinese-trained LLaMA's pattern.","The EU result, where most LLMs prefer inverse scope while humans prefer surface scope, marks a boundary case for using LLMs as models of human sentence processing.","Because BERT models were the only ones showing inverse scope more likely in English than in Chinese, architecture likely modulates how language-specific semantic knowledge is encoded."],"supporting_citations":[{"why":"Supplies the human truth-value judgment dataset for Experiment 2 and the stimulus base that Experiment 1 expands.","marker":"Fang (2023)"},{"why":"Proposes the Jensen-Shannon divergence-based Human Similarity score used to quantify LLM-human alignment.","marker":"Duan et al. (2024)"},{"why":"Prior LLM study of scope with the every...a configuration that supplies the comparison for UE behavior and the direct-query alternative to surprisal.","marker":"Kamath et al. (2024)"},{"why":"Processing Scope Economy principle, used to predict higher processing cost for inverse scope and motivate the human baseline.","marker":"Anderson (2004)"},{"why":"Single Reference Principle, used to explain the human surface-scope bias in EU constructions.","marker":"Kurtzman and MacDonald (1993)"},{"why":"Theoretical claim that Mandarin permits only surface scope, giving the cross-linguistic contrast the paper tests.","marker":"Aoun and Li (1989)"},{"why":"Experimental evidence on cross-linguistic scope that frames how Chinese speakers can accept inverse scope in some contexts.","marker":"Scontras et al. (2017)"}],"fun_headline_variants":["LLMs echo human scope preferences, BERT bucks trend","Surface scope dominates LLMs, mirroring human choices","LLMs favor surface scope like humans; BERT flips it","Scope reading in LLMs: human-like but training matters","Quantifier scope: LLMs align with humans, with gaps"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that surprisal scores computed on the expanded story contexts measure the same scope-interpretation preference as the human truth-value judgments on the original, shorter stories, and the paper itself flags that the items were not identical.","fun_headline_variants_meta":{"raw":{"variants":["LLMs echo human scope preferences, BERT bucks trend","Surface scope dominates LLMs, mirroring human choices","LLMs favor surface scope like humans; BERT flips it","Scope reading in LLMs: human-like but training matters","Quantifier scope: LLMs align with humans, with gaps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000231,"raw_usage":{"total_tokens":1440,"prompt_tokens":853,"completion_tokens":587,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":469,"completion_tokens_details":{"reasoning_tokens":504}},"tokens_in":469,"tokens_out":587,"duration_ms":5235,"temperature":1.0,"reasoning_tokens":504,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:51:46.221982+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same seven LLMs on the exact item set given to human participants and compare surprisal-ranked preferences with human ratings item by item. If the models no longer show inverse scope more likely in English than in Chinese, or if a direct 7-point acceptability probe of Chinese inverse-scope items returns ratings equal to surface-scope ratings, the claim that a subset of LLMs captures human-like cross-linguistic scope contrast fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the human truth-value judgment dataset for Experiment 2 and the stimulus base that Experiment 1 expands."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior LLM study of scope with the every...a configuration that supplies the comparison for UE behavior and the direct-query alternative to surprisal."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Processing Scope Economy principle, used to predict higher processing cost for inverse scope and motivate the human baseline."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Single Reference Principle, used to explain the human surface-scope bias in EU constructions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Theoretical claim that Mandarin permits only surface scope, giving the cross-linguistic contrast the paper tests."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Experimental evidence on cross-linguistic scope that frames how Chinese speakers can accept inverse scope in some contexts."}],"review_version":2}