{"id":"cc25b25e-9fee-4aef-b8e1-ebb898f12ccb","arxiv_id":"2411.10533","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Language models can reinforce Chomsky's generative linguistics through formal generative capacity, discovery procedures, and the Minimalist Program, and generative linguistics can guide LM evaluation.","lead":"This paper argues that neural language models are compatible with Chomsky's generative linguistics and can support it in three ways: as formal generative models, as discovery procedures, and as tools for the Minimalist Program. It is worth reading because it offers a detailed reconciliation of the current divide between language-model research and theoretical linguistics.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central bridge is an unvalidated 'derivational likelihood' quantity: undefined for transformer LMs and, where defined (grammar induction), not shown to track grammaticality rather than model probability.","rationale":"This is a position paper, not an empirical study, so the stress-test must focus on the internal coherence of the argument. The three-level adequacy framing and the historical discussion are careful and useful. The weakest point is exactly the bridge from string probability to grammatical knowledge: the paper's central positive proposal requires a quantity called 'derivational likelihood' that is supposed to exist in LMs and to encode something like grammaticality. The paper honestly flags that extraction is hard for transformers, but the problem is more serious than access. For transformer-based LMs, no explicit grammar or latent derivation is defined, so the proposed quantity has no clear formal meaning. For grammar induction LMs, the quantity is defined relative to an induced grammar, but no evidence is provided that high derivational probability corresponds to grammaticality rather than merely to the model's own prior. The reader's weakest_assumption already names derivational likelihood; I partially agree because I also stress that the quantity is unvalidated even where it is defined, and that the grammar-induction subclass is not shown to represent LMs as a class. The proposed concrete experiment would settle the decisive question: does derivational likelihood, properly operationalized, classify grammatical vs ungrammatical sentences beyond what string likelihood already does? If the experiment fails, the discovery-procedure and descriptive-adequacy arguments collapse, and the paper would need to be substantially revised. If it succeeds, the compatibility claim is considerably strengthened. Given that the construct is currently unvalidated, the reader's CONDITIONAL verdict is appropriate and should remain unchanged; the paper should be accepted only conditional on such a demonstration or on a clear operational definition of derivational likelihood for transformer LMs.","tokens_in":14714,"tokens_out":5209,"duration_ms":62515,"concrete_test":"Train a grammar induction LM (e.g., Kim et al. 2019 compound PCFG or Portelance et al. 2024 style model) on a corpus with gold-standard grammaticality judgments for minimal pairs balanced for surface string probability, including 'Colorless green ideas sleep furiously' vs 'Furiously sleep ideas green colorless' plus a larger set of syntactic violations matched for length and lexical frequency. For each item, compute (a) surface string log-probability and (b) derivational log-likelihood via the inside algorithm over the induced grammar. Measure whether derivational likelihood classifies grammatical vs ungrammatical better than surface likelihood (e.g., AUROC with paired bootstrap). If derivational likelihood does not separate the classes while controlling for string probability, the Section 3 distinction is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing claim in Section 3 is that LMs 'can in principle learn grammars and some decidedly do so, making a clear distinction between string likelihood and derivational likelihood, or grammaticality.' Everything that follows for discovery procedures and descriptive adequacy depends on derivational likelihood. For transformer-based LMs, the paper concedes extraction is not possible, but the deeper problem is that the quantity is not defined: an autoregressive transformer has no latent grammar or derivation variable, so 'derivational likelihood given all relevant model weights and states' has no formal status distinct from the string's probability under the model. For grammar induction LMs, derivational likelihood is well-defined relative to an induced grammar, but the paper never validates that this quantity tracks grammaticality. A PCFG induced from a corpus can assign low probability to rare but grammatical sentences and high probability to ungrammatical ones; grammaticality is not a threshold on one model's derivation score. Thus the compatibility thesis is carried by an unvalidated construct: absent for the models that motivate the debate, and unvalidated for the subclass where it is defined. That subclass is also cited mainly through the authors' own prior work, so the representativeness step is not independently supported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This conceptual position paper argues that neural language models are compatible with, and indeed reinforce, Chomskyan generative linguistics in three ways: (1) LMs are formal generative models in the sense of Chomsky's formal language theory; (2) LMs can serve as discovery procedures in the sense of Syntactic Structures; and (3) LMs can contribute to the Minimalist Program's 'bottom-up' identification of what must be attributed to genetic endowment. The central sub-claim, developed in Section 3, is that LMs can in principle learn grammars and that some 'decidedly do so', thereby distinguishing string likelihood from derivational likelihood, which is equated with grammaticality. The paper is a position piece and contains no new experiments, derivations, or code.","tokens_in":14947,"tokens_out":5452,"duration_ms":54005,"significance":"If the central sub-claim could be made precise and empirically supported, the paper would provide a genuinely useful bridge between computational linguistics and generative linguistics. It is strongest where it connects Chomsky's levels of adequacy to the evaluation of LMs, and where it draws attention to grammar induction models as transparent, underused discovery procedures. The paper does not ship machine-checked proofs, reproducible code, or falsifiable quantitative predictions, so its contribution is conceptual rather than evidential. Those strengths are real, but the argument currently rests on an unformalized construct and on evidence that is not actually presented in the manuscript. As a research agenda the paper is promising; as a demonstration of compatibility it is not yet conclusive.","major_comments":[{"comment":"The load-bearing notion of 'derivational likelihood' is asserted but not formally defined. The text defines string likelihood as a Markov product over next-token distributions and says derivational likelihood is 'the probability that a string be derived given all relevant model weights and states' without the Markov assumption. For a standard autoregressive transformer there is no latent derivation or grammar variable in the inference computation: the sequential next-token distribution fully determines the string's probability, and the paper specifies no additional quantity that would count as 'derivational'. The paper concedes that extraction is not currently possible, but the deeper problem is that the quantity itself is left undefined for the model class that motivates the debate. Without a formal definition, the claim that LMs can distinguish grammaticality from string likelihood is unsupported.","section":"Section 3"},{"comment":"Even for grammar induction LMs, where a derivational likelihood is well-defined relative to an induced grammar, the paper does not validate that this quantity tracks grammaticality. A PCFG can assign low probability to rare but grammatical sentences and high probability to ungrammatical ones; the paper proposes no threshold, normalization, or calibration that would identify grammaticality. The sentence 'making a clear distinction between string likelihood and derivational likelihood, or grammaticality' conflates a model-internal score with a human linguistic judgment. The authors should either provide empirical evidence (for example, correlation with human acceptability judgments on minimal pairs) or explicitly mark the identification of derivational likelihood with grammaticality as a hypothesis to be tested.","section":"Section 3"},{"comment":"The empirical support for the claim that some LMs 'decidedly do' learn grammars is not presented in the manuscript. The only concrete demonstration cited is Portelance et al. (2024), a preprint by the first author; the paper provides no experimental details, no summary of results, and no independent confirmation. Because this claim is needed for the argument that grammar induction LMs already constitute successful discovery procedures, the manuscript should at minimum summarize the supporting evidence and its quality, or weaken the claim to a statement about what these models 'may' achieve.","section":"Section 3"}],"minor_comments":[{"comment":"The abstract contains a typo: 'wonder about the the compatibility' should read 'wonder about the compatibility'.","section":"Abstract"},{"comment":"The phrase 'rules that transformed them into the output s(Chomsky, 1957, 1965)' appears to have a missing word or formatting error; it should likely be 'the output sentences (Chomsky, 1957, 1965)'.","section":"Section 2"},{"comment":"The sentence 'making derivational likelihood easily accessible, see Figure 4' likely refers to the wrong figure; Figure 4 depicts discovery/decision/evaluation procedures, while the comparison of transformer-based and grammar induction LMs is in Figure 6.","section":"Section 3"},{"comment":"Several references contain formatting errors, notably 'The range of aadequacy' (Chomsky 1956b) and the author fields 'Feiman, R., Mody, S., Sanborn, S., & and, S. C.' and 'Liu, Z., & and, M. J.' in which author names appear to be missing.","section":"References"},{"comment":"The 'cognitive duck test' is introduced as a framing device, but its role could be clarified: it is an analogy for levels of adequacy, not itself an argument or a methodological criterion.","section":"Section 2"}],"recommendation":"major_revision","confidential_remarks":"This is a position paper with no empirical content, so its fit with the journal's usual expectations should be considered carefully. The central empirical anchor is a cited preprint by the first author, which is not by itself disqualifying but is a risk for the manuscript's independence. I would not recommend reject if the authors can either formalize derivational likelihood for transformer-based LMs, provide or cite concrete validation for grammar induction LMs, or explicitly reframe the paper as a research agenda rather than a demonstrated compatibility thesis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper is a genuinely useful synthesis of the current debate, but the bridge it relies on is not built. The three-part argument—LMs as formal generative models, LMs as discovery procedures, LMs as Minimalist Program assets—is clearly organized and historically careful. The discovery-procedure framing is the freshest part: it gives grammar induction LMs a concrete scientific role that the \"LMs are just tools\" positions don't. The distinction between string likelihood and derivational likelihood is a helpful way to organize why n-gram style objections don't automatically apply to neural models. The authors also deserve credit for flagging that extraction from black-box LMs is currently impossible.\n\nThe soft spot is exactly where the stress-test note lands. Section 3 asserts that LMs \"can in principle learn grammars and some decidedly do so\" because they can support a derivational likelihood distinct from string probability. For transformer LMs, that quantity is not defined—there is no latent derivation variable in an autoregressive transformer, so \"derivational likelihood given all relevant weights and states\" has no formal content separate from the model's string probability. The paper says the value exists but we can't extract it; the deeper problem is that it's not clear what the value is. For grammar induction LMs, derivational likelihood is well-defined relative to an induced grammar, but nothing in the paper shows that this quantity tracks grammaticality. A PCFG induced from a corpus can assign low probability to rare grammatical sentences and high probability to ungrammatical ones; grammaticality is not a threshold on one model's derivation score. So the compatibility thesis rests on an unvalidated construct: absent for the models that motivate the debate, unvalidated for the subclass where it's defined.\n\nThe reliance on Portelance et al. (2024) for representativeness is a real weakness, though self-citation alone isn't disqualifying. The \"cognitive duck test\" is a framing device, not evidence, which is fine—but it shouldn't be read as doing more than labeling.\n\nFor a position paper, this is solid. The argument is coherent, the literature engagement is fair, and the authors are explicit about key limits. It's not a demonstrated result, and the title overpromises: compatibility is conditional on a derivational-likelihood notion that may not survive contact with actual model internals. Who gets value: anyone in the Piantadosi/Kodner/Katzir debate, and anyone thinking about whether grammar induction models can carry weight in learnability research. It deserves a serious referee—the revision should either define derivational likelihood for a concrete model class or narrow the claim to grammar induction LMs and validate it against grammaticality judgments. I'd take it for reading group.","headline":"A well-organized position paper whose three-part compatibility argument is undercut by an unvalidated 'derivational likelihood' construct—absent for transformers and untested for grammar induction models.","tokens_in":15469,"tokens_out":2524,"would_cite":true,"duration_ms":40947,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that neural language models are not a break from Chomsky's generative linguistics but a continuation of it, because they are formal generative models, can serve as discovery procedures, and can advance the Minimalist…","keywords":["language models","generative linguistics","Chomsky hierarchy","discovery procedures","grammar induction","derivational likelihood","Minimalist Program","language acquisition"],"falsifier":"Train a grammar induction language model on a corpus and test it on the pair 'Colorless green ideas sleep furiously' versus 'Furiously sleep ideas green colorless.' If the model's derivational likelihood ranks the grammatical sentence above the ungrammatical one while its string likelihood does not, the distinction is real; if no accessible internal quantity separates them, the discovery-procedure claim collapses.","tokens_in":14503,"feed_emoji":"💬","tokens_out":6622,"duration_ms":57234,"temperature":0.7,"pith_summary":"This paper argues that the current wave of neural language models is not a break from Chomsky's generative linguistics but a continuation of it. The authors claim that LMs are formal generative models in Chomsky's original sense, that they can serve as discovery procedures which induce grammars from data, and that they can support the Minimalist Program's search for what must be innate in language. If the argument holds, the common opposition between statistical language models and rule-based generative grammar dissolves, and each side becomes a tool for the other. The argument turns on the claim that some LMs track derivational likelihood—a model-internal probability of a sentence given latent grammatical structure—and not just surface string frequency.","feed_headline":"Neural language models fit Chomsky's own definition of grammars","feed_subtitle":"A new paper argues LMs track 'derivational likelihood,' not just string frequency, making them discovery procedures for linguistics.","key_machinery":"The central object is derivational likelihood: the probability that a string is derived given the model's full latent state and weights, as opposed to string likelihood, which is computed by sequential Markovian sampling from the final layer. The distinction lets the authors answer the charge that LMs are only observationally adequate—the paper argues the problem is not that derivational likelihood does not exist but that it is not always easy to extract. The surrounding machinery is Chomsky's hierarchy of formal languages, the three levels of adequacy, and the discovery/decision/evaluation procedure trichotomy, which together locate LMs within generative linguistics rather than outside it.","core_discovery":"The paper's central claim is that LMs are formal generative models as Chomsky originally defined them: devices that generate a set of sentences and are evaluated by observational, descriptive, and explanatory adequacy. Against the objection that LMs only provide string likelihood and therefore cannot distinguish grammaticality from probability, the paper asserts that LMs can in principle learn grammars and some decidedly do—specifically grammar induction LMs, which learn distributions over explicit syntactic rules and make derivational likelihood directly accessible. This makes LMs usable as discovery procedures of the kind Syntactic Structures described, and as computational implementations of the Minimalist Program's third-factor methodology, where whatever cannot be learned from data plus general learning principles is a candidate for innate endowment.","pith_inferences":["If derivational likelihood is real but only readily accessible in grammar induction LMs, the strongest empirical case for the paper rests on that niche class; a natural extension is to probe transformer hidden states for a recoverable analogue.","The paper's framework predicts that adding multimodal grounding and child-scale data to grammar induction LMs should narrow the explanatory-adequacy gap with human learners—a concrete direction for testing the Minimalist claim.","The compatibility argument also recasts benchmark evaluation: if descriptive adequacy is the target, tests of structural ambiguity like 'pancakes or bacon and eggs' become more diagnostic than perplexity or accuracy scores."],"forward_implications":["LMs can be used as discovery procedures that induce grammatical categories and rules directly from corpora, without a pre-specified grammar.","Grammar induction LMs can simulate language acquisition and be checked against children's developmental milestones, giving explanatory adequacy an empirical handle.","LMs' systematic failures under data-driven and third-factor assumptions become hypotheses about which aspects of language are innate.","Chomsky's levels of adequacy provide a scientific evaluation standard for LMs distinct from engineering metrics.","Multimodal versions of these models may bring linguistics closer to the full multimodal experience of human language learning."],"supporting_citations":[{"why":"Supplies the original definition of discovery, decision, and evaluation procedures that the paper reinterprets for LMs.","marker":"Chomsky (1957)"},{"why":"Defines grammars as devices that generate exactly the grammatical sentences, the formalization the paper identifies with LM training.","marker":"Chomsky (1956c)"},{"why":"Introduces observational, descriptive, and explanatory adequacy, the criteria the paper uses to position LMs.","marker":"Chomsky (1965)"},{"why":"Establishes the learnability results that motivate Universal Grammar and the Minimalist Program's search for innate factors.","marker":"Gold (1967)"},{"why":"Introduces compound PCFG grammar induction LMs, the model class whose derivational likelihood is directly accessible.","marker":"Kim, Dyer, and Rush (2019)"},{"why":"Provides another grammar induction architecture that learns latent tree structure, supporting the claim that LMs can learn grammars.","marker":"Drozdov et al. (2019)"},{"why":"Uses grammar induction LMs to model child language acquisition, supporting the claim that these models can approach explanatory adequacy.","marker":"Portelance, Reddy, and O'Donnell (2024)"},{"why":"Tests RNNs and transformer LMs against the Chomsky hierarchy, supplying the empirical base for comparing neural and classical generative models.","marker":"Deletang et al. (2023)"},{"why":"States the three factors of language design that the Minimalist Program uses, which the paper connects to LM induction biases.","marker":"Chomsky (2005)"}],"fun_headline_variants":["Why Chomsky would approve of modern language models","Language models are Chomsky's grammars, literally","AI and Chomsky: a compatible match","Neural LMs meet Chomsky's discovery procedure test","Generative AI and linguistics: not at odds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument stands or falls on the premise that language models contain an internal score for how a sentence is derived grammatically, distinct from how likely the surface string is—a score the paper concedes we often cannot currently extract.","fun_headline_variants_meta":{"raw":{"variants":["Why Chomsky would approve of modern language models","Language models are Chomsky's grammars, literally","AI and Chomsky: a compatible match","Neural LMs meet Chomsky's discovery procedure test","Generative AI and linguistics: not at odds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000159,"raw_usage":{"total_tokens":1206,"prompt_tokens":898,"completion_tokens":308,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":514,"completion_tokens_details":{"reasoning_tokens":235}},"tokens_in":514,"tokens_out":308,"duration_ms":3594,"temperature":1.0,"reasoning_tokens":235,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T19:35:14.626662+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train a grammar induction language model on a corpus and test it on the pair 'Colorless green ideas sleep furiously' versus 'Furiously sleep ideas green colorless.' If the model's derivational likelihood ranks the grammatical sentence above the ungrammatical one while its string likelihood does not, the distinction is real; if no accessible internal quantity separates them, the discovery-procedure claim collapses.","supporting_citations":[{"cited_title":"APACrefauthors \\ 1965","cited_arxiv_id":null,"evidence_quote":"Introduces observational, descriptive, and explanatory adequacy, the criteria the paper uses to position LMs."},{"cited_title":"APACrefauthors \\ 1967","cited_arxiv_id":null,"evidence_quote":"Establishes the learnability results that motivate Universal Grammar and the Minimalist Program's search for innate factors."},{"cited_title":", Dyer, C","cited_arxiv_id":null,"evidence_quote":"Introduces compound PCFG grammar induction LMs, the model class whose derivational likelihood is directly accessible."},{"cited_title":", Verga, P","cited_arxiv_id":null,"evidence_quote":"Provides another grammar induction architecture that learns latent tree structure, supporting the claim that LMs can learn grammars."},{"cited_title":"Reframing linguistic bootstrapping as joint inference using visually-grounded grammar induction models","cited_arxiv_id":"2406.11977","evidence_quote":"Uses grammar induction LMs to model child language acquisition, supporting the claim that these models can approach explanatory adequacy."},{"cited_title":", Ruoss, A","cited_arxiv_id":null,"evidence_quote":"Tests RNNs and transformer LMs against the Chomsky hierarchy, supplying the empirical base for comparing neural and classical generative models."}],"review_version":1}