Pith. sign in

REVIEW 6 cited by

Mission: Impossible Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.06416 v2 pith:UVEXPNRA submitted 2024-01-12 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords impossiblelanguageslearnenglishlanguagemodelsclaimcore
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Chomsky and others have very directly claimed that large language models (LLMs) are equally capable of learning languages that are possible and impossible for humans to learn. However, there is very little published experimental evidence to support such a claim. Here, we develop a set of synthetic impossible languages of differing complexity, each designed by systematically altering English data with unnatural word orders and grammar rules. These languages lie on an impossibility continuum: at one end are languages that are inherently impossible, such as random and irreversible shuffles of English words, and on the other, languages that may not be intuitively impossible but are often considered so in linguistics, particularly those with rules based on counting word positions. We report on a wide range of evaluations to assess the capacity of GPT-2 small models to learn these uncontroversially impossible languages, and crucially, we perform these assessments at various stages throughout training to compare the learning process for each language. Our core finding is that GPT-2 struggles to learn impossible languages when compared to English as a control, challenging the core claim. More importantly, we hope our approach opens up a productive line of inquiry in which different LLM architectures are tested on a variety of impossible languages in an effort to learn more about how LLMs can be used as tools for these cognitive and typological investigations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Tree Transformers are an Ineffective Model of Syntactic Constituency

    cs.CL 2024-11 conditional novelty 7.0 of 10

    Tree Transformers induce constituent structures that diverge from linguistic expectations and provide only marginal gains over standard BERT on hierarchy-sensitive agreement tasks.

  2. Removing Noise, not Finding Gold: Quality Filtering for Large-Scale Pretraining

    cs.LG 2025-10 conditional novelty 6.0 of 10

    Classifier-based quality filtering for LLM pretraining improves downstream tasks by implicitly filtering the reference high-quality set rather than by mimicking it, and its quality scores fail a data-conditioning test.

  3. Why do language models perform worse for morphologically complex languages?

    cs.CL 2024-11 conditional novelty 6.0 of 10

    A language-modeling performance gap between agglutinative and fusional languages largely disappears when training data is measured and scaled in bytes rather than tokens.

  4. Training Bilingual LMs with Data Constraints in the Targeted Language

    cs.CL 2024-11 conditional novelty 6.0 of 10

    Higher-quality auxiliary English pretraining data improves target-language performance for languages close to English (about 2% on translated QA tasks), but not for distant languages, when target-language data is limi...

  5. Beyond Human-Like Processing: Large Language Models Perform Equivalently on Forward and Backward Scientific Text

    cs.CL 2024-11 conditional novelty 5.0 of 10

    GPT-2 models trained on character-reversed neuroscience text perform as well on a neuroscience abstract-selection benchmark as models trained on normal text, despite higher perplexity.

  6. Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models

    cs.LG 2024-12 conditional novelty 4.0 of 10

    This survey organizes LLM synthetic data research around quality, diversity, and complexity, claiming quality mainly helps in-distribution generalization, diversity mainly helps out-of-distribution generalization, and...

Pith tools