Pith. sign in

REVIEW 4 cited by

State-of-the-art generalisation research in NLP: A taxonomy and review

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.03050 v4 pith:K6NAOFTL submitted 2022-10-06 cs.CL cs.AI

classification cs.CLcs.AI
keywords generalisationresearchreviewtaxonomyshiftalongdataresults
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The ability to generalise well is one of the primary desiderata of natural language processing (NLP). Yet, what 'good generalisation' entails and how it should be evaluated is not well understood, nor are there any evaluation standards for generalisation. In this paper, we lay the groundwork to address both of these issues. We present a taxonomy for characterising and understanding generalisation research in NLP. Our taxonomy is based on an extensive literature review of generalisation research, and contains five axes along which studies can differ: their main motivation, the type of generalisation they investigate, the type of data shift they consider, the source of this data shift, and the locus of the shift within the modelling pipeline. We use our taxonomy to classify over 400 papers that test generalisation, for a total of more than 600 individual experiments. Considering the results of this review, we present an in-depth analysis that maps out the current state of generalisation research in NLP, and we make recommendations for which areas might deserve attention in the future. Along with this paper, we release a webpage where the results of our review can be dynamically explored, and which we intend to update as new NLP generalisation studies are published. With this work, we aim to take steps towards making state-of-the-art generalisation testing the new status quo in NLP.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Investigating the (De)Composition Capabilities of Large Language Models in Natural-to-Formal Language Conversion

    cs.CL 2025-01 conditional novelty 7.0 of 10

    LLMs show measurable deficiencies in both decomposition and composition during natural-to-formal conversion, with decomposition errors dominating, under the new DEDC evaluation framework.

  2. Propositional Logic for Probing Generalization in Neural Networks

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Standard neural architectures generalize to unseen variable and operator combinations, but systematically fail when negation is applied to an operator that was hidden during training.

  3. Confidence Estimation for Error Detection in Text-to-SQL Systems

    cs.LG 2025-01 conditional novelty 6.0 of 10

    Entropy-based selective classifiers can detect errors in text-to-SQL systems, and T5 models are better calibrated than GPT-4 and Llama 3 under distribution shift.

  4. A Critical Field Guide for Working with Machine Learning Datasets

    cs.CY 2025-01 unverdicted novelty 2.0 of 10

    A field guide from the Knowing Machines project that turns existing critical dataset studies scholarship into lifecycle questions for practitioners, with no new empirical or formal results.

Pith tools