REVIEW 4 cited by
State-of-the-art generalisation research in NLP: A taxonomy and review
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The ability to generalise well is one of the primary desiderata of natural language processing (NLP). Yet, what 'good generalisation' entails and how it should be evaluated is not well understood, nor are there any evaluation standards for generalisation. In this paper, we lay the groundwork to address both of these issues. We present a taxonomy for characterising and understanding generalisation research in NLP. Our taxonomy is based on an extensive literature review of generalisation research, and contains five axes along which studies can differ: their main motivation, the type of generalisation they investigate, the type of data shift they consider, the source of this data shift, and the locus of the shift within the modelling pipeline. We use our taxonomy to classify over 400 papers that test generalisation, for a total of more than 600 individual experiments. Considering the results of this review, we present an in-depth analysis that maps out the current state of generalisation research in NLP, and we make recommendations for which areas might deserve attention in the future. Along with this paper, we release a webpage where the results of our review can be dynamically explored, and which we intend to update as new NLP generalisation studies are published. With this work, we aim to take steps towards making state-of-the-art generalisation testing the new status quo in NLP.
Forward citations
Cited by 4 Pith papers
-
Investigating the (De)Composition Capabilities of Large Language Models in Natural-to-Formal Language Conversion
LLMs show measurable deficiencies in both decomposition and composition during natural-to-formal conversion, with decomposition errors dominating, under the new DEDC evaluation framework.
-
Propositional Logic for Probing Generalization in Neural Networks
Standard neural architectures generalize to unseen variable and operator combinations, but systematically fail when negation is applied to an operator that was hidden during training.
-
Confidence Estimation for Error Detection in Text-to-SQL Systems
Entropy-based selective classifiers can detect errors in text-to-SQL systems, and T5 models are better calibrated than GPT-4 and Llama 3 under distribution shift.
-
A Critical Field Guide for Working with Machine Learning Datasets
A field guide from the Knowing Machines project that turns existing critical dataset studies scholarship into lifecycle questions for practitioners, with no new empirical or formal results.
Discussion (0). Continue with ORCID to comment.