REVIEW 3 cited by
Masked Language Modeling and the Distributional Hypothesis: Order Word Matters Pre-training for Little
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
A possible explanation for the impressive performance of masked language model (MLM) pre-training is that such models have learned to represent the syntactic structures prevalent in classical NLP pipelines. In this paper, we propose a different explanation: MLMs succeed on downstream tasks almost entirely due to their ability to model higher-order word co-occurrence statistics. To demonstrate this, we pre-train MLMs on sentences with randomly shuffled word order, and show that these models still achieve high accuracy after fine-tuning on many downstream tasks -- including on tasks specifically designed to be challenging for models that ignore word order. Our models perform surprisingly well according to some parametric syntactic probes, indicating possible deficiencies in how we test representations for syntactic information. Overall, our results show that purely distributional information largely explains the success of pre-training, and underscore the importance of curating challenging evaluation datasets that require deeper linguistic knowledge.
Forward citations
Cited by 3 Pith papers
-
Rethinking Addressing in Language Models via Contexualized Equivariant Positional Encoding
TAPE makes positional embeddings content-aware and equivariant, improving Transformer performance on arithmetic and long-context tasks and extending representational power to NC1-complete algorithms.
-
What makes a good metric? Evaluating automatic metrics for text-to-image consistency
None of the four tested text-to-image consistency metrics satisfies all proposed validity criteria, and the VQA-based metrics appear to rely largely on text priors such as yes-bias.
-
Transfer of Structural Knowledge from Synthetic Languages
A new synthetic language, flat_shuffle, transfers more structure to English fine-tuning than earlier synthetic bracket languages, though still far short of training on English from scratch.
Discussion (0). Continue with ORCID to comment.