Pith. sign in

REVIEW 1 cited by

Masked Language Modeling and the Distributional Hypothesis: Order Word Matters Pre-training for Little

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.06644 v2 pith:MZ4KFFVC submitted 2021-04-14 cs.CL cs.LG

classification cs.CLcs.LG
keywords modelswordorderpre-trainingsyntactictaskschallengingdistributional
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A possible explanation for the impressive performance of masked language model (MLM) pre-training is that such models have learned to represent the syntactic structures prevalent in classical NLP pipelines. In this paper, we propose a different explanation: MLMs succeed on downstream tasks almost entirely due to their ability to model higher-order word co-occurrence statistics. To demonstrate this, we pre-train MLMs on sentences with randomly shuffled word order, and show that these models still achieve high accuracy after fine-tuning on many downstream tasks -- including on tasks specifically designed to be challenging for models that ignore word order. Our models perform surprisingly well according to some parametric syntactic probes, indicating possible deficiencies in how we test representations for syntactic information. Overall, our results show that purely distributional information largely explains the success of pre-training, and underscore the importance of curating challenging evaluation datasets that require deeper linguistic knowledge.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Transfer of Structural Knowledge from Synthetic Languages

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A new synthetic language, flat_shuffle, transfers more structure to English fine-tuning than earlier synthetic bracket languages, though still far short of training on English from scratch.

Pith tools