Pith. sign in

REVIEW 2 cited by

Open Sesame: Getting Inside BERT's Linguistic Knowledge

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1906.01698 v1 pith:SB4D37G2 submitted 2019-06-04 cs.CL

classification cs.CL
keywords berthierarchicalinformationrepresentationsstructurelayerslinguisticsensitivity
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

How and to what extent does BERT encode syntactically-sensitive hierarchical information or positionally-sensitive linear information? Recent work has shown that contextual representations like BERT perform well on tasks that require sensitivity to linguistic structure. We present here two studies which aim to provide a better understanding of the nature of BERT's representations. The first of these focuses on the identification of structurally-defined elements using diagnostic classifiers, while the second explores BERT's representation of subject-verb agreement and anaphor-antecedent dependencies through a quantitative assessment of self-attention vectors. In both cases, we find that BERT encodes positional information about word tokens well on its lower layers, but switches to a hierarchically-oriented encoding on higher layers. We conclude then that BERT's representations do indeed model linguistically relevant aspects of hierarchical structure, though they do not appear to show the sharp sensitivity to hierarchical structure that is found in human processing of reflexive anaphora.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VisualBERT: A Simple and Performant Baseline for Vision and Language

    cs.CV 2019-08 conditional novelty 6.0 of 10

    VisualBERT is a Transformer model that implicitly aligns text and image regions through self-attention and achieves competitive or superior results on VQA, VCR, NLVR2, and Flickr30K after pre-training on captions.

  2. Beyond Human-Like Processing: Large Language Models Perform Equivalently on Forward and Backward Scientific Text

    cs.CL 2024-11 conditional novelty 5.0 of 10

    GPT-2 models trained on character-reversed neuroscience text perform as well on a neuroscience abstract-selection benchmark as models trained on normal text, despite higher perplexity.

Pith tools