Pith. sign in

REVIEW 1 cited by

Probing for targeted syntactic knowledge through grammatical error detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.16228 v1 pith:E7QRAXSB submitted 2022-10-28 cs.CL

classification cs.CL
keywords modelsagreementdetectionlanguagewhenencodeerrorinformation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Targeted studies testing knowledge of subject-verb agreement (SVA) indicate that pre-trained language models encode syntactic information. We assert that if models robustly encode subject-verb agreement, they should be able to identify when agreement is correct and when it is incorrect. To that end, we propose grammatical error detection as a diagnostic probe to evaluate token-level contextual representations for their knowledge of SVA. We evaluate contextual representations at each layer from five pre-trained English language models: BERT, XLNet, GPT-2, RoBERTa, and ELECTRA. We leverage public annotated training data from both English second language learners and Wikipedia edits, and report results on manually crafted stimuli for subject-verb agreement. We find that masked language models linearly encode information relevant to the detection of SVA errors, while the autoregressive models perform on par with our baseline. However, we also observe a divergence in performance when probes are trained on different training sets, and when they are evaluated on different syntactic constructions, suggesting the information pertaining to SVA error detection is not robustly encoded.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Enhancing Grammatical Error Detection using BERT with Cleaned Lang-8 Dataset

    cs.CL 2024-11 reject novelty 3.0 of 10

    Fine-tuning BERT-base-uncased on a hand-cleaned Lang-8 subset yields F1 0.91 on that same distribution, but the result is not benchmarked against standard GED tests.

Pith tools