Pith. sign in

REVIEW 1 cited by

Exploring Unsupervised Pretraining and Sentence Structure Modelling for Winograd Schema Challenge

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1904.09705 v1 pith:2QFZ7NWH submitted 2019-04-22 cs.CL cs.AI

classification cs.CLcs.AI
keywords fine-tuningmodellingschemasentencewinogradachievingchallengehelps
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Winograd Schema Challenge (WSC) was proposed as an AI-hard problem in testing computers' intelligence on common sense representation and reasoning. This paper presents the new state-of-theart on WSC, achieving an accuracy of 71.1%. We demonstrate that the leading performance benefits from jointly modelling sentence structures, utilizing knowledge learned from cutting-edge pretraining models, and performing fine-tuning. We conduct detailed analyses, showing that fine-tuning is critical for achieving the performance, but it helps more on the simpler associative problems. Modelling sentence dependency structures, however, consistently helps on the harder non-associative subset of WSC. Analysis also shows that larger fine-tuning datasets yield better performances, suggesting the potential benefit of future work on annotating more Winograd schema sentences.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Align, Mask and Select: A Simple Method for Incorporating Commonsense Knowledge into Language Representation Models

    cs.CL 2019-08 conditional novelty 6.0 of 10

    Pre-training BERT on automatically generated multiple-choice questions from ConceptNet and Wikipedia improves commonsense benchmarks and leaves GLUE performance essentially unchanged.

Pith tools