Pith. sign in

REVIEW

Benchmarking Machine Reading Comprehension: A Psychological Perspective

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.01912 v2 pith:PQ2EEEDY submitted 2020-04-04 cs.CL

classification cs.CL
keywords comprehensiondesignmodelreadingbenchmarkingdatasetsmachinetask
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Machine reading comprehension (MRC) has received considerable attention as a benchmark for natural language understanding. However, the conventional task design of MRC lacks explainability beyond the model interpretation, i.e., reading comprehension by a model cannot be explained in human terms. To this end, this position paper provides a theoretical basis for the design of MRC datasets based on psychology as well as psychometrics, and summarizes it in terms of the prerequisites for benchmarking MRC. We conclude that future datasets should (i) evaluate the capability of the model for constructing a coherent and grounded representation to understand context-dependent situations and (ii) ensure substantive validity by shortcut-proof questions and explanation as a part of the task design.

Discussion (0). Continue with ORCID to comment.

Pith tools