Pith. sign in

REVIEW 2 cited by

Language models and Automated Essay Scoring

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.09482 v1 pith:EMM6DJAC submitted 2019-09-18 cs.CL cs.LGstat.ML

Language models and Automated Essay Scoring

classification cs.CL cs.LGstat.ML
keywords architecturesnetworklanguagemodelsbertcompareessayneural
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

In this paper, we present a new comparative study on automatic essay scoring (AES). The current state-of-the-art natural language processing (NLP) neural network architectures are used in this work to achieve above human-level accuracy on the publicly available Kaggle AES dataset. We compare two powerful language models, BERT and XLNet, and describe all the layers and network architectures in these models. We elucidate the network architectures of BERT and XLNet using clear notation and diagrams and explain the advantages of transformer architectures over traditional recurrent neural network architectures. Linear algebra notation is used to clarify the functions of transformers and attention mechanisms. We compare the results with more traditional methods, such as bag of words (BOW) and long short term memory (LSTM) networks.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Automated Essay Scoring and Language Certification: Assessing Generalizability, Agreement and Validity for French

    cs.CL 2026-06 unverdicted novelty 5.0

    Enhanced ABV framework applied to French AES, comparing 8 models on 27k and 961-essay corpora to assess generalizability, agreement, and validity.

  2. Reconstructing Item Characteristic Curves using Fine-Tuned Large Language Models

    cs.CL 2026-01 conditional novelty 5.0

    Fine-tuned LLMs can reconstruct item characteristic curves from multiple-choice item text, giving useful estimates of IRT difficulty and discrimination without live student response data.