Pith. sign in

REVIEW 3 cited by

Language models and Automated Essay Scoring

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.09482 v1 pith:EMM6DJAC submitted 2019-09-18 cs.CL cs.LGstat.ML

classification cs.CLcs.LGstat.ML
keywords architecturesnetworklanguagemodelsbertcompareessayneural
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we present a new comparative study on automatic essay scoring (AES). The current state-of-the-art natural language processing (NLP) neural network architectures are used in this work to achieve above human-level accuracy on the publicly available Kaggle AES dataset. We compare two powerful language models, BERT and XLNet, and describe all the layers and network architectures in these models. We elucidate the network architectures of BERT and XLNet using clear notation and diagrams and explain the advantages of transformer architectures over traditional recurrent neural network architectures. Linear algebra notation is used to clarify the functions of transformers and attention mechanisms. We compare the results with more traditional methods, such as bag of words (BOW) and long short term memory (LSTM) networks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Adversarial Topic-aware Prompt-tuning for Cross-topic Automated Essay Scoring

    cs.CL 2025-08 reject novelty 6.0 of 10

    ATOP uses shared and topic-specific soft prompts with adversarial training and pseudo-labels to improve cross-topic automated essay scoring, reporting better QWK scores than nine baselines on ASAP++.

  2. Reconstructing Item Characteristic Curves using Fine-Tuned Large Language Models

    cs.CL 2026-01 conditional novelty 5.0 of 10

    Fine-tuned LLMs can reconstruct item characteristic curves from multiple-choice item text, giving useful estimates of IRT difficulty and discrimination without live student response data.

  3. Automated Essay Scoring Incorporating Annotations from Automated Feedback Systems

    cs.CL 2025-05 conditional novelty 4.0 of 10

    On the PERSUADE corpus, adding generated argument-component tags to essay text raised automated scoring agreement from a QWK of 0.860 to 0.868, while error-only tags lowered it.

Pith tools