Pith. sign in

REVIEW 2 cited by

Automated Text Scoring in the Age of Generative AI for the GPU-poor

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.01873 v1 pith:VWOTQDKL submitted 2024-07-02 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords glmsmodelsautomatedefficiencyfeedbackfocusedgenerativeopen-source
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Current research on generative language models (GLMs) for automated text scoring (ATS) has focused almost exclusively on querying proprietary models via Application Programming Interfaces (APIs). Yet such practices raise issues around transparency and security, and these methods offer little in the way of efficiency or customizability. With the recent proliferation of smaller, open-source models, there is the option to explore GLMs with computers equipped with modest, consumer-grade hardware, that is, for the "GPU poor." In this study, we analyze the performance and efficiency of open-source, small-scale GLMs for ATS. Results show that GLMs can be fine-tuned to achieve adequate, though not state-of-the-art, performance. In addition to ATS, we take small steps towards analyzing models' capacity for generating feedback by prompting GLMs to explain their scores. Model-generated feedback shows promise, but requires more rigorous evaluation focused on targeted use cases.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Score Prediction: LLM-Based Essay Scoring and Feedback Generation via Reinforcement Learning with Rubric Rewards

    cs.CL 2026-07 conditional novelty 6.0 of 10

    RLAES-AGFO trains an LLM with GRPO and an LLM-judged 166-item feedback rubric, reaching QWK 0.803 on ASAP while generating feedback rated as well as GPT-5.5.

  2. Automated Essay Scoring Incorporating Annotations from Automated Feedback Systems

    cs.CL 2025-05 conditional novelty 4.0 of 10

    On the PERSUADE corpus, adding generated argument-component tags to essay text raised automated scoring agreement from a QWK of 0.860 to 0.868, while error-only tags lowered it.

Pith tools