Pith. sign in

REVIEW 12 cited by

Human-AI Collaborative Essay Scoring: A Dual-Process Framework with LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.06431 v2 pith:XEP63FEV submitted 2024-01-12 cs.CL cs.AI

classification cs.CLcs.AI
keywords llmsfeedbackgradinghuman-aimodelssystemdual-processessay
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Receiving timely and personalized feedback is essential for second-language learners, especially when human instructors are unavailable. This study explores the effectiveness of Large Language Models (LLMs), including both proprietary and open-source models, for Automated Essay Scoring (AES). Through extensive experiments with public and private datasets, we find that while LLMs do not surpass conventional state-of-the-art (SOTA) grading models in performance, they exhibit notable consistency, generalizability, and explainability. We propose an open-source LLM-based AES system, inspired by the dual-process theory. Our system offers accurate grading and high-quality feedback, at least comparable to that of fine-tuned proprietary LLMs, in addition to its ability to alleviate misgrading. Furthermore, we conduct human-AI co-grading experiments with both novice and expert graders. We find that our system not only automates the grading process but also enhances the performance and efficiency of human graders, particularly for essays where the model has lower confidence. These results highlight the potential of LLMs to facilitate effective human-AI collaboration in the educational context, potentially transforming learning experiences through AI-generated feedback.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Score Prediction: LLM-Based Essay Scoring and Feedback Generation via Reinforcement Learning with Rubric Rewards

    cs.CL 2026-07 conditional novelty 6.0 of 10

    RLAES-AGFO trains an LLM with GRPO and an LLM-judged 166-item feedback rubric, reaching QWK 0.803 on ASAP while generating feedback rated as well as GPT-5.5.

  2. Investigating first-language bias in LLM-based automated essay scoring: A cross-prompt evaluation of an open-weight AI-model on TOEFL essays

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A Gemma-3 essay scorer trained on 480 essays generalizes across eight unseen TOEFL11 prompts but shows a consistent within-band scoring offset favoring European-language over East-Asian-language writers.

  3. Not All Jokes Land: Evaluating Large Language Models Understanding of Workplace Humor

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Five LLMs frequently misclassify the appropriateness of workplace humor, especially offensive and neutral jokes, on a new 304-item industrial humor dataset.

  4. How well can LLMs Grade Essays in Arabic?

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Generative LLMs, including Arabic-specific ones, underperform a fine-tuned BERT model on Arabic essay scoring, with bilingual prompting giving the best LLM results.

  5. Automatic Feedback Generation for Short Answer Questions using Answer Diagnostic Graphs

    cs.CL 2025-01 conditional novelty 6.0 of 10

    A new Answer Diagnostic Graph system generates personalized feedback for short-answer reading comprehension, improving students' self-reported error detection and motivation but not their scores.

  6. CAFES: A Collaborative Multi-Agent Framework for Multi-Granular Multimodal Essay Scoring

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A student-teacher multi-agent pipeline with positive-only feedback improves QWK agreement with human essay scores by 21% on a multimodal benchmark, with gains concentrated in traits where baselines were weakest.

  7. Advancing Student Writing Through Automated Syntax Feedback

    cs.CL 2025-01 conditional novelty 5.0 of 10

    Fine-tuning Llama-2 and Mistral models on a new GPT-generated essay-syntax-feedback dataset improves the quality of automated syntax corrections, according to human ratings.

  8. On the Suitability of pre-trained foundational LLMs for Analysis in German Legal Education

    cs.CL 2024-12 conditional novelty 5.0 of 10

    Pre-trained open LLMs underperform bag-of-words baselines on German Gutachtenstil identification and legal essay grading, but retrieval-based example selection narrows the gap on simpler tasks.

  9. A Benchmark for Math Misconceptions: Bridging Gaps in Middle School Algebra with AI-Supported Instruction

    cs.HC 2024-12 conditional novelty 5.0 of 10

    A new benchmark of 55 algebra misconceptions and 220 examples shows GPT-4-turbo diagnoses around 53% of misconceptions overall, 75% when topic-constrained, and 83.9% when educator feedback is included.

  10. Long Context Automated Essay Scoring with Language Models

    cs.CL 2025-09 conditional novelty 4.0 of 10

    On ASAP 2.0, long-context transformer and state-space models all reached quadratic weighted kappa above the 0.745 human-human baseline, led by Longformer at 0.798 and Mamba at 0.797.

  11. Interactive Sketchpad: A Multimodal Tutoring System for Collaborative, Visual Problem-Solving

    cs.HC 2025-02 reject novelty 4.0 of 10

    Interactive Sketchpad combines code-generated diagrams and an interactive whiteboard for math tutoring, but its learning-benefit claims rest on a small, uncontrolled user study.

  12. A Zero-Shot LLM Framework for Automatic Assignment Grading in Higher Education

    cs.CY 2025-01 conditional novelty 4.0 of 10

    A zero-shot, prompt-engineered GPT-4 system can grade open-ended statistics homework and produce personalized feedback, but the evidence that it improves learning over traditional grading is limited by the survey design.

Pith tools