Pith. sign in

REVIEW 1 cited by

Towards Automated Document Revision: Grammatical Error Correction, Fluency Edits, and Beyond

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.11484 v1 pith:SJLUPPVT submitted 2022-05-23 cs.CL

classification cs.CL
keywords revisionautomateddocumentgrammaticalcorrectiondocument-levelerrorexplore
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Natural language processing technology has rapidly improved automated grammatical error correction tasks, and the community begins to explore document-level revision as one of the next challenges. To go beyond sentence-level automated grammatical error correction to NLP-based document-level revision assistant, there are two major obstacles: (1) there are few public corpora with document-level revisions being annotated by professional editors, and (2) it is not feasible to elicit all possible references and evaluate the quality of revision with such references because there are infinite possibilities of revision. This paper tackles these challenges. First, we introduce a new document-revision corpus, TETRA, where professional editors revised academic papers sampled from the ACL anthology which contain few trivial grammatical errors that enable us to focus more on document- and paragraph-level edits such as coherence and consistency. Second, we explore reference-less and interpretable methods for meta-evaluation that can detect quality improvements by document revision. We show the uniqueness of TETRA compared with existing document revision corpora and demonstrate that a fine-tuned pre-trained language model can discriminate the quality of documents after revision even when the difference is subtle. This promising result will encourage the community to further explore automated document revision models and metrics in future.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ScholaWrite: A Dataset of End-to-End Scholarly Writing Process

    cs.HC 2025-02 conditional novelty 6.0 of 10

    ScholaWrite records over 61,000 keystroke-level edits from five real scholarly preprints, each annotated with one of 15 cognitive writing intentions, and analyzes how writing unfolds non-linearly over months.

Pith tools