Pith. sign in

REVIEW 13 cited by

Is ChatGPT a Highly Fluent Grammatical Error Correction System? A Comprehensive Evaluation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.01746 v1 pith:QRG52T5N submitted 2023-04-04 cs.CL

classification cs.CL
keywords chatgpterrorserrorpotentialcapabilitiescomprehensivecorrectcorrection
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

ChatGPT, a large-scale language model based on the advanced GPT-3.5 architecture, has shown remarkable potential in various Natural Language Processing (NLP) tasks. However, there is currently a dearth of comprehensive study exploring its potential in the area of Grammatical Error Correction (GEC). To showcase its capabilities in GEC, we design zero-shot chain-of-thought (CoT) and few-shot CoT settings using in-context learning for ChatGPT. Our evaluation involves assessing ChatGPT's performance on five official test sets in three different languages, along with three document-level GEC test sets in English. Our experimental results and human evaluations demonstrate that ChatGPT has excellent error detection capabilities and can freely correct errors to make the corrected sentences very fluent, possibly due to its over-correction tendencies and not adhering to the principle of minimal edits. Additionally, its performance in non-English and low-resource settings highlights its potential in multilingual GEC tasks. However, further analysis of various types of errors at the document-level has shown that ChatGPT cannot effectively correct agreement, coreference, tense errors across sentences, and cross-sentence boundary errors.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Patent-CR: A Dataset for Patent Claim Revision

    cs.CL 2024-12 conditional novelty 7.0 of 10

    Patent-CR provides the first English patent claim revision dataset, and benchmark results show current LLMs, including GPT-4, cannot yet revise claims to examination standard.

  2. Benchmarking the Detection of LLMs-Generated Modern Chinese Poetry

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A new modern Chinese poetry detection benchmark shows most current AI-text detectors are unreliable, particularly when LLMs imitate a human style.

  3. APIO: Automatic Prompt Induction and Optimization for Grammatical Error Correction and Text Simplification

    cs.CL 2025-08 conditional novelty 6.0 of 10

    APIO automatically induces and optimizes instruction-list prompts for grammatical error correction and text simplification, reporting improved scores over prior prompt-based methods on BEA-2019 and ASSET.

  4. Adapting LLMs for Minimal-edit Grammatical Error Correction

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A lower learning rate on correct examples after training on errors lets Gemma 2 set a new single-model SOTA on BEA-test, aided by adding unedited pairs.

  5. Which Prompting Technique Should I Use? An Empirical Investigation of Prompting Techniques for Software Engineering Tasks

    cs.SE 2025-06 conditional novelty 6.0 of 10

    Across ten software engineering tasks and four LLMs, no prompting technique wins consistently; ES-KNN is best on many tasks, some techniques underperform the baseline, and USC is best for code QA and code generation.

  6. Beyond Guilt: Legal Judgment Prediction with Trichotomous Reasoning

    cs.CL 2024-12 conditional novelty 6.0 of 10

    The paper introduces LJPIV, the first legal judgment prediction benchmark with innocent verdicts, and shows that trichotomous reasoning prompts and fine-tuning improve prediction of not-guilty outcomes.

  7. Improving Explainability of Sentence-level Metrics via Edit-level Attribution for Grammatical Error Correction

    cs.CL 2024-12 conditional novelty 6.0 of 10

    Sentence-level GEC metric scores are decomposed into per-edit Shapley attributions, enabling edit-level explanation and error-type analysis.

  8. DSGram: Dynamic Weighting Sub-Metrics for Grammatical Error Correction in the Era of Large Language Models

    cs.CL 2024-12 conditional novelty 6.0 of 10

    DSGram is a reference-free GEC evaluation metric that dynamically weights Semantic Coherence, Edit Level, and Fluency using LLM-generated AHP weights, and reports improved correlation with human judgments on the SEEDA...

  9. LLMCL-GEC: Advancing Grammatical Error Correction with LLM-Driven Curriculum Learning

    cs.CL 2024-12 conditional novelty 6.0 of 10

    An LLM-scored easy-to-hard curriculum for training grammatical error correction models yields small but consistent F0.5 gains over one-shot and length-based training.

  10. Harnessing Rule-Based Reinforcement Learning for Enhanced Grammatical Error Correction

    cs.CL 2025-08 conditional novelty 5.0 of 10

    Applying GRPO with a rule-based, reference-match reward to a Qwen3-8B model after reasoning-augmented SFT achieves state-of-the-art F0.5 on Chinese GEC benchmark FCGEC and improves recall.

  11. Predicting Compact Phrasal Rewrites with Large Language Models for ASR Post Editing

    cs.CL 2025-01 conditional novelty 5.0 of 10

    A target-phrase-only edit representation offers the best accuracy-versus-output-length trade-off for LLM-based ASR post editing, closing 50-60% of the WER gap to full rewriting while losing only 10-20% of the length savings.

  12. Trigger$^3$: Refining Query Correction via Adaptive Model Selector

    cs.CL 2024-12 conditional novelty 5.0 of 10

    Trigger3 uses three trained triggers to route Chinese search queries among a small correction model, an LLM, and the original query, improving F0.5 on two datasets while lowering LLM coverage.

  13. CEC-Zero: Chinese Error Correction Solution Based on LLM

    cs.CL 2025-05 reject novelty 4.0 of 10

    The authors claim that reinforcement learning with an embedding-clustering reward improves Chinese spelling correction and cross-domain generalization, but the evidence is missing key baselines and reproducibility artifacts.

Pith tools