Pith. sign in

REVIEW 1 cited by

Improving Explainability of Sentence-level Metrics via Edit-level Attribution for Grammatical Error Correction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.13110 v1 pith:P3RWYJQH submitted 2024-12-17 cs.CL

classification cs.CL
keywords metricsattributionexplainabilitysentence-levelcorrectionediteditserror
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Various evaluation metrics have been proposed for Grammatical Error Correction (GEC), but many, particularly reference-free metrics, lack explainability. This lack of explainability hinders researchers from analyzing the strengths and weaknesses of GEC models and limits the ability to provide detailed feedback for users. To address this issue, we propose attributing sentence-level scores to individual edits, providing insight into how specific corrections contribute to the overall performance. For the attribution method, we use Shapley values, from cooperative game theory, to compute the contribution of each edit. Experiments with existing sentence-level metrics demonstrate high consistency across different edit granularities and show approximately 70\% alignment with human evaluations. In addition, we analyze biases in the metrics based on the attribution results, revealing trends such as the tendency to ignore orthographic edits. Our implementation is available at \url{https://github.com/naist-nlp/gec-attribute}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. gec-metrics: A Unified Library for Grammatical Error Correction Evaluation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A new open-source library unifies ten GEC evaluation metrics and meta-evaluation frameworks, plus new empirical results including an ensemble that reaches 0.984 Spearman on SEEDA-E.

Pith tools