Pith. sign in

REVIEW 1 cited by

CEval: A Benchmark for Evaluating Counterfactual Text Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.17475 v2 pith:GLRDM33F submitted 2024-04-26 cs.CL cs.AI

classification cs.CLcs.AI
keywords counterfactualtextcevalgenerationmethodsmetricsbenchmarkmethod
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Counterfactual text generation aims to minimally change a text, such that it is classified differently. Judging advancements in method development for counterfactual text generation is hindered by a non-uniform usage of data sets and metrics in related work. We propose CEval, a benchmark for comparing counterfactual text generation methods. CEval unifies counterfactual and text quality metrics, includes common counterfactual datasets with human annotations, standard baselines (MICE, GDBA, CREST) and the open-source language model LLAMA-2. Our experiments found no perfect method for generating counterfactual text. Methods that excel at counterfactual metrics often produce lower-quality text while LLMs with simple prompts generate high-quality text but struggle with counterfactual criteria. By making CEval available as an open-source Python library, we encourage the community to contribute more methods and maintain consistent evaluation in future work.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PsyLite Technical Report

    cs.AI 2025-06 conditional novelty 4.0 of 10

    PsyLite fine-tunes InternLM2.5-7B-chat with QLoRA, R1 distillation, and ORPO to improve Chinese psychological counseling quality and dialogue safety, with local deployment in about 5GB memory.

Pith tools