Pith. sign in

REVIEW 1 cited by

A Survey on Natural Language Counterfactual Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.03993 v2 pith:CQI4C7EV submitted 2024-07-04 cs.CL

classification cs.CL
keywords generationcounterfactuallanguagemodelcounterfactualsdifferentfuturemethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Natural language counterfactual generation aims to minimally modify a given text such that the modified text will be classified into a different class. The generated counterfactuals provide insight into the reasoning behind a model's predictions by highlighting which words significantly influence the outcomes. Additionally, they can be used to detect model fairness issues and augment the training data to enhance the model's robustness. A substantial amount of research has been conducted to generate counterfactuals for various NLP tasks, employing different models and methodologies. With the rapid growth of studies in this field, a systematic review is crucial to guide future researchers and developers. To bridge this gap, this survey provides a comprehensive overview of textual counterfactual generation methods, particularly those based on Large Language Models. We propose a new taxonomy that systematically categorizes the generation methods into four groups and summarizes the metrics for evaluating the generation quality. Finally, we discuss ongoing research challenges and outline promising directions for future work.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives

    cs.CL 2024-12 conditional novelty 6.0 of 10

    This paper introduces an automated evaluation framework with extraction-based faithfulness metrics, perplexity for assumptions, and embedding-based human similarity, and shows it can reveal LLM sign self-correction on...

Pith tools