This paper introduces an automated evaluation framework with extraction-based faithfulness metrics, perplexity for assumptions, and embedding-based human similarity, and shows it can reveal LLM sign self-correction on manipulated SHAP tables.
Flemish AI Research Program
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives
This paper introduces an automated evaluation framework with extraction-based faithfulness metrics, perplexity for assumptions, and embedding-based human similarity, and shows it can reveal LLM sign self-correction on manipulated SHAP tables.