REVIEW 2 major objections 1 minor 1 cited by
Quality Without Usefulness: LLM-Generated XAI Narratives as Trust Heuristics Rather Than Decision Aids
T0 review · 2 major / 1 minor · reviewed 2026-06-29 · grok-4.3
Pith's one-line read LLM-generated explanations raise confidence without improving accuracy on forecasting tasks
desk verdict The paper shows high-quality LLM NLEs add no accuracy lift in energy forecasting tasks but raise confidence via text presence, with an OOD task where they mask failures; the placebic control is the key but thinly described element. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The Quality-Usefulness Gap, the observed separation between high scores on text-quality metrics and absence of gains in downstream task performance.
What would settle it
A new experiment in which high-quality NLEs produce a measurable increase in accuracy on one or more of the five tasks, with other variables held constant, would falsify the claim of no accuracy improvement.
Extended reading notes
Core claim
Holding NLE quality constant at high levels, the explanations do not improve task accuracy on any of the five tasks, while inflating self-reported confidence. A placebic control shows the confidence boost stems from text presence rather than content. In an out-of-distribution detection task, NLEs reduce the LLM judge's ability to flag unreliable predictions, providing false reassurance that masks model failure. The findings are characterised as the Quality-Usefulness Gap.
Load-bearing premise
The five experiments each isolate a distinct facet of usefulness from the XAI literature, and the placebic control separates the effect of text presence from explanation content.
Editorial extensions
If this is right
- Evaluation of the XAI-to-NLE pipeline must extend beyond text-quality metrics to include downstream task performance.
- NLEs can reduce detection of model failures by providing false reassurance.
- Confidence effects arise from the presence of text rather than its specific explanatory content.
- Multiple distinct facets of usefulness must be tested separately rather than assumed to follow from text quality.
Reading between the lines
- Comparable gaps may appear when the same NLE generation approach is applied to other domains such as medical or financial decisions.
- Resources might be better allocated to improving underlying model accuracy than to post-hoc explanation generation.
- Training users to discount explanations could be tested as a way to restore proper confidence calibration.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript reports five controlled experiments (2,730 judgments across 60 test instances) in a time-series energy forecasting domain. Holding NLE quality constant at high levels from prior work, the authors find that LLM-generated natural language explanations do not improve accuracy on any of five operationalized usefulness tasks, inflate self-reported confidence (attributed to text presence via a placebic control), and impair out-of-distribution detection by masking model failures. The central claim is the existence of a 'Quality-Usefulness Gap' requiring evaluation beyond text-quality metrics.
Significance. If the results hold after clarification of controls, the work provides empirical evidence that high-quality NLEs can act as trust heuristics without decision-aid value, with direct implications for XAI evaluation practices. The multi-experiment design, placebic control, and OOD task are strengths that make the findings falsifiable and relevant to the field.
major comments (2)
- [Abstract and Methods] Abstract and Methods (placebic control description): The claim that confidence inflation is driven by text presence rather than content rests on the placebic control having zero explanatory value. However, no details are provided on its generation method, length, structure, or lexical properties. If the placebo retains sentence-like form or statistical regularities, the isolation fails and the Quality-Usefulness Gap attribution becomes ambiguous. This is load-bearing for the interpretation of all five experiments' confidence results.
- [Results] Results (five tasks): The abstract states that NLEs 'do not improve task accuracy on any of the five tasks,' but without explicit reporting of effect sizes, confidence intervals, or power analysis for each operationalization, it is difficult to assess whether null results reflect true absence of usefulness or underpowered tests. This directly affects the strength of the central claim.
minor comments (1)
- [Abstract] The abstract refers to 'a prior factorial study' for establishing NLE quality levels; a brief citation or summary of that study's key quality metrics would improve self-containment.
Simulated Author's Rebuttal
We thank the referee for their detailed and constructive review. The two major comments highlight important areas for clarification regarding the placebic control and statistical reporting of null results. We address each point below and commit to revisions that strengthen the manuscript without altering its core claims.
read point-by-point responses
-
Referee: [Abstract and Methods] Abstract and Methods (placebic control description): The claim that confidence inflation is driven by text presence rather than content rests on the placebic control having zero explanatory value. However, no details are provided on its generation method, length, structure, or lexical properties. If the placebo retains sentence-like form or statistical regularities, the isolation fails and the Quality-Usefulness Gap attribution becomes ambiguous. This is load-bearing for the interpretation of all five experiments' confidence results.
Authors: We agree that explicit details on the placebic control are necessary to support the interpretation. The control was generated by sampling non-explanatory, grammatically correct sentences from a general corpus and matching them for length and sentence count to the NLEs, with no domain-specific content or predictive information. We will expand the Methods section in the revision to fully describe the generation procedure, length statistics, structural properties, and any lexical comparisons performed, allowing readers to evaluate the isolation directly. revision: yes
-
Referee: [Results] Results (five tasks): The abstract states that NLEs 'do not improve task accuracy on any of the five tasks,' but without explicit reporting of effect sizes, confidence intervals, or power analysis for each operationalization, it is difficult to assess whether null results reflect true absence of usefulness or underpowered tests. This directly affects the strength of the central claim.
Authors: We acknowledge that additional statistical detail will improve transparency around the null findings. In the revision we will report Cohen's d (or equivalent) effect sizes and 95% confidence intervals for all five tasks. Power analysis was conducted a priori using G*Power based on medium effect sizes from prior XAI studies and our planned sample of 2,730 judgments; we will include these calculations and achieved power estimates per task in a new supplementary table. revision: yes
Circularity Check
No circularity: purely empirical study with no derivations or self-referential reductions
full rationale
The paper reports results from five controlled experiments (2,730 judgments) that test usefulness of NLEs while holding quality constant from a prior study. No equations, fitted parameters, predictions derived from inputs, or mathematical derivations appear in the provided text. Claims rest on experimental outcomes and the placebic control rather than any definitional or self-citation chain that reduces the result to its own inputs by construction. The prior factorial study citation establishes a baseline but does not bear the load of the current empirical findings or create uniqueness/self-definition issues.
Assumptions & free parameters
assumptions (2)
- domain assumption The prior factorial study established high NLE quality levels that can be held constant across conditions.
- standard math Standard principles of controlled experimental design and statistical inference apply to human judgment tasks.
Cite this review
Pith. "Pith review of Quality Without Usefulness: LLM-Generated XAI Narratives as Trust Heuristics Rather Than Decision Aids." pith.science (2026). https://pith.science/paper/F6FWTXZ7
@misc{pith2026260526770,
author = {Pith},
title = {Pith review of: Quality Without Usefulness: LLM-Generated XAI Narratives as Trust Heuristics Rather Than Decision Aids},
year = {2026},
howpublished = {\url{https://pith.science/paper/F6FWTXZ7}},
note = {Machine review of arXiv:2605.26770}
}
read the original abstract
Prior work shows that Large Language Models (LLMs) can transform Explainable AI (XAI) outputs into Natural Language Explanations (NLEs) that score highly on quality metrics such as plausibility, coherence, and comprehensibility. But does explanation quality translate to practical usefulness? We investigate this question in a time-series energy forecasting domain through five controlled experiments (2,730 judgments across 60 test instances), each operationalising a distinct facet of usefulness studied in the XAI literature. Holding NLE quality constant at the high levels established by a prior factorial study, we find that NLEs do not improve task accuracy on any of the five tasks, while inflating self-reported confidence. A placebic control shows that this confidence boost is driven by text presence rather than content. In an out-of-distribution detection task, NLEs reduce the LLM judge's ability to flag unreliable predictions, providing false reassurance that masks model failure. We characterise these findings as the Quality-Usefulness Gap and argue that evaluation of the XAI-to-NLE pipeline must extend beyond text-quality metrics to downstream task performance.
Figures
Forward citations
Cited by 1 Pith paper
-
Scaffold, Not Vocabulary? A Controlled, Two-Tier, Pre-Registered Study of a Popperian Code-Generation Skill
Controlled ablation finds Popperian code-generation skill adds no separable correctness benefit over labels-only scaffold; gains track structure not content.
Reference graph
Works this paper leans on
-
[1]
Zana Bucinca, Phoebe Lin, Krzysztof Z
Evaluating explanations through llms: Be- yond traditional user studies.arXiv preprint arXiv:2410.17781. Zana Bucinca, Phoebe Lin, Krzysztof Z. Gajos, and Elena L. Glassman. 2020. Proxy tasks and subjective measures can be misleading in evaluating explainable ai systems. InProceedings of the 25th International Conference on Intelligent User Interfaces. Za...
-
[2]
InProceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 3681–3688
Interpretation of neural networks is fragile. InProceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 3681–3688. Jiawei Gu, Xuhui Jiang, Zhichao Shi, and 1 others
-
[3]
A survey on llm-as-a-judge.arXiv preprint arXiv:2411.15594. Jiawei Gu, Xuhui Xu, Junyi Ye, Tianyi Zhang, Ming Cheng, and Wenbo Jiao. 2024. A survey on LLM-as- a-judge.Preprint, arXiv:2411.15594. Riccardo Guidotti, Anna Monreale, Salvatore Ruggieri, Franco Turini, Fosca Giannotti, and Dino Pedreschi
work page Pith review arXiv 2024
-
[4]
Daya Guo, Dejian Yang, He Zhang, and 1 others
A survey of methods for explaining black box models.ACM Computing Surveys, 51(5):1–42. Daya Guo, Dejian Yang, He Zhang, and 1 others
-
[5]
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
DeepSeek-R1: Incentivizing reasoning capa- bility in LLMs via reinforcement learning.Preprint, arXiv:2501.12948. Peter Hase and Mohit Bansal. 2020. Evaluating ex- plainable AI: Which algorithmic explanations help users predict model behavior? InProceedings of the 58th Annual Meeting of the Association for Compu- tational Linguistics, pages 5540–5552. Asso...
work page Pith review arXiv 2020
-
[6]
In Conference on Fairness, Accountability, and Trans- parency (FAccT)
How can i choose an explainer? an application- grounded evaluation of post-hoc explanations. In Conference on Fairness, Accountability, and Trans- parency (FAccT). Daniel Kahneman. 2011.Thinking, Fast and Slow. Far- rar, Straus and Giroux. Ellen J. Langer, Arthur Blank, and Benzion Chanowitz
2011
-
[7]
The mindlessness of ostensibly thoughtful ac- tion: The role of “placebic” information in interper- sonal interaction.Journal of Personality and Social Psychology, 36(6):635–642. John D. Lee and Katrina A. See. 2004. Trust in automa- tion: Designing for appropriate reliance.Human Factors, 46(1):50–80. Russell V . Lenth. 2024. emmeans: Estimated marginal m...
-
[8]
From local explanations to global understand- ing with explainable AI for trees.Nature Machine Intelligence, 2(1):56–67. Scott M. Lundberg and Su-In Lee. 2017. A unified approach to interpreting model predictions. InAd- vances in Neural Information Processing Systems, volume 30, pages 4765–4774. Curran Associates, Inc. 11 David Martens, Camille Dams, Jame...
2017
Show all 28 references
-
[9]
Sidra Naveed, Gunnar Stevens, and Dean Robin-Kern
From anecdotal evidence to quantitative eval- uation methods: A systematic review on evaluating explainable ai.ACM Computing Surveys, 55(13s). Sidra Naveed, Gunnar Stevens, and Dean Robin-Kern
-
[10]
Raja Parasuraman and Victor Riley
An overview of the empirical evaluation of ex- plainable AI (XAI).Applied Sciences, 14(23):11288. Raja Parasuraman and Victor Riley. 1997. Humans and automation: Use, misuse, disuse, abuse.Human Factors, 39(2):230–253. Richard E. Petty and John T. Cacioppo. 1986. The elab- ora...
1997
-
[11]
Why Should I Trust You?
Dissenting explanations: Leveraging disagree- ment to reduce model overreliance. InProceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 21537–21544. Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016. “Why Should I Trust You?”: Ex- plaining...
2016
-
[12]
Andreas Theissler, Francesco Spinnato, Udo Schlegel, and Riccardo Guidotti
illuminate: An llm-xai framework leveraging social science explanation theories.arXiv preprint arXiv:2409.08027. Andreas Theissler, Francesco Spinnato, Udo Schlegel, and Riccardo Guidotti. 2022. Explainable AI for time series classification: A review, taxonomy and re- search d...
2022
-
[13]
FR"): df = pd.read_csv(file_path, sep=
Judging LLM-as-a-judge with MT-Bench and chatbot arena. InAdvances in Neural Information Processing Systems, volume 36. Alexandra Zytek, Dongyu Liu, Andrei Vasiliev, and Kalyan Veeramachaneni. 2024a. LLMs for XAI: Fu- ture directions for explaining explanations.Preprint, arXiv...
2006
-
[15]
Compare the prediction to recent lag values – does it follow the recent trend?
-
[16]
If model performance metrics are provided, use them to estimate a typical error range
-
[19]
error_bucket
Weigh all available evidence to determine the most likely error magnitude. Error bucket definitions (based on absolute percentage error): - small:[0%,5%) - medium:[5%,15%) - large:[15%,30%) - very_large:[30%,+∞) You must respond with EXACTLY this JSON format and nothing else: ...
-
[20]
Identify which features are being changed and note their current SHAP values
-
[21]
For each changed feature, determine the direction of change
-
[22]
Consider how the SHAP contribution might change
-
[23]
If multiple features are changed, think about their combined effect
-
[24]
If a natural language explanation is provided, you can use it to better understand the feature–prediction relationships
-
[25]
direction
Estimate the net directional effect and whether it exceeds the 5% threshold. Direction definitions: - higher: prediction increases by more than 5% relative to the original. - lower: prediction decreases by more than 5% relative to the original. - similar: prediction changes by...
-
[26]
Examine the features – consider the lag values, the time of year (season in France), holidays, and whether the prediction seems reasonable given this context
-
[27]
If SHAP values are provided, consider whether each feature’s contribution makes sense given the feature values and context
-
[28]
Consider the model performance metrics to gauge general trustworthiness
-
[29]
If a natural language explanation is provided, you can use it to better understand the model’s reasoning
-
[30]
reliability
Weigh all evidence to decide whether this prediction can be trusted. Reliability: - reliable = inputs look normal and within what the model was trained on. - unreliable = inputs look unusual or far from what the model was trained on. JSON response: { "reliability": "<reliable|...
2009
-
[31]
The direct Real-vs-Placebo equivalence test yields 43.1% inside ROPE: insufficient for formal practical equivalence (≥95%), but consistent with no meaningful difference be- tween real and placebo NLEs. Bayesian ROPE and Real-vs-Placebo Equiva- lence Sensitivity Judge-specific....
Reviewed June 29, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.