Pith. sign in

REVIEW 2 major objections 1 minor 1 cited by

Quality Without Usefulness: LLM-Generated XAI Narratives as Trust Heuristics Rather Than Decision Aids

T0 review · 2 major / 1 minor · reviewed 2026-06-29 · grok-4.3

Pith's one-line read LLM-generated explanations raise confidence without improving accuracy on forecasting tasks

desk verdict The paper shows high-quality LLM NLEs add no accuracy lift in energy forecasting tasks but raise confidence via text presence, with an OOD task where they mask failures; the placebic control is the key but thinly described element. read the letter →

arxiv 2605.26770 v1 pith:F6FWTXZ7 submitted 2026-05-26 cs.CL

classification cs.CL
keywords LLM-generatedexplanationsXAINaturalLanguageQuality-UsefulnessGapTrustHeuristicsDecisionAccuracyOut-of-DistributionDetectionEnergyForecasting
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tests whether high-quality natural language explanations generated by LLMs help users make better decisions in a time-series energy forecasting setting. Five controlled experiments show these explanations produce no accuracy gains across distinct usefulness measures while raising self-reported confidence. The confidence increase occurs with any text present, not because of the specific content. In an out-of-distribution detection task the explanations reduce the ability to spot unreliable predictions. The authors identify a Quality-Usefulness Gap and conclude that XAI evaluation must track downstream task performance rather than text quality alone.

What carries the argument

The Quality-Usefulness Gap, the observed separation between high scores on text-quality metrics and absence of gains in downstream task performance.

What would settle it

A new experiment in which high-quality NLEs produce a measurable increase in accuracy on one or more of the five tasks, with other variables held constant, would falsify the claim of no accuracy improvement.

Watch

Extended reading notes

Core claim

Holding NLE quality constant at high levels, the explanations do not improve task accuracy on any of the five tasks, while inflating self-reported confidence. A placebic control shows the confidence boost stems from text presence rather than content. In an out-of-distribution detection task, NLEs reduce the LLM judge's ability to flag unreliable predictions, providing false reassurance that masks model failure. The findings are characterised as the Quality-Usefulness Gap.

Load-bearing premise

The five experiments each isolate a distinct facet of usefulness from the XAI literature, and the placebic control separates the effect of text presence from explanation content.

Editorial extensions

If this is right

  • Evaluation of the XAI-to-NLE pipeline must extend beyond text-quality metrics to include downstream task performance.
  • NLEs can reduce detection of model failures by providing false reassurance.
  • Confidence effects arise from the presence of text rather than its specific explanatory content.
  • Multiple distinct facets of usefulness must be tested separately rather than assumed to follow from text quality.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Comparable gaps may appear when the same NLE generation approach is applied to other domains such as medical or financial decisions.
  • Resources might be better allocated to improving underlying model accuracy than to post-hoc explanation generation.
  • Training users to discount explanations could be tested as a way to restore proper confidence calibration.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The manuscript reports five controlled experiments (2,730 judgments across 60 test instances) in a time-series energy forecasting domain. Holding NLE quality constant at high levels from prior work, the authors find that LLM-generated natural language explanations do not improve accuracy on any of five operationalized usefulness tasks, inflate self-reported confidence (attributed to text presence via a placebic control), and impair out-of-distribution detection by masking model failures. The central claim is the existence of a 'Quality-Usefulness Gap' requiring evaluation beyond text-quality metrics.

Significance. If the results hold after clarification of controls, the work provides empirical evidence that high-quality NLEs can act as trust heuristics without decision-aid value, with direct implications for XAI evaluation practices. The multi-experiment design, placebic control, and OOD task are strengths that make the findings falsifiable and relevant to the field.

major comments (2)
  1. [Abstract and Methods] Abstract and Methods (placebic control description): The claim that confidence inflation is driven by text presence rather than content rests on the placebic control having zero explanatory value. However, no details are provided on its generation method, length, structure, or lexical properties. If the placebo retains sentence-like form or statistical regularities, the isolation fails and the Quality-Usefulness Gap attribution becomes ambiguous. This is load-bearing for the interpretation of all five experiments' confidence results.
  2. [Results] Results (five tasks): The abstract states that NLEs 'do not improve task accuracy on any of the five tasks,' but without explicit reporting of effect sizes, confidence intervals, or power analysis for each operationalization, it is difficult to assess whether null results reflect true absence of usefulness or underpowered tests. This directly affects the strength of the central claim.
minor comments (1)
  1. [Abstract] The abstract refers to 'a prior factorial study' for establishing NLE quality levels; a brief citation or summary of that study's key quality metrics would improve self-containment.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for their detailed and constructive review. The two major comments highlight important areas for clarification regarding the placebic control and statistical reporting of null results. We address each point below and commit to revisions that strengthen the manuscript without altering its core claims.

read point-by-point responses
  1. Referee: [Abstract and Methods] Abstract and Methods (placebic control description): The claim that confidence inflation is driven by text presence rather than content rests on the placebic control having zero explanatory value. However, no details are provided on its generation method, length, structure, or lexical properties. If the placebo retains sentence-like form or statistical regularities, the isolation fails and the Quality-Usefulness Gap attribution becomes ambiguous. This is load-bearing for the interpretation of all five experiments' confidence results.

    Authors: We agree that explicit details on the placebic control are necessary to support the interpretation. The control was generated by sampling non-explanatory, grammatically correct sentences from a general corpus and matching them for length and sentence count to the NLEs, with no domain-specific content or predictive information. We will expand the Methods section in the revision to fully describe the generation procedure, length statistics, structural properties, and any lexical comparisons performed, allowing readers to evaluate the isolation directly. revision: yes

  2. Referee: [Results] Results (five tasks): The abstract states that NLEs 'do not improve task accuracy on any of the five tasks,' but without explicit reporting of effect sizes, confidence intervals, or power analysis for each operationalization, it is difficult to assess whether null results reflect true absence of usefulness or underpowered tests. This directly affects the strength of the central claim.

    Authors: We acknowledge that additional statistical detail will improve transparency around the null findings. In the revision we will report Cohen's d (or equivalent) effect sizes and 95% confidence intervals for all five tasks. Power analysis was conducted a priori using G*Power based on medium effect sizes from prior XAI studies and our planned sample of 2,730 judgments; we will include these calculations and achieved power estimates per task in a new supplementary table. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: purely empirical study with no derivations or self-referential reductions

full rationale

The paper reports results from five controlled experiments (2,730 judgments) that test usefulness of NLEs while holding quality constant from a prior study. No equations, fitted parameters, predictions derived from inputs, or mathematical derivations appear in the provided text. Claims rest on experimental outcomes and the placebic control rather than any definitional or self-citation chain that reduces the result to its own inputs by construction. The prior factorial study citation establishes a baseline but does not bear the load of the current empirical findings or create uniqueness/self-definition issues.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The central claim rests on the validity of the experimental operationalizations of usefulness facets drawn from prior XAI literature and on the assumption that the placebic control isolates text presence effects; no free parameters or invented entities are introduced.

assumptions (2)
  • domain assumption The prior factorial study established high NLE quality levels that can be held constant across conditions.
    Invoked to isolate usefulness from quality in the current experiments.
  • standard math Standard principles of controlled experimental design and statistical inference apply to human judgment tasks.
    Underpins the five experiments and interpretation of accuracy and confidence measures.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Quality Without Usefulness: LLM-Generated XAI Narratives as Trust Heuristics Rather Than Decision Aids." pith.science (2026). https://pith.science/paper/F6FWTXZ7

@misc{pith2026260526770,
  author       = {Pith},
  title        = {Pith review of: Quality Without Usefulness: LLM-Generated XAI Narratives as Trust Heuristics Rather Than Decision Aids},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/F6FWTXZ7}},
  note         = {Machine review of arXiv:2605.26770}
}
read the original abstract

Prior work shows that Large Language Models (LLMs) can transform Explainable AI (XAI) outputs into Natural Language Explanations (NLEs) that score highly on quality metrics such as plausibility, coherence, and comprehensibility. But does explanation quality translate to practical usefulness? We investigate this question in a time-series energy forecasting domain through five controlled experiments (2,730 judgments across 60 test instances), each operationalising a distinct facet of usefulness studied in the XAI literature. Holding NLE quality constant at the high levels established by a prior factorial study, we find that NLEs do not improve task accuracy on any of the five tasks, while inflating self-reported confidence. A placebic control shows that this confidence boost is driven by text presence rather than content. In an out-of-distribution detection task, NLEs reduce the LLM judge's ability to flag unreliable predictions, providing false reassurance that masks model failure. We characterise these findings as the Quality-Usefulness Gap and argue that evaluation of the XAI-to-NLE pipeline must extend beyond text-quality metrics to downstream task performance.

Figures

Figures reproduced from arXiv: 2605.26770 by the authors.

Figure 1
Figure 1. An LLM narrates a prediction and XAI attributions [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. XGBoost one-step-ahead predictions vs. actual [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Experimental framework. Left: the LLM judge receives up to five information pieces per instance – features and [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Scaffold, Not Vocabulary? A Controlled, Two-Tier, Pre-Registered Study of a Popperian Code-Generation Skill

    cs.SE 2026-06 conditional novelty 7.0 of 10

    Controlled ablation finds Popperian code-generation skill adds no separable correctness benefit over labels-only scaffold; gains track structure not content.

Reference graph

Works this paper leans on

28 extracted references · 8 canonical work pages · cited by 1 Pith paper

  1. [1]

    Zana Bucinca, Phoebe Lin, Krzysztof Z

    Evaluating explanations through llms: Be- yond traditional user studies.arXiv preprint arXiv:2410.17781. Zana Bucinca, Phoebe Lin, Krzysztof Z. Gajos, and Elena L. Glassman. 2020. Proxy tasks and subjective measures can be misleading in evaluating explainable ai systems. InProceedings of the 25th International Conference on Intelligent User Interfaces. Za...

  2. [2]

    InProceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 3681–3688

    Interpretation of neural networks is fragile. InProceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 3681–3688. Jiawei Gu, Xuhui Jiang, Zhichao Shi, and 1 others

  3. [3]

    A Survey on LLM-as-a-Judge

    A survey on llm-as-a-judge.arXiv preprint arXiv:2411.15594. Jiawei Gu, Xuhui Xu, Junyi Ye, Tianyi Zhang, Ming Cheng, and Wenbo Jiao. 2024. A survey on LLM-as- a-judge.Preprint, arXiv:2411.15594. Riccardo Guidotti, Anna Monreale, Salvatore Ruggieri, Franco Turini, Fosca Giannotti, and Dino Pedreschi

  4. [4]

    Daya Guo, Dejian Yang, He Zhang, and 1 others

    A survey of methods for explaining black box models.ACM Computing Surveys, 51(5):1–42. Daya Guo, Dejian Yang, He Zhang, and 1 others

  5. [5]

    DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

    DeepSeek-R1: Incentivizing reasoning capa- bility in LLMs via reinforcement learning.Preprint, arXiv:2501.12948. Peter Hase and Mohit Bansal. 2020. Evaluating ex- plainable AI: Which algorithmic explanations help users predict model behavior? InProceedings of the 58th Annual Meeting of the Association for Compu- tational Linguistics, pages 5540–5552. Asso...

  6. [6]

    In Conference on Fairness, Accountability, and Trans- parency (FAccT)

    How can i choose an explainer? an application- grounded evaluation of post-hoc explanations. In Conference on Fairness, Accountability, and Trans- parency (FAccT). Daniel Kahneman. 2011.Thinking, Fast and Slow. Far- rar, Straus and Giroux. Ellen J. Langer, Arthur Blank, and Benzion Chanowitz

  7. [7]

    placebic

    The mindlessness of ostensibly thoughtful ac- tion: The role of “placebic” information in interper- sonal interaction.Journal of Personality and Social Psychology, 36(6):635–642. John D. Lee and Katrina A. See. 2004. Trust in automa- tion: Designing for appropriate reliance.Human Factors, 46(1):50–80. Russell V . Lenth. 2024. emmeans: Estimated marginal m...

  8. [8]

    From local explanations to global understand- ing with explainable AI for trees.Nature Machine Intelligence, 2(1):56–67. Scott M. Lundberg and Su-In Lee. 2017. A unified approach to interpreting model predictions. InAd- vances in Neural Information Processing Systems, volume 30, pages 4765–4774. Curran Associates, Inc. 11 David Martens, Camille Dams, Jame...

Show all 28 references
  1. [9]

    Sidra Naveed, Gunnar Stevens, and Dean Robin-Kern

    From anecdotal evidence to quantitative eval- uation methods: A systematic review on evaluating explainable ai.ACM Computing Surveys, 55(13s). Sidra Naveed, Gunnar Stevens, and Dean Robin-Kern

  2. [10]

    Raja Parasuraman and Victor Riley

    An overview of the empirical evaluation of ex- plainable AI (XAI).Applied Sciences, 14(23):11288. Raja Parasuraman and Victor Riley. 1997. Humans and automation: Use, misuse, disuse, abuse.Human Factors, 39(2):230–253. Richard E. Petty and John T. Cacioppo. 1986. The elab- ora...

  3. [11]

    Why Should I Trust You?

    Dissenting explanations: Leveraging disagree- ment to reduce model overreliance. InProceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 21537–21544. Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016. “Why Should I Trust You?”: Ex- plaining...

  4. [12]

    Andreas Theissler, Francesco Spinnato, Udo Schlegel, and Riccardo Guidotti

    illuminate: An llm-xai framework leveraging social science explanation theories.arXiv preprint arXiv:2409.08027. Andreas Theissler, Francesco Spinnato, Udo Schlegel, and Riccardo Guidotti. 2022. Explainable AI for time series classification: A review, taxonomy and re- search d...

  5. [13]

    FR"): df = pd.read_csv(file_path, sep=

    Judging LLM-as-a-judge with MT-Bench and chatbot arena. InAdvances in Neural Information Processing Systems, volume 36. Alexandra Zytek, Dongyu Liu, Andrei Vasiliev, and Kalyan Veeramachaneni. 2024a. LLMs for XAI: Fu- ture directions for explaining explanations.Preprint, arXiv...

  6. [15]

    Compare the prediction to recent lag values – does it follow the recent trend?

  7. [16]

    If model performance metrics are provided, use them to estimate a typical error range

  8. [19]

    error_bucket

    Weigh all available evidence to determine the most likely error magnitude. Error bucket definitions (based on absolute percentage error): - small:[0%,5%) - medium:[5%,15%) - large:[15%,30%) - very_large:[30%,+∞) You must respond with EXACTLY this JSON format and nothing else: ...

  9. [20]

    Identify which features are being changed and note their current SHAP values

  10. [21]

    For each changed feature, determine the direction of change

  11. [22]

    Consider how the SHAP contribution might change

  12. [23]

    If multiple features are changed, think about their combined effect

  13. [24]

    If a natural language explanation is provided, you can use it to better understand the feature–prediction relationships

  14. [25]

    direction

    Estimate the net directional effect and whether it exceeds the 5% threshold. Direction definitions: - higher: prediction increases by more than 5% relative to the original. - lower: prediction decreases by more than 5% relative to the original. - similar: prediction changes by...

  15. [26]

    Examine the features – consider the lag values, the time of year (season in France), holidays, and whether the prediction seems reasonable given this context

  16. [27]

    If SHAP values are provided, consider whether each feature’s contribution makes sense given the feature values and context

  17. [28]

    Consider the model performance metrics to gauge general trustworthiness

  18. [29]

    If a natural language explanation is provided, you can use it to better understand the model’s reasoning

  19. [30]

    reliability

    Weigh all evidence to decide whether this prediction can be trusted. Reliability: - reliable = inputs look normal and within what the model was trained on. - unreliable = inputs look unusual or far from what the model was trained on. JSON response: { "reliability": "<reliable|...

  20. [31]

    The direct Real-vs-Placebo equivalence test yields 43.1% inside ROPE: insufficient for formal practical equivalence (≥95%), but consistent with no meaningful difference be- tween real and placebo NLEs. Bayesian ROPE and Real-vs-Placebo Equiva- lence Sensitivity Judge-specific....

Pith tools

Reviewed June 29, 2026 · model on record in the stance chip above.