Pith. sign in

REVIEW 3 major objections 4 minor 116 references

This position paper argues that LLM self-explanations should be evaluated by their actionability—whether they help stakeholders make informed decisions—rather than by plausibility and faithfulness alone.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 21:46 UTC pith:ZT4ESMVE

load-bearing objection A well-argued position paper that usefully diagnoses why faithfulness tests fail for LLM self-explanations, but its actionability proposal stays a slogan—no operational definition, no evidence, and its own cited literature undercuts the safety premise. the 3 major comments →

arxiv 2607.15957 v1 pith:ZT4ESMVE submitted 2026-07-17 cs.CL

From Plausible to Actionable: A Position on LLM Self-Explanations

classification cs.CL
keywords self-explanationsfaithfulnessplausibilityactionabilitylarge language modelsXAI evaluationrationalizationhuman decision-making
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper argues that LLM self-explanations can be highly plausible, questionably faithful, and yet highly actionable, and that the field has been measuring the wrong things. Standard faithfulness evaluation, which uses input perturbations and assumes stable model behavior, breaks down for LLMs because they are nondeterministic, highly sensitive to prompt wording, and often rely on internal knowledge rather than task input. The paper therefore proposes shifting emphasis to actionability: whether an explanation helps diverse stakeholders make better-informed decisions, surface assumptions and uncertainties, and consider alternative perspectives. If this position holds, XAI practice should treat self-explanations less as literal accounts of a model's internal reasoning and more as communicative tools for supporting human judgment.

Core claim

The paper's central claim is that LLM self-explanations occupy a distinct position in explainable AI: they often look convincing (high plausibility), may not reflect the model's actual decision process (questionable faithfulness), and yet can still be useful in practice (high actionability). The authors identify why existing plausibility and faithfulness protocols are inadequate for self-explanations: plausibility is typically measured against a single human reference, ignoring stakeholder expertise and legitimate variation in human explanations; faithfulness evaluation relies on perturbation-based assumptions that fail for LLMs due to nondeterminism, prompt sensitivity, and reliance on inte

What carries the argument

The central mechanism is the plausibility–faithfulness–actionability triad applied to LLM self-explanations, together with the idea of 'LLMs as advocates.' The paper repositions the rationalization capability of LLMs—their ability to generate natural-language justifications for decisions—as a way to produce arguments rather than faithful internal accounts. The key operational move is to evaluate self-explanations by their downstream effect: whether they help stakeholders make informed decisions, communicate uncertainty, and present complementary or countervailing perspectives, rather than whether they mirror the model's hidden reasoning.

Load-bearing premise

The load-bearing premise is that an explanation can be genuinely actionable even when it is not faithful, meaning unfaithful-but-plausible explanations can be identified and used to support good decisions without systematically misleading users.

What would settle it

Run a decision-making study where participants receive LLM self-explanations that are deliberately unfaithful but plausible, and compare their decision accuracy and error detection against participants receiving faithful explanations or no explanation; if unfaithful explanations reduce decision quality or increase unwarranted trust, the paper's central claim that actionability can stand apart from faithfulness would be directly contradicted.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Plausibility evaluation should account for stakeholder expertise and should embrace multiple valid explanations in ambiguous tasks, rather than comparing against a single reference.
  • Faithfulness evaluation for LLMs must address nondeterminism, prompt sensitivity, and the risk that perturbed inputs create exploitable artifacts or trigger reliance on internal knowledge.
  • Self-explanations should be designed to surface assumptions, uncertainties, and alternative arguments, with guardrails that keep the LLM in an argumentation role rather than a final-decision role.
  • Actionability should be treated as a first-class evaluation criterion, alongside plausibility and faithfulness, when assessing LLM self-explanations.
  • In high-stakes domains such as healthcare, self-explanations could present evidence for and against candidate diagnoses, leaving the final decision to human experts.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • An implication the authors leave implicit is that actionability and faithfulness are separable in practice; if unfaithful explanations systematically mislead users, the proposed shift would need to be paired with safeguards to detect and correct such cases.
  • A testable extension is to compare decision quality when stakeholders receive a single explanation versus multiple diverse 'advocate' explanations; the paper's argument predicts the latter improves critical evaluation.
  • Actionability could be measured through downstream behavioral outcomes, such as whether users correctly detect model errors, change their decisions appropriately, or report greater understanding of uncertainty.
  • The paper's framing suggests a further connection to uncertainty communication: self-explanations that explicitly flag their own uncertainty might be highly actionable even when they do not faithfully describe internal reasoning.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This position paper argues that LLM self-explanations can be highly plausible, questionably faithful, and yet highly actionable, and that evaluation of self-explanations should therefore extend beyond plausibility and faithfulness to actionability. Section 2 reviews limitations of current plausibility and faithfulness evaluation, arguing that plausibility evaluations neglect stakeholder expertise and variation, and that perturbation-based faithfulness evaluation rests on assumptions violated by LLMs (nondeterminism, prompt sensitivity, reliance on internal knowledge). Section 3 proposes three applications of rationalization: communicative interfaces for non-experts, decision support, and deliberation via multiple LLMs acting as advocates. The paper concludes by calling for actionability as the primary objective.

Significance. The paper is a timely, well-structured synthesis that raises an important question: whether explanations are for reflecting model internals or for supporting stakeholders' decisions. It provides a useful critique of existing evaluation protocols and offers practical guidelines for plausibility and faithfulness (e.g., accounting for nondeterminism and stakeholders' expertise). The call to consider actionability is a valuable provocation. However, the central recommendation is not yet operationalized, and the paper does not resolve the tension between its own evidence of misleading plausible explanations and its advocacy of actionability. If revised to address these gaps, the paper could serve as a useful agenda-setting piece.

major comments (3)
  1. [§3] The central concept of actionability is not operationalized. The paper defines it as 'the ability to help diverse stakeholders in making better-informed decisions and take effective actions' and cites Orgad et al. (2026), but provides no metrics, evaluation protocol, or criteria for determining when an explanation is actionable. The guidelines in §2 for plausibility and faithfulness are concrete (e.g., account for nondeterminism, prompt sensitivity, stakeholder expertise), but the actionability section consists of examples and a research agenda. Because the paper's main call is to shift evaluation toward actionability, this omission leaves the central recommendation untestable. The authors should provide at least a working definition with constituent criteria (e.g., decision outcome, user's correct use of information) and suggest how it could be measured.
  2. [§2.3, §3] The paper's own §2.3 undercuts the safety of the proposed actionability criterion. It cites Palta et al. (2026), Fan et al. (2026), and Ajwani et al. (2025) showing that highly plausible self-explanations mislead users, reduce error detection, and cause over-reliance. The paper then argues that self-explanations should be treated as 'arguments that human decision-makers can critically evaluate' and mentions safety protocols (e.g., guardrails), but does not explain how these protocols would prevent the demonstrated harmful effects. Without a mechanism to distinguish safe, actionable explanations from those that are merely persuasive, the recommendation risks endorsing the very class of explanations shown to cause automation bias. The authors need to address this tension, e.g., by proposing constraints on explanation content or human-in-the-loop validation.
  3. [§3] The 'LLMs as advocates' paradigm assumes that different LLMs can surface complementary, decision-relevant perspectives. However, the paper itself cites literature on sycophancy (Malmqvist 2025; Wang et al. 2026) and hallucination (Feng et al. 2026) in §2.1–2.2, which suggests LLM outputs may be correlated, sycophantic, or systematically biased. The paper does not argue why the advocate setting would overcome these tendencies, nor how to ensure diversity of perspectives rather than correlated errors. Since the deliberation application is one of the three core applications, this is a load-bearing assumption that requires justification or explicit caveats.
minor comments (4)
  1. [§1, §3] Actionability is introduced with a citation to Orgad et al. (2026) but is not clearly distinguished from related concepts such as 'usefulness' or 'decision support' in XAI (e.g., Hase and Bansal 2020). Clarifying the relationship would help.
  2. [§2.2] The claim that 'masking all words in the task input does not change the predicted label' is presented as a general finding via Fayyaz et al. (2024), but this may be task/model-specific. A more careful statement would note the conditions.
  3. [References] Madsen et al. 2024a and 2024b appear to be the same work (same title) with different entries; please consolidate or clarify.
  4. [§2.1] The guideline to 'embrace variation' is important but the paper does not suggest how to aggregate multiple valid explanations in evaluation (e.g., using agreement metrics or expert adjudication). A concrete suggestion would strengthen the guideline.

Circularity Check

0 steps flagged

No circularity: this is a literature-grounded position paper, not a derivation chain; its only self-citation is peripheral and not load-bearing.

full rationale

The paper makes no formal derivation and contains no equations, fitted parameters, or empirical predictions that could reduce to its inputs. Its central claim—that self-explanations can be highly plausible, questionably faithful, and yet actionable, so evaluation should extend to actionability—is argued from external literature in §2.1–§2.3 and then proposed as a research agenda in §3. The concept of actionability is attributed to Orgad et al. (2026), not to the authors' own prior work. The only author-overlapping citation is Muscato et al. (2026) in §2.1, used to support the secondary recommendation that plausibility evaluation should embrace variation and be validated with domain experts. That recommendation is also grounded in Doshi-Velez and Kim (2017) and Plank (2022), and the §3 shift to actionability does not depend on it. No self-citation is invoked to forbid alternatives, and no quantity is fitted and later renamed as a prediction. Concerns about the safety or operationalization of unfaithful-but-actionable explanations are evidentiary or normative, not circularity. The paper is self-contained as an opinion piece, so the honest finding is no significant circularity.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

The central claim rests on domain assumptions about how LLMs generate explanations, the validity of faithfulness evaluation critiques, and the value of actionability. These are asserted with citations but not proven by the paper. No free parameters or invented physical entities are involved.

axioms (4)
  • domain assumption LLM self-explanations are generated by the same process as prediction and have no access to internal meta-information about the prediction process.
    Invoked in §2.2 (first reason) to argue self-explanations are questionably faithful; relies on Sarkar (2024b).
  • domain assumption Perturbation-based faithfulness evaluation rests on model, prediction, and linearity assumptions that do not hold for LLMs.
    Used in §2.2 to reject standard faithfulness metrics; based on Jacovi & Goldberg (2020) plus cited empirical results.
  • domain assumption The value of a self-explanation lies not in accurately reflecting model reasoning but in surfacing assumptions, uncertainties, and alternative perspectives for human decision-makers.
    Stated in §3; this is the load-bearing premise for the actionability recommendation and is assumed rather than empirically demonstrated.
  • domain assumption Human stakeholders can critically evaluate diverse, argument-like self-explanations and benefit from multiple LLM perspectives.
    Underlies the 'LLMs as advocates' proposal in §3; no evidence that presenting diverse AI arguments improves decisions rather than inducing confusion or automation bias.

pith-pipeline@v1.3.0-alltime-deepseek · 10007 in / 11764 out tokens · 111762 ms · 2026-08-01T21:46:39.092682+00:00 · methodology

0 comments
read the original abstract

Large Language Models (LLMs) can generate natural language explanations that rationalize their own decisions, a phenomenon commonly referred to as self-explanations.Such explanations have emerged as a promising direction for explainable artificial intelligence (XAI), particularly for interpreting LLM behavior.However, while self-explanations often appear plausible, whether they faithfully reflect a model's underlying reasoning process remains an open question. In this opinion paper, we argue that self-explanations can be highly plausible, questionably faithful, and yet highly actionable. From a traditional XAI perspective, we identify the limitations of standard evaluation protocols for LLM-generated self-explanations and propose practical guidelines for assessing their plausibility and faithfulness. Moreover, we argue that evaluation should extend beyond these criteria to actionability, highlighting applications of LLM rationalization capabilities that support informed decision-making and appropriate action across diverse stakeholders.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

116 extracted references · 3 canonical work pages · 1 internal anchor

  1. [1]

    Aho and Jeffrey D

    Alfred V. Aho and Jeffrey D. Ullman , title =. 1972

  2. [2]

    Publications Manual , year = "1983", publisher =

  3. [3]

    Chandra and Dexter C

    Ashok K. Chandra and Dexter C. Kozen and Larry J. Stockmeyer , year = "1981", title =. doi:10.1145/322234.322243

  4. [4]

    2025 , organization=

    Wilde, Katja Mariko and Rugolon, Franco , booktitle=. 2025 , organization=

  5. [5]

    Measuring

    Lanham, Tamera and Chen, Anna and Radhakrishnan, Ansh and Steiner, Benoit and Denison, Carson and Hernandez, Danny and Li, Dustin and Durmus, Esin and Hubinger, Evan and Kernion, Jackson and Luko. Measuring. doi:10.48550/arXiv.2307.13702 , urldate =. arXiv , keywords =:2307.13702 , primaryclass =

  6. [6]

    Paul, Debjit and West, Robert and Bosselut, Antoine and Faltings, Boi , year = 2024, pages =. Making. Findings of the. doi:10.18653/v1/2024.findings-emnlp.882 , urldate =

  7. [7]

    Language

    Turpin, Miles and Michael, Julian and Perez, Ethan and Bowman, Samuel R , journal=. Language. 2023 , url=

  8. [8]

    2024 , url=

    Zhuo, Jingming and Zhang, Songyang and Fang, Xinyu and Duan, Haodong and Lin, Dahua and Chen, Kai , booktitle=. 2024 , url=

  9. [9]

    2025 , url=

    Errica, Federico and Sanvito, Davide and Siracusano, Giuseppe and Bifulco, Roberto , booktitle=. 2025 , url=

  10. [10]

    2023 , url=

    Loya, Manikanta and Sinha, Divya and Futrell, Richard , booktitle=. 2023 , url=

  11. [11]

    2024 , publisher=

    Sarkar, Advait , journal=. 2024 , publisher=

  12. [12]

    2024 , url=

    He, Jia and Rungta, Mukund and Koleczek, David and Sekhon, Arshdeep and Wang, Franklin X and Hasan, Sadid , journal=. 2024 , url=

  13. [13]

    , year = 2026, month = feb, number =

    Mayne, Harry and Kang, Justin Singh and Gould, Dewi and Ramchandran, Kannan and Mahdi, Adam and Siegel, Noah Y. , year = 2026, month = feb, number =. A. doi:10.48550/arXiv.2602.02639 , urldate =. arXiv , keywords =:2602.02639 , primaryclass =

  14. [14]

    Fan, Shutong and Zhang, Lan and Yuan, Xiaoyong , year = 2026, month = may, number =. When. doi:10.48550/arXiv.2602.04003 , urldate =. arXiv , keywords =:2602.04003 , primaryclass =

  15. [15]

    Faithfulness vs

    Agarwal, Chirag and Tanneru, Sree Harsha and Lakkaraju, Himabindu , year = 2024, month = mar, number =. Faithfulness vs. doi:10.48550/arXiv.2402.04614 , urldate =. arXiv , keywords =:2402.04614 , primaryclass =

  16. [16]

    Sherburn, Dane and Chughtai, Bilal and Evans, Owain , journal=

  17. [17]

    doi:10.1162/tacl_a_00367 , abstract =

    Jacovi, Alon and Goldberg, Yoav , year = 2021, month = mar, journal =. doi:10.1162/tacl_a_00367 , abstract =. https://direct.mit.edu/tacl/article-pdf/doi/10.1162/tacl\_a\_00367/1923972/tacl\_a\_00367.pdf , pages =

  18. [18]

    HCXAI workshop , location =

    Large language models cannot explain themselves , author=. HCXAI workshop , location =. 2024 , isbn =

  19. [19]

    Chen, Sirui and Ma, Shuqin and Yu, Shu and Zhang, Hanwang and Zhao, Shengjie and Lu, Chaochao , journal=

  20. [20]

    Heo, Juyeon and Heinze-Deml, Christina and Elachqar, Oussama and Chan, Kwan Ho Ryan and Ren, Shirley and Miller, Andrew and Nallasamy, Udhyakumar and Narain, Jaya , booktitle=

  21. [21]

    2026 , url=

    Feng, Zhaoxin and Chen, Zheng and Ma, Jianfei and Po, Yip Tin and Chersoni, Emmanuele and Li, Bo , booktitle=. 2026 , url=

  22. [22]

    2025 , organization=

    Malmqvist, Lars , booktitle=. 2025 , organization=

  23. [23]

    Implicit bias in

    Lin, Xinru and Li, Luyang , journal=. Implicit bias in

  24. [24]

    2026 , url=

    Wang, Keyu and Li, Jin and Yang, Shu and Zhang, Zhuoran and Wang, Di , booktitle=. 2026 , url=

  25. [25]

    2025 , url=

    Lin, Luyang and Wang, Lingzhi and Guo, Jinsong and Wong, Kam-Fai , booktitle=. 2025 , url=

  26. [26]

    Jacovi, Alon and Goldberg, Yoav , year = 2020, pages =. Towards. Proceedings of the 58th. doi:10.18653/v1/2020.acl-main.386 , urldate =

  27. [27]

    Evaluating

    Fayyaz, Mohsen and Yin, Fan and Sun, Jiao and Peng, Nanyun , year = 2024, month = oct, number =. Evaluating. doi:10.48550/arXiv.2407.00219 , urldate =. arXiv , langid =:2407.00219 , primaryclass =

  28. [28]

    Hsia, Jennifer and Pruthi, Danish and Singh, Aarti and Lipton, Zachary C , booktitle=

  29. [29]

    Findings of the

    Madsen, Andreas and Chandar, Sarath and Reddy, Siva , year = 2024, pages =. Findings of the. doi:10.18653/v1/2024.findings-acl.19 , urldate =

  30. [30]

    Faithful or

    Colegado, Shaun and Jin, Jennifer and Hou, Yunfei , year = 2026, month = feb, pages =. Faithful or. 2026. doi:10.1109/ICSC67292.2026.00018 , urldate =

  31. [31]

    Arguments that Alter Minds: LLM Rationales Sway Human (and LLM) Notions of Plausibility

    Palta, Shramay and Rankel, Peter and Wiegreffe, Sarah and Rudinger, Rachel , year = 2026, month = feb, number =. Everything Is. doi:10.48550/arXiv.2510.08091 , urldate =. arXiv , langid =:2510.08091 , primaryclass =

  32. [32]

    2018 41st International convention on information and communication technology, electronics and microelectronics (MIPRO) , pages=

    Do. 2018 41st International convention on information and communication technology, electronics and microelectronics (MIPRO) , pages=. 2018 , organization=

  33. [33]

    2026 , publisher=

    Jalan, Pratik and Abishethvarman, Vadivel and Chandna, Bhavik and Naseem, Usman , journal=. 2026 , publisher=

  34. [34]

    2026 , url=

    Chen, Xi and Plaat, Aske and van Stein, Niki , booktitle=. 2026 , url=

  35. [35]

    Preprint, alphaXiv , year=

    Barez, Fazl and Wu, Tung-Yu and Arcuschin, Iv. Preprint, alphaXiv , year=

  36. [36]

    Proceedings of the 26th international conference on intelligent user interfaces , pages=

    Visual, textual or hybrid: the effect of user expertise on different explanations , author=. Proceedings of the 26th international conference on intelligent user interfaces , pages=

  37. [37]

    Proceedings of the ACM on human-computer interaction , volume=

    Impact of Explanation Techniques and Representations on Users' Comprehension and Confidence in Explainable AI , author=. Proceedings of the ACM on human-computer interaction , volume=. 2025 , publisher=

  38. [38]

    2023 , isbn =

    Chen, Chacha and Feng, Shi and Sharma, Amit and Tan, Chenhao , title =. 2023 , isbn =. doi:10.1145/3593013.3593970 , booktitle =

  39. [39]

    2022 , url=

    Wei, Jason and Wang, Xuezhi and Schuurmans, Dale and Bosma, Maarten and Xia, Fei and Chi, Ed and Le, Quoc V and Zhou, Denny and others , journal=. 2022 , url=

  40. [40]

    Interpreting

    Hassija, Vikas and Chamola, Vinay and Mahapatra, Atmesh and Singal, Abhinandan and Goel, Divyansh and Huang, Kaizhu and Scardapane, Simone and Spinelli, Indro and Mahmud, Mufti and Hussain, Amir , year = 2024, month = jan, journal =. Interpreting. doi:10.1007/s12559-023-10179-8 , urldate =

  41. [41]

    The good, the bad, and the greedy: Evaluation of llms should not ignore non-determinism , author=. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=

  42. [42]

    Proceedings of the 10th italian conference on computational linguistics (CLiC-it 2024) , pages=

    Is Explanation All You Need? An Expert Survey on LLM-generated Explanations for Abusive Language Detection , author=. Proceedings of the 10th italian conference on computational linguistics (CLiC-it 2024) , pages=. 2024 , url=

  43. [43]

    Properties and

    Kunz, Jenny and Kuhlmann, Marco , year = 2024, pages =. Properties and. Proceedings of the. doi:10.18653/v1/2024.hcinlp-1.2 , urldate =

  44. [44]

    Can LLM s Explain Themselves Counterfactually?

    Dehghanighobadi, Zahra and Fischer, Asja and Zafar, Muhammad Bilal. Can LLM s Explain Themselves Counterfactually?. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. doi:10.18653/v1/2025.emnlp-main.396

  45. [45]

    2023 , url=

    Panigutti, Cecilia and Hamon, Ronan and Hupont, Isabelle and Fernandez Llorca, David and Fano Yela, Delia and Junklewitz, Henrik and Scalzo, Salvatore and Mazzini, Gabriele and Sanchez, Ignacio and Soler Garrido, Josep and others , booktitle=. 2023 , url=

  46. [46]

    Entropy , volume=

    Human-in-the-loop artificial intelligence: A systematic review of concepts, methods, and applications , author=. Entropy , volume=. 2026 , publisher=

  47. [47]

    Transactions on Machine Learning Research , year=

    Open Problems in Mechanistic Interpretability , author=. Transactions on Machine Learning Research , year=

  48. [48]

    2022 , publisher=

    Kostick-Quenet, Kristin M and Gerke, Sara , journal=. 2022 , publisher=

  49. [49]

    2013 IEEE Symposium on visual languages and human centric computing , pages=

    Too much, too little, or just right? Ways explanations impact end users' mental models , author=. 2013 IEEE Symposium on visual languages and human centric computing , pages=. 2013 , organization=

  50. [50]

    2024 , url=

    Albert, Julien and Martin, Balfroid and Doh, Miriam and Bogaert, Jeremie and De Vos, Liesbet and Renard, Bryan and Stragier, Vincent and Jean, Emmanuel and others , booktitle=. 2024 , url=

  51. [51]

    2025 International Conference on Data Science, Agents & Artificial Intelligence (ICDSAAI) , pages=

    Exploring mechanistic interpretability in large language models: Challenges, approaches, and insights , author=. 2025 International Conference on Data Science, Agents & Artificial Intelligence (ICDSAAI) , pages=. 2025 , organization=

  52. [52]

    2024 , url=

    Wataoka, Koki and Takahashi, Tsubasa and Ri, Ryokan , booktitle=. 2024 , url=

  53. [53]

    On the Diversity and Limits of Human Explanations

    Tan, Chenhao. On the Diversity and Limits of Human Explanations. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. doi:10.18653/v1/2022.naacl-main.158

  54. [54]

    Survey on the role of mechanistic interpretability in generative

    Ranaldi, Leonardo , journal=. Survey on the role of mechanistic interpretability in generative. 2025 , publisher=

  55. [55]

    2022 , url=

    Bereska, Leonard F , booktitle=. 2022 , url=

  56. [56]

    2026 , publisher=

    Somvanshi, Shriyank and Islam, Md Monzurul and Rafe, Amir and Tusti, Anannya Ghosh and Chakraborty, Arka and Baitullah, Anika and Chowdhury, Tausif Islam and Alnawmasi, Nawaf and Dutta, Anandi and Das, Subasish , journal=. 2026 , publisher=

  57. [57]

    ACM Transactions on Intelligent Systems and Technology , volume=

    Explainability for large language models: A survey , author=. ACM Transactions on Intelligent Systems and Technology , volume=. 2024 , publisher=

  58. [58]

    arXiv preprint arXiv:2401.12874 , year=

    From understanding to utilization: A survey on explainability for large language models , author=. arXiv preprint arXiv:2401.12874 , year=

  59. [59]

    Human-Centric Intelligent Systems , volume=

    Survey on explainable AI: From approaches, limitations and applications aspects , author=. Human-Centric Intelligent Systems , volume=. 2023 , publisher=

  60. [60]

    F aith LM : Towards Faithful Explanations for Large Language Models

    Chuang, Yu-Neng and Wang, Guanchu and Chang, Chia-Yuan and Tang, Ruixiang and Zhong, Shaochen and Yang, Fan and Wen, Andrew and Du, Mengnan and Cai, Xuanting and Braverman, Vladimir and Hu, Xia. F aith LM : Towards Faithful Explanations for Large Language Models. Proceedings of the 19th Conference of the E uropean Chapter of the A ssociation for C omputat...

  61. [61]

    Proceedings of the 2021 ACM conference on fairness, accountability, and transparency , pages=

    Formalizing trust in artificial intelligence: Prerequisites, causes and goals of human trust in AI , author=. Proceedings of the 2021 ACM conference on fairness, accountability, and transparency , pages=. 2021 , url=

  62. [62]

    2026 , publisher=

    Romeo, Giuseppe and Conti, Daniela , journal=. 2026 , publisher=

  63. [63]

    Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

    On the Complexity-Faithfulness Trade-off of Gradient-Based Explanations , author=. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

  64. [64]

    Nature , volume=

    Detecting hallucinations in large language models using semantic entropy , author=. Nature , volume=. 2024 , publisher=

  65. [65]

    2024 , url=

    Sriramanan, Gaurang and Bharti, Siddhant and Sadasivan, Vinu Sankar and Saha, Shoumik and Kattakinda, Priyatham and Feizi, Soheil , journal=. 2024 , url=

  66. [66]

    and Feng, Shi , title =

    Panickssery, Arjun and Bowman, Samuel R. and Feng, Shi , title =. Proceedings of the 38th International Conference on Neural Information Processing Systems , articleno =. 2024 , isbn =

  67. [67]

    Proceedings of the 41st International Conference on Machine Learning , articleno =

    Chen, Yanda and Zhong, Ruiqi and Ri, Narutatsu and Zhao, Chen and He, He and Steinhardt, Jacob and Yu, Zhou and McKeown, Kathleen , title =. Proceedings of the 41st International Conference on Machine Learning , articleno =. 2024 , publisher =

  68. [68]

    Evaluating

    Hase, Peter and Bansal, Mohit , editor =. Evaluating. Proceedings of the 58th. doi:10.18653/v1/2020.acl-main.491 , urldate =

  69. [69]

    Hong, Pingjun and Roth, Benjamin , year = 2026, month = jan, number =. Do. doi:10.48550/arXiv.2601.03775 , urldate =. arXiv , keywords =:2601.03775 , primaryclass =

  70. [70]

    At. Non-. Proceedings of the 5th Workshop on Evaluation and Comparison of NLP Systems , pages=. 2025 , url=

  71. [71]

    Ribeiro, Marco Tulio and Singh, Sameer and Guestrin, Carlos , booktitle=. ". 2016 , url=

  72. [72]

    doi:10.48550/arXiv.1702.08608 , urldate =

    Towards. doi:10.48550/arXiv.1702.08608 , urldate =. arXiv , langid =:1702.08608 , primaryclass =

  73. [73]

    and Harman, Mark and Wang, Meng , year = 2025, month = feb, journal =

    Ouyang, Shuyin and Zhang, Jie M. and Harman, Mark and Wang, Meng , year = 2025, month = feb, journal =. An. doi:10.1145/3697010 , urldate =

  74. [74]

    , year = 2023, month = oct, number =

    Huang, Shiyuan and Mamidanna, Siddarth and Jangam, Shreedhar and Zhou, Yilun and Gilpin, Leilani H. , year = 2023, month = oct, number =. Can. doi:10.48550/arXiv.2310.11207 , urldate =. arXiv , keywords =:2310.11207 , primaryclass =

  75. [75]

    Astekin, Merve and Hort, Max and Moonen, Leon , year = 2024, month = apr, pages =. An. Proceedings of the. doi:10.1145/3643661.3643952 , urldate =

  76. [76]

    Artificial Intelligence Review , volume=

    Explainable artificial intelligence: a comprehensive review , author=. Artificial Intelligence Review , volume=. 2022 , publisher=

  77. [77]

    Frontiers in Artificial Intelligence , volume=

    Human-annotated rationales and explainable text classification: a survey , author=. Frontiers in Artificial Intelligence , volume=. 2024 , publisher=

  78. [78]

    Explainable Artificial Intelligence: A Comprehensive Review of Techniques, Applications, and Emerging Trends , volume =

    Muia, Charles and Kamiri, Jackson , year =. Explainable Artificial Intelligence: A Comprehensive Review of Techniques, Applications, and Emerging Trends , volume =. International Journal of Scientific Research in Computer Science and Engineering , doi =

  79. [79]

    Proceedings of the 2024 ACM Designing Interactive Systems Conference , pages=

    Requirements and attitudes towards explainable ai in law enforcement , author=. Proceedings of the 2024 ACM Designing Interactive Systems Conference , pages=

  80. [80]

    2025 , eprint=

    Why is plausibility surprisingly problematic as an XAI criterion? , author=. 2025 , eprint=

Showing first 80 references.