Pith. sign in

REVIEW 15 cited by

From Understanding to Utilization: A Survey on Explainability for Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.12874 v2 pith:MLAUHYGD submitted 2024-01-23 cs.CL cs.AI

classification cs.CLcs.AI
keywords explainabilityllmslanguagemodelsunderstandingapplicationsexplanatorylarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Explainability for Large Language Models (LLMs) is a critical yet challenging aspect of natural language processing. As LLMs are increasingly integral to diverse applications, their "black-box" nature sparks significant concerns regarding transparency and ethical use. This survey underscores the imperative for increased explainability in LLMs, delving into both the research on explainability and the various methodologies and tasks that utilize an understanding of these models. Our focus is primarily on pre-trained Transformer-based LLMs, such as LLaMA family, which pose distinctive interpretability challenges due to their scale and complexity. In terms of existing methods, we classify them into local and global analyses, based on their explanatory objectives. When considering the utilization of explainability, we explore several compelling methods that concentrate on model editing, control generation, and model enhancement. Additionally, we examine representative evaluation metrics and datasets, elucidating their advantages and limitations. Our goal is to reconcile theoretical and empirical understanding with practical implementation, proposing exciting avenues for explanatory techniques and their applications in the LLMs era.

Discussion (0). Sign in to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs

    cs.CV 2026-08 conditional novelty 7.0 of 10

    PhyCheck is a 69,825-pair video QA benchmark that tests and improves Video-LLMs' ability to judge whether events obey physical laws, with fine-grained evidence questions and a context-sensitivity pilot.

  2. A Multi-Dimensional Evaluation of Explainability in Media Bias Detection

    cs.CL 2026-07 conditional novelty 6.0 of 10

    In media bias detection, explanation plausibility and mechanistic faithfulness are distinct axes that vary independently across model architectures and finetuning strategies.

  3. Faithfulness Evaluation for Decoder-only LLM Attributions with Controlled Retained Information

    cs.CL 2026-01 conditional novelty 6.0 of 10

    π-Soft-NC/NS matches the expected number of retained tokens across attribution methods, and Grad-ELLM, a gradient-plus-attention attribution method for decoder-only LLMs, achieves top average scores on most evaluated tasks.

  4. CrossEarth-Gate: Fisher-Guided Adaptive Tuning Engine for Efficient Adaptation of Cross-Domain Remote Sensing Semantic Segmentation

    cs.CV 2025-11 conditional novelty 6.0 of 10

    A Fisher-information-guided dynamic selection over a toolbox of LoRA, adapter, and frequency-adapter modules improves cross-domain remote sensing segmentation over static PEFT methods.

  5. Cross-Attention is Half Explanation in Speech-to-Text Models

    cs.CL 2025-09 conditional novelty 6.0 of 10

    Cross-attention in speech-to-text models correlates with saliency-based explanations (Pearson r roughly 0.49-0.75 in the best aggregations) but explains only a minority of the variance, so it should complement, not re...

  6. xInv: Explainable Optimization of Inverse Problems

    cs.LG 2025-05 conditional novelty 6.0 of 10

    An explainability method that instruments differentiable optimizers to emit natural language events and uses a language model to synthesize human-readable explanations of inverse problem optimization.

  7. Exploring LLM-Generated Feedback for Economics Essays: How Teaching Assistants Evaluate and Envision Its Use

    cs.HC 2025-05 conditional novelty 5.0 of 10

    In a think-aloud study, five economics teaching assistants found AI-generated essay feedback useful as suggestions, especially when it included highlighted evidence and intermediate judgments.

  8. From Plausible to Actionable: A Position on LLM Self-Explanations

    cs.CL 2026-07 conditional novelty 4.0 of 10

    Self-explanations from LLMs should be evaluated by their actionability for stakeholders rather than by plausibility or faithfulness alone.

  9. Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations

    cs.CL 2025-10 conditional novelty 4.0 of 10

    LLM explanations split into local and mechanistic tracks; the paper argues they are trustworthy only if they pass causal and contrastive stress tests, adapt to the explainee, and satisfy eight trust principles.

  10. Triadic Fusion of Cognitive, Functional, and Causal Dimensions for Explainable LLMs: The TAXAL Framework

    cs.CL 2025-09 conditional novelty 4.0 of 10

    TAXAL proposes a triadic cognitive-functional-causal framework for role-sensitive explainability in agentic LLMs, demonstrated through cross-domain case studies.

  11. Human-Centered Explainability in Interactive Information Systems: A Survey

    cs.HC 2025-07 conditional novelty 4.0 of 10

    A systematic review of 100 empirical user studies synthesizes explainability research into five conceptual dimensions, a design classification, and six measurement categories.

  12. Composable Building Blocks for Controllable and Transparent Interactive AI Systems

    cs.HC 2025-06 conditional novelty 4.0 of 10

    A framework for making interactive AI systems transparent by exposing their structural and visual building blocks through a shared API.

  13. Transparent AI: The Case for Interpretability and Explainability

    cs.LG 2025-07 unverdicted novelty 3.0 of 10

    The paper is a whitepaper that reviews interpretability concepts, proposes a standardized reporting framework and adoption roadmap, and presents qualitative case studies without new experimental or theoretical claims.

  14. Towards Transparent AI: A Survey on Explainable Large Language Models

    cs.CL 2025-06 conditional novelty 3.0 of 10

    A review that groups LLM explainability methods by transformer architecture and discusses their evaluation and applications.

  15. Large Language Models as Computable Approximations to Solomonoff Induction

    cs.LG 2025-05 reject novelty 2.0 of 10

    The paper argues LLMs are computable approximations of Solomonoff induction, but its central derivation recovers the model's own probabilities by construction.

Pith tools