REVIEW 15 cited by
From Understanding to Utilization: A Survey on Explainability for Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Explainability for Large Language Models (LLMs) is a critical yet challenging aspect of natural language processing. As LLMs are increasingly integral to diverse applications, their "black-box" nature sparks significant concerns regarding transparency and ethical use. This survey underscores the imperative for increased explainability in LLMs, delving into both the research on explainability and the various methodologies and tasks that utilize an understanding of these models. Our focus is primarily on pre-trained Transformer-based LLMs, such as LLaMA family, which pose distinctive interpretability challenges due to their scale and complexity. In terms of existing methods, we classify them into local and global analyses, based on their explanatory objectives. When considering the utilization of explainability, we explore several compelling methods that concentrate on model editing, control generation, and model enhancement. Additionally, we examine representative evaluation metrics and datasets, elucidating their advantages and limitations. Our goal is to reconcile theoretical and empirical understanding with practical implementation, proposing exciting avenues for explanatory techniques and their applications in the LLMs era.
Forward citations
Cited by 15 Pith papers
-
PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs
PhyCheck is a 69,825-pair video QA benchmark that tests and improves Video-LLMs' ability to judge whether events obey physical laws, with fine-grained evidence questions and a context-sensitivity pilot.
-
A Multi-Dimensional Evaluation of Explainability in Media Bias Detection
In media bias detection, explanation plausibility and mechanistic faithfulness are distinct axes that vary independently across model architectures and finetuning strategies.
-
Faithfulness Evaluation for Decoder-only LLM Attributions with Controlled Retained Information
π-Soft-NC/NS matches the expected number of retained tokens across attribution methods, and Grad-ELLM, a gradient-plus-attention attribution method for decoder-only LLMs, achieves top average scores on most evaluated tasks.
-
CrossEarth-Gate: Fisher-Guided Adaptive Tuning Engine for Efficient Adaptation of Cross-Domain Remote Sensing Semantic Segmentation
A Fisher-information-guided dynamic selection over a toolbox of LoRA, adapter, and frequency-adapter modules improves cross-domain remote sensing segmentation over static PEFT methods.
-
Cross-Attention is Half Explanation in Speech-to-Text Models
Cross-attention in speech-to-text models correlates with saliency-based explanations (Pearson r roughly 0.49-0.75 in the best aggregations) but explains only a minority of the variance, so it should complement, not re...
-
xInv: Explainable Optimization of Inverse Problems
An explainability method that instruments differentiable optimizers to emit natural language events and uses a language model to synthesize human-readable explanations of inverse problem optimization.
-
Exploring LLM-Generated Feedback for Economics Essays: How Teaching Assistants Evaluate and Envision Its Use
In a think-aloud study, five economics teaching assistants found AI-generated essay feedback useful as suggestions, especially when it included highlighted evidence and intermediate judgments.
-
From Plausible to Actionable: A Position on LLM Self-Explanations
Self-explanations from LLMs should be evaluated by their actionability for stakeholders rather than by plausibility or faithfulness alone.
-
Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations
LLM explanations split into local and mechanistic tracks; the paper argues they are trustworthy only if they pass causal and contrastive stress tests, adapt to the explainee, and satisfy eight trust principles.
-
Triadic Fusion of Cognitive, Functional, and Causal Dimensions for Explainable LLMs: The TAXAL Framework
TAXAL proposes a triadic cognitive-functional-causal framework for role-sensitive explainability in agentic LLMs, demonstrated through cross-domain case studies.
-
Human-Centered Explainability in Interactive Information Systems: A Survey
A systematic review of 100 empirical user studies synthesizes explainability research into five conceptual dimensions, a design classification, and six measurement categories.
-
Composable Building Blocks for Controllable and Transparent Interactive AI Systems
A framework for making interactive AI systems transparent by exposing their structural and visual building blocks through a shared API.
-
Transparent AI: The Case for Interpretability and Explainability
The paper is a whitepaper that reviews interpretability concepts, proposes a standardized reporting framework and adoption roadmap, and presents qualitative case studies without new experimental or theoretical claims.
-
Towards Transparent AI: A Survey on Explainable Large Language Models
A review that groups LLM explainability methods by transformer architecture and discusses their evaluation and applications.
-
Large Language Models as Computable Approximations to Solomonoff Induction
The paper argues LLMs are computable approximations of Solomonoff induction, but its central derivation recovers the model's own probabilities by construction.
Discussion (0). Sign in to comment.