REVIEW 8 cited by
Rethinking Interpretability in the Era of Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Interpretable machine learning has exploded as an area of interest over the last decade, sparked by the rise of increasingly large datasets and deep neural networks. Simultaneously, large language models (LLMs) have demonstrated remarkable capabilities across a wide array of tasks, offering a chance to rethink opportunities in interpretable machine learning. Notably, the capability to explain in natural language allows LLMs to expand the scale and complexity of patterns that can be given to a human. However, these new capabilities raise new challenges, such as hallucinated explanations and immense computational costs. In this position paper, we start by reviewing existing methods to evaluate the emerging field of LLM interpretation (both interpreting LLMs and using LLMs for explanation). We contend that, despite their limitations, LLMs hold the opportunity to redefine interpretability with a more ambitious scope across many applications, including in auditing LLMs themselves. We highlight two emerging research priorities for LLM interpretation: using LLMs to directly analyze new datasets and to generate interactive explanations.
Forward citations
Cited by 8 Pith papers
-
DrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis?
A new five-level medical imaging benchmark, DrVD-Bench, shows that vision-language models lose accuracy sharply as reasoning complexity grows and often diagnose without grounding in lesion evidence.
-
Controllable LLM Reasoning via Sparse Autoencoder-Based Steering
SAE-Steering finds, via keyword-logit recall plus effectiveness ranking, sparse-autoencoder features that steer a reasoning model into a chosen reasoning strategy, beating baseline steering by ~15% on a judge-based me...
-
Localizing Persona Representations in LLMs
Persona information is most separable in the final third of LLM layers, and in Llama3's last layer ethical personas share 17.6% of salient activations while political personas have 2.1% to 5.5% unique activations.
-
Response Uncertainty and Probe Modeling: Two Sides of the Same Coin in LLM Interpretability?
LLM response uncertainty and linear probe performance are strongly negatively correlated across six fact-based datasets and six models, with high-uncertainty responses associated with more spread-out feature importance.
-
Transforming Remanufacturing Automation with Large Language Models: A Forward-Looking Analysis with Case Studies
The authors propose ReManGPT, a conceptual orchestration framework for applying LLMs to remanufacturing, and illustrate it with case studies in disassembly planning, repair guidance, and robotic execution.
-
Benchmarking Foundation Models with Multimodal Public Electronic Health Records
A standardized MIMIC-IV benchmark comparing eight unimodal and multimodal foundation models shows multimodal inputs improve predictive performance without adding bias, while medical LVLMs underperform on length-of-sta...
-
Interpretable Adaptive Sampling for LLM Test-Time Scaling
A fuzzy controller that allocates a per-prompt sampling budget keeps LLM accuracy near a fixed full-budget baseline while reducing the average number of candidate answers on some datasets.
-
Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations
LLM explanations split into local and mechanistic tracks; the paper argues they are trustworthy only if they pass causal and contrastive stress tests, adapt to the explainee, and satisfy eight trust principles.
Discussion (0). Continue with ORCID to comment.