Towards Intrinsic Interpretability of Large Language Models:A Survey of Design Principles and Architectures

· 2026 · cs.CL · arXiv 2604.16042

2 Pith papers cite this work. Polarity classification is still indexing.

2 Pith papers citing it

open full Pith review browse 2 citing papers arXiv PDF

abstract

While Large Language Models (LLMs) have achieved strong performance across many NLP tasks, their opaque internal mechanisms hinder trustworthiness and safe deployment. Existing surveys in explainable AI largely focus on post-hoc explanation methods that interpret trained models through external approximations. In contrast, intrinsic interpretability, which builds transparency directly into model architectures and computations, has recently emerged as a promising alternative. This paper presents a systematic review of the recent advances in intrinsic interpretability for LLMs, categorizing existing approaches into five design paradigms: functional transparency, concept alignment, representational decomposability, explicit modularization, and latent sparsity induction. We further discuss open challenges and outline future research directions in this emerging field. The paper list is available at: https://github.com/PKU-PILLAR-Group/Survey-Intrinsic-Interpretability-of-LLMs.

representative citing papers

Sparsely gated tiny linear experts

cs.LG · 2026-06-05 · unverdicted · novelty 6.0

Sgatlin replaces transformer FF layers with sparse single linear neurons, improving perplexity across compute budgets and enabling direct interpretation of semantically clustered circuits for factual recall.

Faithful by Definition: Emotion Analysis via Natural Semantic Metalanguage Explications

cs.CL · 2026-07-01 · unverdicted · novelty 5.0

An NSM-based explication parser with fixed semantic rules produces emotion labels for events, achieving 0.33 accuracy on held-out crowd-sourced data while shifting empirical risk to an inspectable parser.

citing papers explorer

Showing 1 of 1 citing paper after filters.

Faithful by Definition: Emotion Analysis via Natural Semantic Metalanguage Explications cs.CL · 2026-07-01 · unverdicted · none · ref 2 · internal anchor
An NSM-based explication parser with fixed semantic rules produces emotion labels for events, achieving 0.33 accuracy on held-out crowd-sourced data while shifting empirical risk to an inspectable parser.

Towards Intrinsic Interpretability of Large Language Models:A Survey of Design Principles and Architectures

fields

years

verdicts

representative citing papers

citing papers explorer