Pith. sign in

REVIEW 3 cited by

On Identifiability in Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1908.04211 v4 pith:MXKNR2CR submitted 2019-08-12 cs.CL cs.LG

classification cs.CLcs.LG
keywords attentionembeddingscontextualidentifiabilityidentityinformationinputself-attention
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper we delve deep in the Transformer architecture by investigating two of its core components: self-attention and contextual embeddings. In particular, we study the identifiability of attention weights and token embeddings, and the aggregation of context into hidden tokens. We show that, for sequences longer than the attention head dimension, attention weights are not identifiable. We propose effective attention as a complementary tool for improving explanatory interpretations based on attention. Furthermore, we show that input tokens retain to a large degree their identity across the model. We also find evidence suggesting that identity information is mainly encoded in the angle of the embeddings and gradually decreases with depth. Finally, we demonstrate strong mixing of input information in the generation of contextual embeddings by means of a novel quantification method based on gradient attribution. Overall, we show that self-attention distributions are not directly interpretable and present tools to better understand and further investigate Transformer models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Provably Learning Multi-Head Attention with Queries

    cs.LG 2026-08 accept novelty 8.0 of 10

    A randomized algorithm recovers all (W,v) pairs of multi-head softmax attention from final-token scalar outputs with 4Hd^2 queries, no subspace assumptions, exact up to permutation.

  2. PLEX: Perturbation-free Local Explanations for LLM-Based Text Classification

    cs.CL 2025-07 conditional novelty 6.0 of 10

    PLEX learns a mapping from BERT or RoBERTa token embeddings to word importance scores, reproducing LIME and SHAP style explanations without per-sentence perturbations.

  3. Towards Transparent AI: A Survey on Explainable Large Language Models

    cs.CL 2025-06 conditional novelty 3.0 of 10

    A review that groups LLM explainability methods by transformer architecture and discusses their evaluation and applications.

Pith tools