Pith. sign in

REVIEW 4 cited by

A Close Look at Decomposition-based XAI-Methods for Transformer Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.15886 v1 pith:IMWCV675 submitted 2025-02-21 cs.CL

classification cs.CL
keywords methodslanguagemodelsxai-methodsalti-logitattributionbeenbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Various XAI attribution methods have been recently proposed for the transformer architecture, allowing for insights into the decision-making process of large language models by assigning importance scores to input tokens and intermediate representations. One class of methods that seems very promising in this direction includes decomposition-based approaches, i.e., XAI-methods that redistribute the model's prediction logit through the network, as this value is directly related to the prediction. In the previous literature we note though that two prominent methods of this category, namely ALTI-Logit and LRP, have not yet been analyzed in juxtaposition and hence we propose to close this gap by conducting a careful quantitative evaluation w.r.t. ground truth annotations on a subject-verb agreement task, as well as various qualitative inspections, using BERT, GPT-2 and LLaMA-3 as a testbed. Along the way we compare and extend the ALTI-Logit and LRP methods, including the recently proposed AttnLRP variant, from an algorithmic and implementation perspective. We further incorporate in our benchmark two widely-used gradient-based attribution techniques. Finally, we make our carefullly constructed benchmark dataset for evaluating attributions on language models, as well as our code, publicly available in order to foster evaluation of XAI-methods on a well-defined common ground.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. From Clever Hans to Scientific Discovery: Interpreting EEG Foundational Transformers with LRP

    cs.AI 2026-05 unverdicted novelty 6.0 of 10

    LRP on EEG transformers reveals Clever Hans artifacts in motor imagery tasks and a recurring central electrode cluster as a candidate sensorimotor signature of arousal.

  2. Contrastive Attribution in the Wild: An Interpretability Analysis of LLM Failures on Realistic Benchmarks

    cs.AI 2026-04 conditional novelty 6.0 of 10

    Token-level contrastive attribution yields informative signals for some LLM benchmark failures but is not universally applicable across datasets and models.

  3. Evaluating Post-hoc Explanations of the Transformer-based Genome Language Model DNABERT-2

    cs.LG 2026-04 unverdicted novelty 5.0 of 10

    AttnLRP explanations of DNABERT-2 reliably capture known biological patterns in genomic sequences, showing that transformer-based genome language models can yield biologically meaningful insights comparable to CNNs.

  4. Fast & Faithful Function Vectors

    cs.CL 2026-06 unverdicted novelty 4.0 of 10

    LRP-based attention head selection and distributed application improve the efficiency and accuracy of function vectors for steering LLMs compared to prior choices.

Pith tools