Pith. sign in

REVIEW 2 cited by

Is Sparse Attention more Interpretable?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.01087 v2 pith:MF57I23I submitted 2021-06-02 cs.CL

classification cs.CL
keywords attentionsparseinputsmodelassumptioninfluentialinterpretabilityplausible
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Sparse attention has been claimed to increase model interpretability under the assumption that it highlights influential inputs. Yet the attention distribution is typically over representations internal to the model rather than the inputs themselves, suggesting this assumption may not have merit. We build on the recent work exploring the interpretability of attention; we design a set of experiments to help us understand how sparsity affects our ability to use attention as an explainability tool. On three text classification tasks, we verify that only a weak relationship between inputs and co-indexed intermediate representations exists -- under sparse attention and otherwise. Further, we do not find any plausible mappings from sparse attention distributions to a sparse set of influential inputs through other avenues. Rather, we observe in this setting that inducing sparsity may make it less plausible that attention can be used as a tool for understanding model behavior.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Membership Inference Attack against Long-Context Large Language Models

    cs.CL 2024-11 conditional novelty 6.0 of 10

    Long-context language models leak membership of documents in their input context, detectable via snippet-based generation loss and semantic similarity attacks.

  2. New Faithfulness-Centric Interpretability Paradigms for Natural Language Processing

    cs.CL 2024-11 conditional novelty 4.0 of 10

    The thesis shows that randomly masking input tokens during fine-tuning makes post-hoc explanations of NLP models consistently faithful under an erasure-based faithfulness metric.

Pith tools