Pith. sign in

REVIEW 1 cited by

Rethinking Self-Attention: Towards Interpretability in Neural Parsing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.03875 v3 pith:PKNHJTZV submitted 2019-11-10 cs.CL cs.LG

classification cs.CLcs.LG
keywords attentionself-attentionmodelheadsinterpretabilitylabellayerparsing
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Attention mechanisms have improved the performance of NLP tasks while allowing models to remain explainable. Self-attention is currently widely used, however interpretability is difficult due to the numerous attention distributions. Recent work has shown that model representations can benefit from label-specific information, while facilitating interpretation of predictions. We introduce the Label Attention Layer: a new form of self-attention where attention heads represent labels. We test our novel layer by running constituency and dependency parsing experiments and show our new model obtains new state-of-the-art results for both tasks on both the Penn Treebank (PTB) and Chinese Treebank. Additionally, our model requires fewer self-attention layers compared to existing work. Finally, we find that the Label Attention heads learn relations between syntactic categories and show pathways to analyze errors.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Your Attention Matters: to Improve Model Robustness to Noise and Spurious Correlations

    cs.LG 2025-07 conditional novelty 5.0 of 10

    Doubly stochastic attention is the most robust of five ViT attention mechanisms to fog corruption in relative accuracy, based on single-seed experiments on CIFAR-10, CIFAR-100, and Imagenette.

Pith tools