Pith. sign in

REVIEW 6 cited by

Talking-Heads Attention

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.02436 v1 pith:D7VLL33G submitted 2020-03-05 cs.LG cs.NEcs.SDeess.ASstat.ML

Talking-Heads Attention

classification cs.LG cs.NEcs.SDeess.ASstat.ML
keywords attentiontalking-headsbetterlanguagetasksacrossadditionaladditionalcomputation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We introduce "talking-heads attention" - a variation on multi-head attention which includes linearprojections across the attention-heads dimension, immediately before and after the softmax operation.While inserting only a small number of additional parameters and a moderate amount of additionalcomputation, talking-heads attention leads to better perplexities on masked language modeling tasks, aswell as better quality when transfer-learning to language comprehension and question answering tasks.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Mosaic: A Fleet of User Embedding Specialists for Recommendation at Meta

    cs.IR 2026-07 conditional novelty 6.0

    Mosaic shows that a fleet of four heterogeneous user-embedding specialists, trained with redundancy-reduction and composite-label losses, improves downstream recommendation quality at Meta.

  2. Motion-Compensated Weight Compression

    cs.CV 2026-05 unverdicted novelty 6.0

    MCWC aligns permutation-symmetric blocks across layers to enable sequential prediction and residual entropy coding, improving rate-accuracy tradeoffs versus quantization and prior codecs on language and vision models.

  3. Position-Agnostic Pre-Projection for Transformer Attention: Nonlinear Feature Construction and Content Skip Before Q/K/V

    cs.CL 2026-04 unverdicted novelty 6.0

    A position-agnostic nonlinear pre-projection MLP plus content skip connection in transformer attention improves LAMBADA accuracy by 40.6% and reduces perplexity by 39% on 160M-scale models.

  4. TalkLoRA: Communication-Aware Mixture of Low-Rank Adaptation for Large Language Models

    cs.LG 2026-04 unverdicted novelty 6.0

    TalkLoRA equips MoE-LoRA experts with a communication module that smooths routing dynamics and improves performance on language tasks under similar parameter budgets.

  5. LoSA-Net: A Localized and Scale-Adaptive Network for Boundary-Sensitive Prediction of Perineural Invasion in 3D MRI

    cs.CV 2026-07 conditional novelty 5.0

    A localized, scale-adaptive 3D encoder (TNA+SAFM+CSRA) predicts perineural invasion from contrast-enhanced MRI with AUC 0.7567, outperforming matched CNN and transformer baselines on 168 patients.

  6. Nexusformer: Nonlinear Attention Expansion for Stable and Inheritable Transformer Scaling

    cs.LG 2026-04 unverdicted novelty 5.0

    Nexusformer uses a three-stage nonlinear mapping in attention to enable stable, inheritable scaling of transformers, matching baseline perplexity with up to 41.5% less compute when growing from 240M to 440M parameters.