Pith. sign in

REVIEW

Influence Patterns for Explaining Information Flow in BERT

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.00740 v3 pith:UEXXGMYE submitted 2020-11-02 cs.CL

classification cs.CL
keywords patternsbertinformationflowmodelattentionattention-basedinfluence
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While attention is all you need may be proving true, we do not know why: attention-based transformer models such as BERT are superior but how information flows from input tokens to output predictions are unclear. We introduce influence patterns, abstractions of sets of paths through a transformer model. Patterns quantify and localize the flow of information to paths passing through a sequence of model nodes. Experimentally, we find that significant portion of information flow in BERT goes through skip connections instead of attention heads. We further show that consistency of patterns across instances is an indicator of BERT's performance. Finally, We demonstrate that patterns account for far more model performance than previous attention-based and layer-based methods.

Discussion (0). Sign in to comment.

Pith tools