Pith. sign in

REVIEW 3 cited by

Unveiling Transformers with LEGO: a synthetic reasoning task

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.04301 v3 pith:ZTIQGL6G submitted 2022-06-09 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords attentionlegoproposereasoningtaskchainpretrainingarchitectural
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a synthetic reasoning task, LEGO (Learning Equality and Group Operations), that encapsulates the problem of following a chain of reasoning, and we study how the Transformer architectures learn this task. We pay special attention to data effects such as pretraining (on seemingly unrelated NLP tasks) and dataset composition (e.g., differing chain length at training and test time), as well as architectural variants such as weight-tied layers or adding convolutional components. We study how the trained models eventually succeed at the task, and in particular, we manage to understand some of the attention heads as well as how the information flows in the network. In particular, we have identified a novel \emph{association} pattern that globally attends only to identical tokens. Based on these observations we propose a hypothesis that here pretraining helps for LEGO tasks due to certain structured attention patterns, and we experimentally verify this hypothesis. We also observe that in some data regime the trained transformer finds ``shortcut" solutions to follow the chain of reasoning, which impedes the model's robustness, and moreover we propose ways to prevent it. Motivated by our findings on structured attention patterns, we propose the LEGO attention module, a drop-in replacement for vanilla attention heads. This architectural change significantly reduces Flops and maintains or even \emph{improves} the model's performance at large-scale pretraining.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Small transformer architectures for task switching

    cs.LG 2025-08 reject novelty 6.0 of 10

    On a new task-switching benchmark, a cisformer with expressive attention reaches about 95% accuracy, while standard transformers, LSTMs, and MLPs all stay near or below 60%.

  2. Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence

    cs.AI 2025-06 conditional novelty 5.0 of 10

    Reliable deductive reasoning in AI requires replacing average-case statistical objectives with the exact learning criterion of universal correctness, a thesis supported by sample-complexity lower bounds showing statis...

  3. Bayesian Inference of Discretization Error Means in ODEs via Ensemble Kalman Filtering

    math.NA 2026-07 conditional novelty 4.0 of 10

    A Bayesian state-space model with an Ensemble Kalman Filter infers the mean of ODE discretization errors from noisy observations, using a step-size-dependent Markov prior whose convergence is proven.

Pith tools