Pith. sign in

REVIEW 2 cited by

Making Transformers Solve Compositional Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2108.04378 v2 pith:BUWXIBYI submitted 2021-08-09 cs.AI cs.CL

classification cs.AIcs.CL
keywords compositionalgeneralizationtaskstransformerbenchmarkcompositionallydesigngeneralize
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Several studies have reported the inability of Transformer models to generalize compositionally, a key type of generalization in many NLP tasks such as semantic parsing. In this paper we explore the design space of Transformer models showing that the inductive biases given to the model by several design decisions significantly impact compositional generalization. Through this exploration, we identified Transformer configurations that generalize compositionally significantly better than previously reported in the literature in a diverse set of compositional tasks, and that achieve state-of-the-art results in a semantic parsing compositional generalization benchmark (COGS), and a string edit operation composition benchmark (PCFG).

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Generalized Locomotion in Out-of-distribution Conditions with Robust Transformer

    cs.RO 2025-07 conditional novelty 6.0 of 10

    A transformer with body tokenization and consistent dropout generalizes to unseen leg damages and sensor noise while trained on limited dynamics and clean observations.

  2. Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models

    cs.CV 2025-05 conditional novelty 5.0 of 10

    VLMs score far below adult humans on a new 13,188-question benchmark of 36 atomic 2D geometry perception skills.

Pith tools