Pith. sign in

REVIEW 2 cited by

Compositional generalization in a deep seq2seq model by separating syntax and semantics

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1904.09708 v3 pith:HZLNEAHK submitted 2019-04-22 cs.LG cs.CLstat.ML

classification cs.LGcs.CLstat.ML
keywords compositionaldeepgeneralizationlearningstandardsyntactichumanlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Standard methods in deep learning for natural language processing fail to capture the compositional structure of human language that allows for systematic generalization outside of the training distribution. However, human learners readily generalize in this way, e.g. by applying known grammatical rules to novel words. Inspired by work in neuroscience suggesting separate brain systems for syntactic and semantic processing, we implement a modification to standard approaches in neural machine translation, imposing an analogous separation. The novel model, which we call Syntactic Attention, substantially outperforms standard methods in deep learning on the SCAN dataset, a compositional generalization task, without any hand-engineered features or additional supervision. Our work suggests that separating syntactic from semantic learning may be a useful heuristic for capturing compositional structure.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GXJoin: Generalized Cell Transformations for Explainable Joinability

    cs.DB 2025-05 conditional novelty 6.0 of 10

    Adding relative indexes, reusable and optional rule parts, bidirectional search, and simplicity tie-breaking raises the coverage of the best discovered transformation by up to about 10% on the authors' benchmarks.

  2. Cognitive Decision Routing in Large Language Models: When to Think Fast, When to Think Slow

    cs.CL 2025-08 reject novelty 3.0 of 10

    A four-feature routing rule is claimed to cut token use by 34% and improve accuracy, but the paper provides no code, no data, and only hand-wavy details.

Pith tools