Pith. sign in

REVIEW 1 cited by

Learning Syntax Without Planting Trees: Understanding Hierarchical Generalization in Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.16367 v3 pith:42KLFC6W submitted 2024-04-25 cs.CL cs.LG

classification cs.CLcs.LG
keywords hierarchicalgeneralizationtransformerslanguagemodelingtrainedgeneralizemodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Transformers trained on natural language data have been shown to learn its hierarchical structure and generalize to sentences with unseen syntactic structures without explicitly encoding any structural bias. In this work, we investigate sources of inductive bias in transformer models and their training that could cause such generalization behavior to emerge. We extensively experiment with transformer models trained on multiple synthetic datasets and with different training objectives and show that while other objectives e.g. sequence-to-sequence modeling, prefix language modeling, often failed to lead to hierarchical generalization, models trained with the language modeling objective consistently learned to generalize hierarchically. We then conduct pruning experiments to study how transformers trained with the language modeling objective encode hierarchical structure. When pruned, we find joint existence of subnetworks within the model with different generalization behaviors (subnetworks corresponding to hierarchical structure and linear order). Finally, we take a Bayesian perspective to further uncover transformers' preference for hierarchical generalization: We establish a correlation between whether transformers generalize hierarchically on a dataset and whether the simplest explanation of that dataset is provided by a hierarchical grammar compared to regular grammars exhibiting linear generalization.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Information Locality as an Inductive Bias for Neural Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Neural LMs learn languages with lower m-local entropy more easily, suggesting a shared sensitivity to local statistical structure with human learners.

Pith tools