Pith. sign in

REVIEW 1 cited by

A Theory for Length Generalization in Learning to Reason

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.00560 v1 pith:GXWSLBM2 submitted 2024-03-31 cs.AI

classification cs.AI
keywords problemslearningreasonreasoningchallenginggeneralizationlengthlengths
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Length generalization (LG) is a challenging problem in learning to reason. It refers to the phenomenon that when trained on reasoning problems of smaller lengths or sizes, the resulting model struggles with problems of larger sizes or lengths. Although LG has been studied by many researchers, the challenge remains. This paper proposes a theoretical study of LG for problems whose reasoning processes can be modeled as DAGs (directed acyclic graphs). The paper first identifies and proves the conditions under which LG can be achieved in learning to reason. It then designs problem representations based on the theory to learn to solve challenging reasoning problems like parity, addition, and multiplication, using a Transformer to achieve perfect LG.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Saving for the future: Enhancing generalization via partial logic regularization

    cs.LG 2025-08 reject novelty 4.0 of 10

    PL-Reg adds a trainable mask and a defined/undefined classification loss to logic-based regularization, improving unknown-class accuracy across GCD, mDG+GCD, and CIL benchmarks.

Pith tools