Pith. sign in

REVIEW 6 cited by

Grokking Modular Polynomials

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.03495 v1 pith:FNONOHGW submitted 2024-06-05 cs.LG cond-mat.dis-nnhep-thmath.NTstat.ML

classification cs.LGcond-mat.dis-nnhep-thmath.NTstat.ML
keywords modularnetworksgeneralizepolynomialssolutionsadditionanalyticalgrokking
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Neural networks readily learn a subset of the modular arithmetic tasks, while failing to generalize on the rest. This limitation remains unmoved by the choice of architecture and training strategies. On the other hand, an analytical solution for the weights of Multi-layer Perceptron (MLP) networks that generalize on the modular addition task is known in the literature. In this work, we (i) extend the class of analytical solutions to include modular multiplication as well as modular addition with many terms. Additionally, we show that real networks trained on these datasets learn similar solutions upon generalization (grokking). (ii) We combine these "expert" solutions to construct networks that generalize on arbitrary modular polynomials. (iii) We hypothesize a classification of modular polynomials into learnable and non-learnable via neural networks training; and provide experimental evidence supporting our claims.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning words in groups: fusion algebras, tensor ranks and grokking

    cs.LG 2025-09 conditional novelty 8.0 of 10

    Group word operations can be learned by small two-layer networks because the associated word tensor has low rank, decomposable through the fusion algebra of the group's self-conjugate representations.

  2. (How) Can Transformers Predict Pseudo-Random Numbers?

    cs.LG 2025-02 conditional novelty 7.0 of 10

    Transformers predict LCG sequences in-context for fixed moduli up to 2^32 and unseen moduli up to 2^16 by learning the modulus factorization and digit-wise periodic structure.

  3. Algebraic Representability as the Limiting Regime of Grokking: An Exactly Solvable Model with Holomorphic Activations

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A task ma+nb mod p is representable by a z^k holomorphic network iff m+n=k; non-representable tasks cannot be memorised at any width.

  4. Grokking Is Conditional and Fragile: A Fully-Tractable, Multi-Seed Study at 12K Parameters

    cs.LG 2026-07 accept novelty 6.0 of 10

    In a fully tractable 12K Llama-style model, grokking is a conditional fragile phase transition gated by coverage (tracking modulus more than structure), weight decay, and floating-point reduction order, so evidence mu...

  5. Uncovering a Universal Abstract Algorithm for Modular Addition in Neural Networks

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Trained MLPs and transformers solving modular addition can be unified under an approximate Chinese Remainder Theorem, and deep or embedding-based networks learn only O(log n) frequency features.

  6. Mechanistic Insights into Grokking from the Embedding Layer

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Trainable embeddings in a simple MLP cause delayed generalization (grokking) on modular arithmetic, and a higher embedding learning rate plus balanced sampling accelerates it.

Pith tools