Pith. sign in

REVIEW 3 cited by

How Feature Learning Can Improve Neural Scaling Laws

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.17858 v2 pith:YAPQSSBD submitted 2024-09-26 stat.ML cond-mat.dis-nncs.LG

classification stat.MLcond-mat.dis-nncs.LG
keywords learningscalingtasksfeatureeasyhardkernelmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We develop a solvable model of neural scaling laws beyond the kernel limit. Theoretical analysis of this model shows how performance scales with model size, training time, and the total amount of available data. We identify three scaling regimes corresponding to varying task difficulties: hard, easy, and super easy tasks. For easy and super-easy target functions, which lie in the reproducing kernel Hilbert space (RKHS) defined by the initial infinite-width Neural Tangent Kernel (NTK), the scaling exponents remain unchanged between feature learning and kernel regime models. For hard tasks, defined as those outside the RKHS of the initial NTK, we demonstrate both analytically and empirically that feature learning can improve scaling with training time and compute, nearly doubling the exponent for hard tasks. This leads to a different compute optimal strategy to scale parameters and training time in the feature learning regime. We support our finding that feature learning improves the scaling law for hard tasks but not for easy and super-easy tasks with experiments of nonlinear MLPs fitting functions with power-law Fourier spectra on the circle and CNNs learning vision tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks

    cs.LG 2025-07 conditional novelty 7.0 of 10

    Compute-optimally trained networks of different sizes show loss curves that collapse onto one universal curve after normalization; with learning rate decay, the collapse is tighter than seed-to-seed noise, providing a...

  2. Predictable Scale: Part II, Farseer: A Refined Scaling Law in Large Language Models

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A newly fitted scaling law with model-size-dependent data exponents predicts LLM loss more accurately than Chinchilla, including at a held-out 25.1B model.

  3. Feature learning is decoupled from generalization in high capacity neural networks

    cs.LG 2025-07 conditional novelty 4.0 of 10

    Current feature learning measures quantify the magnitude of representation change, which the authors argue is decoupled from the generalization benefit that neural networks show over their neural tangent kernel.

Pith tools