Pith. sign in

Neural scaling laws rooted in the data distribution

4 Pith papers cite this work. Polarity classification is still indexing.

4 Pith papers citing it

citation-role summary

background 2

citation-polarity summary

fields

cs.LG 3 cs.CL 1

years

2026 3 2025 1

roles

background 2

polarities

background 2

representative citing papers

Smooth Scaling Laws Hide Stepwise Token Learning

cs.CL · 2026-06-29 · conditional · novelty 7.0

Power-law LLM scaling laws are largely the aggregate of stepwise token learning events whose heavy-tailed learning-time spectrum reconstructs loss derivatives along step, data, and model axes.

Critical Percolation as a Synthetic Data Model for Interpretability

cs.LG · 2026-06-18 · unverdicted · novelty 6.0

Critical percolation clusters embedded in high dimensions, combined with taxonomic latent variables, form an analytically tractable synthetic data model whose ground-truth hierarchy can be linearly decoded from network activations.

Superposition Yields Robust Neural Scaling

cs.LG · 2025-05-15 · conditional · novelty 6.0

Strong superposition causes neural loss to scale as the inverse of model dimension due to geometric feature overlaps, explaining scaling laws for broad frequency distributions.

citing papers explorer

Showing 4 of 4 citing papers.

  • Smooth Scaling Laws Hide Stepwise Token Learning cs.CL · 2026-06-29 · conditional · none · ref 7

    Power-law LLM scaling laws are largely the aggregate of stepwise token learning events whose heavy-tailed learning-time spectrum reconstructs loss derivatives along step, data, and model axes.

  • Critical Percolation as a Synthetic Data Model for Interpretability cs.LG · 2026-06-18 · unverdicted · none · ref 8

    Critical percolation clusters embedded in high dimensions, combined with taxonomic latent variables, form an analytically tractable synthetic data model whose ground-truth hierarchy can be linearly decoded from network activations.

  • Practical Scaling Laws: Converting Compute into Performance in a Data-Constrained World cs.LG · 2026-05-09 · conditional · none · ref 11

    A new scaling law L(N, D, T) = E + (L0 - E) h/(1+h) with h = a/N^α + b/T^β + c N^γ/D^δ that decomposes loss into undercapacity, undertraining, and overfitting terms and saturates between E and L0.

  • Superposition Yields Robust Neural Scaling cs.LG · 2025-05-15 · conditional · none · ref 23

    Strong superposition causes neural loss to scale as the inverse of model dimension due to geometric feature overlaps, explaining scaling laws for broad frequency distributions.