Pith. sign in

How transformers learn structured data: insights from hierarchical filtering

6 Pith papers cite this work. Polarity classification is still indexing.

6 Pith papers citing it

citation-role summary

background 1

citation-polarity summary

years

2026 6

verdicts

UNVERDICTED 6

roles

background 1

polarities

background 1

representative citing papers

Phase structure of the Random Language Model

cond-mat.dis-nn · 2026-06-26 · unverdicted · novelty 7.0

The Random Language Model exhibits a hierarchy of phase transitions in the double-scaling limit ε̃_d → 0, N → ∞ at fixed x = ε̃_d log N, with symbol correlations, non-uniform marginals, and glassy freezing, yielding scaling laws consistent with large language models.

Critical Percolation as a Synthetic Data Model for Interpretability

cs.LG · 2026-06-18 · unverdicted · novelty 6.0

Critical percolation clusters embedded in high dimensions, combined with taxonomic latent variables, form an analytically tractable synthetic data model whose ground-truth hierarchy can be linearly decoded from network activations.

Sampling Data with Chains of Forward-Backward Diffusion Steps

cs.LG · 2026-05-26 · unverdicted · novelty 5.0

U-turn chains are Markov chains formed by short forward-backward diffusion steps that remain on the learned manifold and, with Metropolis-Hastings, sample from energy-modified targets, exhibiting an ergodicity-breaking transition on fragmented manifolds.

citing papers explorer

Showing 6 of 6 citing papers.

  • Phase structure of the Random Language Model cond-mat.dis-nn · 2026-06-26 · unverdicted · none · ref 12

    The Random Language Model exhibits a hierarchy of phase transitions in the double-scaling limit ε̃_d → 0, N → ∞ at fixed x = ε̃_d log N, with symbol correlations, non-uniform marginals, and glassy freezing, yielding scaling laws consistent with large language models.

  • Learn from your own latents and not from tokens: A sample-complexity theory cs.LG · 2026-05-26 · unverdicted · none · ref 45

    Latent prediction SSL recovers latent trees from PCFG data with sample complexity constant in hierarchy depth L (up to logs), unlike exponential for token-level or supervised methods.

  • Critical Percolation as a Synthetic Data Model for Interpretability cs.LG · 2026-06-18 · unverdicted · none · ref 22

    Critical percolation clusters embedded in high dimensions, combined with taxonomic latent variables, form an analytically tractable synthetic data model whose ground-truth hierarchy can be linearly decoded from network activations.

  • Distributional simplicity bias and effective convexity in Energy Based Models cs.LG · 2026-05-08 · unverdicted · none · ref 6

    Gradient flow in energy-based models for strictly positive binary distributions produces stable data-consistent fixed points and a learning hierarchy that favors lower-order interactions first, mechanistically explaining distributional simplicity bias.

  • Sampling Data with Chains of Forward-Backward Diffusion Steps cs.LG · 2026-05-26 · unverdicted · none · ref 57

    U-turn chains are Markov chains formed by short forward-backward diffusion steps that remain on the learned manifold and, with Metropolis-Hastings, sample from energy-modified targets, exhibiting an ergodicity-breaking transition on fragmented manifolds.

  • Statistical Properties of Training & Generalization stat.ML · 2026-06-18 · unverdicted · none · ref 212 · 2 links

    Review of neural scaling laws and their relation to constraints and inductive biases when applying machine learning to physics problems.