Pith. sign in

hub

Neural Tangent Kernel: Convergence and Generalization in Neural Networks

24 Pith papers cite this work. Polarity classification is still indexing.

24 Pith papers citing it

hub tools

citation-role summary

method 2 background 1 other 1

citation-polarity summary

representative citing papers

Editing Models with Task Arithmetic

cs.LG · 2022-12-08 · accept · novelty 8.0

Task vectors from weight differences allow arithmetic operations to edit pre-trained models, improving multiple tasks simultaneously and enabling analogical inference on unseen tasks.

Dimensional Criticality at Grokking Across MLPs and Transformers

cs.LG · 2026-04-06 · unverdicted · novelty 7.0

Effective cascade dimension D(t) crosses D=1 at the grokking transition in MLPs and Transformers, with opposite directions for modular addition versus XOR, consistent with attraction to a shared critical manifold.

Does Weight Decay Enhance Training Stability?

cs.LG · 2026-05-15 · conditional · novelty 6.0

Weight decay slows progressive sharpening at the edge of stability, inducing damped oscillations in CNNs and a phase transition to sub-2/η sharpness in MLPs driven by parameter-sharpness gradient alignment, yielding more stable NTK dynamics.

Grokking as Dimensional Phase Transition in Neural Networks

cs.LG · 2026-04-06 · unverdicted · novelty 6.0

Grokking occurs as the effective dimensionality of the gradient field transitions from sub-diffusive to super-diffusive at the onset of generalization, exhibiting self-organized criticality.

Bayesian Inference with Shaped Deep Non-linear MLPs

math.ST · 2026-05-29 · unverdicted · novelty 5.0

In the LP/N = Θ(1) regime, Bayesian predictive posteriors for deep MLPs equal those of data-dependent kernels to first order, with a criterion identifying data processes that benefit from larger effective depth.

The Thermodynamic Costs of Simple Linear Regression

cond-mat.stat-mech · 2026-05-18 · unverdicted · novelty 5.0

Thermodynamic lower bounds are approximated for exact and SGD linear regression, producing energy-aware scaling laws for optimal training dataset size given a target generalization error.

Lectures on Semiclassical Methods for Composite Operators

hep-th · 2026-06-09 · unverdicted · novelty 3.0

Lecture notes develop semiclassical methods to compute large-n scaling dimensions of composite operators in CFTs, recovering known results in free theory and deriving one-loop corrections at the Wilson-Fisher fixed point.

Some Inverse Problems in Particle Physics

hep-lat · 2026-06-06 · unverdicted · novelty 2.0

Lectures reviewing three established numerical methods for inverse problems in extracting PDFs and spectral functions from lattice QCD and experimental data.

citing papers explorer

Showing 24 of 24 citing papers.