One training example via RLVR boosts LLM math reasoning from 17.6% to 35.7% average across six benchmarks.
Deep double descent: Where bigger models and more data hurt.Journal of Statistical Mechanics: Theory and Experiment, 2021(12):124003
6 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
Overparameterized two-qubit SU(4) QNNs exhibit grokking and epoch-wise double descent; depth raises generalization success, and weak L2 regularization anchors the post-grokking state against weight-norm drift.
Small transformers learn to forecast unseen dynamical systems in-context by using delay embeddings to recover the manifold and forecasting its invariant sets via a transfer-operator strategy.
λ-Orthogonality regularization enables distribution-specific adaptation of representations via affine transformations while retaining original learned structures.
Negative-capable ridge regression uses controlled negative regularization as anti-shrinkage to increase effective complexity along weak eigendirections and mitigate underfitting in small-data regression.
Experiments on modular arithmetic with heavy label noise show that over-parameterized networks form a distributed internal generalization structure that can be extracted via frequency methods to achieve high accuracy despite 80% noise.
citing papers explorer
-
Reinforcement Learning for Reasoning in Large Language Models with One Training Example
One training example via RLVR boosts LLM math reasoning from 17.6% to 35.7% average across six benchmarks.
-
Grokking and epoch-wise double descent in quantum neural networks
Overparameterized two-qubit SU(4) QNNs exhibit grokking and epoch-wise double descent; depth raises generalization success, and weak L2 regularization anchors the post-grokking state against weight-norm drift.
-
Transformers for dynamical systems learn transfer operators in-context
Small transformers learn to forecast unseen dynamical systems in-context by using delay embeddings to recover the manifold and forecasting its invariant sets via a transfer-operator strategy.
-
$\boldsymbol{\lambda}$-Orthogonality Regularization for Compatible Representation Learning
λ-Orthogonality regularization enables distribution-specific adaptation of representations via affine transformations while retaining original learned structures.
-
A Ridge Too Far: Correcting Over-Shrinkage via Negative Regularization
Negative-capable ridge regression uses controlled negative regularization as anti-shrinkage to increase effective complexity along weak eigendirections and mitigate underfitting in small-data regression.
-
Unveiling Memorization-Generalization Coexistence: A Case Study on Arithmetic Tasks with Label Noise
Experiments on modular arithmetic with heavy label noise show that over-parameterized networks form a distributed internal generalization structure that can be extracted via frequency methods to achieve high accuracy despite 80% noise.