Pith. sign in

Canonical reference

Title resolution pending

Canonical reference. 71% of citing Pith papers cite this work as background.

25 Pith papers citing it
Background 71% of classified citations

citation-role summary

background 5 baseline 1 method 1

citation-polarity summary

years

2026 24 2025 1

representative citing papers

Simply Stabilizing the Loop via Fully Looped Transformer

cs.LG · 2026-05-11 · unverdicted · novelty 7.0

Fully Looped Transformer stabilizes looped training up to 12 iterations via distributed inter-loop signals and attention injection, improving downstream performance by up to 13.2%.

A Rod Flow Model for Adam at the Edge of Stability

cs.LG · 2026-05-07 · unverdicted · novelty 7.0

Rod flow models for Adam and related optimizers track discrete iterates at the edge of stability more accurately than standard stable flows across tested ML architectures.

Architecture Generalization with MetaNCA

cs.LG · 2026-07-08 · conditional · novelty 6.0

A learned local rule (Weight Transformer) iteratively self-organizes task-network weights from local graph neighborhoods and generalizes across unseen MLP, CNN, and ResNet architectures up to ~2M parameters.

AMUSE: Anytime Muon with Stable Gradient Evaluation

cs.LG · 2026-05-21 · accept · novelty 6.0

AMUSE stabilizes Muon with time-varying schedule-free gradient evaluation, improving the performance-iteration Pareto frontier without learning-rate schedules.

Budget-aware Auto Optimizer Configurator

cs.AI · 2026-05-06 · unverdicted · novelty 6.0

BAOC samples gradient streams to compute per-block risk metrics for cheap optimizer configs then solves a constrained optimization to minimize total risk under memory and time budgets while preserving training quality.

Muon is Scalable for LLM Training

cs.LG · 2025-02-24 · unverdicted · novelty 6.0

Muon optimizer with weight decay and update scaling achieves ~2x efficiency over AdamW for large LLMs, shown via the Moonlight 3B/16B MoE model trained on 5.7T tokens.

GPC: Large-Scale Generative Pretraining for Transferable Motor Control

cs.CV · 2026-06-28 · unverdicted · novelty 5.0

GPC learns a motion vocabulary via Finite Scalar Quantization and end-to-end RL, then trains an autoregressive transformer for next-token control generation, achieving 99.98% motion reproduction success with emergent robustness.

Personalized Federated Learning for Gradient Alignment

cs.LG · 2026-05-04 · unverdicted · novelty 5.0

pFLAlign uses two gradient alignment mechanisms derived from PAC-Bayesian analysis to reduce variance in local training and distortion in aggregation, yielding state-of-the-art personalization in federated learning.

Can Muon Fine-tune Adam-Pretrained Models?

cs.LG · 2026-05-11 · unverdicted · novelty 4.0

Constraining fine-tuning updates with LoRA mitigates performance degradation when switching from Adam to Muon on pretrained models.

citing papers explorer

Showing 25 of 25 citing papers.