Pion modifies Muon's Newton-Schulz iterations into a controllable high-pass filter that anchors dominant singular values at 1 while suppressing noisy tails, outperforming Muon and AdamW in VLA and RLVR regimes.
Back to basics: Revisiting exploration in reinforcement learning for llm reasoning via generative probabilities.arXiv preprint arXiv:2602.05281
4 Pith papers cite this work. Polarity classification is still indexing.
years
2026 4representative citing papers
Replacing the importance sampling ratio with a stop-gradient self-anchored ratio for positive advantages yields unclipped, REINFORCE-equivalent gradients that improve exploration without training instability.
ISPO densifies GRPO rewards with sequence-level informativeness and token-level directional signals from policy probabilities to reduce zero-advantage collapse and hallucinated certainty on math benchmarks.
Entrocraft uses rejection sampling to enforce precise entropy schedules in LLM RL by biasing advantages, enabling longer training, better generalization, and higher performance than baselines.
citing papers explorer
-
Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR
Pion modifies Muon's Newton-Schulz iterations into a controllable high-pass filter that anchors dominant singular values at 1 while suppressing noisy tails, outperforming Muon and AdamW in VLA and RLVR regimes.
-
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma
Replacing the importance sampling ratio with a stop-gradient self-anchored ratio for positive advantages yields unclipped, REINFORCE-equivalent gradients that improve exploration without training instability.
-
Momentum for Reasoning: Dense Intrinsic Signals in Policy Optimization
ISPO densifies GRPO rewards with sequence-level informativeness and token-level directional signals from policy probabilities to reduce zero-advantage collapse and hallucinated certainty on math benchmarks.
-
Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control
Entrocraft uses rejection sampling to enforce precise entropy schedules in LLM RL by biasing advantages, enabling longer training, better generalization, and higher performance than baselines.