PWO is a trust-region optimizer for autoregressive NQS that improves stability over Adam and stochastic reconfiguration methods while scaling to 1.5B-parameter models on spin systems.
arXiv preprint arXiv:2511.16652 , year =
7 Pith papers cite this work. Polarity classification is still indexing.
years
2026 7representative citing papers
EGGROLL applies low-rank evolution strategies to train leaky integrate-and-fire spiking neural networks, reaching 79.21% accuracy on N-MNIST with 2.23 times lower per-generation time than full-rank ES.
Elitist (1+M) genetic algorithms follow the loss gradient via mutation-selection, slowed only by noise in the effective-rank directions of the Hessian rather than the full parameter count.
PopuLoRA shows that co-evolving populations of LoRA adapters through cross-evaluated self-play can outperform compute-matched single-agent baselines on multiple code and math reasoning benchmarks.
Frozen random backbones with low-rank LoRA adapters recover 96-100% of fully trained performance on diverse architectures while training only 0.5-40% of parameters.
A screening rule skips evolutionary outer loops when the ratio of best single-shot gain to best cheap gain meets or exceeds 90%, validated on pre-registered lab cases where the gate fired and loops were abandoned.
SMT trains nonlinear RNNs by imitating one-step memory-transition labels generated by a Transformer, replacing BPTT's unrolled credit assignment with time-parallel supervised learning.
citing papers explorer
-
One More Time: Revisiting Neural Quantum States from a Reinforcement Learning Perspective
PWO is a trust-region optimizer for autoregressive NQS that improves stability over Adam and stochastic reconfiguration methods while scaling to 1.5B-parameter models on spin systems.
-
Gradient-Free Training of Spiking Neural Networks via Low-Rank Evolution Strategies
EGGROLL applies low-rank evolution strategies to train leaky integrate-and-fire spiking neural networks, reaching 79.21% accuracy on N-MNIST with 2.23 times lower per-generation time than full-rank ES.
-
Why can genetic algorithms work in high-dimensional search spaces?
Elitist (1+M) genetic algorithms follow the loss gradient via mutation-selection, slowed only by noise in the effective-rank directions of the Hessian rather than the full parameter count.
-
PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play
PopuLoRA shows that co-evolving populations of LoRA adapters through cross-evaluated self-play can outperform compute-matched single-agent baselines on multiple code and math reasoning benchmarks.
-
A Little Rank Goes a Long Way: Random Scaffolds with LoRA Adapters Are All You Need
Frozen random backbones with low-rank LoRA adapters recover 96-100% of fully trained performance on diverse architectures while training only 0.5-40% of parameters.
-
Knowing in Advance When an Evolutionary Outer Loop Will Not Help: A Pre-Registered Cheap-Baseline Screening Rule
A screening rule skips evolutionary outer loops when the ratio of best single-shot gain to best cheap gain meets or exceeds 90%, validated on pre-registered lab cases where the gate fired and loops were abandoned.
-
Pretraining Recurrent Networks without Recurrence
SMT trains nonlinear RNNs by imitating one-step memory-transition labels generated by a Transformer, replacing BPTT's unrolled credit assignment with time-parallel supervised learning.