Additive regret is governed by the active weighted-mass exponent p near cutoffs: every policy incurs at least T^{1/2-1/(2p)} when p>1, and a sample-path marginal policy matches this (up to logs) while attaining O((log T)^2) when p=1, without non-degeneracy assumptions.
Online resource allocation with stochastic resource consumption
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3representative citing papers
Establishes Ω((log T)^2) lower bound on regret for multi-secretary problem with gapped distributions via Bellman certificates, showing prior O((log T)^2) upper bounds are tight.
CERO uses Beta posteriors and Fenchel-dual online optimization to adaptively allocate a fixed rollout budget across prompts and epochs in LLM RL, outperforming fixed-allocation GRPO on math reasoning benchmarks.
citing papers explorer
-
Online Resource Allocation with Continuous Random Consumption: Regret under Degeneracy
Additive regret is governed by the active weighted-mass exponent p near cutoffs: every policy incurs at least T^{1/2-1/(2p)} when p>1, and a sample-path marginal policy matches this (up to logs) while attaining O((log T)^2) when p=1, without non-degeneracy assumptions.
-
Tight Lower Bounds for the Multi-Secretary Problem via Bellman Certificates
Establishes Ω((log T)^2) lower bound on regret for multi-secretary problem with gapped distributions via Bellman certificates, showing prior O((log T)^2) upper bounds are tight.
-
Cross-Epoch Adaptive Rollout Optimization for RL Post-Training
CERO uses Beta posteriors and Fenchel-dual online optimization to adaptively allocate a fixed rollout budget across prompts and epochs in LLM RL, outperforming fixed-allocation GRPO on math reasoning benchmarks.