A myopic MINMPC framework learns a value function offline via inverse optimization from expert data, allowing short horizons with near-optimal performance and strict integer feasibility online for hybrid systems.
Advances in neural information processing systems , volume=
6 Pith papers cite this work. Polarity classification is still indexing.
years
2026 6verdicts
UNVERDICTED 6representative citing papers
Wahkon unifies Kolmogorov superposition with RKHS regularization to produce a deep network whose penalized estimator is exactly the MAP under a hierarchical GP prior and achieves minimax-optimal rates.
EDRBO uses ensemble surrogates and Wasserstein ambiguity sets to robustify BO acquisition functions against context distribution mismatch, with sublinear regret O(γ_T √T) and SOTA empirical results on continuous contexts.
ERPPO adds a DSA-based ambiguity estimator to MAPPO and switches between L1 and L2 entropy regularization to improve exploration and stability in non-stationary multi-dimensional observations.
RASP-Tuner matches or beats GP-UCB and CMA-ES regret on seven of nine synthetic non-stationary tasks while running 8-12 times faster per step.
On five tabular security datasets at 10% labels, tuning only the classifier with Bayesian optimization recovers a median 86% of the gains from full joint SSL-classifier optimization.
citing papers explorer
-
Learning myopic mixed-integer nonlinear model predictive control from expert demonstrations
A myopic MINMPC framework learns a value function offline via inverse optimization from expert data, allowing short horizons with near-optimal performance and strict integer feasibility online for hybrid systems.
-
Wahkon: A Statistically Principled Deep RKHS Superposition Network
Wahkon unifies Kolmogorov superposition with RKHS regularization to produce a deep network whose penalized estimator is exactly the MAP under a hierarchical GP prior and achieves minimax-optimal rates.
-
Ensemble Distributionally Robust Bayesian Optimisation with Continuous Context
EDRBO uses ensemble surrogates and Wasserstein ambiguity sets to robustify BO acquisition functions against context distribution mismatch, with sublinear regret O(γ_T √T) and SOTA empirical results on continuous contexts.
-
ERPPO: Entropy Regularization-based Proximal Policy Optimization
ERPPO adds a DSA-based ambiguity estimator to MAPPO and switches between L1 and L2 entropy regularization to improve exploration and stability in non-stationary multi-dimensional observations.
-
RASP-Tuner: Retrieval-Augmented Soft Prompts for Context-Aware Black-Box Optimization in Non-Stationary Environments
RASP-Tuner matches or beats GP-UCB and CMA-ES regret on seven of nine synthetic non-stationary tasks while running 8-12 times faster per step.
-
SemiScope: Disentangling Classifier Tuning and Joint Optimization in Semi-Supervised Security Classification
On five tabular security datasets at 10% labels, tuning only the classifier with Bayesian optimization recovers a median 86% of the gains from full joint SSL-classifier optimization.