REVIEW 4 major objections 6 minor 111 references
Learning-Augmented Online Control for Decarbonizing Water Infrastructures
T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read A safe action set with a future-risk hedge lets LAOC guarantee per-round safety for ML-based water pump control, no matter how bad the predictions are.
desk verdict A genuine per-round safety guarantee relative to a control prior, built on a reservation-based safe action set, but the abstract overstates it as absolute safety and the experiments need tighter reporting. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The named object is the safe action set $U_{\lambda,h}$ with reservation function. The reservation $\phi_h(u)$ is a quadratic penalty on the distance between the state that action would produce and the state the control prior would produce, scaled by a coefficient $q_h$ that grows with the remaining horizon. This hedges worst-case future risk differences: if the chosen action drives the state away from the prior's virtual trajectory, the reservation consumes part of the safety budget now, so that the set remains feasible later. The proof of non-emptiness is an induction: since $u^\dagger_h$ lies in $U_{\lambda,h}$ at every round, LAOC can always fall back to the prior action, which is what converts a relative safety constraint into a strict anytime guarantee.
What would settle it
Run LAOC on a horizon with linear dynamics, adversarial demand, and a risk function satisfying Assumption 3.2, using a prior whose cumulative risk $R^\dagger_h$ is exactly zero at some intermediate round; the non-emptiness argument must still place an action in $U_{\lambda,h}$ that keeps $R_h + \phi_h(u) \le 0$, and any violation of (4) at that round falsifies Theorem 4.3. A simpler check is to search over synthetic sequences for any round where LAOC's cumulative risk exceeds $(1+\lambda)R^\dagger_h$, which the theorem says never happens.
Extended reading notes
Core claim
The paper's central discovery is a sufficient design for safe action sets in online control: at round $h$ the set $U_{\lambda,h} = \{ u_h : R_h + \phi_h(u_h) \le (1+\lambda) R^{\pi^\dagger}_h \}$, with reservation $\phi_h(u) = q_h \|f_h(x_h,u) - f_h(x^\dagger_h,u^\dagger_h)\|^2$, where $(x^\dagger_h,u^\dagger_h)$ is the state-action of the control prior and $q_h$ scales with the remaining horizon. Proposition 4.2 shows by induction that if the previous action was in the safe set, the next safe set is never empty and always contains the prior action $u^\dagger_h$. Because LAOC always chooses an action in $U_{\lambda,h}$—the ML action when it is admissible, otherwise a projection or linear interpolation onto the set—Theorem 4.3 follows: the $(1+\lambda)$-safety constraint holds per round, for every problem sequence $y_{1:H}$ and every round $h$, regardless of ML prediction quality. The companion performance bound (Theorem 4.4) shows the expected loss of LAOC is at most that of the pure ML policy plus a term that vanishes as $\lambda$ grows, quantifying the safety-versus-efficiency tradeoff.
Load-bearing premise
The load-bearing premise is that the control prior $\pi^\dagger$ is genuinely safe in the absolute sense; LAOC's guarantee is relative, so if the prior can violate absolute safety in some scenario, LAOC inherits that failure.
Editorial extensions
If this is right
- For any water supply system with a trusted controller as prior, LAOC strictly satisfies $(1+\lambda)$-safety per round, even against adversarial demand and arbitrarily bad ML predictions (Theorem 4.3).
- The expected loss of LAOC is bounded by the pure ML policy's expected loss plus a safety-induced gap that shrinks as $\lambda$ grows; at large $\lambda$ it recovers pure-ML average performance (Theorem 4.4).
- Safety-aware finetuning of the ML model converges at $O(\sqrt{1/n})$ to the unconstrained optimal expected loss plus the same $\lambda$-dependent gap (Theorem 4.5).
- On building water supply data, LAOC reduces average carbon and energy costs compared to OGD, ROBD, and MPC priors while keeping the maximum risk ratio low, and it never violates the safety constraint in distribution or out-of-distribution tests.
- The same safe-action-set machinery transfers to battery management of EV charging stations and cooling control for data centers by replacing the dynamics, cost, and risk functions.
Reading between the lines
- Because the guarantee is relative to the prior, real-world absolute safety still depends on the prior's own worst-case behavior; LAOC inherits any scenario where the prior violates an absolute safety limit, so the algorithm should be paired with a certified prior in deployment.
- The reservation idea is a general recipe: maintain a risk budget with a lookahead hedge proportional to the state discrepancy, which could be adapted to other safety metrics such as control barrier functions or chance constraints.
- A testable extension is to set $\lambda$ adaptively online, using the Theorem 4.4 bound to choose the loosest safety requirement that still meets a given worst-case risk target, improving average cost without violating the constraint.
- The per-round (anytime) guarantee is stronger than the cumulative guarantees typical in learning-augmented control; one could investigate whether the reservation double-counting can be tightened to improve the constant in Theorem 4.4.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes LAOC, a learning-augmented online control algorithm for water supply systems. The controller takes action recommendations from a machine-learned policy and combines them with a 'control prior', a trusted baseline policy, through a per-round safe action set that reserves risk budget for future rounds. The main theoretical claim is Theorem 4.3: for any problem sequence and any round, LAOC satisfies the (1+λ)-safety constraint R^π_h ≤ (1+λ)R^{π†}_h, where R^{π†}_h is the cumulative safety risk of the control prior. The paper also provides an average-cost bound (Theorem 4.4) and a generalization bound for a safety-aware finetuning variant (Theorem 4.5). Experiments on a building water supply case study report lower average energy/carbon costs than traditional priors while keeping the risk ratio low. The core design idea, a reservation function φ_h that hedges against worst-case future risk differences, is technically interesting and the induction in Proposition 4.2 appears coherent.
Significance. If the results stand, the paper makes a useful contribution to learning-augmented online control: it moves beyond expected or high-probability safety by giving a per-round, worst-case guarantee relative to a trusted baseline, and it supports the guarantee with a constructive non-emptiness argument. The explicit tradeoff between the safety parameter λ and the average-cost bound is also valuable, and the experimental study is on a realistic water supply problem with public data. However, the guarantee is relative to the control prior's risk; the paper does not construct a prior with an absolute physical-safety certificate, and the experimental risk-ratio metric is not clearly defined. These issues affect the interpretation of the central safety claim and the empirical evidence.
major comments (4)
- [Section 3.1 and Theorem 4.3] The safety guarantee in Theorem 4.3 is relative: it bounds R^π_h by (1+λ)R^{π†}_h, and the paper assumes in Section 3.1 that a 'genuinely safe control prior' exists, but no such prior is constructed. The risk function in Eq. (3) is a soft quadratic penalty, so a finite cumulative risk bound does not imply that the water level stays above any absolute emergency threshold. The abstract's statement that LAOC 'can provably guarantee safety constraints' is therefore stronger than what Theorem 4.3 delivers. This is a load-bearing limitation because the water-supply safety problem is ultimately about hard physical constraints. Please either (i) construct or identify a prior with a per-round hard safety certificate and state the resulting absolute guarantee, or (ii) qualify the abstract and introduction to say the guarantee is relative to a trusted control prior, and add an explicit discussion of what properties the prior must have for the guarantee to imply physical safety.
- [Section 5.1.2, Tables 1 and 2] The 'Max risk ratio' metric is ambiguous. The text defines it as max_{y∈D_test} R^π_H / R^{π†}_H, but it does not specify which control prior π† is used for the rows corresponding to the priors themselves. If OGD is the prior, then OGD's ratio to itself should be 1.0, yet Table 1 reports 2.04; similar remarks apply to ROBD (reported 1.14). This makes the empirical safety comparisons difficult to interpret. Please state the reference prior for each column, or use a common reference (e.g., ROBD) and label it clearly.
- [Theorem 4.4 and Proposition 4.2] The 'optimal' choice of C2 in Theorem 4.4 is given as arg min_{c≥1} { c/(c-1) σ_u² (1-(cσ_x²)^{H-h})/(1-cσ_x²) }, which depends on the round h. However, Proposition 4.2 and the proof of Lemma D.2 use a single constant C2 in the reservation function q_h. If C2 must vary by round, the notation and proofs should be changed accordingly (e.g., to C2,h); if C2 is fixed, the optimization objective should not contain H-h. As written, the parameter choice in the theorem is not well-defined, although the proof of the safety guarantee appears to hold for any admissible fixed C2≥1.
- [Appendix B, Proposition 4.1] The statement of Proposition 4.1 defines the quality of the pure ML policy as ∥ũ-u*∥²/J*_H, but the proof bounds ∥ũ-u*∥²/R*_H (the risk of the offline optimal) and uses the prior's competitive ratio for the risk objective. The proposition and its proof are therefore mismatched. Please either change the proposition's quality measure to the risk-based one, or revise the proof to use the cost J*_H. This is a negative result motivating the safe action set, but it should still be stated correctly.
minor comments (6)
- [Section 1] There is a typo in the third paragraph: 'no worse than a the safety performance benchmark' should read 'no worse than the safety performance benchmark'.
- [Example 4.1] In the sentence 'the true loss c_{h+1}(x_{h+1},u) is lager than the scaled prior loss', 'lager' should be 'larger'.
- [Theorem 4.4 statement] The theorem statement says the bound depends on 'A is the size of the state-action set', but A does not appear in inequality (11). Either remove this mention or clarify where A enters the bound.
- [Theorem 4.5] The O(·) term is described only as 'scaling with the loss upper bound P, the horizon H, and the size of action-state space X×U'. Please state the explicit dependence or give a more precise order notation, since the current form makes the convergence rate hard to verify.
- [Section 5.1.4] The text says 'C1 and C2 are chosen based on Theorem 4.4' but does not specify how the Lipschitz constants σ_x, σ_u are obtained for the water supply dynamics in Eq. (1), nor what values of C1, C2 were used. Please provide these details (or state that the bounds are implemented with conservative estimates) so the experiments are reproducible.
- [Appendix D, Lemma D.2 proof] In the proof of Lemma D.2, 'Since U_{λ,h} is a close set' should read 'closed set'.
Circularity Check
No significant circularity: the safety guarantee is a design invariant of the safe action set with a nontrivial non-emptiness proof, and the cost bounds are analytic.
full rationale
The derivation chain is self-contained. The safety guarantee (Theorem 4.3) is obtained by defining the safe action set U_{λ,h} in (7) so that membership implies R_h + φ_h(u_h) ≤ (1+λ)R†_h, i.e. the same inequality as the safety constraint (4) up to a nonnegative reservation; Theorem 4.3 is therefore a corollary of this definition together with Proposition 4.2, which supplies the substantive content by proving U_{λ,h} is always non-empty (it contains u†_h) under Assumptions 3.1 and 3.2. This is a standard invariant-set design, not a hidden tautology: the non-emptiness proof uses Lipschitz dynamics, strong convexity/smoothness, and the choice q_h = C1(1+1/λ)β/2 Σ (C2 σ_x^2)^{h'}, with no fitted parameter relabeled as a prediction. The expected-cost bounds in Theorems 4.4 and 4.5 are derived analytically from Lipschitz continuity of the ML policy and cost, the reservation design, and a standard covering-number generalization bound; the constants C1 and C2 are chosen by optimization, not fit to data. The paper cites related work by its own authors (e.g., [41], [54], [59], [79], [100]) but only as background or as instantiations of control priors; the core proof depends on generic lemmas (Lemma C.1 from [45], the generalization theorem [15]), not on a self-citation chain. The main caveat is semantic rather than circular: the guarantee in (4) is relative to an assumed-safe control prior π†, and the paper does not certify absolute water-level thresholds; this is an assumption/scope concern, not a circularity. Hence no significant circularity; score 2 reflects the mild by-construction nature of the safety invariant, not a fitted or self-citational derivation.
Assumptions & free parameters
free parameters (4)
- lambda (safety requirement) =
0.4 or 0.8 in experiments
- C1 =
1+1/sqrt(1+lambda)
- C2 =
argmin expression in Theorem 4.4
- gamma1, gamma2, gamma3, gamma_w, gamma_b =
not reported
assumptions (5)
- domain assumption Assumption 3.1: each f_h is Lipschitz continuous in x and u with constants σ_x and σ_u.
- domain assumption Assumption 3.2: each risk function r_h is non-negative, α-strongly convex, and β-smooth.
- domain assumption The control prior π† is safe and can serve as a safety benchmark.
- domain assumption For Theorem 4.4, the ML policy is L_π-Lipschitz and the cost functions are L_c-Lipschitz.
- ad hoc to paper C1≥1 and C2≥1 are constants; C2 is chosen such that the geometric sums in q_h are finite.
Cite this review
Pith. "Pith review of Learning-Augmented Online Control for Decarbonizing Water Infrastructures." pith.science (2026). https://pith.science/paper/LPTTPMVW
@misc{pith2026250114232,
author = {Pith},
title = {Pith review of: Learning-Augmented Online Control for Decarbonizing Water Infrastructures},
year = {2026},
howpublished = {\url{https://pith.science/paper/LPTTPMVW}},
note = {Machine review of arXiv:2501.14232}
}
read the original abstract
Water infrastructures are essential for drinking water supply, irrigation, fire protection, and other critical applications. However, water pumping systems, which are key to transporting water to the point of use, consume significant amounts of energy and emit millions of tons of greenhouse gases annually. With the wide deployment of digital water meters and sensors in these infrastructures, Machine Learning (ML) has the potential to optimize water supply control and reduce greenhouse gas emissions. Nevertheless, the inherent vulnerability of ML methods in terms of worst-case performance raises safety concerns when deployed in critical water infrastructures. To address this challenge, we propose a learning-augmented online control algorithm, termed LAOC, designed to dynamically schedule the activation and/or speed of water pumps. To ensure safety, we introduce a novel design of safe action sets for online control problems. By leveraging these safe action sets, LAOC can provably guarantee safety constraints while utilizing ML predictions to reduce energy and environmental costs. Our analysis reveals the tradeoff between safety requirements and average energy/environmental cost performance. Additionally, we conduct an experimental study on a building water supply system to demonstrate the empirical performance of LAOC. The results indicate that LAOC can effectively reduce environmental and energy costs while guaranteeing safety constraints.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Planning for Sustainable Water Infrastructure
2023. Planning for Sustainable Water Infrastructure. https://www.epa.gov/ sustainable-water-infrastructure/planning-sustainable-water-infrastructure
2023
-
[2]
Energy Efficiency for Water Utilities
2024. Energy Efficiency for Water Utilities. https://www.epa.gov/sustainable- water-infrastructure/energy-efficiency-water-utilities
2024
-
[3]
Exploring the interdependence of two critical resources Energy and Water
2024. Exploring the interdependence of two critical resources Energy and Water. https://www.iea.org/topics/energy-and-water
2024
-
[4]
Horsepower required to pump water
2024. Horsepower required to pump water. https://www.engineeringtoolbox. com/pumping-water-horsepower-d_753.html
2024
-
[5]
Real-time Price by CAISO
2024. Real-time Price by CAISO. http://www.energyonline.com/Data/ GenericData.aspx?DataId=19&CAISO___Real-time_Price/
2024
-
[6]
STRATEGIES FOR SAVING ENERGY AT PUBLIC WATER SYSTEMS
2024. STRATEGIES FOR SAVING ENERGY AT PUBLIC WATER SYSTEMS. https://www.epa.gov/sites/default/files/2015-04/documents/epa816f13004.pdf
2024
-
[7]
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel. 2017. Constrained Policy Optimization. In International conference on machine learning . PMLR, 22–31
2017
-
[8]
Akshay Agrawal, Brandon Amos, Shane Barratt, Stephen Boyd, Steven Diamond, and J Zico Kolter. 2019. Differentiable convex optimization layers. Advances in neural information processing systems 32 (2019)
2019
Show all 111 references
-
[9]
Mohammad Ali Alomrani, Reza Moravej, and Elias Boutros Khalil. 2022. Deep Policies for Online Bipartite Matching: A Reinforcement Learning Approach. Transactions on Machine Learning Research (2022). https://openreview.net/ forum?id=mbwm7NdkpO
2022
-
[10]
Sanae Amani, Christos Thrampoulidis, and Lin Yang. 2021. Safe Reinforce- ment Learning with Linear Function Approximation. In Proceedings of the 38th International Conference on Machine Learning (Proceedings of Machine Learn- ing Research, Vol. 139), Marina Meila and Tong Zhan...
2021
-
[11]
BD Anderson and John B Moore. 2012. Optimal Filtering. Courier Corporation. Courier Corporation (2012)
2012
-
[12]
Antonios Antoniadis, Christian Coester, Marek Elias, Adam Polak, and Bertrand Simon. 2020. Online Metric Algorithms with Untrusted Predictions. In ICML
2020
-
[13]
Anil Aswani, Humberto Gonzalez, S Shankar Sastry, and Claire Tomlin. 2013. Provably safe and robust learning-based model predictive control. Automatica 49, 5 (2013), 1216–1226
2013
-
[14]
Gissella Bejarano, Adita Kulkarni, Raushan Raushan, Anand Seetharam, and Arti Ramesh. 2019. Swap: Probabilistic graphical and deep learning models for water consumption prediction. In Proceedings of the 6th ACM International Conference on Systems for Energy-Efficient Buildings...
2019
-
[15]
Olivier Bousquet, Stéphane Boucheron, and Gábor Lugosi. 2004. Introduction to statistical learning theory. Advanced Lectures on Machine Learning: ML Summer Schools 2003, Canberra, Australia, February 2-14, 2003, Tübingen, Germany, August 4-16, 2003, Revised Lectures (2004), 169–207
2004
-
[16]
Lukas Brunke, Melissa Greeff, Adam W Hall, Zhaocong Yuan, Siqi Zhou, Jacopo Panerati, and Angela P Schoellig. 2022. Safe learning in robotics: From learning- based control to safe reinforcement learning. Annual Review of Control, Robotics, and Autonomous Systems 5, 1 (2022), 411–444
2022
-
[17]
Eduardo F Camacho and Carlos Bordons Alba. 2013. Model predictive control. Springer
2013
-
[18]
Agustin Castellano, Hancheng Min, Juan Bazerque, and Enrique Mallada. 2022. Reinforcement Learning with Almost Sure Constraints. InLearning for Dynamics and Control
2022
-
[19]
Carmen Cheh, Justin Albrethsen, Zhen Wei Ng, Binbin Chen, Xin Lou, Zaki Masood, and David KY Yau. 2024. Water Pump Operation Optimization under Dynamic Market and Consumer Behaviour. (2024), 335–346
2024
-
[20]
Bingqing Chen, Priya L Donti, Kyri Baker, J Zico Kolter, and Mario Bergés. 2021. Enforcing policy feasibility constraints through differentiable projection for energy optimization. In Proceedings of the Twelfth ACM International Conference on Future Energy Systems . 199–210
2021
-
[22]
Yuri Chervonyi, Praneet Dutta, Piotr Trochim, Octavian Voicu, Cosmin Padu- raru, Crystal Qian, Emre Karagozler, Jared Quincy Davis, Richard Chippendale, Gautam Bajaj, et al. 2022. Semi-analytical industrial cooling system model for reinforcement learning. arXiv preprint arXiv:...
2022 arXiv
-
[23]
Nicolas Christianson, Junxuan Shen, and Adam Wierman. 2023. Optimal robustness-consistency tradeoffs for learning-augmented metrical task systems. In AI STATS
2023
-
[24]
Joshua Comden, Sijie Yao, Niangjun Chen, Haipeng Xing, and Zhenhua Liu
-
[25]
Dongsheng Ding, Xiaohan Wei, Zhuoran Yang, Zhaoran Wang, and Mihailo Jovanovic. 2021. Provably efficient safe exploration via primal-dual policy optimization. In International Conference on Artificial Intelligence and Statistics . PMLR, 3304–3312
2021
-
[26]
Dongsheng Ding, Kaiqing Zhang, Tamer Basar, and Mihailo Jovanovic. 2020. Natural policy gradient primal-dual method for constrained markov decision processes. Advances in Neural Information Processing Systems 33 (2020), 8378– 8390
2020
-
[27]
Bingqian Du, Chuan Wu, and Zhiyi Huang. 2019. Learning Resource Al- location and Pricing for Cloud Profit Maximization. In Proceedings of the Thirty-Third AAAI Conference on Artificial Intelligence and Thirty-First Inno- vative Applications of Artificial Intelligence Conferenc...
2019 doi
-
[28]
Claudia D’Ambrosio, Andrea Lodi, Sven Wiese, and Cristiana Bragalli. 2015. Mathematical programming techniques in water network optimization. Euro- pean Journal of Operational Research 243, 3 (2015), 774–788
2015
-
[29]
Yonathan Efroni, Shie Mannor, and Matteo Pirotta. 2020. Exploration- exploitation in constrained mdps. arXiv preprint arXiv:2003.02189 (2020)
2020 arXiv
-
[30]
David D Fan, Ali-akbar Agha-mohammadi, and Evangelos A Theodorou. 2020. Deep learning tubes for tube mpc. arXiv preprint arXiv:2002.01587 (2020)
2020 arXiv
-
[31]
Ana Lùcia D Franco, Henri Bourlès, Edson R De Pieri, and Herve Guillard
-
[32]
Randy Freeman and Petar V Kokotovic. 2008. Robust nonlinear control design: state-space and Lyapunov techniques . Springer Science & Business Media
2008
-
[33]
Aditya Gahlawat, Pan Zhao, Andrew Patterson, Naira Hovakimyan, and Evan- gelos Theodorou. 2020. L1-GP: L1 adaptive control with Bayesian learning. In Learning for dynamics and control . PMLR, 826–837
2020
-
[34]
Bissan Ghaddar, Joe Naoum-Sawaya, Akihiro Kishimoto, Nicole Taheri, and Bradley Eck. 2015. A Lagrangian decomposition approach for the pump sched- uling problem in water networks. European Journal of Operational Research 241, 2 (2015), 490–501
2015
-
[35]
Zabih Ghelichi, Javad Tajik, and Mir Saman Pishvaee. 2018. A novel robust optimization approach for an integrated municipal water distribution system design under uncertainty: A case study of Mashhad. Computers & Chemical Engineering 110 (2018), 13–34
2018
-
[36]
Arnob Ghosh, Xingyu Zhou, and Ness Shroff. 2022. Provably efficient model-free constrained rl with linear function approximation. arXiv preprint arXiv:2206.11889 (2022)
2022 arXiv
-
[37]
Arnob Ghosh, Xingyu Zhou, and Ness Shroff. 2022. Provably efficient model- free constrained rl with linear function approximation. Advances in Neural Information Processing Systems 35 (2022), 13303–13315
2022
-
[38]
Gautam Goel, Naman Agarwal, Karan Singh, and Elad Hazan. 2022. Best of Both Worlds in Online Control: Competitive Ratio and Policy Regret. arXiv preprint arXiv:2211.11219 (2022)
2022 arXiv
- [39]
-
[40]
Gautam Goel and Babak Hassibi. 2022. Competitive control.IEEE Trans. Automat. Control (2022)
2022
-
[41]
Gautam Goel, Yiheng Lin, Haoyuan Sun, and Adam Wierman. 2019. Beyond Online Balanced Descent: An Optimal Algorithm for Smoothed Online Opti- mization. In NeurIPS, Vol. 32. https://proceedings.neurips.cc/paper/2019/file/ 9f36407ead0629fc166f14dde7970f68-Paper.pdf
2019
-
[42]
Gautam Goel, Yiheng Lin, Haoyuan Sun, and Adam Wierman. 2019. Beyond online balanced descent: an optimal algorithm for smoothed online optimiza- tion. In Proceedings of the 33rd International Conference on Neural Information Processing Systems. Curran Associates Inc., Red Hook...
2019
-
[43]
Gautam Goel and Adam Wierman. 2019. An Online Algorithm for Smoothed Online Convex Optimization. SIGMETRICS Perform. Eval. Rev. 47, 2 (Dec. 2019), 6–8
2019
-
[44]
Alexander P Goryashko and Arkadi S Nemirovski. 2014. Robust energy cost optimization of water distribution system with uncertain demand. Automation and Remote Control 75 (2014), 1754–1769
2014
-
[45]
Moritz Hardt and Max Simchowitz. 2018. Convex Optimization and Approxi- mation. https://ee227c.github.io/notes/ee227c-notes.pdf
2018
-
[46]
Lukas Hewing, Kim P Wabersich, Marcel Menner, and Melanie N Zeilinger
-
[47]
Keke Huang, Ke Wei, Fanbiao Li, Chunhua Yang, and Weihua Gui. 2022. LSTM- MPC: A deep learning based predictive control method for multimode process control. IEEE Transactions on Industrial Electronics 70, 11 (2022), 11544–11554
2022
-
[48]
Girish Joshi and Girish Chowdhary. 2019. Deep model reference adaptive control. In 2019 IEEE 58th Conference on Decision and Control (CDC) . IEEE, 4601–4608
2019
-
[49]
Girish Joshi, Jasvir Virdi, and Girish Chowdhary. 2021. Asynchronous deep model reference adaptive control. In Conference on Robot Learning . PMLR, 984– 1000
2021
-
[50]
Marvin Jung, Paulo Renato da Costa Mendes, Magnus Önnheim, and Emil Gustavsson. 2023. Model Predictive Control when utilizing LSTM as dynamic models. Engineering Applications of Artificial Intelligence 123 (2023), 106226
2023
-
[51]
IS Khalil, JC Doyle, and K Glover. 1996. Robust and optimal control . Prentice hall
1996
-
[52]
Rüdiger Kiesel and Michael Kusterman. 2016. Structural models for coupled electricity markets. Journal of Commodity Markets 3, 1 (2016), 16–38
2016
-
[53]
Zachary J Lee, Tongxin Li, and Steven H Low. 2019. ACN-data: Analysis and applications of an open EV charging dataset. In Proceedings of the tenth ACM international conference on future energy systems . 139–149
2019
-
[54]
Pengfei Li, Jianyi Yang, and Shaolei Ren. 2022. Expert-Calibrated Learning for Online Optimization with Switching Costs. In SIGMETRICS
2022
-
[55]
Pengfei Li, Jianyi Yang, and Shaolei Ren. 2022. Expert-Calibrated Learning for Online Optimization with Switching Costs. Proc. ACM Meas. Anal. Comput. Syst. 6, 2, Article 28 (Jun 2022), 35 pages
2022
-
[56]
Pengfei Li, Jianyi Yang, and Shaolei Ren. 2023. Learning for Edge-Weighted Online Bipartite Matching with Robustness Guarantees. ICML (2023)
2023
-
[57]
Pengfei Li, Jianyi Yang, and Shaolei Ren. 2023. Robustified Learning for Online Optimization with Memory Costs. INDOCOM (2023)
2023
-
[58]
Tongxin Li, Bo Sun, Yue Chen, Zixin Ye, Steven H Low, and Adam Wierman
-
[59]
Tongxin Li, Ruixiao Yang, Guannan Qu, Guanya Shi, Chenkai Yu, Adam Wier- man, and Steven Low. 2022. Robustness and Consistency in Linear Quadratic Control with Untrusted Predictions. Proc. ACM Meas. Anal. Comput. Syst. 6, 1, Article 18 (feb 2022), 35 pages. https://doi.org/10....
2022 doi
-
[60]
Yingying Li, Xin Chen, and Na Li. 2019. Online optimal control with linear dynamics and predictions: Algorithms and regret analysis. Advances in Neural Information Processing Systems 32 (2019)
2019
-
[61]
Yingying Li and Na Li. 2020. Leveraging Predictions in Smoothed Online Convex Optimization via Gradient-based Algorithms. In NeurIPS, Vol. 33. https: //proceedings.neurips.cc/paper/2020/file/a6e4f250fb5c56aaf215a236c64e5b0a- Paper.pdf
2020
-
[62]
Yingying Li, Jing Yu, Lauren Conger, Taylan Kargin, and Adam Wierman. [n. d.]. Learning the Uncertainty Sets of Linear Control Systems via Set Membership: A Non-asymptotic Analysis. In Forty-first International Conference on Machine Learning
-
[63]
Enming Liang, Minghua Chen, and Steven H. Low. 2023. Low Complexity Homeomorphic Projection to Ensure Neural-Network Solution Feasibility for Optimization over (Non-)Convex Set. In ICML
2023
-
[64]
Yiheng Lin, Judy Gan, Guannan Qu, Yash Kanoria, and Adam Wierman. 2022. De- centralized Online Convex Optimization in Networked Systems. InInternational Conference on Machine Learning . PMLR, 13356–13393
2022
-
[65]
Yiheng Lin, Yang Hu, Guanya Shi, Haoyuan Sun, Guannan Qu, and Adam Wierman. 2021. Perturbation-based Regret Analysis of Predictive Control in Linear Time Varying Systems. In Advances in Neural Information Processing Systems, M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, ...
2021
-
[66]
Tao Liu, Ruida Zhou, Dileep Kalathil, Panganamala Kumar, and Chao Tian. 2021. Learning policies with zero or bounded constraint violation for constrained mdps. Advances in Neural Information Processing Systems 34 (2021), 17183– 17193
2021
-
[67]
Brett T Lopez, Jean-Jacques E Slotine, and Jonathan P How. 2019. Dynamic tube MPC for nonlinear systems. In 2019 American Control Conference (ACC). IEEE, 1655–1662
2019
-
[68]
Tiago Luna, João Ribau, David Figueiredo, and Rita Alves. 2019. Improving energy efficiency in water supply systems with pump scheduling optimization. Journal of cleaner production 213 (2019), 342–356
2019
-
[69]
Jerry Luo, Cosmin Paduraru, Octavian Voicu, Yuri Chervonyi, Scott Munns, Jerry Li, Crystal Qian, Praneet Dutta, Jared Quincy Davis, Ningjia Wu, et al
-
[70]
Electricity Maps. 2024. Carbon Intensity Data (Version January 17, 2024). Elec- tricity Maps Data Portal (2024). https://www.electricitymaps.com/data-portal
2024
-
[71]
Konstantinos Oikonomou and Masood Parvania. 2018. Optimal coordination of water distribution energy flexibility with power systems operation. IEEE Transactions on Smart Grid 10, 1 (2018), 1101–1110
2018
-
[72]
Weici Pan, Guanya Shi, Yiheng Lin, and Adam Wierman. 2022. Online Opti- mization with Feedback Delay and Nonlinear Switching Cost. Proc. ACM Meas. Anal. Comput. Syst. 6, 1, Article 17 (Feb 2022), 34 pages. https://doi.org/10.1145/ 3508037
2022
-
[73]
Santiago Paternain, Luiz Chamon, Miguel Calvo-Fullana, and Alejandro Ribeiro
-
[74]
Marios M Polycarpou and Petros A Ioannou. 1993. A robust adaptive nonlinear control design. In 1993 American control conference. IEEE, 1365–1369
1993
-
[75]
Claudia Quintiliani and Enrico Creaco. 2019. Using additional time slots for improving pump control optimization based on trigger levels. Water Resources Management 33 (2019), 3175–3186
2019
-
[76]
Daan Rutten, Nico Christianson, Debankur Mukherjee, and Adam Wierman
-
[77]
Mohammad A Salahuddin, Ala Al-Fuqaha, and Mohsen Guizani. 2016. Reinforce- ment learning for resource provisioning in the vehicular cloud. IEEE Wireless Communications 23, 4 (2016), 128–135
2016
-
[78]
Guanya Shi. 2021. Competitive Control via Online Optimization with Memory, Delayed Feedback, and Inexact Predictions. In 2021 55th Annual Conference on Information Sciences and Systems (CISS)
2021
-
[79]
Advances in Neural Information Processing Systems 32 (2019)
Constrained reinforcement learning has zero duality gap. Advances in Neural Information Processing Systems 32 (2019)
2019
-
[80]
Jerome Sieber, Samir Bennani, and Melanie N Zeilinger. 2021. A system level approach to tube-based model predictive control. IEEE Control Systems Letters 6 (2021), 776–781
2021
-
[81]
Manish K Singh and Vassilis Kekatos. 2019. Optimal scheduling of water dis- tribution systems. IEEE Transactions on Control of Network Systems 7, 2 (2019), 711–723
2019
-
[82]
Aivar Sootla, Alexander I Cowen-Rivers, Taher Jafferjee, Ziyan Wang, David H Mguni, Jun Wang, and Haitham Ammar. 2022. Sauté rl: Almost surely safe reinforcement learning using state augmentation. In International Conference on Machine Learning. PMLR, 20423–20443
2022
-
[83]
arXiv preprint arXiv:2202.03519 (2022)
Online Optimization with Untrusted Predictions. arXiv preprint arXiv:2202.03519 (2022)
2022 arXiv
-
[84]
Anna Stuhlmacher and Johanna L Mathieu. 2020. Chance-constrained wa- ter pumping to manage water and power demand uncertainty in distribution networks. Proc. IEEE 108, 9 (2020), 1640–1655
2020
-
[85]
Anna Stuhlmacher and Johanna L Mathieu. 2020. Water distribution networks as flexible loads: A chance-constrained programming approach. Electric Power Systems Research 188 (2020), 106570
2020
-
[86]
Guanya Shi, Yiheng Lin, Soon-Jo Chung, Yisong Yue, and Adam Wierman. 2020. Online optimization with memory and competitive control. Advances in Neural Information Processing Systems 33 (2020), 20636–20647
2020
-
[87]
Wei Sun and Chenchen Huang. 2022. Predictions of carbon emission intensity based on factor analysis and an improved extreme learning machine from the perspective of carbon emission efficiency. Journal of Cleaner Production 338 (2022), 130414
2022
-
[88]
Wentao Tang and Prodromos Daoutidis. 2022. Data-driven control: Overview and perspectives. In 2022 American Control Conference (ACC). IEEE, 1048–1064
2022
-
[89]
Andrew Taylor, Andrew Singletary, Yisong Yue, and Aaron Ames. 2020. Learn- ing for safety-critical control with control barrier functions. In Learning for Dynamics and Control. PMLR, 708–717. Learning-Augmented Online Control for Decarbonizing Water Infrastructures Conference ...
2020
-
[90]
Pantelis Sopasakis, Ajay K Sampathirao, Alberto Bemporad, and Panagiotis Patrinos. 2018. Uncertainty-aware demand management of water distribution networks in deregulated energy markets. Environmental modelling & software 101 (2018), 10–22
2018
-
[91]
Michael Volk. 2013. Pump characteristics and applications . CRC Press
2013
-
[92]
Kim P Wabersich, Lukas Hewing, Andrea Carron, and Melanie N Zeilinger. 2021. Probabilistic model predictive safety certification for learning-based control. IEEE Trans. Automat. Control 67, 1 (2021), 176–188
2021
-
[93]
Chenxi Sun, Tongxin Li, and Xiaoying Tang. 2021. Data-driven Electric Vehicle Charging Station Placement for Incentivizing Potential Demand. In2021 IEEE In- ternational Conference on Communications, Control, and Computing Technologies for Smart Grids (SmartGridComm) . IEEE, 27–32
2021
-
[94]
Evan M Wanjiru, Lijun Zhang, and Xiaohua Xia. 2016. Model predictive control strategy of energy-water management in urban households. Applied Energy 179 (2016), 821–831
2016
-
[95]
Honghao Wei, Xin Liu, and Lei Ying. 2021. A provably-efficient model-free algo- rithm for constrained markov decision processes.arXiv preprint arXiv:2106.01577 (2021)
2021 arXiv
-
[96]
William Wong, Praneet Dutta, Octavian Voicu, Yuri Chervonyi, Cosmin Padu- raru, and Jerry Luo. 2022. Optimizing industrial hvac systems with hierarchical reinforcement learning. arXiv preprint arXiv:2209.08112 (2022)
2022 arXiv
-
[97]
Vestal Tutterow and Aimee T McKane. 2004. Variable speed pumping: A guide to successful applications. US DOE. Lawrence Berkeley National Laboratory. https://escholarship.org/uc/item/4691d71q
2004
-
[98]
Tsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, and Peter J. Ramadge
-
[99]
Yunchang Yang, Tianhao Wu, Han Zhong, Evrard Garcelon, Matteo Pirotta, Alessandro Lazaric, Liwei Wang, and Simon Shaolei Du. 2022. A Reduction- Based Framework for Conservative Bandits and Reinforcement Learning. In International Conference on Learning Representations . https:...
2022
-
[100]
Ye Wang, Kevin Too Yok, Wenyan Wu, Angus R Simpson, Erik Weyer, and Chris Manzie. 2021. Minimizing pumping energy cost in real-time operations of water distribution systems using economic model predictive control. Journal of Water Resources Planning and Management 147, 7 (2021...
2021
-
[101]
Chenkai Yu, Guanya Shi, Soon-Jo Chung, Yisong Yue, and Adam Wierman. 2022. Competitive control with delayed imperfect information. In 2022 American Control Conference (ACC). IEEE, 2604–2610
2022
-
[102]
Krzysztof Zarzycki and Maciej Ławryńczuk. 2024. LSTM for Modelling and Predictive Control of Multivariable Processes. In International Conference on Innovative Techniques and Applications of Artificial Intelligence. Springer, 74–87
2024
-
[103]
𝐻∑︁ ℎ=0 𝑐ℎ 𝑥ℎ,𝑚( ˜𝜋(𝑠ℎ),U𝜆,ℎ) −𝑐ℎ( ˜𝑥ℎ, ˜𝜋( ˜𝑠ℎ)) # . (41) We can bound this difference as E𝑦0:𝐻 𝐽𝜋𝜆 𝐻 (𝑦0:𝐻) − E𝑦0:𝐻 h 𝐽 ˜𝜋 𝐻(𝑦0:𝐻) i =E𝑦0:𝐻
Runyu Zhang, Yingying Li, and Na Li. 2021. On the regret analysis of online LQR control with predictions. In 2021 American Control Conference (ACC). IEEE, 697–703. Conference acronym ’XX, June 17 - 20, 2025, Rotterdam, Netherlands Yang et al. Appendix A Additional Numerical Re...
2021
-
[104]
Jianyi Yang and Shaolei Ren. 2022. Learning-Assisted Algorithm Unrolling for Online Optimization with Budget Constraints. AAAI (2022)
2022
-
[106]
InInternational Confer- ence on Learning Representations
Projection-Based Constrained Policy Optimization. InInternational Confer- ence on Learning Representations . https://openreview.net/forum?id=rke3TJrtPS
-
[108]
Chenkai Yu, Guanya Shi, Soon-Jo Chung, Yisong Yue, and Adam Wierman. 2020. The power of predictions in online control. Advances in Neural Information Processing Systems 33 (2020), 1994–2004
2020
-
[112]
is the𝜖−covering number of the competitive policy space Π𝜆 with𝐿1−norm as the distance measure: the distance of two functions 𝜋 and 𝜋′ is∥𝜋− 𝜋′∥ ˆ𝐿𝑛 1 = 1 𝑛 Í𝑛 𝑡 =1∥𝜋(𝑠(𝑡))− 𝜋′(𝑠(𝑡))∥ 1. By Eqn. (10), we have 1 𝑛 Í𝑛 𝑡 =1𝐽 𝜋(𝑛) 𝜆 𝐻 (𝑦(𝑡) 0:𝐻)≤ 1 𝑛 Í𝑛 𝑡 =1𝐽 𝜋∗ 𝜆 𝐻 (𝑦(𝑡) 0:𝐻). Th...
-
[2006]
IEEE transactions on automatic control 51, 7 (2006), 1200–1207
Robust nonlinear control associating robust feedback linearization and H/sub/spl infin//control. IEEE transactions on automatic control 51, 7 (2006), 1200–1207
2006
-
[2019]
Online Optimization in Cloud Resource Provisioning: Predictions, Regrets, and Algorithms. Proc. ACM Meas. Anal. Comput. Syst. 3, 1, Article 16 (March 2019), 30 pages. https://doi.org/10.1145/3322205.3311087
2019
-
[2020]
Annual Review of Control, Robotics, and Autonomous Systems3, 1 (2020), 269–296
Learning-based model predictive control: Toward safe learning in control. Annual Review of Control, Robotics, and Autonomous Systems3, 1 (2020), 269–296
2020
-
[2021]
IEEE Transactions on Smart Grid 12, 6 (2021), 4897–4913
Learning-based Predictive Control via Real-time Aggregate Flexibility. IEEE Transactions on Smart Grid 12, 6 (2021), 4897–4913
2021
-
[2022]
arXiv preprint arXiv:2211.07357 (2022)
Controlling Commercial Cooling Systems Using Reinforcement Learning. arXiv preprint arXiv:2211.07357 (2022)
2022 arXiv
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.