A hierarchical RL-OC method uses inverse optimization to derive structured lower-level policies from demonstrations, claiming superior efficiency and quality over end-to-end RL and existing hierarchical baselines on two control tasks.
arXiv preprint arXiv:2302.03122 , year=
2 Pith papers cite this work. Polarity classification is still indexing.
years
2026 2verdicts
UNVERDICTED 2representative citing papers
A DRL-based heating controller with a real-time adaptive safety filter guarantees flexibility compliance, achieves up to 50% energy savings over rule-based methods, and outperforms plain DRL with only minor comfort violations.
citing papers explorer
-
Hierarchical Decision Making with Structured Policies: A Principled Design via Inverse Optimization
A hierarchical RL-OC method uses inverse optimization to derive structured lower-level policies from demonstrations, claiming superior efficiency and quality over end-to-end RL and existing hierarchical baselines on two control tasks.
-
Safe Deep Reinforcement Learning for Building Heating Control and Demand-side Flexibility
A DRL-based heating controller with a real-time adaptive safety filter guarantees flexibility compliance, achieves up to 50% energy savings over rule-based methods, and outperforms plain DRL with only minor comfort violations.