Pith. sign in

REVIEW 1 cited by

Model-predictive control and reinforcement learning in multi-energy system case studies

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.09785 v2 pith:MNMYYRL7 submitted 2021-04-20 eess.SY cs.AIcs.LGcs.SYmath.OC

classification eess.SYcs.AIcs.LGcs.SYmath.OC
keywords lmpcsystemcontrolmulti-energycaselearningoptimalrealistic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Model-predictive-control (MPC) offers an optimal control technique to establish and ensure that the total operation cost of multi-energy systems remains at a minimum while fulfilling all system constraints. However, this method presumes an adequate model of the underlying system dynamics, which is prone to modelling errors and is not necessarily adaptive. This has an associated initial and ongoing project-specific engineering cost. In this paper, we present an on- and off-policy multi-objective reinforcement learning (RL) approach, that does not assume a model a priori, benchmarking this against a linear MPC (LMPC - to reflect current practice, though non-linear MPC performs better) - both derived from the general optimal control problem, highlighting their differences and similarities. In a simple multi-energy system (MES) configuration case study, we show that a twin delayed deep deterministic policy gradient (TD3) RL agent offers potential to match and outperform the perfect foresight LMPC benchmark (101.5%). This while the realistic LMPC, i.e. imperfect predictions, only achieves 98%. While in a more complex MES system configuration, the RL agent's performance is generally lower (94.6%), yet still better than the realistic LMPC (88.9%). In both case studies, the RL agents outperformed the realistic LMPC after a training period of 2 years using quarterly interactions with the environment. We conclude that reinforcement learning is a viable optimal control technique for multi-energy systems given adequate constraint handling and pre-training, to avoid unsafe interactions and long training periods, as is proposed in fundamental future work.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Generalising Battery Control in Net-Zero Buildings via Personalised Federated RL

    cs.LG 2024-12 reject novelty 4.0 of 10

    In a simplified net-zero microgrid, untuned federated TRPO learns useful battery policies, but tuned PPO gets much closer to the known optimal policy; personal encoding and feature grouping sometimes shrink the gap.

Pith tools