Pith. sign in

REVIEW 2 cited by

Multi-agent reinforcement learning for the control of three-dimensional Rayleigh-B\'enard convection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.21565 v2 pith:MPVLJX7C submitted 2024-07-31 physics.flu-dyn cs.LG

classification physics.flu-dyncs.LG
keywords controlconvectionmarlmathrmcontrollerenardflowlearned
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Deep reinforcement learning (DRL) has found application in numerous use-cases pertaining to flow control. Multi-agent RL (MARL), a variant of DRL, has shown to be more effective than single-agent RL in controlling flows exhibiting locality and translational invariance. We present, for the first time, an implementation of MARL-based control of three-dimensional Rayleigh-B\'enard convection (RBC). Control is executed by modifying the temperature distribution along the bottom wall divided into multiple control segments, each of which acts as an independent agent. Two regimes of RBC are considered at Rayleigh numbers $\mathrm{Ra}=500$ and $750$. Evaluation of the learned control policy reveals a reduction in convection intensity by $23.5\%$ and $8.7\%$ at $\mathrm{Ra}=500$ and $750$, respectively. The MARL controller converts irregularly shaped convective patterns to regular straight rolls with lower convection that resemble flow in a relatively more stable regime. We draw comparisons with proportional control at both $\mathrm{Ra}$ and show that MARL is able to outperform the proportional controller. The learned control strategy is complex, featuring different non-linear segment-wise actuator delays and actuation magnitudes. We also perform successful evaluations on a larger domain than used for training, demonstrating that the invariant property of MARL allows direct transfer of the learnt policy.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Physics-enhanced reinforcement learning for real-time optimal control of dynamical systems

    cs.LG 2026-07 conditional novelty 6.0 of 10

    PEARL trains a policy by differentiating a short horizon of the simulator and using a learned adjoint network to account for long-term return gradients, beating PPO, TD3, BPTT, and SHAC on two double-gyre navigation tasks.

  2. Latent feedback control of distributed systems in multiple scenarios through deep learning-based reduced order models

    math.OC 2024-12 conditional novelty 6.0 of 10

    A learned reduced-order feedback controller computes near-optimal distributed controls for parametrized PDEs in real time, with a latent loop that works even without online state measurements.

Pith tools