Pith. sign in

REVIEW 1 cited by

Complementary Meta-Reinforcement Learning for Fault-Adaptive Control

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2009.12634 v1 pith:XJ6LAPGR submitted 2020-09-26 cs.LG cs.SYeess.SYstat.ML

classification cs.LGcs.SYeess.SYstat.ML
keywords faultssystemapproachcontrollearningpolicysystemsabrupt
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Faults are endemic to all systems. Adaptive fault-tolerant control maintains degraded performance when faults occur as opposed to unsafe conditions or catastrophic events. In systems with abrupt faults and strict time constraints, it is imperative for control to adapt quickly to system changes to maintain system operations. We present a meta-reinforcement learning approach that quickly adapts its control policy to changing conditions. The approach builds upon model-agnostic meta learning (MAML). The controller maintains a complement of prior policies learned under system faults. This "library" is evaluated on a system after a new fault to initialize the new policy. This contrasts with MAML, where the controller derives intermediate policies anew, sampled from a distribution of similar systems, to initialize a new policy. Our approach improves sample efficiency of the reinforcement learning process. We evaluate our approach on an aircraft fuel transfer system under abrupt faults.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Study of the Efficacy of Generative Flow Networks for Robotics and Machine Fault-Adaptation

    cs.RO 2025-01 conditional novelty 5.0 of 10

    On a simulated Reacher arm with four injected faults, CFlowNets matches or beats DDPG, TD3, PPO, and SAC on adaptation speed and asymptotic reward, while using far more GPU memory and wall-clock time.

Pith tools