Pith. sign in

REVIEW 1 cited by

Data-Driven Inverse Optimal Control for Continuous-Time Nonlinear Systems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.09090 v2 pith:WL4VOCWR submitted 2025-03-12 eess.SY cs.SY

classification eess.SYcs.SY
keywords model-freecontrolsystemsalgorithmalgorithmsinverseoptimalcontinuous-time
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper introduces a novel model-free and a partially model-free algorithm for inverse optimal control (IOC), also known as inverse reinforcement learning (IRL), aimed at estimating the cost function of continuous-time nonlinear deterministic systems. Using the input-state trajectories of an expert agent, the proposed algorithms separately utilize control policy information and the Hamilton-Jacobi-Bellman equation to estimate different sets of cost function parameters. This approach allows the algorithms to achieve broader applicability while maintaining a model-free framework. Also, the model-free algorithm reduces complexity compared to existing methods, as it requires solving a forward optimal control problem only once during initialization. Furthermore, in our partially model-free algorithm, this step can be bypassed entirely for systems with known input dynamics. Simulation results demonstrate the effectiveness and efficiency of our algorithms, highlighting their potential for real-world deployment in autonomous systems and robotics.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Online Learning Control Strategies for Industrial Processes with Application for Loosening and Conditioning

    eess.SY 2025-06 reject novelty 4.0 of 10

    An adaptive Koopman MPC with historical safety constraints is proposed for tobacco conditioning, but the claimed Cpk improvements come from model-generated advisor-mode trajectories, not real closed-loop control.

Pith tools