Pith. sign in

REVIEW 4 major objections 5 minor 28 references

Offline and Online Nonlinear Inverse Differential Games with Known and Approximated Cost and Value Function Structures

T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Inverse differential games can be solved by computing the full set of equivalent cost parameters, with online convergence to one element.

desk verdict Real progress on nonlinear inverse differential games, but the solution-set claim rests on an uncharacterized condition and a finite-sample assumption; referee it with a demanded revision. read the letter →

arxiv 2411.10297 v1 pith:WXQ44RIG submitted 2024-11-15 math.OC

classification math.OC MSC 49N7049N4591A2393B3093C10
keywords inversedifferentialgamesoptimalcontrolHamilton-Jacobi-BellmanequationsfeedbackNashequilibriumnon-uniquenesssolutionsetonlinelearningapproximationerrornonlinearsystems
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper aims to solve the inverse differential game problem for nonlinear multi-player systems: from observed state and control trajectories that constitute a feedback Nash equilibrium, recover the cost functions of all players. Its central claim is that an offline HJB-based method computes the full set of equivalent cost function parameters that reproduce the observed trajectories, and that an online gradient-descent counterpart is guaranteed to converge to one element of that set. A separate approximation analysis shows that errors in value-function approximation produce bounded trajectory errors, whereas errors in cost-function approximation generally produce unbounded or unquantifiable errors unless cost and value structures are aligned so the coupled Hamilton-Jacobi-Bellman equations can be fulfilled. A two-player numerical example illustrates the results under both known and violated structural assumptions.

What carries the argument

The central machinery is the reformulation of the coupled HJB equations into N decoupled linear-in-parameters equations after replacing the true Nash strategies with identified ones. Solving these equations as least-squares problems gives parameter sets parametrized by the null-space vectors of the stacked regression matrices, displayed as equations (25) and (26). The online method uses the same regression structure in gradient-descent laws, with persistent excitation assumptions guaranteeing exponential convergence. The approximation analysis uses the Stone-Weierstrass theorem to represent value and cost functions as basis expansions plus bounded residuals, and compares how the two types of residuals propagate.

What would settle it

Compute the solution set for a nonlinear game with two distinct free-vector choices that both satisfy the paper's conditions, then solve the forward differential game for each parameter pair and compare trajectories to the observed ones; if any parameter inside the set produces a trajectory that deviates from the observations, the claim that the set contains all equivalent parameters fails. Alternatively, construct a game with value-function approximation error where no parameterized cost function can realize the identified strategy, which would show that Assumption 7 and Lemma 4 do not apply.

Watch

Extended reading notes

Core claim

The core discovery is that the coupled HJB equations, which characterize a feedback Nash equilibrium, can be decoupled after first identifying the equilibrium strategies from data, and then solved as convex quadratic programs whose null spaces describe all equivalent cost parameters. When the value functions are known up to parameters, the offline method returns the complete solution set of equivalent parameters; the online method, under persistent excitation, converges to one element of this set. When value-function structures are approximated, the identified strategies stay within a bounded error of the true Nash strategies, but this only transfers to cost parameters if an additional existence assumption holds. When cost functions are approximated, the residual approximation error biases the parameter estimates and the resulting Nash equilibrium generally differs from the observed one, with no guaranteed bound.

Load-bearing premise

The load-bearing premise is that for every parameter in the computed sets one can choose the free null-space vector so that the decoupled optimal control problem has a unique solution, and that cost parameters exist making approximated value functions exact; neither choice is given constructively.

Editorial extensions

If this is right

  • For a given dataset, all cost parameters consistent with the data can be enumerated or sampled from the solution set, making non-uniqueness explicit rather than hidden by a scaling assumption.
  • The online method is the first nonlinear multi-player inverse differential game algorithm with a convergence guarantee to the offline solution set, enabling real-time identification from streaming trajectories.
  • Value-function approximation errors lead to bounded errors in the identified Nash strategies and, under an additional existence assumption, in the final trajectories.
  • Cost-function approximation errors can bias parameters and break trajectory matching, so the paper's alignment condition on cost and value basis functions is necessary for bounded final errors.
  • When the structural assumptions hold, the estimated cost parameters produce a forward Nash equilibrium whose trajectories match the observed ones, as demonstrated by the numerical example.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper leaves open the problem of constructively characterizing the free null-space vectors that make the decoupled optimal control problem well-posed; a practical extension would be to project the solution set onto parameters satisfying positive definiteness or other regularity constraints.
  • The asymmetry between value and cost approximation suggests a design rule: choose value-function basis functions first, then construct cost basis functions via a converse HJB argument so the coupled equations are exactly solvable; the paper hints at this in Remark 2 but does not turn it into an algorithm.
  • Because the online estimator converges to an arbitrary element of the solution set, the converged parameters are not individually identifiable; set-valued tracking or additional selection criteria would be needed for interpretations in applications like human-robot interaction.
  • A concrete testable extension for linear-quadratic games would be to check whether the implicit free-vector condition reduces to positive-definiteness constraints, which would give a fully constructive version of Theorem 1 in that setting.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes offline and online inverse differential game (IDG) methods for nonlinear differential games, aiming to recover all cost-function parameters that reproduce given ground-truth trajectories of a feedback Nash equilibrium. The offline method (Section III) uses the observed trajectories to first identify the Nash equilibrium strategies and parts of the value functions (Lemmas 1 and 2), then reformulates the coupled Hamilton-Jacobi-Bellman equations as decoupled linear equations, and solves quadratic programs to obtain affine solution sets (Theorem 1, Eqs. (25) and (26)). The online method (Section IV) uses gradient-descent updates for the same parameters and proves convergence to a fixed element of the offline solution set (Theorem 2). The paper also analyzes the effect of approximating the value functions (bounded strategy error) and the cost functions (in general, unbounded or biased parameter estimates). A two-player numerical example illustrates the offline and online methods and the approximation-error cases.

Significance. If the central equivalence claim of Theorem 1 is established, the paper would make a substantial contribution: it would provide a formal treatment of non-uniqueness in nonlinear multi-player inverse differential games beyond the known scaling ambiguity, give the first online nonlinear IDG method with a convergence guarantee to the offline solution set, and provide a separate approximation-error analysis for value and cost functions that is genuinely informative for practitioners. The paper's strengths include the clean derivation of the decoupled HJB reformulation under the stated parametric assumptions, the explicit handling of per-player scaling non-uniqueness in Lemma 2, the machine-checkable linear-algebra formulation of the solution sets, and an honest negative result for cost-function approximation (Lemma 5). The numerical example is reproducible and demonstrates the proposed algorithms on a nontrivial two-player nonlinear example.

major comments (4)
  1. [Theorem 1, Eqs. (25)-(26)] The central claim that every parameter in the sets (25) and (26) solves Problem 1 is not fully proven. Equations (21) and (22) are required to hold for every x in X, but the QP (28) only enforces them at the K-bar sampled states x-bar_k. Assumption 6 ensures that further data points do not increase the column rank of the sampled regressor, but it does not imply that the null space of the finite-sample matrix equals the null space of the operator on all of X. A parameter satisfying the finite-grid equations can therefore have nonzero HJB residual at unsampled states and does not necessarily define the same optimal-control solution. To support the claim, the authors need either a genericity or analyticity condition that forces equality of the solution sets, or a revised statement that the computed sets are only the finite-data versions and the global equivalence holds under an additional verifiable condition.
  2. [Theorem 1, w-bar conditions] The restriction on the free vectors w-bar_i (or w-bar_i^(r)) in (25) and (26) is non-constructive: the condition that the decoupled HJB equation be 'necessary and sufficient for a unique OC solution' is not characterized by any checkable condition, such as positive definiteness of R-hat_ii and Q-hat_i or closed-loop stability of the resulting feedback law. Without such a characterization, the algorithm cannot actually compute the claimed 'set of all equivalent cost function parameters', because one cannot decide which affine parameters belong to the set. The numerical example implicitly uses positive definiteness of Q_i to define the set, but this criterion is not connected to the theorem; the claim should be either proved or made conditional on an explicit, verifiable condition.
  3. [Lemma 4 and Assumption 7] Assumption 7 presupposes exactly the alignment that the paper argues is necessary: it assumes that there exist cost parameters R-tilde_ij and beta-tilde_i such that the approximated value functions theta*_i^T phi_i(x) are the exact value functions of a differential game with those cost functions. Lemma 4 then establishes only that the parameters from (25)/(26) yield a FNE equal to the identified control laws mu-tilde*_i, under this assumption. This is a conditional result, and Section V-B shows that Assumption 10 (the online analogue) can fail. The paper should state in the main text, not only in the numerical example, that the approximation-error bound for the offline method rests on an unverifiable alignment assumption and that when the assumption fails the HJB-based identification may not produce a FNE matching the identified strategies.
  4. [Theorem 2, proof after Eq. (41)] The proof of Theorem 2 introduces an additional condition that is not listed among the theorem's assumptions: it is stated that 'we can assume that the GT trajectory x*(t0 to infinity) is excited such that Assumption 6 is fulfilled as well'. Assumption 6 is defined for the offline grid (23)-(24) and is not a standing assumption on the online trajectory or the probing signal. Furthermore, the proof uses m_HJB(t) evaluated along the probing signal x_HJB(t) of Remark 3, but the theorem does not state that this signal stays inside X or that the continuous-time rank condition holds. These points need to be clarified and made part of the hypotheses, or the proof needs to be revised to derive them from Assumptions 8 and 9.
minor comments (5)
  1. [Lemma 3, Eq. (29)] The bound in (29) uses the norm of the gradient of the approximation error function epsilon-bar_i, but the Stone-Weierstrass theorem only provides boundedness of the function itself. If X is not compact, boundedness of the gradient does not follow; the assumption should state that X is compact or that the derivative of epsilon-bar_i is bounded.
  2. [Section IV, stopping criterion (35)] The stopping criterion (35) depends on a time window T and a threshold that are not specified; it would be helpful to state how these are chosen in practice and whether the stopping time affects any of the theoretical guarantees, since Lemmas 6-8 discuss exponential convergence but the algorithm terminates at a finite time.
  3. [Section V-B, NSAE values] The NSAE values in the approximation-error examples (e.g., delta_x ≈ 404.8 and delta_u ≈ 920.3) exceed 1, meaning the normalized error is larger than the maximum absolute value of the signal; while this is consistent with the theory's guarantee of only boundedness, a remark explaining that the theoretical bound is not tight and that these values indicate a large practical error would improve the presentation.
  4. [Assumptions 5 and 6] The 'highest possible column rank' conditions in Assumptions 5 and 6 are phrased in terms of saturation with respect to data points, but no finite procedure to verify them from data is described; for nonlinear basis functions, this condition is not directly checkable and deserves a comment on how a practitioner would confirm it.
  5. [Figure 1 caption] The caption of Figure 1 refers to 'the last two initial state resets' but the time axis begins at 12 s; the initial reset times should be identified to make the comparison between ground-truth and estimated trajectories easy to follow.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the inverse-game parameters are fitted to HJB residuals constructed from externally observed GT trajectories, and the claimed solution-set equivalence rests on an uncharacterized sufficiency condition rather than on a definitional reduction.

full rationale

The derivation chain is not circular. The inputs are the GT trajectories; Lemma 2 fits the feedback strategies and the reduced value-function weights to those external data, and Theorem 1 then computes cost parameters from the HJB residual equations. The numerical validation simulates forward the FNE from the estimated parameters and compares it with the GT trajectories, which is a genuine consistency check. The online method minimizes the same least-squares HJB residual that defines the offline set, so convergence to that set is consistency with the method's own objective, not a hidden refitting of the claimed output. Citation practice is normal: [23] is used only to run a PI solver in the numerical example, and [5] is background prior work; neither carries the theoretical load. The real weakness is a correctness gap, not circularity. The sets (25)/(26) are defined by satisfying (21)/(22) at a finite grid of points, while the proof of Theorem 1 needs the HJB residual to vanish on all of X and needs the free vector w-bar to make the decoupled HJB equation necessary and sufficient for a unique optimal-control solution. These conditions are not characterized, and the finite-grid rank assumption (Assumption 6) does not by itself imply the global HJB equality; so the phrase 'all equivalent cost function parameters that yield the observed trajectories' may overstate what the computed affine nullspace actually contains. That is an unproven inference about the forward problem, not a reduction of the output to the input by definition.

Assumptions & free parameters 4 free parameters · 7 assumptions · 0 invented entities

The central claim rests on standard optimal control and game theory results (HJB sufficiency, Stone-Weierstrass) plus a set of domain assumptions about data richness and the existence of parameterized cost/value structures. The most delicate additions are the implicit 'unique OC solution' condition in Theorem 1 and Assumption 7, which essentially assume the alignment that the paper argues is necessary. The method itself introduces no new physical entities or forces.

free parameters (4)
  • Per-player scaling factor c_i = unknown (depends on ||theta_i^{(r)*}||)
    Introduced in Lemma 2, Eq. (16), through normalization of the identified value-function weights. The inverse problem is solvable only up to this per-player scale, and the solution set (25) is parameterized by it.
  • Arbitrary vectors \bar{w}_i^{(r)} and \bar{w}_i = arbitrary
    These parameterize the non-uniqueness of the HJB least-squares solution in (25) and (26). Any value satisfying the 'unique OC solution' condition yields an equivalent cost parameter set, but that condition is not characterized.
  • Learning rates tau_i and kappa_i = not specified
    Used in the online updates (34) and (37). They are chosen by hand, affect convergence speed, and are not identified by the method. They are implementation parameters rather than scientific outputs.
  • Probing sine frequencies f_i^{(1)}, f_i^{(2)} = random, unspecified
    Used in the numerical example to satisfy Assumption 9. The paper states they are randomly chosen from a uniform distribution between 0.5 Hz and 5 Hz but does not report the actual values, making exact reproduction impossible.
assumptions (7)
  • domain assumption There exists a unique feedback Nash equilibrium for the differential game under study.
    Stated in Section II before Definition 1. The inverse problem presupposes the observed trajectories come from a unique FNE; without uniqueness, the identification is ambiguous.
  • standard math The coupled HJB equations are necessary and sufficient for a feedback Nash equilibrium.
    Invoked in Section II with citations to Basar and Olsder [21, Theorem 6.16] and Anderson and Moore [22]. The entire method reformulates the inverse problem as solving HJB residuals, so this verification theorem is load-bearing.
  • domain assumption Assumptions 1-4: cost and value functions are linear combinations of known basis functions with known system dynamics.
    Stated in Section II. These assumptions define the parameterized structure within which identification occurs; if they fail, approximation errors arise and the theoretical guarantees weaken.
  • domain assumption Assumptions 5 and 6: rank/excitation conditions on the data matrices M_ui and M_HJBi.
    Assumed in Section III-A. The least-squares identification and the HJB parameter sets require the data to be sufficiently rich; otherwise the solution set is not minimal and the results degrade.
  • ad hoc to paper Assumption 7: there exist cost parameters \tilde{R}_ij, \tilde{\beta}_i such that the approximated value functions are the exact value functions of a DG with those costs.
    Introduced in Section III-B. This is a strong converse-HJB assumption that presumes the alignment of cost and value structures, which is the paper's central message. It is needed for Lemma 4 to guarantee bounded trajectory errors.
  • domain assumption Assumptions 8 and 9: persistence of excitation of the regression matrices M_ui(t) and m_HJBi(t).
    Used in Section IV. Standard adaptive-control conditions required for exponential convergence of the online gradient descent updates.
  • standard math Stone-Weierstrass theorem and subalgebra conditions on the basis functions.
    Invoked in Lemma 3 and Lemma 5 to guarantee the existence of bounded approximation errors for value and cost functions on the compact set X.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Offline and Online Nonlinear Inverse Differential Games with Known and Approximated Cost and Value Function Structures." pith.science (2026). https://pith.science/paper/WXQ44RIG

@misc{pith2026241110297,
  author       = {Pith},
  title        = {Pith review of: Offline and Online Nonlinear Inverse Differential Games with Known and Approximated Cost and Value Function Structures},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WXQ44RIG}},
  note         = {Machine review of arXiv:2411.10297}
}
read the original abstract

In this work, we propose novel offline and online Inverse Differential Game (IDG) methods for nonlinear Differential Games (DG), which identify the cost functions of all players from control and state trajectories constituting a feedback Nash equilibrium. The offline approach computes the sets of all equivalent cost function parameters that yield the observed trajectories. Our online method is guaranteed to converge to cost function parameters of the offline calculated sets. For both methods, we additionally analyze the case where the cost and value functions are not given by known parameterized structures and approximation structures, like polynomial basis functions, need to be chosen. Here, we found that for guaranteeing a bounded error between the trajectories resulting from the offline and online IDG solutions and the observed trajectories an appropriate selection of the cost function structures is required. They must be aligned to assumed value function structures such that the coupled Hamilton-Jacobi-Bellman equations can be fulfilled. Finally, the theoretical results and the effectiveness of our new methods are illustrated with a numerical example.

Figures

Figures reproduced from arXiv: 2411.10297 by the authors.

Figure 1
Figure 1. Comparison of GT trajectories (solid lines) and trajec [PITH_FULL_IMAGE:figures/full_fig_p010_1.png] view at source ↗
Figure 2
Figure 2. Parameter trajectories of the online IDG method in case [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 3
Figure 3. Comparison of GT trajectories (solid lines), trajecto [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Comparison of GT trajectories (solid lines) and trajec [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

28 extracted references · 28 canonical work pages

  1. [1]

    When is a linear control system optimal?

    R. E. Kalman, “When is a linear control system optimal?” J. Basic Eng., vol. 86, no. 1, pp. 51–60, 1964

  2. [2]

    From human to humanoid locomotion—an inverse optimal control approach,

    K. Mombaur, A. Truong, and J.-P. Laumond, “From human to humanoid locomotion—an inverse optimal control approach,” Auton. Robots , vol. 28, no. 3, pp. 369–383, 2010

  3. [3]

    Trustworthi- ness of optimality condition violation in inverse dynamic game methods based on the minimum principle,

    P. Karg, A. Kienzle, J. Kaub, B. Varga, and S. Hohmann, “Trustworthi- ness of optimality condition violation in inverse dynamic game methods based on the minimum principle,” 8th IEEE Conf. Control Tech. App. (CCTA), 2024

  4. [4]

    T. L. Molloy, J. Inga, S. Hohmann, and T. Perez, Inverse Optimal Control and Inverse Noncooperative Dynamic Game Theory. A Minimum Principle Approach. Cham: Springer Nature, 2022

  5. [5]

    Solution sets for inverse non-cooperative linear-quadratic differential games,

    J. Inga, E. Bischoff, T. L. Molloy, M. Flad, and S. Hohmann, “Solution sets for inverse non-cooperative linear-quadratic differential games,” IEEE Control Syst. Lett. , vol. 3, no. 4, pp. 871–876, 2019

  6. [6]

    Inverse reinforcement learning for multi-player noncooperative apprentice games,

    B. Lian, W. Xue, F. L. Lewis, and T. Chai, “Inverse reinforcement learning for multi-player noncooperative apprentice games,” Automatica, vol. 145, 2022

  7. [7]

    Reinforcement learning for inverse linear- quadratic dynamic non-cooperative games,

    E. Martirosyan and M. Cao, “Reinforcement learning for inverse linear- quadratic dynamic non-cooperative games,” Syst. Control Lett., vol. 191, 2024

  8. [8]

    Inverse optimal control with polynomial optimization,

    E. Pauwels, D. Henrion, and J.-B. Lasserre, “Inverse optimal control with polynomial optimization,” 53rd IEEE Conf. Decis. Control (CDC) , pp. 5581–5586, 2014

Show all 28 references
  1. [9]

    Linear inverse reinforcement learning in continuous time and space,

    R. Kamalapurkar, “Linear inverse reinforcement learning in continuous time and space,” 2018 Am. Control Conf. (ACC) , pp. 1683–1688, 2018

  2. [10]

    Online observer- based inverse reinforcement learning,

    R. Self, K. Coleman, H. Bai, and R. Kamalapurkar, “Online observer- based inverse reinforcement learning,” IEEE Control Syst. Lett. , vol. 5, no. 6, pp. 1922–1927, 2021

  3. [11]

    Model- based inverse reinforcement learning for deterministic systems,

    R. Self, M. Abudia, S. M. N. Mahmud, and R. Kamalapurkar, “Model- based inverse reinforcement learning for deterministic systems,” Auto- matica, vol. 140, 2022

  4. [12]

    Nonuniqueness and con- vergence to equivalent solutions in observer-based inverse reinforcement learning,

    J. Town, Z. Morrison, and R. Kamalapurkar, “Nonuniqueness and con- vergence to equivalent solutions in observer-based inverse reinforcement learning,” 2023 Am. Control Conf. (ACC) , pp. 3989–3994, 2023

  5. [13]

    Inverse q-learning using input-output data,

    B. Lian, W. Xue, F. L. Lewis, and A. Davoudi, “Inverse q-learning using input-output data,” IEEE Trans. Cybern. , vol. 54, no. 2, 2024

  6. [14]

    Adaptive inverse nonlinear optimal control based on finite-time concurrent learning and semidefinite programming,

    H.-N. Wu and J. Lin, “Adaptive inverse nonlinear optimal control based on finite-time concurrent learning and semidefinite programming,” IEEE Trans. Cybern., vol. 54, no. 10, 2024

  7. [15]

    Online inverse linear-quadratic differential games applied to human behavior identification in shared control,

    J. Inga, A. Creutz, and S. Hohmann, “Online inverse linear-quadratic differential games applied to human behavior identification in shared control,” 2021 European Control Conf. (ECC) , 2021

  8. [16]

    Data-driven inverse reinforcement learning control for linear multiplayer games,

    B. Lian, V . S. Donge, F. L. Lewis, T. Chai, and A. Davoudi, “Data-driven inverse reinforcement learning control for linear multiplayer games,” IEEE Trans. Neural Netw. Learning Syst., vol. 35, no. 2, pp. 2028–2041, 2024

  9. [17]

    Model-based online adaptive inverse nonco- operative linear-quadratic differential games via finite-time concurrent learning,

    J. Lin and H.-N. Wu, “Model-based online adaptive inverse nonco- operative linear-quadratic differential games via finite-time concurrent learning,” IEEE Trans. Artif. Intell., vol. 5, no. 8, pp. 4247–4257, 2024

  10. [18]

    Learning human behavior in shared control: Adaptive inverse differential game approach,

    H.-N. Wu and M. Wang, “Learning human behavior in shared control: Adaptive inverse differential game approach,” IEEE Trans. Cybern. , vol. 54, no. 6, 2024

  11. [19]

    Multi-player non-zero-sum games: Online adaptive learning solution of coupled hamilton-jacobi equations,

    K. G. Vamvoudakis and F. L. Lewis, “Multi-player non-zero-sum games: Online adaptive learning solution of coupled hamilton-jacobi equations,” Automatica, vol. 47, 2011

  12. [20]

    K. S. Narendra and A. M. Annaswamy, Stable Adaptive Systems . Mineola, New York: Dover Publications, Inc., 2005

  13. [21]

    Bas ¸ar and G

    T. Bas ¸ar and G. J. Olsder, Dynamic Noncooperative Game Theory . Philadelphia: SIAM, 1999

  14. [22]

    B. D. O. Anderson and J. B. Moore, Optimal Control: Linear Quadratic Methods. Englewood Cliffs, New Jersey: Prentice-Hall, Inc., 1989

  15. [23]

    Excitation for adaptive optimal control of nonlinear systems in differential games,

    P. Karg, F. K ¨opf, C. A. Braun, and S. Hohmann, “Excitation for adaptive optimal control of nonlinear systems in differential games,” IEEE Trans. Autom. Control, vol. 68, no. 1, pp. 596–603, 2023

  16. [24]

    Online synchronous approximate optimal learning algorithm for multiplayer nonzero-sum games with unknown dynamics,

    D. Liu, H. Li, and D. Wang, “Online synchronous approximate optimal learning algorithm for multiplayer nonzero-sum games with unknown dynamics,” IEEE Trans. Syst. Man. Cybern., Syst. , vol. 44, no. 8, pp. 1015–1027, 2014

  17. [25]

    H. N. Mhaskar and D. V . Pai, Fundamentals of Approximation Theory . Narosa Publishing House, 2000

  18. [26]

    Constrained nonlinear optimal control: A converse hjb approach,

    V . Nevistic and J. A. Primbs, “Constrained nonlinear optimal control: A converse hjb approach,” 1996

  19. [27]

    Convergence properties of adaptive systems and the definition of exponential stability,

    B. M. Jenkins, A. M. Annaswamy, E. Lavretsky, and T. E. Gibson, “Convergence properties of adaptive systems and the definition of exponential stability,” SIAM J. Control Optim., vol. 56, no. 4, pp. 2463– 2484, 2018

  20. [28]

    Online actor-critic algorithm to solve the continuous-time infinite horizon optimal control problem,

    K. G. Vamvoudakis and F. L. Lewis, “Online actor-critic algorithm to solve the continuous-time infinite horizon optimal control problem,” Automatica, vol. 46, 2010. Philipp Karg studied Electrical Engineering and Information Technologies at the Karlsruhe Institute of Technolog...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.