Pith. sign in

REVIEW 2 major objections 5 minor 62 references

Entropy-regularized equilibria become exponentially sensitive to model error when the control graph has a positive feedback cycle, changing the statistical resolution boundary from a power law to order 1/log n.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 15:56 UTC pith:B2KOH3NF

load-bearing objection A serious, honestly hedged theory paper: real new results on temperature-explicit susceptibility for EEHJB, with the main soft spot being that exponential amplification is proven for exact models plus a conditional abstract theorem, not for general state-dependent EEHJB. the 2 major comments →

arxiv 2607.18128 v1 pith:B2KOH3NF submitted 2026-07-20 math.OC cs.NAmath.NAmath.PR

Feedback Cycles in Exploratory Equilibria

classification math.OC cs.NAmath.NAmath.PR MSC 93E2049L2068T0560H3065M1262M05
keywords time-inconsistent stochastic controlentropy regularizationexploratory equilibrium HJB equationVolterra operatorMittag-Leffler functionfeedback cyclesstatistical conditioningLambert W function
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper asks how errors in learned rewards and dynamics propagate through entropy-regularized equilibria of time-inconsistent stochastic control, where each temporal self optimizes against the future policies of later selves. It establishes that the derivative of the equilibrium policy with respect to model parameters is governed by a backward Volterra resolvent whose low-temperature growth is decided by the causal graph of the system: on a directed acyclic graph the sensitivity is a power of 1/τ, while a single positive feedback cycle can produce exponential growth e^{C/τ}. If true, cooling an entropy-regularized system with feedback cycles is dramatically more fragile than the local Gibbs factor 1/τ would suggest, and the statistical boundary for root-n estimation drops from a power law to order 1/log n. The paper also proves that at fixed temperature the equilibrium branch is twice differentiable in finite-dimensional model parameters, yielding a function-valued delta method, and it constructs explicit models—an affine two-action model and a bounded uniformly elliptic diffusion—that realize both regimes exactly. The same machinery shows that along a cyclic Perron mode, right-endpoint time discretization is relatively consistent exactly when N τ^2 → ∞.

Core claim

The paper claims that the derivative of an entropy-regularized equilibrium policy with respect to model parameters is governed by a backward Volterra resolvent u = τ^{-1}S(c + Ku), and that its low-temperature size is set by the causal graph of the system. On a directed acyclic graph the response is exactly polynomial, Θ(τ^{-(L+1)}); a positive feedback cycle produces an exponential factor E_β(g(T-t)^β/τ^q) under an aligned cone condition. In the bounded uniformly elliptic model the susceptibility matrix is χ_τ(t) = (v/τ) exp{(-νI + vK/τ)(T-t)}, and the paper shows that closing one positive cycle changes the root-n linear-response boundary from a power law to order 1/log n, with an exact Lam

What carries the argument

The load-bearing object is the causal Volterra resolvent: the policy tangent solves u = τ^{-1}S(c + Ku), where S is the centered Gibbs covariance operator and K is the future-policy-to-current-score derivative. The influence kernel is measured by a fractional integral I^α_{T-} of order α ∈ (0,2], so the resolvent series is bounded by the Mittag-Leffler function E_α(κ(T-t)^α/τ). In block form, powers of A = SK count directed walks: an acyclic graph satisfies A^{L+1}=0, terminating the series after the longest path, while a positive cycle contributes the factor E_β(g(T-t)^β/τ^q). This same machinery powers the fixed-temperature delta method.

Load-bearing premise

The exponential amplification claim rests on an assumed 'aligned positive mode'—a cone preserved by the feedback operator with a temperature-independent test path—which the paper verifies only for special bounded-diffusion models, so without that condition negative feedback could damp the response and no exponential growth need occur.

What would settle it

In the exact affine model with β=T=1 and ξ_n = Z/√n, integrate (5.5)–(5.7) at the Lambert-W scale τ*_n = βT/W(βT√n) for n ≥ 10^8; Theorem 5.2 predicts m(0) has a nondegenerate random limit with mean about 0.52, while m(t) for fixed t>0 converges to 0. A distribution converging to a point mass, or a nonvanishing later-time response, would refute it.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • On a directed acyclic graph of longest path L, the equilibrium susceptibility is Θ(τ^{-(L+1)}); along a positive cycle it is at least of exponential order, so closing one edge can change the small-temperature scaling.
  • The root-n statistical resolution boundary for learned-model noise is of order (log n)^{-α} in general and order 1/log n when a positive cycle with α=1 is present.
  • At fixed τ>0, the local equilibrium branch is twice differentiable in finite-dimensional model parameters, so function-valued delta-method inference for equilibrium policies is justified.
  • In the bounded diffusion model, right-endpoint time discretization along a cyclic Perron mode is relatively consistent exactly when Nτ^2→∞; for a DAG, N→∞ suffices without coupling to τ.
  • The affine model has an exact phase transition at τ*_n = βT/W(βT√n) ~ 2βT/log n, at which the initial policy has a nondegenerate random limit while every fixed later-time policy converges to the reference mixture.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the exponential amplification holds generally, then entropy-annealing in cyclic control environments will not achieve root-n accuracy: n^{-1/2} model errors, harmless on a DAG, leave order-one policy errors on a cycle. A practical corollary is to estimate the causal return structure before cooling.
  • The Lambert-W scale implies that for a cyclic system, achieving a fixed statistical precision at temperature τ requires a sample size exponential in βT/τ—a quantitative prediction testable in tabular or linear-quadratic experiments.
  • The lower bound's dependence on the invariant-cone condition (assumed, not derived) suggests that negative feedback can mask cycle amplification; characterising the largest class of state-dependent models in which the exponential rate is actually attained remains an open problem, and the paper's e^{C/τ^2} general bound hints the true rate may be even worse in some models.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This paper studies the sensitivity of entropy-regularized equilibria in time-inconsistent stochastic control to perturbations of the model parameters, with emphasis on low-temperature amplification. The core objects are a backward Volterra resolvent for the derivative of the equilibrium policy and a graph-theoretic decomposition of the causal influence operator. Under an abstract fractional-causal-influence assumption (Assumption 3.2), the paper proves an upper bound of Mittag-Leffler type; under additional cone/order hypotheses (3.12)-(3.14) it proves a matching lower bound. The block decomposition yields a dichotomy: on acyclic influence graphs the susceptibility is polynomial in 1/tau, whereas a positive cycle can produce exponential growth. The paper then proves fixed-temperature C^2 differentiability of the EEHJB equilibrium branch and a function-valued delta method, with an e^{C/tau^2} general stability certificate. Two explicit models - an affine two-action model and a bounded trigonometric diffusion - realize the rates, including a Lambert-W critical temperature for root-n reward noise and a discrete-time mesh-stiffness threshold N tau^2 -> infinity along a cyclic Perron mode. The numerical section reports reproducible experiments with deterministic verification.

Significance. If the results hold, they identify a new and concrete mechanism: the topology of causal feedback cycles, not just the 1/tau Gibbs factor, controls the conditioning of exploratory equilibria at low temperature. The path-cycle distinction is sharp and quantitatively explicit (Theta(tau^{-(L+1)}) vs e^{C/tau}), with a matching nonlinear selection phenomenon in a bounded uniformly elliptic model. The paper is unusually careful: the abstract theorems are stated with explicit hypotheses, the exact models are solved in closed form, the proofs are self-contained relative to stated assumptions, and the numerical code is archived and deterministic, with no fitted constants. The main limitation is that the exponential lower bound is conditional on cone conditions verified only for the exact models; this does not invalidate the exact-model results but restricts the generality of the advertised cycle-induced amplification.

major comments (2)
  1. [§3, Thm 3.4–Cor 3.7; §4, Assump. 4.1–4.2; §5.2] The exponential lower bound (3.15) and the resolution boundary (Cor 3.5, Cor 3.7) are conditional on the cone/order conditions (3.12)–(3.14). These are assumed, not derived from the EEHJB hypotheses of Section 4; the only verification is for the constant-coefficient bounded-diffusion model in §5.2 (Perron cone with ψ(t)=e^{νt}r). For the state-dependent parabolic setting of §4 the natural causal order is α=1/2 from (4.11), and no ψ, cone, or κ_- is constructed. The scalar example following Thm 3.3 (A=-κ∫) shows that without the sign condition the response is damped, so the cone condition is essential. Consequently the paper establishes the e^{C/τ} amplification and the 1/log n boundary for an explicit model class plus a conditional abstract statement, not for general state-dependent EEHJB. Please either verify (3.12)–(3.14) for a nontrivial state-dependent class, or state this limitation
  2. [Abstract; §7] The abstract's sentence 'Closing one positive cycle changes the root-n linear-response boundary from a power law to order 1/log n' is not qualified by the aligned-mode/cone assumptions. Since §7 itself leaves open whether a state-dependent model attains even the e^{C/τ} lower rate (the general parabolic upper bound is e^{C/τ²}), the unqualified phrasing overstates the proven scope. Please amend the abstract and the analogous introduction sentence so that the exact models and the conditional abstract theorem are presented as the proven statements.
minor comments (5)
  1. [§3, Eq. (3.18)] The 'asymptotic solution' of the canonical balance contains an extra log α in the second-order term. Since only the {1+o(1)} form is claimed, please label it as a leading-order asymptotic solution rather than the exact solution of the displayed equation.
  2. [§5.1, Eq. (5.3)] The actionwise reward r(y,s,x,a) is not written explicitly; the integrand uses the policy-averaged m^π. State explicitly that r(y,s,x,a)=β(x−y)a, so that the averaged term in (5.3) is the mean reward under π.
  3. [§6, Prop. 6.1] In the DAG part, the assumption N≥L appears only in the text after the proposition. Move it into the statement for clarity.
  4. [Title page] 'EXPLORA TOR Y' contains a stray space in the title; please check the typesetting.
  5. [§6.1 (numerics)] Specify that the 16.2% and 1.32% errors at n=1014 refer to τ_{n,1} and τ_{n,2} in (5.13).

Circularity Check

0 steps flagged

No significant circularity: the central bounds, graph dichotomy, and statistical scales are derived from explicit assumptions or exact models, and the numerical claims are not fitted to the theory.

full rationale

The derivation chain is self-contained relative to its stated hypotheses. The abstract susceptibility estimates (Thms 3.3–3.4) follow directly from Assumption 3.2 and the explicitly stated aligned-cone conditions (3.12)–(3.14); these are hypotheses, not conclusions forced by the claimed result. The paper does not present the general parabolic upper bound e^{C/τ^2} as a sharp prediction, and it openly states that whether state-dependent models attain the larger rate is open. The exponential amplification and the 1/log n boundary for the bounded-diffusion model come from the exact susceptibility formula (5.19) and its spectral analysis in Cor 5.4, not from fitting any constant. The affine model's Lambert-W critical temperature is obtained by solving the exact equation a_n=1 in Theorem 5.2. The numerical section explicitly states that no constant in a theoretical curve is fitted to the output. The dependence on [14] is background well-posedness for EEHJB systems, and the paper constructs its local equilibrium branch and influence equations directly rather than importing a uniqueness theorem from its own prior work. The skeptical concern that cone conditions (3.12)–(3.14) are verified only for the constant-coefficient bounded-diffusion model is a scope limitation about generality, not a circular reduction: the theorem that uses them states them as assumptions, and the paper does not claim those conditions are consequences of the general well-posedness hypotheses. No fitted-input-called-prediction, self-citation load-bearing, ansatz-smuggling, or renaming pattern is present.

Axiom & Free-Parameter Ledger

0 free parameters · 10 axioms · 0 invented entities

The analysis introduces no fitted constants; model parameters in the exact models (β, σ, K, v, T) are fixed inputs, and the abstract constants (κ, α, C_c) are hypotheses of the bounds, not fit values. The main extra assumptions are the fractional causal influence bound (Assumption 3.2), the uniform comparison envelope (Assumption 4.2), and the cone/order conditions (3.12)–(3.14) that produce the exponential lower bounds. The latter are verified only in the explicit bounded diffusion model, not for a general state-dependent EEHJB.

axioms (10)
  • domain assumption Assumption 3.2: ∃α∈(0,2], κ≥0, C_c such that ∥S_t K v(t)∥ ≤ κ (I^α_{T−}∥v∥)(t) and ∥S_t c[h](t)∥ ≤ C_c ∥h∥.
    Load-bearing for all abstract susceptibility bounds (Thms 3.3–3.6). The paper asserts α=1/2 from the parabolic D_x singularity but does not verify this fractional bound for a general closed-loop EEHJB tangent system.
  • domain assumption Assumption 4.1: uniform ellipticity, heat-kernel gradient bounds (s−t)^{−1/2}, C² finite-dimensional parameterization.
    Underpins Lemma 4.3, Theorem 4.4 and the delta method; standard PDE regularity, but nontrivial for the controlled drift.
  • domain assumption Assumption 4.2: uniform comparison envelope ∥V*_M,τ∥_V ≤ B_V and uniform mild-clause bounds over τ∈I.
    Needed for temperature-dependent stability (4.17)–(4.19); fixed-temperature well-posedness does not supply it, as the paper notes. Verified for the trigonometric model but not for general data.
  • domain assumption Theorem 4.7 assumes existence of one mild equilibrium (V0,π0) at θ0.
    The delta method constructs a local branch around an existing equilibrium; for the general class, existence is taken from background theory or constructed only for a symmetric base model in Cor 4.8.
  • ad hoc to paper Cone/order conditions (3.12)–(3.14): AX+⊂X+, w_h⪰c_−ψ, A(fψ)⪰κ_− I^α f ψ.
    Without these positivity hypotheses the lower bound of Theorem 3.4 fails; the paper's own negative-feedback scalar example damps the response. Verified only in the explicit bounded-diffusion model via the Perron cone, not for a general state-dependent EEHJB.
  • domain assumption Cor 5.5: K≥0 with positive row sums, sup supp μ0 = 1, T < π/2.
    Positive row sums and the short-horizon cosine positivity are required for the nonlinear cooperative-selection conclusion; the paper notes longer horizons can change the sign of the return kernel.
  • domain assumption Statistical input: √n(θ̂_n−θ0)⇒Ξ and P(θ̂_n∈U)→1.
    The delta-method output (4.40) is only as good as the finite-dimensional estimator convergence; the paper explicitly does not claim a nonparametric coefficient-process delta method.
  • standard math Mittag-Leffler positive-axis asymptotics log E_α(z) ∼ z^{1/α} (proved in SM1).
    Used to convert Volterra bounds into exponential rates; the supplement supplies a Stirling-based proof.
  • standard math Fractional integral semigroup identities I^α I^γ = I^{α+γ} and gamma denominators.
    Core to summing the Neumann series into Mittag-Leffler functions.
  • standard math Perron-Frobenius and nonnegative-matrix facts: acyclic nilpotence, ρ(K)>0 on cycles, norm comparisons.
    Used in Theorem 3.6 and Cor 5.4 to turn block structure into path/cycle rates.

pith-pipeline@v1.3.0-alltime-deepseek · 30354 in / 24577 out tokens · 265349 ms · 2026-08-01T15:56:21.147568+00:00 · methodology

0 comments
read the original abstract

Entropy regularization smooths equilibrium policies in time-inconsistent stochastic control. At low temperature, the same Gibbs response can strongly amplify errors in learned rewards and dynamics. We show that the derivative of an exploratory equilibrium is governed by a backward Volterra-parabolic resolvent. Along an aligned positive mode, a lower bound has the same exponential order. A block decomposition identifies the source of the amplification: causal paths contribute powers of 1/tau, whereas a positive feedback cycle can produce exponential growth. At fixed temperature, a local equilibrium branch is twice differentiable with respect to finite-dimensional model parameters, which yields a function-valued delta method. A bounded uniformly elliptic diffusion realizes this path-cycle distinction in every finite dimension. Closing one positive cycle changes the root-n linear-response boundary from a power law to order 1/log n; along the cyclic Perron mode, right-endpoint discretization is relatively consistent exactly when N tau^2 -> infinity. An affine model also gives an exact nonlinear transition at the Lambert-W temperature beta T / W(beta T sqrt(n)). Numerical calculations illustrate these rates.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

62 extracted references · 12 canonical work pages

  1. [1]

    Bayraktar, Y.-J

    E. Bayraktar, Y.-J. Huang, Z. Wang, and Z. Zhou , Relaxed equilibria for time-inconsistent Markov decision processes , Math. Oper. Res., 50 (2025), pp. 2666--2687, https://doi.org/10.1287/moor.2023.0209. Published online October 23, 2024

  2. [2]

    Bayraktar, Z

    E. Bayraktar, Z. Wang, and Z. Zhou , Short communication: Stability of time-inconsistent stopping for one-dimensional diffusions , SIAM J. Financial Math., 13 (2022), pp. SC123--SC135, https://doi.org/10.1137/22M1510005

  3. [3]

    Bayraktar, Z

    E. Bayraktar, Z. Wang, and Z. Zhou , Stability of equilibria in time-inconsistent stopping problems , SIAM J. Control Optim., 61 (2023), pp. 674--696, https://doi.org/10.1137/22M1496955

  4. [4]

    Berman and R

    A. Berman and R. J. Plemmons , Nonnegative Matrices in the Mathematical Sciences , vol. 9 of Classics in Applied Mathematics, Society for Industrial and Applied Mathematics, Philadelphia, 1994, https://doi.org/10.1137/1.9781611971262

  5. [5]

    Bj \"o rk, M

    T. Bj \"o rk, M. Khapko, and A. Murgoci , On time-inconsistent stochastic control in continuous time , Finance Stoch., 21 (2017), pp. 331--360, https://doi.org/10.1007/s00780-017-0327-5

  6. [6]

    Bj \"o rk, M

    T. Bj \"o rk, M. Khapko, and A. Murgoci , Time-Inconsistent Control Theory with Finance Applications , Springer Finance, Springer, Cham, 2021, https://doi.org/10.1007/978-3-030-81843-2

  7. [8]

    M. Dai, Y. Dong, and Y. Jia , Learning equilibrium mean--variance strategy , Math. Finance, 33 (2023), pp. 1166--1212, https://doi.org/10.1111/mafi.12402

  8. [9]

    Diethelm , The Analysis of Fractional Differential Equations , vol

    K. Diethelm , The Analysis of Fractional Differential Equations , vol. 2004 of Lecture Notes in Mathematics, Springer, Berlin, 2010, https://doi.org/10.1007/978-3-642-14574-2

  9. [10]

    Gripenberg, S.-O

    G. Gripenberg, S.-O. Londen, and O. Staffans , Volterra Integral and Functional Equations , vol. 34 of Encyclopedia of Mathematics and its Applications, Cambridge University Press, Cambridge, 1990, https://doi.org/10.1017/CBO9780511662805

  10. [11]

    X. Guo, Y. Huang, and X. Yu , Deterministic policy gradient for learning equilibrium in time-inconsistent control problems , 2026, https://arxiv.org/abs/2606.11798

  11. [12]

    Hajek , Cooling schedules for optimal annealing , Math

    B. Hajek , Cooling schedules for optimal annealing , Math. Oper. Res., 13 (1988), pp. 311--329, https://doi.org/10.1287/moor.13.2.311

  12. [13]

    Huang, Z

    Y.-J. Huang, Z. Wang, and Z. Zhou , Convergence of policy iteration for entropy-regularized stochastic control problems , SIAM J. Control Optim., 63 (2025), pp. 752--777, https://doi.org/10.1137/24M1638744

  13. [14]

    Huang, X

    Y.-J. Huang, X. Yu, and K. Zhang , Policy iteration achieves regularized equilibrium under time inconsistency , 2026, https://arxiv.org/abs/2603.06145

  14. [15]

    R. M. Karp , A characterization of the minimum cycle mean in a digraph , Discrete Mathematics, 23 (1978), pp. 309--311, https://doi.org/10.1016/0012-365X(78)90011-0

  15. [16]

    Lei and C

    Q. Lei and C. S. Pun , Nonlocal fully nonlinear parabolic differential equations arising in time-inconsistent problems , J. Differential Equations, 358 (2023), pp. 339--385, https://doi.org/10.1016/j.jde.2023.02.025

  16. [17]

    Lei and C

    Q. Lei and C. S. Pun , On the well-posedness of Hamilton--Jacobi--Bellman equations of the equilibrium type , 2023, https://arxiv.org/abs/2307.01986. Revised May 2026

  17. [18]

    Lei and C

    Q. Lei and C. S. Pun , Nonlocality, nonlinearity, and time inconsistency in stochastic differential games , Math. Finance, 34 (2024), pp. 190--256, https://doi.org/10.1111/mafi.12420

  18. [19]

    Leonidov, A

    A. Leonidov, A. Savvateev, and A. G. Semenov , Quantal response equilibria in binary choice games on graphs , 2019, https://arxiv.org/abs/1912.09584

  19. [20]

    N. S. Lesmana and C. S. Pun , A subgame perfect equilibrium reinforcement learning approach to time-inconsistent problems , SIAM J. Financial Math., 16 (2025), pp. 68--122, https://doi.org/10.1137/23M1594510

  20. [21]

    J. Ma, G. Wang, and J. Zhang , Convergence analysis for entropy-regularized control problems: A probabilistic approach , SIAM J. Control Optim., 64 (2026), pp. 816--842, https://doi.org/10.1137/24M1680039

  21. [22]

    R. D. McKelvey and T. R. Palfrey , Quantal response equilibria for extensive form games , Experimental Economics, 1 (1998), pp. 9--41, https://doi.org/10.1023/A:1009905800005

  22. [23]

    Podlubny , Fractional Differential Equations , vol

    I. Podlubny , Fractional Differential Equations , vol. 198 of Mathematics in Science and Engineering, Academic Press, San Diego, 1999

  23. [25]

    Reisinger and Y

    C. Reisinger and Y. Zhang , Regularity and stability of feedback relaxed controls , SIAM J. Control Optim., 59 (2021), pp. 3118--3151, https://doi.org/10.1137/20M1312435

  24. [26]

    Sethi, D

    D. Sethi, D. S i s ka, and Y. Zhang , Entropy annealing for policy mirror descent in continuous time and space , SIAM J. Control Optim., 63 (2025), pp. 3006--3041, https://doi.org/10.1137/24M166591X

  25. [27]

    W. Tang, Y. P. Zhang, and X. Y. Zhou , Exploratory HJB equations and their convergence , SIAM J. Control Optim., 60 (2022), pp. 3191--3216, https://doi.org/10.1137/21M1448185

  26. [28]

    H. V. Tran, Z. Wang, and Y. P. Zhang , Policy iteration for exploratory Hamilton--Jacobi--Bellman equations , Appl. Math. Optim., 91 (2025), 50, https://doi.org/10.1007/s00245-025-10249-3

  27. [29]

    A. W. van der Vaart , Asymptotic Statistics , vol. 3 of Cambridge Series in Statistical and Probabilistic Mathematics, Cambridge University Press, Cambridge, 1998, https://doi.org/10.1017/CBO9780511802256

  28. [30]

    Z. Wang, X. Yu, J. Zhang, and Z. Zhou , Equilibrium under time-inconsistency: A new existence theory by vanishing entropy regularization , 2026, https://arxiv.org/abs/2603.10321

  29. [31]

    Yong , Time-inconsistent optimal control problems and the equilibrium HJB equation , Math

    J. Yong , Time-inconsistent optimal control problems and the equilibrium HJB equation , Math. Control Relat. Fields, 2 (2012), pp. 271--329, https://doi.org/10.3934/mcrf.2012.2.271

  30. [32]

    N. V. Krylov , Lectures on Elliptic and Parabolic Equations in H \"o lder Spaces , vol. 12 of Graduate Studies in Mathematics, American Mathematical Society, Providence, RI, 1996, https://doi.org/10.1090/gsm/012

  31. [33]

    Tang, Wenpin and Zhang, Yuming Paul and Zhou, Xun Yu , title =. SIAM J. Control Optim. , volume =. 2022 , doi =

  32. [34]

    Huang, Yu-Jui and Wang, Zhenhua and Zhou, Zhou , title =. SIAM J. Control Optim. , volume =. 2025 , doi =

  33. [35]

    Ma, Jin and Wang, Gaozhan and Zhang, Jianfeng , title =. SIAM J. Control Optim. , volume =. 2026 , doi =

  34. [36]

    Tran, Hung Vinh and Wang, Zhenhua and Zhang, Yuming Paul , title =. Appl. Math. Optim. , volume =. 2025 , doi =

  35. [37]

    Reisinger, Christoph and Zhang, Yufei , title =. SIAM J. Control Optim. , volume =. 2021 , doi =

  36. [38]

    Bayraktar, Erhan and Wang, Zhenhua and Zhou, Zhou , title =. SIAM J. Control Optim. , volume =. 2023 , doi =

  37. [39]

    Bayraktar, Erhan and Wang, Zhenhua and Zhou, Zhou , title =. SIAM J. Financial Math. , volume =. 2022 , doi =

  38. [40]

    and Palfrey, Thomas R

    McKelvey, Richard D. and Palfrey, Thomas R. , title =. Experimental Economics , volume =. 1998 , doi =

  39. [41]

    , title =

    Leonidov, Andrey and Savvateev, Alexey and Semenov, Andrew G. , title =. 2019 , eprint =

  40. [42]

    Here, There and Everywhere: State-Dependent Time-Inconsistent Stochastic Control , year =

    Possama. Here, There and Everywhere: State-Dependent Time-Inconsistent Stochastic Control , year =. 2603.22022 , archivePrefix =

  41. [43]

    Hajek, Bruce , title =. Math. Oper. Res. , volume =. 1988 , doi =

  42. [44]

    Entropy Annealing for Policy Mirror Descent in Continuous Time and Space , journal =

    Sethi, Deven and. Entropy Annealing for Policy Mirror Descent in Continuous Time and Space , journal =. 2025 , doi =

  43. [45]

    Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning , year =

    Cao, Jialun and Acero, Fernando and. Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning , year =. 2607.03168 , archivePrefix =

  44. [46]

    Yong, Jiongmin , title =. Math. Control Relat. Fields , volume =. 2012 , doi =

  45. [47]

    On Time-Inconsistent Stochastic Control in Continuous Time , journal =

    Bj. On Time-Inconsistent Stochastic Control in Continuous Time , journal =. 2017 , doi =

  46. [48]

    Time-Inconsistent Control Theory with Finance Applications , series =

    Bj. Time-Inconsistent Control Theory with Finance Applications , series =. 2021 , doi =

  47. [49]

    Lei, Qian and Pun, Chi Seng , title =. J. Differential Equations , volume =. 2023 , doi =

  48. [50]

    Lei, Qian and Pun, Chi Seng , title =. Math. Finance , volume =. 2024 , doi =

  49. [51]

    2023 , eprint =

    Lei, Qian and Pun, Chi Seng , title =. 2023 , eprint =

  50. [52]

    Dai, Min and Dong, Yuchao and Jia, Yanwei , title =. Math. Finance , volume =. 2023 , doi =

  51. [53]

    Lesmana, Nixie Sapphira and Pun, Chi Seng , title =. SIAM J. Financial Math. , volume =. 2025 , doi =

  52. [54]

    Bayraktar, Erhan and Huang, Yu-Jui and Wang, Zhenhua and Zhou, Zhou , title =. Math. Oper. Res. , volume =. 2025 , doi =

  53. [55]

    2026 , eprint =

    Huang, Yu-Jui and Yu, Xiang and Zhang, Keyu , title =. 2026 , eprint =

  54. [56]

    2026 , eprint =

    Wang, Zhenhua and Yu, Xiang and Zhang, Jingjie and Zhou, Zhou , title =. 2026 , eprint =

  55. [57]

    2026 , eprint =

    Guo, Xin and Huang, Yijie and Yu, Xiang , title =. 2026 , eprint =

  56. [58]

    2010 , doi =

    Diethelm, Kai , title =. 2010 , doi =

  57. [59]

    Podlubny, Igor , title =

  58. [60]

    1990 , doi =

    Gripenberg, Gustaf and Londen, Stig-Olof and Staffans, Olof , title =. 1990 , doi =

  59. [61]

    , title =

    Berman, Abraham and Plemmons, Robert J. , title =. 1994 , doi =

  60. [62]

    , title =

    Karp, Richard M. , title =. Discrete Mathematics , volume =. 1978 , doi =

  61. [63]

    , title =

    van der Vaart, Aad W. , title =. 1998 , doi =

  62. [64]

    Krylov, N. V. , title =. 1996 , doi =