Pith. sign in

REVIEW 3 major objections 4 minor 2 cited by

In a two-investor portfolio game, the informed leader's optimal strategy is Gaussian randomization.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

In a two-investor mean-variance Stackelberg game with asymmetric information, the leader's equilibrium randomized strategy is Gaussian and the follower's strategy depends linearly on the leader's observed trades.

T0 review reviewed 2026-08-05 challenge →

load-bearing objection The continuous-time exploratory equilibrium is a genuine, mostly verified contribution; the discrete-sampling ε-Stackelberg theorem is not proven as written. the 3 major comments →

arxiv 2509.03669 v1 pith:54ISB4N5 submitted 2025-09-03 q-fin.MF math.OC

Mean-Variance Stackelberg Games with Asymmetric Information

classification q-fin.MF math.OC MSC 91G1591A6593E11
keywords asymmetric informationmean-variance portfolio selectionStackelberg gamerandomized strategyintra-personal equilibriumentropy regularizationrelative performanceGaussian policy
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that when a fully informed investor acts as a Stackelberg leader and a partially informed investor as follower, both caring about relative performance, an equilibrium exists in which the leader deliberately randomizes her trades to hide information. In that equilibrium the leader samples her portfolio positions from a Gaussian distribution, and the follower's best response is a linear function of the leader's observed positions. Because wealth dynamics are linear, the follower's optimal combined portfolio is deterministic, so his equilibrium value function does not depend on the leader's realized random actions. The paper proves this as an exact intra-personal equilibrium under idealized continuous action sampling, and as a time-consistent ε-Stackelberg equilibrium when actions are sampled on a discrete grid. If correct, this gives a tractable description of information-hiding in financial markets: camouflage is Gaussian, its variance is set by an entropy weight and risk aversion, and its mean is unaffected by that weight.

Core claim

On the paper's own terms, the central discovery is the explicit equilibrium profile (Π*, u₂*). The leader's randomized policy is Gaussian with mean (θ(p)-r)/(σ²l) - β(p)/(χσ)(∂_p a₁ + (1-χ)∂_p a₂) and constant variance λ0/(γ₁σ²χ²); the follower responds linearly, u₂* = [(θ(p)-r)/(σ²γ₂) - β(p)∂_p a₂/σ]/(1-λ₂/2) + u₁λ₂/(2-λ₂). The follower's linear response cancels the leader's sampled noise in the combined portfolio u*, so the follower's equilibrium value function is the deterministic affine function (1-λ₂/2)x₂ - (λ₂/2)x₁ + A₂(t,p). Under continuous sampling this is an intra-personal Stackelberg equilibrium; under discrete sampling it becomes an ε-Stackelberg equilibrium with ε governed by th

What carries the argument

The argument rests on three pieces. (1) A change of variables turns the pair's relative wealth into a single process Z driven by a combined portfolio u = (1-λ/2)u₂ - (λ/2)u₁; linearity means the follower can pick u* independent of the leader's realized action, and this is why the follower's value function becomes deterministic. (2) An extended Hamilton-Jacobi-Bellman system—two coupled PDEs for the value function and an auxiliary expected-terminal-wealth function—handles the time inconsistency of mean-variance preferences and produces the semi-analytic formulas. (3) The leader's problem is cast in exploratory dynamics, where sampled random actions are replaced by equivalent Brownian noise; m

Load-bearing premise

The load-bearing premise is that the follower never updates his estimate of the stock's drift from the leader's observed trades, justified only by practical infeasibility; if a statistical follower did learn from trades, the equilibrium formulas would no longer be guaranteed.

What would settle it

Re-solve the follower's problem with the posterior P(t) computed from both the stock path and the history of the leader's sampled actions u₁(δ(t)), instead of from the stock path alone; if the resulting best response differs from (3.9), the claimed Stackelberg equilibrium does not survive inference. A simulation can make this concrete: generate leader actions from (4.10), let an econometrician estimate μ from (S, u₁), and check whether the follower's optimal position still matches (3.9).

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • The leader's equilibrium randomization has constant variance λ0/(γ₁σ²χ²): more randomization weight or less risk aversion means noisier camouflage, while the mean trade depends only on fundamentals and the posterior p.
  • The follower's value is deterministic despite the leader's random actions, because his linear response cancels exactly the leader's action noise; random leadership does not impose welfare risk on a rational relative-performance follower.
  • On a discrete trading grid, sampling from the Gaussian policy delivers a time-consistent ε-Stackelberg equilibrium, with ε shrinking as the grid is refined; implementation only requires sampling at sufficiently frequent times.
  • As the entropy weight λ0 goes to zero, the Gaussian policy degenerates to the deterministic strategy of a Stackelberg leader under partial information, recovering known single-investor and full-information benchmarks.
  • The leader's Gaussian mean is independent of λ0, so the choice of how much to randomize separates from the choice of average position—randomization can be tuned without distorting expected trades.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the follower were allowed to update his posterior with the leader's observed trades, the Gaussian equilibrium would not be expected to survive; a natural extension would replace pure randomization with a signalling equilibrium, where the leader's trades convey information and the follower's response changes.
  • The same linear-cancellation mechanism suggests a general result: in any linear wealth dynamics with relative performance, a follower's best response can neutralize a leader's action noise, so informed-player randomization may be welfare-neutral for followers.
  • The exploratory-to-sampled approximation with linear error in grid mesh gives a practical prescription: the leader can quantify the trade-off between sampling frequency and equilibrium accuracy, choosing grid size from the desired ε.
  • The entropy-regularized exploratory method could be re-run for a Nash version of the same asymmetric-information game; the expectation is a Gaussian equilibrium again, but with coupled first-order conditions instead of the leader's Stackelberg optimality.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper studies a two-investor mean-variance portfolio choice problem with asymmetric information and relative performance. Investor 1 knows the true drift μ; investor 2 observes stock prices and filters μ from the price path. The interaction is modeled as a Stackelberg game: the informed leader randomizes her strategy under an entropy-regularized objective to protect her informational advantage, while the follower responds to the sampled actions. Section 3 derives a semi-analytic intra-personal equilibrium for the follower under a filtration G_t that includes the full future path of the leader's sampled actions; Section 4 introduces exploratory dynamics and derives the leader's Gaussian equilibrium strategy. Section 4.3 claims that discrete sampling of the leader's Gaussian policy yields an ε-Stackelberg equilibrium, using an approximation result of Jia et al. (2025).

Significance. If the continuous-time results are correct, the paper makes a useful contribution: it extends entropy-regularized mean-variance portfolio selection to a Stackelberg game with asymmetric information, provides explicit Gaussian equilibrium strategies, and identifies a deterministic follower value function despite the leader's random actions. The verification proofs for Theorems 3.1 and 4.1 are written out in detail. However, the paper's advertised 'realistic case' result, Theorem 4.2, is not established by the proof, and the information assumptions in Section 3 weaken the economic motivation. The value of the paper depends on whether the discrete-sampling bridge can be repaired.

major comments (3)
  1. [Section 4.3, Theorem 4.2 proof] The proof does not establish the ε-intra-personal equilibrium of Definition 4.2. Definition 4.2 requires a supremum over arbitrary one-point deviations π at a fixed grid point t_i in the sampled objective J^D_1. The proof instead starts from inequality (4.11), which controls an infinitesimal exploratory perturbation Π^{h,eπ} on [t,t+h] in eJ_1, and then in (4.15) equates J^D_1(Π^π)+ε(D2,Π^π) with eJ_1(Π^{h,eπ}). These are different objects: a sampled deviation acts on an interval of length |D|, not an infinitesimal h, and the policy in J^D_1 is not the exploratory policy. The assertion lim_{h↓0} ε(D2,Π^π)=ε(D2,Π*) is not justified, and no argument shows this limit is uniform over π∈A1. Lemma 4.1's constant C depends on Π*, so the weak-convergence estimate (4.12) gives no control over arbitrary π (e.g., π with very large variance). Consequently the displayed inequality after (4.15) does n
  2. [Section 3, Definition 3.1 and Remark 3.2] The follower's information structure is load-bearing but economically questionable. Definition 3.1 defines A2 as progressively measurable with respect to G_t = F^S_t ⊗ F^ξ_T, so the follower's strategy u2(t) may depend on the entire future path of the leader's sampled actions. This is not a causal observation structure and is never justified as a limit of discrete-time updating. Remark 3.2 simultaneously bars the follower from using the observed actions u1 to update his belief P(t) about μ. Under these assumptions, the leader's randomization does not function as information protection: the follower cannot learn from trades by fiat, so randomization only injects entropy into the leader's objective. The equilibrium formulas (3.9) and (3.10) are derived under this filtration and this no-learning assumption, so they do not address the information-leakage concern that motivates the model. The
  3. [Lemma 4.1 and Section 4.1] The approximation result that underpins Theorem 4.2 is not actually verified. The proof of Lemma 4.1 asserts that the coefficients of the exploratory dynamics (4.2), namely ebt(θ(Pt))-r, ebt σ, and σ eσt, belong to C^4_p with bounded derivatives, based on the regularity of a1 and a2 derived in Theorem 4.1. However, the cited regularity (Lemma 3.3 of Huang and Sun 2023) only provides C^{1,2} solutions that are bounded on [0,T]×(0,1); no C^4_p estimates are given. Moreover, the exploratory dynamics (4.2) are derived in Section 4.1 only through an LLN/CLT heuristic, with no rigorous statement of the convergence from the sampled dynamics (3.1). Since (4.12)–(4.13) are the sole bridge from exploratory to sampled equilibrium, the missing verification of the conditions of Jia et al. (2025) is a substantive gap.
minor comments (4)
  1. [Theorem 4.1 and Appendix A.2, Eq. (A.23)] The entropy term in the A1 PDE is written as λ0/2 log(2πλ0/(γ1χ^2)). Given the variance in (4.10), λ0/(γ1σ^2χ^2), Shannon's differential entropy is 1/2 log(2πeλ0/(γ1σ^2χ^2)). The expression in (A.30) correctly uses σ^2 in the denominator, so (A.23) and Theorem 4.1 are internally inconsistent. The value function A1 should be corrected.
  2. [Definition 4.2] The definition of the one-point perturbation Π^π_t = π(t) for t = t_i is imprecise for continuous-time policies: a deviation at a single instant has no effect on the sampled dynamics over a positive-length interval. The perturbation should be defined on the interval [t_i, t_{i+1}) or in the discrete-time sampling protocol. This imprecision contributes to the gap in the proof of Theorem 4.2.
  3. [Equation (3.9)] In (3.9), the notation u1 appears in the third term without an explicit argument; since the follower's strategy is evaluated at time t, it would be clearer to write u1(δ(t)) or u1(t_i) to match the sampled dynamics in (3.1).
  4. [Throughout] Minor typos: 'equalibrium' in the proof of Theorem 4.1; 'we show that p1(·) can be characterized' is a slightly awkward construction; the paper would benefit from a table of notation for θ, β, Γ, χ, and l, which appear in several places.

Circularity Check

0 steps flagged

No significant circularity: the follower and leader equilibria are derived in-paper; self-citations supply only ancillary filtering and PDE facts.

full rationale

The paper's central derivation chain is self-contained rather than circular. The follower's equilibrium (3.8)-(3.9) is obtained in Appendix A.1 by the standard extended-HJB verification procedure: an ansatz (A.3), a maximizer computation (A.6), and a martingale verification that the candidate indeed equals the conditional expectation objective. The citations to Huang and Sun (2023) are ancillary: Lemma 3.2 gives the standard Wonham-filter SDE (2.3), and Lemma 3.3/Corollary 3.1 give existence-uniqueness of classical solutions to the linear parabolic Cauchy problems (A.8)/(A.9). These results are parameter-free, do not contain the equilibrium claim, and are not used to forbid alternative equilibria; indeed, the paper explicitly says uniqueness of equilibrium remains open. The remark after Theorem 3.1 that (3.9) is 'consistent with' Theorem 3.2 of Huang and Sun is a comment, not a load-bearing step. The leader's Gaussian equilibrium (4.10) is derived in Appendix A.2 by maximizing the entropy-regularized HJB, with the Gaussian form following from the classical maximum-entropy property, again not from a self-citation. The deterministic value function (3.10) is a consequence of the linear combination u* being independent of u1, and is verified by Itô-martingale arguments, not assumed as the conclusion. The only notable weakness is in the proof of Theorem 4.2 (the epsilon-Stackelberg claim): after (4.15), the assertion lim_{h↓0} epsilon(D2,Pi^pi) = epsilon(D2,Pi*) is not justified, the o(h) bound from (4.11) is not uniform over pi in A1, and (4.11) controls infinitesimal exploratory perturbations while Definition 4.2 perturbs a single sampled grid point. This is a correctness gap, not a circular reduction: the continuous-time equilibrium derivation does not presuppose the discrete-sampling result. Therefore the circularity score is 0, with the above caveat noted under correctness risk rather than circularity.

Axiom & Free-Parameter Ledger

1 free parameters · 6 axioms · 0 invented entities

The central claim rests on a chain of external results: filtering from the authors' prior paper, PDE well-posedness from the authors' prior paper, and an approximation theorem from an unpublished preprint. The only truly free modeling parameter is λ0; the more consequential load-bearing premises are the two behavioral assumptions on the follower (no learning from leader's actions, and conditioning on future actions).

free parameters (1)
  • λ0 (entropy regularization weight)
    Controls the variance λ0/(γ1σ²χ²) of the leader's Gaussian policy; it is chosen ad hoc to represent the leader's preference for randomization and is not derived from an information-leakage constraint or calibrated to data. The value-function entropy term also depends on it.
axioms (6)
  • domain assumption The posterior P(t) for μ=μ1 evolves as dP = ((μ1-μ2)/σ) P(1-P) dW̃ (Eq. 2.3), from Lemma 3.2 of Huang and Sun (2023).
    Standard two-point Wonham filter; the paper cites its own prior result for existence/uniqueness of this SDE.
  • standard math The Cauchy problems for a2 and A2 (and a1, A1) have unique classical solutions with bounded derivatives, taken from Lemma 3.3 and Corollary 3.1 of Huang and Sun (2023).
    Standard parabolic PDE results, but cited to the authors' own preprint rather than proved.
  • domain assumption The exploratory dynamics (4.2)-(4.4) obtained by replacing sampled actions with e_b + e_σ ε and passing to the LLN/CLT limit are a valid idealization of the sampled dynamics.
    This is the RL exploratory-control ansatz of Wang et al. (2020); used without a self-contained derivation of the limit.
  • domain assumption The approximation theorem of Jia et al. (2025) applies, i.e., the sampled dynamics converge to the exploratory dynamics with rate C|D| and the conditions (C4_p regularity, bounded derivatives) hold for the Gaussian equilibrium policy.
    The paper asserts the conditions hold by regularity of a1,a2 without full verification; Jia et al. is an unpublished preprint.
  • ad hoc to paper The follower conditions his mean-variance objective on the full path of the leader's sampled actions (filtration G_t = F^S_t ⊗ F^ξ_T, Section 3), i.e., he knows the leader's future actions when forming his objective.
    This pathwise/random-field formulation is borrowed from Buckdahn and Ma (2007); it is a strong modeling assumption because the follower in reality only observes actions as they occur.
  • ad hoc to paper The follower does not use the leader's sampled actions to update his belief about μ (Remark 3.2).
    Explicitly restricts the follower to the stock-price-only posterior P(t). Without this, the leader's randomization would interact with the follower's learning and the Stackelberg equilibrium would be a signaling game.

reviewed 2026-08-05 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Mean-Variance Stackelberg Games with Asymmetric Information." pith.science (2026). https://pith.science/paper/54ISB4N5

@misc{pith2026250903669,
  author       = {Pith},
  title        = {Pith review of: Mean-Variance Stackelberg Games with Asymmetric Information},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/54ISB4N5}},
  note         = {Machine review of arXiv:2509.03669}
}
Share X Bluesky LinkedIn Reddit HN
abstract

This paper considers two investors who perform mean-variance portfolio selection with asymmetric information: one knows the true stock dynamics, while the other has to infer the true dynamics from observed stock evolution. Their portfolio selection is interconnected through relative performance concerns, i.e., each investor is concerned about not only her terminal wealth, but how it compares to the average terminal wealth of both investors. We model this as Stackelberg competition: the partially-informed investor (the "follower") observes the trading behavior of the fully-informed investor (the "leader") and decides her trading strategy accordingly; the leader, anticipating the follower's response, in turn selects a trading strategy that best suits her objective. To prevent information leakage, the leader adopts a randomized strategy selected under an entropy-regularized mean-variance objective, where the entropy regularizer quantifies the randomness of a chosen strategy. The follower, on the other hand, observes only the actual trading actions of the leader (sampled from the randomized strategy), but not the randomized strategy itself. Her mean-variance objective is thus a random field, in the form of an expectation conditioned on a realized path of the leader's trading actions. In the idealized case of continuous sampling of the leader's trading actions, we derive a Stackelberg equilibrium where the follower's trading strategy depends linearly on the actual trading actions of the leader and the leader samples her trading actions from Gaussian distributions. In the realistic case of discrete sampling of the leader's trading actions, the above becomes an $\epsilon$-Stackelberg equilibrium.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Mean-field game of mean-variance portfolio optimization with peer-based risk aversion

    q-fin.MF 2026-05 accept novelty 7.0

    A mean-field equilibrium exists for a mean-variance portfolio game in which each agent's risk aversion switches discontinuously as wealth crosses the population average.

  2. Mean-field game of mean-variance portfolio optimization with peer-based risk aversion

    q-fin.MF 2026-05 unverdicted novelty 6.0

    Existence of mean-field equilibrium is shown for a time-inconsistent mean-variance portfolio game with piecewise peer-based relative risk aversion via regularization of discontinuous FBSDEs.

Reference graph

Works this paper leans on

36 extracted references · 33 canonical work pages · cited by 1 Pith paper · 1 internal anchor

  1. [1]

    Amendinger, P

    J. Amendinger, P. Imkeller, and M. Schweizer. Additional logarithmic utility of an insider. Stochastic Processes and their Applications, 75 0 (2): 0 263--286, 1998

  2. [2]

    Back and S

    K. Back and S. Baruch. Information in securities markets: Kyle meets G losten and M ilgrom. Econometrica, 72 0 (2): 0 433--465, 2004

  3. [3]

    Basak and G

    S. Basak and G. Chabakauri. Dynamic mean-variance asset allocation. The Review of Financial Studies, 23 0 (8): 0 2970--3016, 2010

  4. [4]

    Bender and N

    C. Bender and N. T. Thuan. On the grid-sampling limit SDE . arXiv preprint arXiv:2410.07778, 2024

  5. [5]

    Bj \"o rk, M

    T. Bj \"o rk, M. Khapko, and A. Murgoci. On time-inconsistent stochastic control in continuous time. Finance and Stochastics, 21: 0 331--360, 2017

  6. [6]

    L. Bo, J. Wang, X. Wei, and X. Yu. Mean field control with poissonian common noise: A pathwise compactification approach. arXiv preprint arXiv:2505.23441, 2025

  7. [7]

    Buckdahn and J

    R. Buckdahn and J. Ma. Stochastic viscosity solutions for nonlinear stochastic partial differential equations. P art I . Stochastic Processes and their Applications, 93 0 (2): 0 181--204, 2001 a

  8. [8]

    Buckdahn and J

    R. Buckdahn and J. Ma. Stochastic viscosity solutions for nonlinear stochastic partial differential equations. P art II . Stochastic Processes and their Applications, 93 0 (2): 0 205--228, 2001 b

  9. [9]

    Buckdahn and J

    R. Buckdahn and J. Ma. Pathwise stochastic control problems and stochastic HJB equations. SIAM Journal on Control and Optimization, 45 0 (6): 0 2224--2256, 2007

  10. [10]

    Cardaliaguet

    P. Cardaliaguet. Differential games with asymmetric information. SIAM Journal on Control and Optimization, 46 0 (3): 0 816--838, 2007

  11. [11]

    Cardaliaguet and C

    P. Cardaliaguet and C. Rainer. Stochastic differential games with asymmetric information. Applied Mathematics and Optimization, 59 0 (1): 0 1--36, 2009

  12. [12]

    Carmona, F

    R. Carmona, F. Delarue, and D. Lacker. Mean field games with common noise. Annals of Probability, 44 0 (6): 0 3740--3803, 2016

  13. [13]

    J. M. Corcuera, P. Imkeller, A. Kohatsu-Higa, and D. Nualart. Additional utility of insiders with imperfect dynamical information. Finance and Stochastics, 8 0 (3): 0 437--450, 2004

  14. [14]

    T. M. Cover and J. A. Thomas. Elements of information theory, volume 1. John Wiley & Sons, 2006

  15. [15]

    M. Dai, Y. Dong, and Y. Jia. Learning equilibrium mean-variance strategy. Mathematical Finance, 33 0 (4): 0 1166--1212, 2023

  16. [16]

    De Angelis, E

    T. De Angelis, E. Ekstr \"o m, and K. Glover. Dynkin games with incomplete and asymmetric information. Mathematics of Operations Research, 47 0 (1): 0 560--586, 2022

  17. [17]

    Espinosa and N

    G.-E. Espinosa and N. Touzi. Optimal investment under relative performance concerns. Mathematical Finance, 25 0 (2): 0 221--257, 2015

  18. [18]

    W. H. Fleming and M. Nisio. On stochastic relaxed control for partially observed diffusions. Nagoya Mathematical Journal, 93: 0 71--108, 1984

  19. [19]

    G \^a rleanu and L

    N. G \^a rleanu and L. H. Pedersen. Dynamic trading with predictable returns and transaction costs. The Journal of Finance, 68 0 (6): 0 2309--2340, 2013

  20. [20]

    G \^a rleanu and L

    N. G \^a rleanu and L. H. Pedersen. Dynamic portfolio choice with frictions. Journal of Economic Theory, 165: 0 487--516, 2016

  21. [21]

    Graewe, U

    P. Graewe, U. Horst, and J. Qiu. A non- M arkovian liquidation problem and backward SPDE s with singular terminal conditions. SIAM Journal on Control and Optimization, 53 0 (2): 0 690--711, 2015

  22. [22]

    Gr \"u n

    C. Gr \"u n. On D ynkin games with incomplete information. SIAM Journal on Control and Optimization, 51 0 (5): 0 4039--4065, 2013

  23. [23]

    P. Guasoni. Asymmetric information in fads models. Finance and Stochastics, 10 0 (2): 0 159--177, 2006

  24. [24]

    J. Han, X. Li, G. Ma, and A. P. Kennedy. Strategic trading with information acquisition and long-memory stochastic liquidity. European Journal of Operational Research, 308 0 (1): 0 480--495, 2023

  25. [25]

    X. D. He and Z. L. Jiang. On the equilibrium strategies for time-inconsistent problems in continuous time. SIAM Journal on Control and Optimization, 59 0 (5): 0 3860--3886, 2021

  26. [26]

    Partial Information in a Mean-Variance Portfolio Selection Game

    Y.-J. Huang and L.-H. Sun. Partial information breeds systemic risk. arXiv preprint arXiv:2312.04045, 2023

  27. [27]

    Huang and Z

    Y.-J. Huang and Z. Zhou. Strong and weak equilibria for time-inconsistent stochastic control in continuous time. Mathematics of Operations Research, 46 0 (2): 0 428--451, 2021

  28. [28]

    Huang and Z

    Y.-J. Huang and Z. Zhou. A time-inconsistent D ynkin game: from intra-personal to inter-personal equilibria. Finance and Stochastics, 26 0 (2): 0 301--334, 2022

  29. [29]

    Y. Jia, D. Ouyang, and Y. Zhang. Accuracy of discretely sampled stochastic policies in continuous-time reinforcement learning. arXiv preprint arXiv:2503.09981, 2025

  30. [30]

    Karatzas and S

    I. Karatzas and S. E. Shreve. Methods of mathematical finance, volume 39. Springer, 1998

  31. [31]

    Lacker and T

    D. Lacker and T. Zariphopoulou. Mean field and n-agent games for optimal investment under relative performance criteria. Mathematical Finance, 29 0 (4): 0 1003--1038, 2019

  32. [32]

    Pikovsky and I

    I. Pikovsky and I. Karatzas. Anticipative portfolio optimization. Advances in Applied Probability, 28 0 (4): 0 1095--1122, 1996

  33. [33]

    Szpruch, T

    L. Szpruch, T. Treetanthiploet, and Y. Zhang. Optimal scheduling of entropy regularizer for continuous-time linear-quadratic reinforcement learning. SIAM Journal on Control and Optimization, 62 0 (1): 0 135--166, 2024

  34. [34]

    Wang and X

    H. Wang and X. Y. Zhou. Continuous-time mean--variance portfolio selection: A reinforcement learning framework. Mathematical Finance, 30 0 (4): 0 1273--1308, 2020

  35. [35]

    H. Wang, T. Zariphopoulou, and X. Y. Zhou. Reinforcement learning in continuous time and space: A stochastic control approach. Journal of Machine Learning Research, 21 0 (198): 0 1--34, 2020

  36. [36]

    X. Y. Zhou. On the existence of optimal relaxed controls of stochastic partial differential equations. SIAM Journal on Control and Optimization, 30 0 (2): 0 247--261, 1992

This paper was first reviewed by deepseek-v4-flash on August 5, 2026.