{"id":"79ff174b-5c5f-43c2-8168-d7c957c81fa1","arxiv_id":"2502.04788","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A two-agent non-zero-sum mean-variance game with Choquet-regularized exploration has a time-consistent Nash equilibrium, explicit in a Gaussian market, and a policy iteration scheme that is claimed to converge uniformly to it.","lead":"Two competing investors who care about both their own wealth and the gap to a rival can learn a time-consistent equilibrium investment strategy in a continuous-time market with unknown parameters. The paper gives explicit equilibrium distributions for a Gaussian market and a policy iteration scheme that is claimed to converge uniformly to the Nash equilibrium.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Lemma 1 requires h' ∈ L^2; the paper's class H allows h(p)=√p−p, for which the equilibrium quantile formula (3.16) is not even defined.","rationale":"The reader's weakest assumption identifies exactly the same load-bearing issue: Lemma 1 is imported without the regularity condition needed for its conclusion. This is the correct concern because every subsequent equilibrium formula—Proposition 1, Theorem 2's quantile expression (3.16), the value-function coefficients in (3.15), and Proposition 2's explicit solution (3.17)—depends on writing the optimal exploratory distribution as m + s h'(1−p)/‖h'‖₂ and its Choquet value as s‖h'‖₂. If h' is not square-integrable, the term ‖h'‖₂ diverges and the whole derivation collapses. The paper's Section 2.2 defines H too broadly, so the central theorem is not valid as stated. The gap is fixable by an explicit absolute-continuity plus L² assumption on h′, which the numerical examples (normal and Gini distortion) already satisfy, so it does not warrant rejection. The rest of the paper has genuine independent support: the verification theorem follows the standard extended-HJB route, the Gaussian-model algebra is coherent under the stronger regularity assumption, and the policy-iteration uniform convergence proof is a reasonable contraction argument, though the joint convergence of the means and value coefficients in Theorem 4 is written a little loosely. The numerical section also contains an internal inconsistency (Table 1 sets ρ=−0.93 while Section 2.1 assumes ρ∈[0,1], and the uniform-distribution support has a spurious factor 3), but these are secondary and do not affect the main theoretical claim. Since the reader already gave CONDITIONAL and this concern sharpens rather than reverses that verdict, no adjustment is needed.","tokens_in":24724,"tokens_out":4481,"duration_ms":53467,"concrete_test":"Take h(p)=√p−p and solve the inner maximization problem (3.9) directly for fixed finite m and s, without invoking Lemma 1. First verify that ‖h'‖₂ is infinite, so the paper's proposed quantile formula (3.8) is undefined. Then test existence of a finite-variance maximizer: for a two-point distribution with mean m and variance s², compute Φ_h explicitly and see whether a finite supremum is attained. If no such maximizer exists, or if the maximizer differs from the asserted formula, then Proposition 1 and Theorem 2 fail for h∈H, and the paper must either restrict H to functions with h'∈L² or supply an alternative derivation for functions with unbounded derivative.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central derivation rests on Lemma 1 (attributed to Liu et al. 2020), which says that a maximizer of Φ_h subject to fixed mean m and variance s² has quantile Q(p) = m + s·h'(1−p)/‖h'‖₂ and value s‖h'‖₂. This formula is only meaningful when h' exists and is square-integrable on [0,1]. The paper, however, defines H in Section 2.2 as the set of concave functions of bounded variation with h(0)=h(1)=0, and Proposition 1 and Theorem 2 are stated for this full class. A legitimate member is h(p)=√p−p, which is concave, of bounded variation, and satisfies h(0)=h(1)=0, but h'(p)=1/(2√p)−1 is not in L²[0,1] because ∫₀¹ (h')² dp = ∞. For such h, Equation (3.8) contains a division by an infinite norm, and the claimed equilibrium variance σ_i(t)=λ_i(t)‖h'_i‖₂/(γ_i b²) is infinite. The variance terms λ_i²‖h'_i‖₂²/(2γ_i b²) in (3.15) are likewise undefined. Consequently, Proposition 1, Theorem 2, and Proposition 2 are not established for the full class H claimed; at best they hold when H is restricted to absolutely continuous h with h'∈L². This is a genuine regularity gap in the main equilibrium construction, not merely a stylistic issue.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies a two-agent non-zero-sum differential game in a continuous-time reinforcement-learning setting. Each agent has a Choquet-regularized, time-inconsistent mean-variance objective that includes a concern for relative wealth, and the market is incomplete with a stochastic factor. The paper derives a time-consistent Nash equilibrium by dynamic programming, obtains an explicit equilibrium in the Gaussian mean-return model, proves uniform convergence of an exact policy iteration, and proposes an actor-critic algorithm with numerical illustrations.","tokens_in":25080,"tokens_out":9513,"duration_ms":102273,"significance":"If the results are correct, the paper gives a useful extension of single-agent exploratory mean-variance RL to a competitive multi-agent setting, with explicit equilibrium policies up to coupled one-dimensional ODEs and a policy iteration that converges without a monotone improvement property. The paper is careful to state the time-inconsistency issue and the non-uniqueness caveat, and the equilibrium characterization is not circular: the target equilibrium solves the PDE system (3.14)-(3.15) independently of the iteration. The main contribution is the closed-form equilibrium plus the uniform-convergence analysis of the exact iteration; the numerical experiments are illustrative rather than a substitute for a convergence theorem for the sampled algorithm. The value of the paper is reduced by several proof gaps and an overly broad regularity class, but these are repairable.","major_comments":[{"comment":"The equilibrium construction relies on Lemma 1, whose quantile formula Q(p)=m + s h'(1-p)/||h'||_2 and value s||h'||_2 are defined only when h' exists and is square-integrable on [0,1]. The class H is introduced as concave functions of bounded variation with h(0)=h(1)=0, which admits h(p)=sqrt(p)-p; for this h, h'(p)=1/(2 sqrt(p))-1 is not in L^2, so ||h'||_2 is infinite and Eqs. (3.8), (3.11), (3.15), and (3.16) are undefined. The statements of Proposition 1, Theorem 2, Proposition 2, Corollary 1, and the policy-iteration theorems should be restricted to absolutely continuous, nonconstant h with h' in L^2 and h' not identically zero, or the proof of Lemma 1 under the stated hypotheses should be supplied. As written, the main equilibrium theorem is not established for the claimed class H; this is a load-bearing regularity gap, though it is fixable by adding the standard assumption.","section":"Section 2.2, Lemma 1, Eq. (3.8)"},{"comment":"Theorem 1 is the verification theorem that justifies all subsequent equilibrium derivations, but it is stated without proof. The text says the result is analogous to Björk et al. (2017), yet the present objective contains a Choquet regularizer and a two-agent game structure, so the verification argument is not literally a special case. A proof or a precise statement of the hypotheses under which the cited result applies should be provided, including the integrability and smoothness conditions needed for the generator computations in Eqs. (3.2)-(3.4).","section":"Section 3.1, Theorem 1"},{"comment":"The proof of Theorem 4 replaces the coefficients a^{in}_1(t), a^{in}_2(t) in the policy update (4.5) by the limiting coefficients a_i^1(t), a_i^2(t) inside the recursion for the means. However, (4.5) and (4.6) define a coupled, time-varying iteration: the a-coefficients themselves converge only as n tends to infinity, so the constant-coefficient contraction argument written in the proof does not apply directly. The theorem may be true, but the proof should address the joint convergence of (mu^n_1, mu^n_2, a^{1n}, a^{2n}) or provide a perturbation argument that controls the difference between the time-varying and the limiting recursions.","section":"Section 4, Theorem 4"},{"comment":"The displayed closed form for a_i^0(t) contains terms such as e^{2(iota+rho v)T} and e^{(iota+rho v)T} that do not involve (T-t), and it does not satisfy the terminal condition a_i^0(T)=0 implied by d_i(T,y)=0. As printed, Eq. (3.20) is not the solution of the ODE system (3.22). The equilibrium policy (3.17) does not use a_i^0, but the value function (3.18)-(3.19) is part of the claimed analytical solution, so this formula needs to be corrected.","section":"Proposition 2, Eq. (3.20)"},{"comment":"The abstract states that 'the proposed algorithm achieves uniform convergence,' but Theorems 3 and 4 analyze the exact policy iteration with full knowledge of the model and exact updates (4.5)-(4.6), not Algorithm 1, which uses function approximation, stochastic gradients, finite samples, and a smoothed functional gradient. The manuscript itself, near the end of Section 6, lists factors that can prevent the sampled algorithm from converging to the true equilibrium. The claim in the abstract should be qualified or a convergence theorem for Algorithm 1 under the stated idealizations should be proved.","section":"Abstract, Sections 4-5"}],"minor_comments":[{"comment":"The model assumes rho in [0,1] at the start of Section 2.1, but Table 1 uses rho = -0.93; either the assumption should be rho in [-1,1] or the numerical values should be changed to satisfy the stated condition.","section":"Section 2.1, Table 1"},{"comment":"The norm on R^2 is defined as ||x|| = max{x1,x2}, which is not a norm because it can be negative; it should be max{|x1|,|x2|} or the proof should use the max-norm on absolute values.","section":"Section 4, Theorem 4 proof"},{"comment":"The initialization block contains a repeated assignment 'xi(tn) <- xi' and the tilde notation for the perturbed wealth path is introduced inconsistently; also the sentence 'Use u to generate ui(tn) and ui(tn)' does not distinguish the unperturbed and perturbed actions. Please clean up the pseudocode.","section":"Algorithm 1"},{"comment":"The heading 'Gauss mean return model' should read 'Gaussian mean return model' for consistency with the text.","section":"Section 3.3 title"},{"comment":"The sign of the term involving gamma_i k_i b^2 sigma_j^2 appears to differ between Eq. (3.15) and the equation for b_i^0'(t) in (3.21); please check whether the signs are consistent after substitution.","section":"Eqs. (3.15) and (3.21)"},{"comment":"The label 'Uniform Distributiion' in the second panel contains a typo; it should be 'Uniform Distribution'.","section":"Figure 2"},{"comment":"The standing assumptions on h are stated twice in different ways: first 'Given a concave function h : [0,1] -> R of bounded variation with h(0)=h(1)=0' and then 'We denote the set of h : [0,1] -> R by H.' Please define H once, with the concavity and normalization requirements, and use that definition throughout.","section":"Section 2.2, definition of H"}],"recommendation":"major_revision","confidential_remarks":"The paper relies heavily on the same group's earlier work (Han et al. 2023, Guo et al. 2023) for the Choquet-regularizer machinery; the authors should make explicit which parts of the equilibrium characterization are new beyond those papers. The regularity gap and the missing proof of Theorem 1 are fixable, but the proof of Theorem 4 as written is not correct because it uses the limiting coefficients inside the iterative recursion. I would be willing to review a careful revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe short version: this is a genuine extension, not a repackaging. The authors construct the first non-zero-sum differential game with continuous-time RL and Choquet regularizers, and in the Gaussian mean return model they get a closed-form Nash equilibrium up to one-dimensional ODEs. If you work in continuous-time RL or time-inconsistent portfolio choice, this is worth knowing.\n\nWhat is actually new: the two-agent mean-variance objective with relative wealth concerns, the Choquet-based exploration, and the observation that policy iteration converges uniformly even though the usual policy improvement theorem fails. The single-agent policy iteration argument in Theorem 3 is sound, and the contraction argument for the means is clean. Decomposing the game into two independent learning problems is a sensible way to avoid multi-agent dimensionality issues.\n\nNow the soft spots, roughly in increasing order.\n\nMild: Theorem 1 is a verification theorem quoted without proof. Standard in this literature, but it should be sketched.\n\nMedium: the abstract claims the RL algorithm achieves uniform convergence. What is actually proved is uniform convergence of the idealized policy iteration under known parameters. The implemented algorithm uses parameterized critics and a smoothed gradient, and no convergence proof is given for that. The claim needs scaling back.\n\nMore serious: the main equilibrium construction overclaims the class H. Lemma 1 needs h' in L^2 for the quantile formula Q = m + s h'/||h'||_2, but H is defined in Section 2.2 as concave functions of bounded variation with h(0)=h(1)=0. That admits h(p)=sqrt(p)-p, for which ||h'||_2 diverges. For such h, Eqs. (3.8), (3.11), and (3.16) are undefined. So Proposition 1, Theorem 2, and Proposition 2 are not established for the full class H; they hold only for a restricted class, say absolutely continuous h with h' in L^2 and not identically zero. This must be stated.\n\nAlso serious: the proof of Theorem 4 is incomplete. The simultaneous iteration recursion for the means uses the limiting coefficient functions a_i(t) instead of the current iterate a_i^{(n)}(t). The fixed-point idea is fine once the a_i^{(n)} have converged, but the proof as written does not establish joint convergence. It needs a real argument.\n\nFinally, the numerical section cannot be checked as posted. Table 1 sets rho = -0.93 even though the model assumes rho in [0,1]. The normal variance and uniform support shown in Section 6 are inconsistent with Section 3 by factors involving lambda^2/(gamma^2 sigma^4), because the paper mixes standard deviation and variance. No code or error bars are provided. These are fixable but they undermine the numerical claims.\n\nBottom line: the equilibrium algebra and the policy iteration idea are legitimate and useful, but the paper as posted overclaims in three places. I would send it to peer review with a request for major revision: restrict H properly, repair the Theorem 4 proof, and redo the numerics. It deserves referee time. For a reading group, maybe; read it knowing the gaps.","headline":"Real extension of Choquet-regularized mean-variance RL to non-zero-sum games; the equilibrium algebra holds up, but the paper overclaims the class of h, leaves Theorem 4 incomplete, and the numerics are inconsistent.","tokens_in":25621,"tokens_out":5390,"would_cite":false,"duration_ms":57110,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91A23","91G10","93E35"],"pacs":[],"model":"deepseek-v4-flash","headline":"Two competing investors can learn a time-consistent Nash equilibrium in closed form under a mean-variance objective, even when market parameters are unknown.","keywords":["mean-variance","time-consistent Nash equilibrium","Choquet regularizer","reinforcement learning","non-zero-sum differential game","exploratory policy","Gaussian mean return model","policy iteration"],"falsifier":"Take the distortion $h(p)=\\sqrt p - p$, which is continuous, of bounded variation, and satisfies $h(0)=h(1)=0$. Its derivative $h'(p)=\\frac{1}{2\\sqrt p}-1$ has divergent $\\|h'\\|_2$ over $[0,1]$, so the right-hand side of the equilibrium quantile formula (3.16) is undefined; computing this single example would show that Proposition 2 and Theorem 2 require a smoothness condition on $h$ that the paper never states.","tokens_in":24482,"feed_emoji":"📈","tokens_out":3505,"duration_ms":40735,"temperature":0.7,"pith_summary":"This paper studies two agents who invest competitively, each caring both about their own terminal wealth and about beating the other agent's wealth. Each agent maximizes a mean-variance objective that is regularized by a Choquet penalty on exploration, in a market that can be incomplete and whose parameters may be unknown. The paper's central claim is that this non-zero-sum differential game has a time-consistent Nash equilibrium, and that under a Gaussian mean-return model the equilibrium is given explicitly: each agent's randomized investment policy is a distribution whose quantile function is a known affine function of the market state plus a regularizer-dependent shape. If correct, the result gives a practical reinforcement-learning algorithm with uniform convergence to the equilibrium, even though the standard policy-improvement theorem does not hold for time-inconsistent problems.","feed_headline":"Two rival investors get a closed-form Nash equilibrium","feed_subtitle":"Reinforcement learning finds each agent's optimal exploratory strategy despite mean-variance time inconsistency, with uniform convergence.","key_machinery":"The load-bearing object is the Choquet regularizer $\\Phi_h(\\Pi) = \\int h\\circ\\Pi([x,\\infty))\\,dx$, which measures the randomness of an exploratory distribution through a concave distortion function $h$. Its quantile representation $\\Phi_h(\\Pi)=\\int_0^1 Q_\\Pi(1-p)\\,dh(p)$ reduces the infinite-dimensional maximization over distributions to a mean-variance constraint, and the imported Lemma 1 states that the maximizer has quantile $m + s\\,h'(1-p)/\\|h'\\|_2$ with maximum value $s\\|h'\\|_2$. This identity lets the extended Hamilton-Jacobi-Bellman verification theorem for time-inconsistent control be solved explicitly, turning the two-agent game into a coupled pair of single-agent optimality conditions for the means while each variance is set independently.","core_discovery":"The authors establish that the game admits a specific time-consistent Nash equilibrium in which each agent's exploratory policy has a quantile function of the form $Q_{\\Pi_i^*(t)}(p) = \\frac{1}{1-k_1k_2}\\big[\\frac{\\theta(t,y)}{b(t,y)}(\\frac{1}{\\gamma_i}+\\frac{k_i}{\\gamma_j}) - \\frac{\\rho v(t,y)}{b(t,y)}(\\frac{\\partial d_i(t,y)}{\\partial y}+k_i\\frac{\\partial d_j(t,y)}{\\partial y})\\big] + \\frac{\\lambda_i(t)}{\\gamma_i b^2(t,y)}h_i'(1-p)$. Here $k_i$ is the agent's sensitivity to the opponent's wealth, $\\gamma_i$ is risk aversion, $\\lambda_i$ is the exploration weight, and $h_i$ is the Choquet distortion function that measures randomness. The equilibrium value function has the quadratic form $x + \\frac12 b_i^2(t)y^2 + b_i^1(t)y + b_i^0(t)$, with coefficient functions determined by one-dimensional ODEs. The paper also proves that a policy-iteration scheme, where the two agents update simultaneously, converges uniformly to this equilibrium despite the absence of monotone policy improvement.","pith_inferences":["The mean-variance decomposition suggests an immediate testable extension: if the same model is run with three or more agents, the quantile formula should generalize with the coupling matrix $k_i k_j$ replaced by a Perron-Frobenius condition, and convergence of policy iteration should depend on the spectral radius of that matrix; the paper does not state this extension.","The authors' convergence theorem relies on the Gaussian assumption for the market-state dynamics; an editorially plausible conjecture is that the same uniform convergence holds for any affine mean-reverting state process, because the ODE arguments depend only on linearity and not on Gaussianity beyond the state equation.","A practical test the paper leaves implicit: the Algorithm's empirical convergence should degrade predictably when the Choquet distortion $h$ is chosen so that $h'$ is only of bounded variation but not square-integrable, because the quantile formula itself becomes undefined; a numerical experiment with $h(p)=\\sqrt p - p$ would expose this boundary.","The separation of variance from opponent parameters implies that in a competitive setting an agent's exploration intensity is a private decision, not a strategic response; this could be used to design a decentralized multi-agent RL protocol that avoids the usual dimensionality explosion in centralized critics."],"forward_implications":["If the central claim is correct, a competitive two-agent mean-variance market can be solved analytically up to one-dimensional ODEs, so no numerical PDE or game-tree search is needed for the Gaussian mean-return model.","The variance of each agent's exploratory distribution depends only on that agent's own risk aversion, exploration weight, and Choquet regularizer, while the mean couples the two agents; this separation explains why the multi-agent problem decomposes into two independent learning tasks.","The uniform convergence of the policy-iteration scheme means that reinforcement learning can be applied to a time-inconsistent equilibrium without relying on monotone improvement, extending RL to a class of problems where policy improvement fails.","Since the quantile formula works for any admissible Choquet regularizer, agents can choose different exploration styles (e.g., Gaussian-like for one agent and Gini/uniform-like for the other) and still share the same equilibrium structure.","In the Black-Scholes complete-market limit, the equilibrium distribution collapses to a formula showing that higher sensitivity to an opponent pushes an agent to take riskier positions while higher risk aversion reduces the mean investment."],"supporting_citations":[{"why":"Supplies the crucial Lemma 1 that characterizes the quantile function of the maximally random distribution with fixed mean and variance, the engine behind the closed-form equilibrium.","marker":"Liu et al. (2020)"},{"why":"Introduces Choquet regularizers for continuous-time reinforcement learning and establishes the concavity and quantile representation used throughout the paper.","marker":"Han et al. (2023)"},{"why":"Provides the extended Hamilton-Jacobi-Bellman verification framework for time-inconsistent stochastic control, which the paper adapts to the two-agent game.","marker":"Björk et al. (2017)"},{"why":"Establishes the continuous-time exploratory control framework with distributional controls, which is the starting point for the randomized wealth process.","marker":"Wang et al. (2020a)"},{"why":"Gives the incomplete-market learning framework and the equilibrium mean-variance strategy whose policy-iteration method the paper extends to the non-zero-sum game.","marker":"Dai et al. (2023)"},{"why":"Provides the dynamic mean-variance formulation for incomplete markets and the time-consistency background that motivates the equilibrium concept.","marker":"Basak and Chabakauri (2010)"},{"why":"Supplies the Gaussian mean-return model with mean-reverting market price of risk, the concrete setting where the paper derives explicit ODEs and numerical parameters.","marker":"Wachter (2002)"}],"fun_headline_variants":["Rival investors' RL shows uniform convergence in MV game","Closed-form equilibrium for two-agent mean-variance RL","Non-zero-sum RL game: explicit Nash strategy","Time-consistent equilibrium despite mean-variance inconsistency","Two investors, one equilibrium: RL converges uniformly"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole closed-form equilibrium rests on a lemma saying that the most random distribution with given mean and variance has quantiles of the shape $m + s\\, h'(1-p)/\\|h'\\|_2$; this lemma requires $h'$ to exist and have finite squared integral, while the paper only assumes $h$ is continuous with bounded variation, so the formula is not defined for legitimate distortion functions such as $h(p)=\\sqrt p - p$ unless an extra regularity condition is added.","fun_headline_variants_meta":{"raw":{"variants":["Rival investors' RL shows uniform convergence in MV game","Closed-form equilibrium for two-agent mean-variance RL","Non-zero-sum RL game: explicit Nash strategy","Time-consistent equilibrium despite mean-variance inconsistency","Two investors, one equilibrium: RL converges uniformly"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000258,"raw_usage":{"total_tokens":1596,"prompt_tokens":970,"completion_tokens":626,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":586,"completion_tokens_details":{"reasoning_tokens":552}},"tokens_in":586,"tokens_out":626,"duration_ms":7379,"temperature":1.0,"reasoning_tokens":552,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T21:28:56.822791+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the distortion $h(p)=\\sqrt p - p$, which is continuous, of bounded variation, and satisfies $h(0)=h(1)=0$. Its derivative $h'(p)=\\frac{1}{2\\sqrt p}-1$ has divergent $\\|h'\\|_2$ over $[0,1]$, so the right-hand side of the equilibrium quantile formula (3.16) is undefined; computing this single example would show that Proposition 2 and Theorem 2 require a smoothness condition on $h$ that the paper never states.","supporting_citations":[{"cited_title":"and Wang, R","cited_arxiv_id":null,"evidence_quote":"Supplies the crucial Lemma 1 that characterizes the quantile function of the maximally random distribution with fixed mean and variance, the engine behind the closed-form equilibrium."},{"cited_title":"and Zhou, X","cited_arxiv_id":null,"evidence_quote":"Introduces Choquet regularizers for continuous-time reinforcement learning and establishes the concavity and quantile representation used throughout the paper."},{"cited_title":"and Jia, Y","cited_arxiv_id":null,"evidence_quote":"Gives the incomplete-market learning framework and the equilibrium mean-variance strategy whose policy-iteration method the paper extends to the non-zero-sum game."},{"cited_title":"and Chabakauri, G","cited_arxiv_id":null,"evidence_quote":"Provides the dynamic mean-variance formulation for incomplete markets and the time-consistency background that motivates the equilibrium concept."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Gaussian mean-return model with mean-reverting market price of risk, the concrete setting where the paper derives explicit ODEs and numerical parameters."}],"review_version":1}