{"id":"7df66a0e-550d-4d3d-8488-089f2ae79d00","arxiv_id":"2509.03669","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"In a two-investor mean-variance Stackelberg game with asymmetric information, the leader's equilibrium randomized strategy is Gaussian and the follower's strategy depends linearly on the leader's observed trades.","lead":"Two investors with different information about a stock's true drift play a mean-variance portfolio game where one leads and the other follows, and the leader randomizes her trades to protect her information. The paper derives a closed-form equilibrium in which the leader's random trades follow a Gaussian distribution and the follower's best response is linear in the observed trades, then extends it to discrete-time sampling.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 4.2's ε-Stackelberg claim is not proven: the proof swaps an infinitesimal exploratory perturbation for a fixed grid-point deviation and asserts uniformity over all deviations without a bound.","rationale":"The reader's CONDITIONAL verdict is appropriate, but the most concrete defect is in Theorem 4.2, not (only) in the modeling assumptions of Remark 3.2. The theorem is a headline result, and its proof is a sketch with an unjustified uniformity step. The continuous-time construction in Theorems 3.1 and 4.1 appears internally consistent, and I do not see an algebraic error in the Gaussian policy derivation. However, the paper's advertised ε-Stackelberg equilibrium under discrete sampling depends on a limit interchange that is not justified. This is a correctness risk: if the supremum over one-point deviations is not controlled, the sampled equilibrium may fail for large-variance deviations even when the continuous-time equilibrium is valid. The motivational concern in Remark 3.2 about the follower not learning from the leader's trades is real but secondary; it weakens the economic narrative that randomization protects information, not the conditional mathematical statement. Thus the verdict should remain CONDITIONAL (UNCHANGED), pending a rigorous proof of Theorem 4.2 or an explicit, uniform ε bound.","tokens_in":26137,"tokens_out":9457,"duration_ms":105934,"concrete_test":"Analytical/numerical check: replace the proof's h-limit by the actual Definition 4.2 deviation. Fix a parameter set (e.g., μ1=0.2, μ2=0.05, σ=0.2, γ1=γ2=1, λ1=λ2=0.5, λ0=0.1, T=1). For a grid D with mesh |D|, compute J^D_1(Π^π)−J^D_1(Π*) for Π^π a Gaussian one-point deviation at t_i with variance v, using (3.9), (4.10), and (4.8). Check whether the supremum over v decays as |D|→0. Analytically, attempt to bound this supremum directly via (4.12) and determine whether the constant is independent of v; if the constant grows with v, the claimed uniform ε bound fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Theorem 4.2, the basis for the abstract's 'realistic case' claim, is not established by the proof. Definition 4.2 requires a supremum over arbitrary one-point deviations π at a fixed grid point t_i in the sampled objective J^D_1. The proof instead starts from the exploratory inequality (4.11), which controls an infinitesimal perturbation Π^{h,eπ} on [t,t+h] in eJ_1. These are different objects: a sampled deviation acts on an interval of length |D|, not on an infinitesimal h. The proof then asserts lim_{h↓0} ε(D2,Π^π) = ε(D2,Π*) and chooses a finer grid D so that ε(D1,Π*)−ε(D2,Π*) ≤ ε, but no argument shows this limit is uniform over π∈A1. Lemma 4.1's constant C depends on Π*, so the weak-convergence estimate (4.12) gives no uniform control over arbitrary π (e.g., π with very large variance). Consequently, the displayed inequality after (4.15) does not imply the required sup_π bound, and no concrete ε is produced. The continuous-time exploratory equilibrium of Theorem 4.1 may well be correct; the discrete-sampling ε-Stackelberg theorem is unsupported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies a two-investor mean-variance portfolio choice problem with asymmetric information and relative performance. Investor 1 knows the true drift μ; investor 2 observes stock prices and filters μ from the price path. The interaction is modeled as a Stackelberg game: the informed leader randomizes her strategy under an entropy-regularized objective to protect her informational advantage, while the follower responds to the sampled actions. Section 3 derives a semi-analytic intra-personal equilibrium for the follower under a filtration G_t that includes the full future path of the leader's sampled actions; Section 4 introduces exploratory dynamics and derives the leader's Gaussian equilibrium strategy. Section 4.3 claims that discrete sampling of the leader's Gaussian policy yields an ε-Stackelberg equilibrium, using an approximation result of Jia et al. (2025).","tokens_in":26527,"tokens_out":6650,"duration_ms":70096,"significance":"If the continuous-time results are correct, the paper makes a useful contribution: it extends entropy-regularized mean-variance portfolio selection to a Stackelberg game with asymmetric information, provides explicit Gaussian equilibrium strategies, and identifies a deterministic follower value function despite the leader's random actions. The verification proofs for Theorems 3.1 and 4.1 are written out in detail. However, the paper's advertised 'realistic case' result, Theorem 4.2, is not established by the proof, and the information assumptions in Section 3 weaken the economic motivation. The value of the paper depends on whether the discrete-sampling bridge can be repaired.","major_comments":[{"comment":"The proof does not establish the ε-intra-personal equilibrium of Definition 4.2. Definition 4.2 requires a supremum over arbitrary one-point deviations π at a fixed grid point t_i in the sampled objective J^D_1. The proof instead starts from inequality (4.11), which controls an infinitesimal exploratory perturbation Π^{h,eπ} on [t,t+h] in eJ_1, and then in (4.15) equates J^D_1(Π^π)+ε(D2,Π^π) with eJ_1(Π^{h,eπ}). These are different objects: a sampled deviation acts on an interval of length |D|, not an infinitesimal h, and the policy in J^D_1 is not the exploratory policy. The assertion lim_{h↓0} ε(D2,Π^π)=ε(D2,Π*) is not justified, and no argument shows this limit is uniform over π∈A1. Lemma 4.1's constant C depends on Π*, so the weak-convergence estimate (4.12) gives no control over arbitrary π (e.g., π with very large variance). Consequently the displayed inequality after (4.15) does n","section":"Section 4.3, Theorem 4.2 proof"},{"comment":"The follower's information structure is load-bearing but economically questionable. Definition 3.1 defines A2 as progressively measurable with respect to G_t = F^S_t ⊗ F^ξ_T, so the follower's strategy u2(t) may depend on the entire future path of the leader's sampled actions. This is not a causal observation structure and is never justified as a limit of discrete-time updating. Remark 3.2 simultaneously bars the follower from using the observed actions u1 to update his belief P(t) about μ. Under these assumptions, the leader's randomization does not function as information protection: the follower cannot learn from trades by fiat, so randomization only injects entropy into the leader's objective. The equilibrium formulas (3.9) and (3.10) are derived under this filtration and this no-learning assumption, so they do not address the information-leakage concern that motivates the model. The","section":"Section 3, Definition 3.1 and Remark 3.2"},{"comment":"The approximation result that underpins Theorem 4.2 is not actually verified. The proof of Lemma 4.1 asserts that the coefficients of the exploratory dynamics (4.2), namely ebt(θ(Pt))-r, ebt σ, and σ eσt, belong to C^4_p with bounded derivatives, based on the regularity of a1 and a2 derived in Theorem 4.1. However, the cited regularity (Lemma 3.3 of Huang and Sun 2023) only provides C^{1,2} solutions that are bounded on [0,T]×(0,1); no C^4_p estimates are given. Moreover, the exploratory dynamics (4.2) are derived in Section 4.1 only through an LLN/CLT heuristic, with no rigorous statement of the convergence from the sampled dynamics (3.1). Since (4.12)–(4.13) are the sole bridge from exploratory to sampled equilibrium, the missing verification of the conditions of Jia et al. (2025) is a substantive gap.","section":"Lemma 4.1 and Section 4.1"}],"minor_comments":[{"comment":"The entropy term in the A1 PDE is written as λ0/2 log(2πλ0/(γ1χ^2)). Given the variance in (4.10), λ0/(γ1σ^2χ^2), Shannon's differential entropy is 1/2 log(2πeλ0/(γ1σ^2χ^2)). The expression in (A.30) correctly uses σ^2 in the denominator, so (A.23) and Theorem 4.1 are internally inconsistent. The value function A1 should be corrected.","section":"Theorem 4.1 and Appendix A.2, Eq. (A.23)"},{"comment":"The definition of the one-point perturbation Π^π_t = π(t) for t = t_i is imprecise for continuous-time policies: a deviation at a single instant has no effect on the sampled dynamics over a positive-length interval. The perturbation should be defined on the interval [t_i, t_{i+1}) or in the discrete-time sampling protocol. This imprecision contributes to the gap in the proof of Theorem 4.2.","section":"Definition 4.2"},{"comment":"In (3.9), the notation u1 appears in the third term without an explicit argument; since the follower's strategy is evaluated at time t, it would be clearer to write u1(δ(t)) or u1(t_i) to match the sampled dynamics in (3.1).","section":"Equation (3.9)"},{"comment":"Minor typos: 'equalibrium' in the proof of Theorem 4.1; 'we show that p1(·) can be characterized' is a slightly awkward construction; the paper would benefit from a table of notation for θ, β, Γ, χ, and l, which appear in several places.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The main question for the editor is whether Theorem 4.2 can be repaired within the scope of a revision. The proof as written does not merely omit an estimate; it conflates a sampled one-point deviation with an exploratory interval perturbation and asserts uniformity without a bound. If the authors cannot supply a rigorous argument, the abstract's central claim about the 'realistic case' should be withdrawn or substantially weakened. Additionally, the paper leans heavily on the authors' own prior work (Huang and Sun 2023) for the filtering SDE, PDE existence-uniqueness, and even the structure of the follower's strategy; while the cited results are standard, the dependence is extensive enough that the paper should state the needed theorems explicitly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take. The main result — the Gaussian equilibrium for the leader and the linear follower response that makes the follower's value function deterministic — is real and worth attention. The verification theorems are written out, the algebra is consistent, and the reduction of a random-field objective to a deterministic value function is a nice byproduct of the linear structure. The paper also credits the RL lineage and Huang-Sun for the filtering and PDE inputs; that's honest and appropriate.\n\nThe soft spots. First, Theorem 4.2, which carries the abstract's \"realistic case\" claim, is not established. The proof borrows the infinitesimal perturbation inequality (4.11) and applies it to a fixed grid-point deviation in Definition 4.2. Those are different objects. The limit argument in (4.14)-(4.15) needs uniformity over π ∈ A1, and nothing provides it: the constant in Lemma 4.1 depends on Π*, and the assertion that ε(D2,Π^π) converges to ε(D2,Π*) is not justified. As written, the theorem is unsupported. The exploratory equilibrium in Theorem 4.1 may well be correct, but the paper should either prove the discrete-sampling version carefully, with a concrete ε, or drop the claim.\n\nSecond, the model's two information assumptions are strong and explicitly acknowledged: the follower does not update his belief from the leader's trades (Remark 3.2), and he conditions on the full future path of the leader's actions (Remark 3.1). The first undercuts the stated motivation that randomization protects information; the second is a heavy cognitive assumption. These are limitations, not fatal flaws, but they deserve more discussion than they get.\n\nThird, there is a small internal typo in the entropy constant in Theorem 4.1: the log term is missing σ² in the denominator compared with the entropy calculation in the proof. Easy to fix, but it should be corrected.\n\nThe exploratory dynamics in Section 4.1 rest on an LLN/CLT heuristic rather than a rigorous derivation; that's standard in this literature and acceptable for a first pass, but it is worth flagging.\n\nMy recommendation: the core continuous-time equilibrium is a solid contribution and deserves a serious referee. But Theorem 4.2 needs real work before the claims are taken as established. If I were the editor, I'd send it out, and ask the referee to focus on Theorem 4.2 and the information assumptions.","headline":"The continuous-time exploratory equilibrium is a genuine, mostly verified contribution; the discrete-sampling ε-Stackelberg theorem is not proven as written.","tokens_in":26920,"tokens_out":3669,"would_cite":true,"duration_ms":37046,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91G15","91A65","93E11"],"pacs":[],"model":"deepseek-v4-flash","headline":"In a two-investor portfolio game, the informed leader's optimal strategy is Gaussian randomization.","keywords":["asymmetric information","mean-variance portfolio selection","Stackelberg game","randomized strategy","intra-personal equilibrium","entropy regularization","relative performance","Gaussian policy"],"falsifier":"Re-solve the follower's problem with the posterior P(t) computed from both the stock path and the history of the leader's sampled actions u₁(δ(t)), instead of from the stock path alone; if the resulting best response differs from (3.9), the claimed Stackelberg equilibrium does not survive inference. A simulation can make this concrete: generate leader actions from (4.10), let an econometrician estimate μ from (S, u₁), and check whether the follower's optimal position still matches (3.9).","tokens_in":26017,"feed_emoji":"📈","tokens_out":9429,"duration_ms":89412,"temperature":0.7,"pith_summary":"The paper claims that when a fully informed investor acts as a Stackelberg leader and a partially informed investor as follower, both caring about relative performance, an equilibrium exists in which the leader deliberately randomizes her trades to hide information. In that equilibrium the leader samples her portfolio positions from a Gaussian distribution, and the follower's best response is a linear function of the leader's observed positions. Because wealth dynamics are linear, the follower's optimal combined portfolio is deterministic, so his equilibrium value function does not depend on the leader's realized random actions. The paper proves this as an exact intra-personal equilibrium under idealized continuous action sampling, and as a time-consistent ε-Stackelberg equilibrium when actions are sampled on a discrete grid. If correct, this gives a tractable description of information-hiding in financial markets: camouflage is Gaussian, its variance is set by an entropy weight and risk aversion, and its mean is unaffected by that weight.","feed_headline":"Informed leader's best camouflage: Gaussian trading","feed_subtitle":"A two-investor mean-variance model finds a Stackelberg equilibrium where the leader randomizes trades and the follower cancels the noise.","key_machinery":"The argument rests on three pieces. (1) A change of variables turns the pair's relative wealth into a single process Z driven by a combined portfolio u = (1-λ/2)u₂ - (λ/2)u₁; linearity means the follower can pick u* independent of the leader's realized action, and this is why the follower's value function becomes deterministic. (2) An extended Hamilton-Jacobi-Bellman system—two coupled PDEs for the value function and an auxiliary expected-terminal-wealth function—handles the time inconsistency of mean-variance preferences and produces the semi-analytic formulas. (3) The leader's problem is cast in exploratory dynamics, where sampled random actions are replaced by equivalent Brownian noise; m","core_discovery":"On the paper's own terms, the central discovery is the explicit equilibrium profile (Π*, u₂*). The leader's randomized policy is Gaussian with mean (θ(p)-r)/(σ²l) - β(p)/(χσ)(∂_p a₁ + (1-χ)∂_p a₂) and constant variance λ0/(γ₁σ²χ²); the follower responds linearly, u₂* = [(θ(p)-r)/(σ²γ₂) - β(p)∂_p a₂/σ]/(1-λ₂/2) + u₁λ₂/(2-λ₂). The follower's linear response cancels the leader's sampled noise in the combined portfolio u*, so the follower's equilibrium value function is the deterministic affine function (1-λ₂/2)x₂ - (λ₂/2)x₁ + A₂(t,p). Under continuous sampling this is an intra-personal Stackelberg equilibrium; under discrete sampling it becomes an ε-Stackelberg equilibrium with ε governed by th","pith_inferences":["If the follower were allowed to update his posterior with the leader's observed trades, the Gaussian equilibrium would not be expected to survive; a natural extension would replace pure randomization with a signalling equilibrium, where the leader's trades convey information and the follower's response changes.","The same linear-cancellation mechanism suggests a general result: in any linear wealth dynamics with relative performance, a follower's best response can neutralize a leader's action noise, so informed-player randomization may be welfare-neutral for followers.","The exploratory-to-sampled approximation with linear error in grid mesh gives a practical prescription: the leader can quantify the trade-off between sampling frequency and equilibrium accuracy, choosing grid size from the desired ε.","The entropy-regularized exploratory method could be re-run for a Nash version of the same asymmetric-information game; the expectation is a Gaussian equilibrium again, but with coupled first-order conditions instead of the leader's Stackelberg optimality."],"forward_implications":["The leader's equilibrium randomization has constant variance λ0/(γ₁σ²χ²): more randomization weight or less risk aversion means noisier camouflage, while the mean trade depends only on fundamentals and the posterior p.","The follower's value is deterministic despite the leader's random actions, because his linear response cancels exactly the leader's action noise; random leadership does not impose welfare risk on a rational relative-performance follower.","On a discrete trading grid, sampling from the Gaussian policy delivers a time-consistent ε-Stackelberg equilibrium, with ε shrinking as the grid is refined; implementation only requires sampling at sufficiently frequent times.","As the entropy weight λ0 goes to zero, the Gaussian policy degenerates to the deterministic strategy of a Stackelberg leader under partial information, recovering known single-investor and full-information benchmarks.","The leader's Gaussian mean is independent of λ0, so the choice of how much to randomize separates from the choice of average position—randomization can be tuned without distorting expected trades."],"supporting_citations":[{"why":"Supplies the nonlinear filtering equation for the posterior P and the partial-information mean-variance equilibrium formulas that the follower's and leader's strategies extend.","marker":"Huang and Sun (2023)"},{"why":"Provides the extended HJB system used to define and verify intra-personal equilibria for time-inconsistent mean-variance objectives.","marker":"Björk et al. (2017)"},{"why":"Establishes the entropy-regularized mean-variance formulation and its Gaussian optimal strategy, the template for the leader's exploratory problem.","marker":"Wang and Zhou (2020)"},{"why":"Introduces exploratory dynamics for entropy-regularized stochastic control, from which the leader's exploratory wealth process is taken.","marker":"Wang et al. (2020)"},{"why":"Derives an equilibrium Gaussian policy for entropy-regularized mean-variance control, the closest single-agent antecedent for the leader's equilibrium.","marker":"Dai et al. (2023)"},{"why":"Provides the discretization estimate that converts the exploratory equilibrium into an ε-equilibrium under sampled actions.","marker":"Jia et al. (2025)"},{"why":"Supplies the pathwise/random-field formulation that the follower's conditional mean-variance problem uses before linearity reduces it to a deterministic value function.","marker":"Buckdahn and Ma (2007)"},{"why":"Sets the relative-performance objective (own wealth minus a share of the average wealth) that defines both investors' payoffs.","marker":"Espinosa and Touzi (2015)"}],"fun_headline_variants":["Leader hides info with Gaussian trades, follower cancels noise","Randomized leader strategy yields explicit Stackelberg equilibrium","Follower's linear response cancels leader's noise in mean-variance game","Two investors, one informed: leader randomizes, follower cancels noise","Closed-form equilibrium in asymmetric-information Stackelberg game"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise is that the follower never updates his estimate of the stock's drift from the leader's observed trades, justified only by practical infeasibility; if a statistical follower did learn from trades, the equilibrium formulas would no longer be guaranteed.","fun_headline_variants_meta":{"raw":{"variants":["Leader hides info with Gaussian trades, follower cancels noise","Randomized leader strategy yields explicit Stackelberg equilibrium","Follower's linear response cancels leader's noise in mean-variance game","Two investors, one informed: leader randomizes, follower cancels noise","Closed-form equilibrium in asymmetric-information Stackelberg game"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001087,"raw_usage":{"total_tokens":4449,"prompt_tokens":880,"completion_tokens":3569,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":624,"completion_tokens_details":{"reasoning_tokens":3482}},"tokens_in":624,"tokens_out":3569,"duration_ms":28467,"temperature":1.0,"reasoning_tokens":3482,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T10:46:43.247508+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-solve the follower's problem with the posterior P(t) computed from both the stock path and the history of the leader's sampled actions u₁(δ(t)), instead of from the stock path alone; if the resulting best response differs from (3.9), the claimed Stackelberg equilibrium does not survive inference. A simulation can make this concrete: generate leader actions from (4.10), let an econometrician estimate μ from (S, u₁), and check whether the follower's optimal position still matches (3.9).","supporting_citations":[],"review_version":1}