{"id":"7fd80e5f-e07c-4468-8d5e-3fc94e524844","arxiv_id":"1908.03728","paper_version":5,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A Nash-type fictitious game between a real and an auxiliary player yields a new class of open-loop equilibrium policies for time-inconsistent linear-quadratic control, characterized by Riccati-like equations and applied to mean-variance portfolio selection.","lead":"This paper proposes a fictitious game between a real decision maker and an auxiliary player to handle time-inconsistent optimal control, and derives conditions for an open-loop equilibrium. The framework provides a tunable approach to balance globally optimal precommitted policies and time-consistent policies, with an application to multi-period mean-variance portfolio selection.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The advertised full characterization of the open-loop self-coordination control for the original LQ problem is not actually stated; the LQ specialization is deferred with 'Due to space limitations', so the paper's central claim rests on unstated theorems.","rationale":"Read in good faith: the paper's main technical object is Theorem 2.2 for the generalized nonzero-sum game, and the proof structure (convex variation, Riccati-like equations, pseudo-inverse selection) is standard and largely coherent. The mean-variance application is concrete, and the genericity result Theorem 3.3 gives a substantive existence statement. The reader's identified weakest assumption, that punishment parameters must satisfy range and semidefiniteness conditions, is a legitimate limitation but not a defect in a necessary-and-sufficient characterization: such theorems are conditional by design, and the generic-existence result mitigates the concern. The more load-bearing issue is that the paper's central advertised deliverable for Problem (LQ), the 'full characterization' of open-loop self-coordination control, is explicitly omitted from the manuscript. Since the LQ problem is the title subject and the stated motivation for introducing Problem (GLQ), the absence of the specialized theorem means the central claim cannot be verified from the text. This justifies the reader's CONDITIONAL verdict: the GLQ theory may be sound, but the promised LQ characterization must be supplied or the claims narrowed. No evidence of fraud or internal inconsistency is suggested.","tokens_in":46370,"tokens_out":15934,"duration_ms":167757,"concrete_test":"Instantiate Theorem 2.2 with the embedding matrices displayed after Theorem 2.5, writing the specialized conditions (2.17)-(2.18), the semidefiniteness conditions on O_{t,k}, O_{t,k}, O_{k,k}, and the all-u range conditions (2.21)-(2.22) solely in terms of the original LQ data and punishment parameters. If these specialized necessary and sufficient conditions cannot be stated explicitly, or if they do not reduce to finitely many matrix equations, then the advertised 'full characterization' is unsupported and should be narrowed in the abstract and introduction.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract and Section 1.3 promise that 'open-loop self-coordination control of Problem (LQ) is fully characterized,' and Theorem 2.2 is the advertised engine. However, after embedding Problem (LQ) into Problem (GLQ), the paper states only 'Combining (1.16) and (1.17), we can get results that are parallel to Theorem 2.2, Theorem 2.4 and Theorem 2.5 ... Due to space limitations, the results are not presented here.' No theorem in the manuscript gives necessary and sufficient conditions for existence of the open-loop self-coordination control in terms of the original data (A0, B0, Q0, R0, G0, and the punishment parameters). This is not a cosmetic gap: condition c) of Theorem 2.2 is a condition over all controls u, and it is not transparent whether the LQ specialization reduces to a finite set of matrix range/rank equations. The mean-variance example is worked out, but it is a single application, not the promised full characterization. A reader cannot check the central claimed result; the manuscript must either include the specialized theorems and proofs or narrow the advertised claim. The GLQ theory itself and the MV derivation appear internally consistent, so this is a completeness concern rather than a discovered contradiction.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a 'fictitious game' framework for time-inconsistent discrete-time stochastic linear-quadratic (LQ) control. A real player seeks a time-consistent policy while an auxiliary fictitious player seeks a precommitted optimal policy; the two play a Nash-type game with quadratic punishment terms, and the real player's equilibrium policy is called an open-loop self-coordination control. The paper embeds this problem in a general nonzero-sum stochastic LQ game (Problem (GLQ)) and derives necessary and sufficient conditions for an open-loop equilibrium via stationary conditions, convexity conditions, a set of Riccati-like equations, and linear equations (Theorems 2.1, 2.2, 2.4, 2.5). It then specializes these results to multi-period mean-variance portfolio selection (Theorems 3.1–3.3) and presents numerical examples comparing self-coordination controls with precommitted and time-consistent policies and with the planner-doer framework of [11].","tokens_in":46619,"tokens_out":5278,"duration_ms":55849,"significance":"Should the advertised claims be fully established, the paper would contribute a useful auxiliary-variable mechanism for interpolating between precommitted and time-consistent policies in a class of genuinely time-inconsistent stochastic LQ problems. The strength of the paper is its detailed derivation of the GLQ equilibrium characterization by discrete-time convex variation, including explicit formulas for the equilibrium controls and a clear separation of stationary and convexity conditions. The mean-variance application is nontrivial and the numerical section gives a concrete comparison with the earlier planner-doer framework of [11], including the observation that self-coordination controls can outperform both extreme policies at late instants. However, as discussed in Major Comment 1, the paper's central advertised contribution—the full characterization of the open-loop self-coordination control of the original Problem (LQ)—is not actually stated or proved in the manuscript, so the significance of the paper in its current form is substantially reduced.","major_comments":[{"comment":"The abstract and Section 1.3 promise that the open-loop self-coordination control of Problem (LQ) is fully characterized. However, after embedding Problem (LQ) into Problem (GLQ), the manuscript only says: 'Combining (1.16) and (1.17), we can get results that are parallel to Theorem 2.2, Theorem 2.4 and Theorem 2.5 ... Due to space limitations, the results are not presented here.' No theorem in the manuscript states necessary and sufficient conditions for existence of the open-loop self-coordination control in terms of the original data (A0, B0, Q0, R0, G0 and the punishment parameters). In particular, condition c) of Theorem 2.2 is a condition over all controls u, and it is not demonstrated that the LQ specialization reduces to finite-dimensional matrix range/rank conditions. This is load-bearing because a reader cannot check the paper's central claimed result. The manuscript must either include the specialized theorems (with proofs or at least precise statements) or narrow the advertised claim to the GLQ theory plus the mean-variance example.","section":"Section 2 (after equations (1.16)–(1.17))"},{"comment":"Theorem 3.1 states conditions only on Wk and ~Wk (equation (3.12)), yet an open-loop equilibrium also requires the semidefiniteness conditions and the invariance conditions of Theorem 2.2. The proof of Theorem 3.1 delegates the entire convexity half to 'Theorem 4.3 of [33]' after introducing an auxiliary static mean-field problem (5.29)–(5.30). As written, the reduction is too terse: it is not shown explicitly which matrices in (5.29)–(5.30) correspond to Ok, Ok, Ok and Mt,k, Mt,k, nor why all hypotheses of Theorem 4.3 of [33] are exactly satisfied at every step. This is a gap in the proof of a key application. Please expand the argument or state the correspondence explicitly.","section":"Section 3, Theorem 3.1 and its proof"}],"minor_comments":[{"comment":"The line 'Let Ot,kuk = ...' appears to be a typo for a decomposition of Ft,kuk, and the displayed construction of c1 and c2 in the contradiction argument has several apparent typographical errors (e.g., indicator functions with overlapping or missing cases). This part of the proof is very hard to follow and should be rewritten.","section":"Section 5.2, proof of Proposition 5.2"},{"comment":"The phrase 'for any initial pair (t,y)×R~n' is nonstandard; it should presumably read 'for any (t,y)∈T×R~n'.","section":"Theorem 2.4"},{"comment":"In the proof, 'For ζ0∈Ξk' should be 'For ζ0∈Ξc_k', since Ξk was defined as a set of matrices while Ξc_k was defined as the set of vectors satisfying Cov(Θk)ζ = EΘk.","section":"Section 3, proof of Theorem 3.2(iii)"},{"comment":"The caption of Figure 1 says 'k = 3' although the figure contains four subfigures for k = 0, 1, 2, 3; also Table 3 reports the minimizer for k = 0 as 99953 while the text and Figure 3 refer to μ = 99954. These values should be reconciled.","section":"Section 4, Figures and Tables"},{"comment":"Reference [31] is a previous arXiv version of this same manuscript (arXiv:1908.03728v3). The authors should indicate how the present version differs or should cite the published version if one exists; self-citation to an earlier arXiv version is confusing.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The core GLQ theory appears internally consistent, and the mean-variance derivation is mostly convincing, but the paper's headline claim about the LQ problem is effectively deferred to 'space limitations.' If the authors can add the missing specialized theorems and proofs, or explicitly reframe the contribution as the GLQ theory with an MV application, the paper would be publishable. I also note that the proof of Theorem 3.1 leans heavily on an external theorem from [33]; the editor may wish to ask for a self-contained verification or a clear statement that this external result is applied verbatim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know going in: the GLQ theory here is a genuine advance on [11], but the abstract's promise to \"fully characterize\" the original LQ problem is not kept. The authors derive a Nash-type fictitious game between a real time-consistent player and an auxiliary precommitted player, then say that combining (1.16)-(1.17) gives results parallel to Theorems 2.2-2.5 and defer them \"due to space limitations.\" So the central advertised result for Problem (LQ) never appears. That is a real completeness gap, since the motivation of the entire paper is the LQ problem, not the GLQ game as an end in itself.\n\nWhat is genuinely new: the fictitious-game construction itself, which differs structurally from the leader-follower planner-doer of [11] and is clearly described as such. The necessary-and-sufficient conditions in Theorem 2.2 — range conditions, semidefiniteness, and the invariance conditions (2.21)-(2.22) — are not in the prior cited literature, and the claim that the convexity characterization is the first for mean-field LQ is plausible. The mean-variance section is worked out in real detail: the zero-punishment limit correctly recovers the known open-loop time-consistent control, which is a good sanity check, and the \"sequently generic\" uniqueness result (Theorem 3.3) is a nice way of saying that the conditions hold for all but a measure-zero set of punishment intensities. The GLQ derivation itself looks sound to me; the convex-variation and Riccati steps follow the standard pattern.\n\nSoft spots, in proportion. The missing LQ theorems are the big one. A reader cannot check the paper's headline claim. This is fixable: either include the specialized theorems and proofs, or explicitly narrow the advertisement. The proof of Proposition 5.2 is also hard to verify; the indicator-function construction around (5.25)-(5.26) is intricate, and the manuscript has enough typos (\"nonsigular,\" duplicated \"Hence,\" notational slips) that I spent real time just parsing assertions. I do not think the argument is wrong, but it needs a cleaner rewrite. The dependence of existence on user-chosen punishment parameters is honestly handled: the N&S conditions make it explicit that some choices fail, and the authors do not hide that. Self-citations to [33] and [30] are prior results used as tools, not as inputs containing the conclusion, so the circularity burden is low. The numerical examples are illustrative rather than empirical validation, which is fine for a theory paper.\n\nBottom line: this deserves a serious referee, but with a demand for major revision — either deliver the LQ theorems or scale back the claim. The paper is valuable for people working on time-inconsistent LQ control and multi-period mean-variance portfolio selection; the GLQ framework and the MV application are solid enough to build on, and with the LQ statements included it would be a much stronger piece than the current version.","headline":"A real extension of the planner-doer idea with a solid GLQ core, but the advertised full characterization of the LQ self-coordination control is not actually stated in the manuscript.","tokens_in":47157,"tokens_out":4065,"would_cite":true,"duration_ms":47855,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["49N10","49N70","91A15","93E20"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proves that a fictitious two-player game yields an explicit open-loop equilibrium for time-inconsistent linear-quadratic control.","keywords":["time-inconsistent control","stochastic linear-quadratic problem","fictitious game","open-loop equilibrium","Riccati equations","mean-variance portfolio selection","Nash equilibrium","precommitted policy"],"falsifier":"Take a scalar one-step LQ instance with a punishment direction $\\Psi$ and intensity $\\mu$ for which (2.17) fails, for example by choosing $R_k$ and $B_k$ so that $\\tilde H_{t,k}(\\mathbb{E}_t X_k,\\mathbb{E}_t X_k)^\\top + h_{t,k}$ is not in $\\mathrm{Ran}(\\tilde W_{t,k})$, and verify numerically that no pair $(u^*,v^*)$ satisfies the stationary conditions (2.1); this would show the range condition is genuinely necessary, or conversely identify the true boundary of existence.","tokens_in":46109,"feed_emoji":"🎮","tokens_out":5461,"duration_ms":50942,"temperature":0.7,"pith_summary":"This paper tries to establish that a time-inconsistent linear-quadratic control problem can be solved by pitting two players against each other: a real player who wants a time-consistent policy and a fictitious player who wants a globally precommitted optimal policy. The two players are coupled through punishment terms chosen by the modeler, and their Nash equilibrium, called an open-loop self-coordination control, balances local and global optimality. The paper proves necessary and sufficient conditions for this equilibrium to exist, expressed through Riccati-like equations, range conditions, and semidefiniteness of certain matrices, and it shows the construction works for multi-period mean-variance portfolio selection.","feed_headline":"Two-player game yields a self-coordination control for time-inconsistent LQ","feed_subtitle":"A real player seeks a time-consistent policy; a fictitious player seeks a precommitted one. Their equilibrium is the solution.","key_machinery":"The central object is a fictitious Nash-type game between a real player and an auxiliary fictitious player, formalized as Problem (LQ)$^g$ and generalized to a nonzero-sum game Problem (GLQ). The work-horse is the set of Riccati-like equations (2.9)--(2.10), (2.14)--(2.15) whose solutions organize the stationary and convexity conditions, together with Moore-Penrose inverse range conditions that turn solvability of the equilibrium equations into checkable linear-algebra conditions on the state mean and fluctuation.","core_discovery":"The central claim is that a time-inconsistent stochastic LQ problem admits an open-loop self-coordination control exactly when a set of algebraic conditions holds: the range conditions (2.17)--(2.18) on the Moore-Penrose inverses of the assembled coefficient matrices, the semidefiniteness $O_{t,k}$, $\\bar O_{t,k}$, $O_{k,k} \\succeq 0$, and the invariance conditions (2.21)--(2.22). Under these conditions the equilibrium is selected by the explicit feedback formula (2.25), driven by the state trajectory (2.19). For the mean-variance portfolio selection case, existence is guaranteed for generically chosen punishment intensities, and when the punishment is zero the scheme recovers the open-loop time-consistent equilibrium control.","pith_inferences":["Beyond the paper, the same fictitious-game construction could be applied to closed-loop or feedback time-consistent policies, where the equilibrium would likely be characterized by coupled Riccati equations with a similar range-condition structure.","Beyond the paper, the dependence on the user-chosen punishment parameters suggests a design procedure: search over punishment matrices to optimize a secondary criterion at intermediate times, since the numerical example shows late-time objectives can beat both precommitted and time-consistent policies.","Beyond the paper, if the invariance conditions (2.21)--(2.22) fail, the equilibrium may still exist but the paper's feedback selection formula would not be valid; testing that failure on simple scalar counterexamples could delineate the true boundary of the framework."],"forward_implications":["For any LQ problem whose parameters satisfy the range and semidefiniteness conditions, the paper gives an explicit open-loop equilibrium control rather than a merely existential statement.","Adjusting the punishment direction and intensity yields a family of self-coordination controls interpolating between precommitted and time-consistent behavior; at zero punishment the scheme reduces to the time-consistent equilibrium.","The necessary-and-sufficient characterization provides a finite check for existence before computing the control.","In multi-period mean-variance portfolio selection, existence and uniqueness hold for generically chosen punishment intensities, with an explicit Riccati recursion for the equilibrium.","The nonzero-sum formulation makes the method applicable when one agent commits and the other behaves time-consistently."],"supporting_citations":[{"why":"Introduces the planner-doer self-coordination game that motivates this paper's fictitious-game construction.","marker":"[11]"},{"why":"Supplies the indefinite mean-field LQ convexity theory used to verify the semidefiniteness conditions in the mean-variance application.","marker":"[33]"},{"why":"Provides the Moore-Penrose inverse lemma used to turn stationary conditions into range conditions.","marker":"[2]"},{"why":"Establishes open-loop time-consistent mean-variance policies that the zero-punishment limit should recover.","marker":"[9]"},{"why":"Gives the multi-period mean-variance portfolio formulation that the paper embeds in its LQ framework.","marker":"[28]"},{"why":"Defines the time-consistent versus precommitted dichotomy that the fictitious game is designed to balance.","marker":"[37]"}],"fun_headline_variants":["Fictitious game solves time-inconsistent stochastic control","Self-coordination via a Nash fictitious game","Fictitious player yields time-consistent control policy","Nash fictitious game yields open-loop equilibrium"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The existence result stands on the requirement that, for the modeler-chosen punishment parameters, the range conditions (2.17)--(2.18) and the semidefiniteness of $O_{t,k}$, $\\bar O_{t,k}$, $O_{k,k}$ hold; if a chosen punishment fails these, no open-loop self-coordination control is guaranteed.","fun_headline_variants_meta":{"raw":{"variants":["Fictitious game solves time-inconsistent stochastic control","Self-coordination via a Nash fictitious game","Fictitious player yields time-consistent control policy","Nash fictitious game yields open-loop equilibrium"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000902,"raw_usage":{"total_tokens":3896,"prompt_tokens":971,"completion_tokens":2925,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":587,"completion_tokens_details":{"reasoning_tokens":2866}},"tokens_in":587,"tokens_out":2925,"duration_ms":22242,"temperature":1.0,"reasoning_tokens":2866,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:03:46.960984+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a scalar one-step LQ instance with a punishment direction $\\Psi$ and intensity $\\mu$ for which (2.17) fails, for example by choosing $R_k$ and $B_k$ so that $\\tilde H_{t,k}(\\mathbb{E}_t X_k,\\mathbb{E}_t X_k)^\\top + h_{t,k}$ is not in $\\mathrm{Ran}(\\tilde W_{t,k})$, and verify numerically that no pair $(u^*,v^*)$ satisfies the stationary conditions (2.1); this would show the range condition is genuinely necessary, or conversely identify the true boundary of existence.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the planner-doer self-coordination game that motivates this paper's fictitious-game construction."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the indefinite mean-field LQ convexity theory used to verify the semidefiniteness conditions in the mean-variance application."},{"cited_title":"Ait Rami, X","cited_arxiv_id":null,"evidence_quote":"Provides the Moore-Penrose inverse lemma used to turn stationary conditions into range conditions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes open-loop time-consistent mean-variance policies that the zero-punishment limit should recover."},{"cited_title":"Li and W.L","cited_arxiv_id":null,"evidence_quote":"Gives the multi-period mean-variance portfolio formulation that the paper embeds in its LQ framework."},{"cited_title":"Strotz, Myopia and inconsistency in dynamic utility maximization, The Review of Economic Studies, 1955-1956, vol.23, pp.165-180","cited_arxiv_id":null,"evidence_quote":"Defines the time-consistent versus precommitted dichotomy that the fictitious game is designed to balance."}],"review_version":1}