{"id":"3c62b668-f339-4b38-b5c8-9215f31385d9","arxiv_id":"2502.06356","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":1.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of the randomization method proving that the value of an optimal control problem equals the value of a randomized problem and is represented by a constrained BSDE, with a complete tour of applications.","lead":"This survey explains the randomization method, which represents the value of a stochastic optimal control problem through a backward stochastic differential equation, and gives detailed proofs of the main equivalence and representation results. It is a self-contained entry point for researchers who want to learn, apply, or extend a method that has spread across many control settings.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Proof of υ0 ≤ υR0 rests on unproved Lemma 5.6; appendix also contains a false bound after (A.7).","rationale":"The reader correctly identified Lemma 5.6 as the externally quoted load-bearing premise in the proof of υ0 ≤ υR0. My stress-test confirms that this is the main structural dependency: without density of deterministic-grid simple controls in AW, Proposition A.2 cannot produce the approximating randomized trajectory needed to pass from an arbitrary control α to randomized controls. The paper explicitly declines to prove this lemma, making the proof of the value equality incomplete as written. In addition, I found a concrete internal error in the appendix: the estimate leading to (A.6) bounds cumulative sums of the exponential parameters by their total sum, which is false for N>1; the left side has multiplicities N−j. This is not fatal because the corrected O(N/m) bound still tends to zero and the choice of m can be made after fixing N, but it means the appendix does not currently contain a valid proof of (A.6). Both issues are in the same step of the central theorem, so they reinforce the reader's CONDITIONAL verdict rather than overturning it. The survey is honest that it contains no new results, and the BSDE part (Theorems 6.2 and 6.3) is essentially self-contained once the randomized problem is accepted; the main gap is in the value-equivalence proof. I therefore recommend keeping the CONDITIONAL verdict, with conditions being to supply a proof or precise reference for Lemma 5.6 and to correct the appendix estimate.","tokens_in":69176,"tokens_out":28114,"duration_ms":231679,"concrete_test":"Independently prove Lemma 5.6 for FW-progressive controls under assumption (A1) with the deterministic-grid finitely-valued step controls used in Proposition A.2; if the approximation fails for non-right-continuous progressive processes, the inequality υ0 ≤ υR0 and hence Theorem 4.8 is not established. In the same re-derivation, recompute the bound after Eq. (A.7) with the true cumulative sums and confirm that a corrected N-dependent bound still allows choosing m so that tilde-rho(α, I) < δ.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The proof of υ0 ≤ υR0 in Section 5.2 depends on Proposition A.2, which constructs, for every admissible control α, a step process I^k within ρ-distance 1/k of α and with compensator density bounded above and away from zero. Two fragile points gate this construction. First, Proposition A.2 reduces α to a deterministic-grid finitely-valued step process using Lemma 5.6, quoted from Krylov [49] with the proof explicitly omitted; if Lemma 5.6 does not hold for the class AW of FW-progressive controls under (A1), the inequality υ0 ≤ υR0 has no proof and the value equality is unsupported. Second, the appendix's own proof of (A.6) contains a false estimate: immediately after Eq. (A.7) the cumulative sums ∑_{n=0}^{N−1}(λ_{1m}^{−1}+...+λ_{nm}^{−1}) are incorrectly bounded by ∑_{n≥1} λ_{nm}^{−1}=1/m. The left side contains N−j copies of λ_{jm}^{−1}, so the correct bound is O(N/m), not O(1/m). Because N is fixed before choosing m, the conclusion (A.6) is repairable by taking m large, but the displayed proof is invalid as written. Both issues sit at the load-bearing step of the central theorem.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a survey of the randomization method for stochastic optimal control. It studies a finite-horizon controlled diffusion with Borel control space A under assumptions (A1), constructs an auxiliary randomized control problem driven by an independent marked Poisson process with intensity measure λ, proves the equality of the two values (Theorem 4.8), and shows that the common value is represented by the first component at time 0 of the unique minimal solution to a constrained BSDE with jumps (Theorems 6.2 and 6.3). The proof of the inequality υ0^R ≤ υ0 uses a point-process construction and a Girsanov argument in Section 5.1, while the reverse inequality in Section 5.2 uses a density lemma imported from Krylov, a stability lemma, and a construction deferred to the Appendix. The final sections survey applications to switching, stopping, impulse control, partial observation, McKean-Vlasov equations, infinite-dimensional systems, and other problems, and list open research directions. The manuscript explicitly states that it contains no new results and positions itself as a synthetic exposition.","tokens_in":69388,"tokens_out":6854,"duration_ms":61346,"significance":"If the results are correct, the paper provides a useful rigorous survey of a method that yields BSDE representations for fully nonlinear Hamilton-Jacobi-Bellman equations and for a broad class of stochastic control problems, including non-Markovian and partially observed cases. A particular strength is the detailed treatment of the basic case, including the marked-point-process constructions in Section 3 and the Appendix, which collects material not previously published in this form. The paper is also honest about its scope: it explicitly identifies where proofs are omitted or deferred. However, because the central equality of values rests on load-bearing arguments that are either imported without proof or contain an invalid estimate, the manuscript cannot currently be used as a fully reliable self-contained reference and needs revision.","major_comments":[{"comment":"Lemma 5.6 is the density result quoted from Krylov [49] that AW_0 is dense in AW with respect to the metric tilde-rho. The paper explicitly says the proof will not be reported. This lemma is load-bearing: it is used in Proposition A.2, which in turn underpins the proof of the inequality υ0 ≤ υ0^R in Section 5.2. If the lemma does not apply to the class of FW-progressive controls under assumption (A1), the proof of Theorem 4.8 has a gap. The author should either provide a proof, or state a precise version of the lemma with all hypotheses and a full reference that verifies its applicability to this setting. For a survey that aims at a complete exposition, an unproved external lemma at this critical juncture is a significant omission.","section":"Section 5.2, Lemma 5.6"},{"comment":"The displayed estimate following (A.7) is false. The cumulative sums ∑_{n=0}^{N−1}(λ_{1m}^{−1}+...+λ_{nm}^{−1}) contain N−j copies of λ_{jm}^{−1}, so the correct upper bound is of order N/m, not the claimed ∑_{n≥1} λ_{nm}^{−1} = 1/m. As written, the proof of the claim tilde-rho(bar-alpha, alpha-hat^m) → 0 in (A.6) is invalid. The conclusion is repairable: N is fixed before choosing m, so taking m large with N/m < δ/3 would suffice, but the displayed argument must be corrected. This is not a cosmetic issue, since Proposition A.2 is essential to the proof of υ0 ≤ υ0^R.","section":"Appendix, after Eq. (A.7)"},{"comment":"Lemma 5.7 is the stability result used to conclude that J^R(ν^k) → J(α) from tilde-rho(I-hat^k, alpha-hat) → 0. Its proof is omitted with the remark that it is entirely analogous to Lemma 5.5. However, the lemma is stated with different filtrations G^k and control processes γ^k that are only G^k-progressive, and it is applied to controls that are adapted to different enlarged filtrations. Because this is a load-bearing step in the final convergence argument, a proof or a precise reference should be supplied rather than left as an analogy.","section":"Section 5.2, Lemma 5.7"}],"minor_comments":[{"comment":"The statement refers to the 'partially observed control problem', but the problem formulated in Section 4.2 is fully observed; this should be 'original control problem' or 'classical control problem'.","section":"Section 4.5, Theorem 4.8"},{"comment":"In the product term of κ^{hat-nu}, the symbol appears as ν_{hat-S_n}(hat-eta_n) without the hat; it should be hat-nu_{hat-S_n}(hat-eta_n).","section":"Section 4.3, Eq. (4.13)"},{"comment":"In the final displayed formula, the conditional survival probability is written as an integral without the exponential factor; the right-hand side should be exp(−∫∫ ν λ da ds), consistent with the earlier proof and with (3.9).","section":"Section 3.2, Proposition 3.10"},{"comment":"There is a typo 'Fron now on' that should read 'From now on'. Also, in the proof of Lemma A.4 the expression 'H_t = F_t ∨ F^κ_t' should be 'H_t = hat-G_t ∨ F^κ_t' to match the definition of H.","section":"Appendix, proof of Proposition A.2"},{"comment":"In the display of the convergence of the reward functional, the running cost term should read E^Q[∫_0^T f(...) dt + g(...)]; the 'dt' is missing in the displayed formula.","section":"Section 5.2, Lemma 5.7"}],"recommendation":"major_revision","confidential_remarks":"The paper is a survey by an author who is a main contributor to the randomization method, and the bibliography is appropriately thorough. The manuscript explicitly disclaims new results, so novelty is not the issue. The main concern is the reliability of the central proof: the false estimate in the Appendix is a clear, fixable error, and the reliance on an unproved external lemma at a load-bearing point weakens the survey's claim to completeness. I would suggest revisions that correct the Appendix estimate and either prove Lemma 5.6 in a form sufficient for the present setting or give a precise statement with a complete reference. The paper is not beyond repair, but it is not ready for acceptance in its current form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: this is a solid survey that does exactly what it says—no new results, but full proofs of the core equivalence and BSDE representation theorems, plus an appendix proof of a result that was previously only in a preprint. It deserves referee time and, after a minor revision, publication. It will be a standard entry point for people who want to learn the method.\n\nWhat's genuinely useful: the exposition in Sections 4–6 is careful and self-contained up to the quoted density lemma. The survey gathers a scattered literature into one place, and the Appendix finally gives the details of Proposition A.2 (the approximation of an admissible control by a step process with bounded, bounded-away-from-zero compensator density). That alone is a service. The discussion of limitations in Section 2.4 is honest, and the bibliography is comprehensive.\n\nThe soft spots are both in the proof of Theorem 4.8. First, Lemma 5.6—the density of simple step controls in the class of progressive controls under the metric ρ̃—is imported from Krylov [49] without proof. It is genuinely load-bearing for the inequality υ0 ≤ υR0, because Proposition A.2 relies on it to reduce an arbitrary admissible control to a finitely-valued, deterministic-grid step process. This is a dependency, not a fatal gap, but for a paper claiming to be a complete exposition it would be better to include a proof or at least a precise statement of the hypotheses.\n\nSecond, the appendix contains a wrong estimate. After (A.7), the cumulative sums ∑_{n=0}^{N−1}(λ_{1m}^{−1}+...+λ_{nm}^{−1}) are bounded as if each cumulative sum were just λ_{nm}^{−1}; the actual sum is O(N/m), not O(1/m). The conclusion (A.6) is still repairable—fix N and choose m large enough—but the displayed proof is invalid as written. A referee should ask for that line to be corrected.\n\nThe central argument holds up in the sense that both issues are fixable and the underlying theorems are already published. This paper is for readers who want the method in one place, and for researchers who need the appendix proof. I'd send it to review and accept after minor revision.","headline":"A useful, honest survey of the randomization method; the central theorem is credible, but the proof has two fixable gaps in the appendix.","tokens_in":69929,"tokens_out":2976,"would_cite":true,"duration_ms":25510,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60H10","93E20"],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey shows that the value of a general stochastic optimal control problem is reproduced by a randomized control problem and is given by the time-zero component of the unique minimal solution to a constrained backward SDE.","keywords":["randomization of controls","backward stochastic differential equations","constrained BSDEs","stochastic optimal control","Hamilton-Jacobi-Bellman equations","fully nonlinear PDEs","minimal solutions","Girsanov transformation"],"falsifier":"A concrete way to test the central claim is to check the omitted density lemma for the stated class of progressive controls on an arbitrary Borel action space: find an admissible control, for instance one with irregular path dependence on the Brownian motion, that cannot be approximated in dt-times-probability measure by any sequence of finite-valued controls measurable at deterministic times. If such a control exists under the paper's assumptions, the equality of the two values would lack its proof and could fail.","tokens_in":68936,"feed_emoji":"🎲","tokens_out":4544,"duration_ms":44616,"temperature":0.7,"pith_summary":"This survey establishes that a large class of stochastic optimal control problems can be solved by first randomizing the control, replacing the chosen action process with a piecewise-constant process driven by an independent Poisson random measure, and then letting the control act through Girsanov changes of probability. Its main theorem states that the randomized problem has exactly the same value as the original one, and that this common value is the time-zero component of the unique minimal solution of a constrained backward stochastic differential equation. If correct, this provides a probabilistic representation for general control problems and for fully nonlinear Hamilton-Jacobi-Bellman equations, extending the nonlinear Feynman-Kac formula beyond the semilinear setting.","feed_headline":"Control values match a constrained BSDE at time zero","feed_subtitle":"Randomizing the control with a Poisson clock preserves the optimum and yields a backward SDE representation.","key_machinery":"The machinery is the randomized control problem: the original control process is replaced by a piecewise-constant process built from an independent Poisson random measure on the action space, and admissible controls become bounded positive intensity fields that reshape the law of the randomizing process via Girsanov transformation while leaving the Brownian motion untouched. The value of this auxiliary problem is then shown to be the first component of the unique minimal solution of a constrained BSDE in which one martingale integrand is required to be nonpositive, with minimality supplying uniqueness and a penalization procedure supplying existence.","core_discovery":"The central claim is that, under standard Lipschitz and growth assumptions on the coefficients plus a full-support intensity measure on the action space, the value of the original control problem equals the value of the randomized problem, and this common value is represented by the first component at time zero of the unique minimal solution to the constrained BSDE with jumps. The equality is proved by showing the randomized value is no larger than the original value through a pathwise law-matching argument, and no smaller through a density approximation of arbitrary admissible controls by randomized step processes. The survey therefore presents a complete route from a fully nonlinear stochastic optimization problem to a well-posed backward equation, and identifies the constrained BSDE as the common object representing both the control value and the solution of the associated Hamilton-Jacobi-Bellman equation.","pith_inferences":["Beyond the paper: the identity suggests a practical duality certificate for optimal control, where approximate minimal BSDE solutions from one side and admissible controls from the other bracket the true value.","Beyond the paper: because the method adds a Poisson clock independent of the original noise, it may yield Monte Carlo schemes for the value function even when the value function has low regularity, replacing PDE solvers with BSDE simulation.","Beyond the paper: applying the same randomization to deterministic optimal control would produce a genuinely stochastic BSDE whose first component at time zero is the value of a deterministic problem, potentially enabling stochastic numerical methods for deterministic control."],"forward_implications":["The value of a general stochastic control problem can be computed as the time-zero component of a unique minimal constrained BSDE solution, with no nondegeneracy assumption on the diffusion coefficient.","Fully nonlinear Hamilton-Jacobi-Bellman equations obtain a probabilistic representation analogous to the Feynman-Kac formula, a case previously treated by second-order BSDEs or G-expectation.","The same scheme applies to switching, impulse, stopping, partially observed, mean-field, infinite-dimensional, and jump-diffusion problems, each with its own constrained BSDE.","A randomized dynamic programming principle holds, providing an alternative route to viscosity solutions of the fully nonlinear HJB equation."],"supporting_citations":[{"why":"Supplies the density lemma that simple step controls are dense in the admissible controls under the metric used to prove the inequality from the original value to the randomized value.","marker":"[49]"},{"why":"Provides the randomized and backward SDE representation for non-Markovian control problems whose proof structure Section 6 follows.","marker":"[36]"},{"why":"Provides the point process construction stated as Proposition A.2, used to approximate any control by a randomized step process.","marker":"[5]"},{"why":"Introduces the randomization method in the context of optimal switching and is the earliest precise formulation of the technique.","marker":"[14]"},{"why":"Introduces constrained BSDEs with jumps and connects them to quasi-variational inequalities, a central object the survey builds on.","marker":"[47]"},{"why":"Establishes Feynman-Kac representations for Hamilton-Jacobi-Bellman integro-differential equations, motivating the extension to fully nonlinear settings.","marker":"[48]"}],"fun_headline_variants":["Poisson clock randomization pins control value to BSDE","Optimal control value equals minimal constrained BSDE solution","Randomizing actions turns control problem into a BSDE","Survey: randomization method for control-to-BSDE reduction"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The proof that the original value cannot exceed the randomized value rests on a quoted density lemma asserting that simple finite-valued step controls are dense in the space of all admissible progressive controls under an L1-type metric, and the paper does not report that lemma's proof.","fun_headline_variants_meta":{"raw":{"variants":["Poisson clock randomization pins control value to BSDE","Optimal control value equals minimal constrained BSDE solution","Randomizing actions turns control problem into a BSDE","Survey: randomization method for control-to-BSDE reduction"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000676,"raw_usage":{"total_tokens":3000,"prompt_tokens":798,"completion_tokens":2202,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":414,"completion_tokens_details":{"reasoning_tokens":2138}},"tokens_in":414,"tokens_out":2202,"duration_ms":16230,"temperature":1.0,"reasoning_tokens":2138,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T15:41:54.594551+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete way to test the central claim is to check the omitted density lemma for the stated class of progressive controls on an arbitrary Borel action space: find an admissible control, for instance one with irregular path dependence on the Brownian motion, that cannot be approximated in dt-times-probability measure by any sequence of finite-valued controls measurable at deterministic times. If such a control exists under the paper's assumptions, the equality of the two values would lack its proof and could fail.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the density lemma that simple step controls are dense in the admissible controls under the metric used to prove the inequality from the original value to the randomized value."},{"cited_title":"Fuhrman and H","cited_arxiv_id":null,"evidence_quote":"Provides the randomized and backward SDE representation for non-Markovian control problems whose proof structure Section 6 follows."},{"cited_title":"Backward SDEs for optimal control of partially observed path-dependent stochastic systems: a control randomization approach","cited_arxiv_id":null,"evidence_quote":"Provides the point process construction stated as Proposition A.2, used to approximate any control by a randomized step process."},{"cited_title":"Bouchard","cited_arxiv_id":null,"evidence_quote":"Introduces the randomization method in the context of optimal switching and is the earliest precise formulation of the technique."},{"cited_title":"Kharroubi, J","cited_arxiv_id":null,"evidence_quote":"Introduces constrained BSDEs with jumps and connects them to quasi-variational inequalities, a central object the survey builds on."},{"cited_title":"Kharroubi and H","cited_arxiv_id":null,"evidence_quote":"Establishes Feynman-Kac representations for Hamilton-Jacobi-Bellman integro-differential equations, motivating the extension to fully nonlinear settings."}],"review_version":1}