{"id":"3619add2-4795-4582-b853-571739e7d4c4","arxiv_id":"2606.01058","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":1,"one_line_summary":"Entropy-regularized relaxed control for infinite-horizon Itô stochastic systems with input delay yields Gaussian optimal distributions that converge to Dirac measures as the exploration weight vanishes.","lead":"The paper reformulates stochastic optimal control problems with input delay using entropy regularization on relaxed controls to derive Gaussian optimal controllers. A generalist might read it for insight into handling delays and uncertainty in control systems via probabilistic methods that recover deterministic solutions in the limit.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Relaxed-control reformulation for input-delay Itô systems may fail to preserve dynamics/cost when entropy term is added","rationale":"The reader's weakest_assumption directly names the load-bearing step; the delay structure makes that step non-obvious and the abstract supplies no explicit verification. Because the full manuscript was not reviewed by the reader, the present pass does not alter the UNVERDICTED status.","tokens_in":1595,"tokens_out":371,"duration_ms":13662,"concrete_test":"Write the relaxed dynamics explicitly (state segment X_t(·) driven by ∫ u(t-τ) μ_t(du) where μ_t is the relaxed control measure) and substitute into the original cost; verify whether the entropy term ∫ log(dμ_t) dμ_t appears only as an additive regularizer or whether it forces a modification of the delay integral. If the latter occurs, recompute the HJB equation for the regularized value function and check whether the optimizer remains Gaussian.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that a relaxed-control version of the delayed stochastic system exists such that the entropy-regularized objective can be introduced while the original Itô dynamics (including the delayed input operator) and running cost remain exactly the same. In the presence of input delay the state is an infinite-dimensional segment; the relaxed control is a measure-valued process whose action on the delayed input must be integrated against the delay kernel. Nothing in the abstract or the stated weakest assumption shows that this measure-valued lifting commutes with the delay operator without introducing an extra integral term or altering the quadratic variation of the Itô integral. If that commutation fails, the subsequent derivation of the Gaussian optimizer and the zero-temperature limit both rest on an altered problem.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript investigates the infinite-horizon stochastic optimal control problem for Itô systems with input delay under an entropy-regularized relaxed-control framework. It asserts that constructing a relaxed system and adding an entropy term reformulates the classical problem, yields an optimal controller whose distribution is Gaussian, and that this Gaussian distribution converges to the optimal Dirac measure as the exploration weight tends to zero; numerical simulations are supplied for validation.","tokens_in":1761,"tokens_out":300,"duration_ms":19960,"significance":"If the relaxed-control reformulation preserves the original Itô dynamics and cost exactly (including the action of the delayed input operator on the measure-valued control), the result would extend entropy-regularized methods to delayed stochastic systems and supply both a tractable Gaussian policy and a rigorous zero-temperature limit. The numerical simulations constitute a concrete strength that can be assessed independently of the analytic claims.","major_comments":[{"comment":"Abstract (and the weakest assumption stated therein): the claim that a relaxed system can be constructed so that the entropy-regularized objective is introduced while the original Itô dynamics and running cost remain exactly the same is load-bearing for the subsequent Gaussian derivation and convergence result, yet no equation, integral representation, or sketch is supplied showing how the measure-valued control integrates against the delay kernel without altering the quadratic variation or introducing an extra term.","section":"Abstract"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the detailed reading and the constructive comment on the abstract. We address the point below.","responses":[{"response":"We acknowledge that the abstract is brief and does not contain the supporting equation or sketch. In the body of the manuscript (Section 2), the relaxed dynamics are defined by replacing the control with a probability measure μ on the admissible set, with the delayed input entering via the Bochner integral ∫ K(t,s) u(s) μ(du) against the delay kernel; because this is a deterministic integral with respect to the measure (no additional Itô integral is introduced), the quadratic variation of the state process remains exactly that of the original system, and the running cost is unchanged. The entropy term is added only to the objective. To make this explicit for readers, we will revise the abstract to include a one-sentence reference to this integral representation and its preservation properties.","revision_made":"yes","referee_comment":"[Abstract] Abstract (and the weakest assumption stated therein): the claim that a relaxed system can be constructed so that the entropy-regularized objective is introduced while the original Itô dynamics and running cost remain exactly the same is load-bearing for the subsequent Gaussian derivation and convergence result, yet no equation, integral representation, or sketch is supplied showing how the measure-valued control integrates against the delay kernel without altering the quadratic variation or introducing an extra term."}],"tokens_in":1162,"tokens_out":318,"duration_ms":16030,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper applies entropy regularization to relaxed control for infinite-horizon Itô systems with input delay. It claims to derive an optimal Gaussian controller and shows that this distribution converges to the deterministic optimum as the regularization parameter goes to zero. A numerical simulation is included to illustrate the approach.\n\nThis combination for the delayed case is the new element. Prior work on entropy-regularized control exists, but handling input delay requires lifting the control to measures while keeping the delayed input operator intact. The paper positions this as the contribution.\n\nThe numerical example is a positive point, as it provides some concrete validation beyond the theory.\n\nThe soft spot is the reformulation of the delayed system into relaxed form. The state is infinite-dimensional due to the delay, and the measure-valued control must be applied through the delay kernel. It is not obvious that this preserves the original Itô dynamics and quadratic variation exactly. The stress-test concern about commutation with the delay operator is worth examining in the proofs. If that holds, the Gaussian result and convergence follow from standard arguments. If not, the claims apply to an altered problem.\n\nThe paper engages with the literature on relaxed control and entropy terms in a direct way. The math appears to be formal, though the abstract alone does not show the equations.\n\nThis work is for specialists in stochastic optimal control, particularly those interested in delays or regularized problems. A reader in that area would see the explicit Gaussian form as potentially useful for computation or further analysis.\n\nIt deserves a serious referee. The claims are checkable against the derivations, and the topic has practical relevance in control with delays.\n\nI recommend sending it to peer review, with attention to the delay-relaxation step.","headline":"The paper extends entropy-regularized relaxed control to Itô systems with input delay, deriving an explicit Gaussian controller and its convergence to the Dirac optimum, but the reformulation step that preserves the delayed dynamics needs verification.","tokens_in":2254,"tokens_out":437,"would_cite":false,"duration_ms":26388,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Entropy regularization turns the optimal control problem for delayed stochastic systems into one whose solution is a Gaussian distribution that approaches the classical optimum as the regularization weight vanishes.","keywords":["stochastic optimal control","input delay","entropy regularization","relaxed control","Gaussian distribution","infinite horizon","Itô processes"],"falsifier":"For a simple scalar linear delayed system whose classical optimum is known in closed form, compute the Gaussian controller at successively smaller entropy weights; if the mean does not approach the known deterministic control and the variance does not approach zero, the convergence claim is false.","tokens_in":2499,"feed_emoji":"","tokens_out":615,"duration_ms":18678,"temperature":0.7,"pith_summary":"The paper shows how to reformulate an infinite-horizon stochastic optimal control problem with input delay by using relaxed controls together with an entropy regularization term. This produces an optimal controller whose distribution is Gaussian. The Gaussian solution is proved to converge to the classical deterministic optimum (a Dirac measure) when the regularization weight is driven to zero. A sympathetic reader would care because the construction supplies an explicit, smoothed controller for a class of delayed systems while still recovering the original solution in the limit.","feed_headline":"Entropy term yields Gaussian controller for delayed stochastic systems","feed_subtitle":"The Gaussian distribution converges to the classical Dirac optimum as the regularization weight tends to zero.","key_machinery":"The entropy-regularized relaxed-control formulation, which replaces deterministic controls by probability measures over controls and augments the cost with an entropy penalty while preserving the original dynamics and running cost.","core_discovery":"By constructing a relaxed system and introducing an entropy regularization term, the classical optimal control problem is reformulated into an entropy regularized formulation, and the optimal controller is shown to follow a Gaussian distribution. The optimal Gaussian control distribution converges to the optimal Dirac measure as the exploration weight tends to zero.","pith_inferences":["The same relaxed-plus-entropy construction could be tested on finite-horizon versions of the problem to check whether the Gaussian limit still holds.","Adaptive choice of the entropy weight might turn the method into a practical algorithm that gradually reduces exploration while approaching optimality.","The convergence result suggests entropy regularization could serve as a theoretical device for connecting stochastic and deterministic control in other delay or partial-observation settings."],"forward_implications":["The optimal controller under the regularized criterion is explicitly Gaussian.","As the exploration weight tends to zero the Gaussian distribution concentrates on the classical deterministic optimum.","The reformulation applies directly to infinite-horizon Itô systems with input delay.","Numerical simulation on concrete examples confirms that the Gaussian controller behaves as predicted."],"fun_headline_variants":["Entropy regularization yields Gaussian controller for delayed Ito systems","Gaussian controller from entropy regularization in stochastic delay systems","Relaxed control produces Gaussian optima via entropy term with input delay","Entropy regularization shapes Gaussian control in Itô systems with delay"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The stochastic system with input delay admits a relaxed-control reformulation that permits the entropy term to be introduced while preserving the original dynamics and cost structure.","fun_headline_variants_meta":{"raw":{"variants":["Entropy regularization yields Gaussian controller for delayed Ito systems","Gaussian controller from entropy regularization in stochastic delay systems","Relaxed control produces Gaussian optima via entropy term with input delay","Entropy regularization shapes Gaussian control in Itô systems with delay"]},"model":"grok-4.3","cost_usd":0.005275,"raw_usage":{"total_tokens":2470,"prompt_tokens":505,"num_sources_used":0,"completion_tokens":56,"cost_in_usd_ticks":52749500,"prompt_tokens_details":{"text_tokens":505,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1909,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":505,"tokens_out":56,"duration_ms":15100,"temperature":1.0,"reasoning_tokens":1909,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-28T16:52:44.133149+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"For a simple scalar linear delayed system whose classical optimum is known in closed form, compute the Gaussian controller at successively smaller entropy weights; if the mean does not approach the known deterministic control and the variance does not approach zero, the convergence claim is false.","supporting_citations":[],"review_version":1}