{"id":"e52608d7-0fc7-400f-9ae3-b5887b7672dd","arxiv_id":"2605.20718","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Derives an explicit Gâteaux policy-gradient formula for entropy-regularized continuous-time mean-field control using the value function and cylindrical representations, then builds a model-based actor-critic scheme with PDE well-posedness analysis.","lead":"This paper develops a policy gradient method for entropy-regularized mean-field control in continuous time. It derives an explicit gradient formula directly from the value function on the state-population space, enabling a model-based actor-critic algorithm without solving an extra equation for the gradient.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"The reader's weakest assumption (polynomial-growth well-posedness of the linear HJB) is the necessary prerequisite for the value function to exist in the class where the Gâteaux formula applies. No additional internal gap appears in the gradient derivation itself; the distinction from prior methods that require an extra equation is supported by the advantage-based construction.","tokens_in":1832,"tokens_out":245,"duration_ms":57702,"concrete_test":"On the LQR example, recompute the policy gradient via the explicit formula (value + advantage) and via centered finite differences on the objective under small policy-parameter perturbations (which include induced μ variation); agreement within 5% confirms the formula captures the full variation without hidden terms.","verdict_should_be":"UNCHANGED","load_bearing_attack":"No load-bearing concern identified for the central claim. The explicit Gâteaux formula expresses the first-order objective variation directly via the value function and instantaneous advantage after the linear stationary HJB is solved; the derivation uses the HJB residual to absorb future and measure-coupling effects without an auxiliary equation, consistent with the coupled representative-particle/population setup and cylindrical representation.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper develops a policy gradient method for entropy-regularized mean-field control in the continuous-time, discounted infinite-horizon setting. It considers randomized feedback policies acting on a coupled representative-particle and population system whose law evolves according to a McKean-Vlasov equation. The central contribution is an explicit Gâteaux formula expressing the first-order variation of the objective directly in terms of the value function and an instantaneous advantage function, obtained after solving only the linear stationary HJB equation; this yields a model-based actor-critic scheme in which the critic uses cylindrical functions to represent measure dependence and the actor is updated via the derived ascent direction. Well-posedness of the PDE is established in a polynomial-growth function class, and the method is illustrated numerically on an LQR example and a crowd-motion problem.","tokens_in":1876,"tokens_out":514,"duration_ms":27842,"significance":"If the Gâteaux formula is valid, the approach offers a genuine simplification by eliminating the need for an auxiliary equation to compute the policy gradient, which is a practical advantage in mean-field settings. The combination of the explicit formula, cylindrical-function representation, and well-posedness result in a polynomial-growth class supplies a coherent theoretical framework. The numerical experiments on standard benchmark problems provide concrete evidence of implementability, although their scope remains limited to linear-quadratic and simple interaction models.","major_comments":[],"minor_comments":[{"comment":"The abstract states that the value function is defined on R^d × P_2(R^d) but does not indicate whether the cylindrical-function representation is introduced before or after the Gâteaux derivation; a brief forward reference would improve readability.","section":"Abstract"},{"comment":"In the well-posedness analysis, the precise polynomial-growth exponents and the constants appearing in the a-priori estimates are not restated in the main theorem statement; adding them would make the result self-contained.","section":"Well-posedness analysis"},{"comment":"The numerical section reports results for the LQR and crowd-motion examples but does not specify the discretization parameters (time step, number of particles, or basis size for the cylindrical functions); these details are needed for reproducibility.","section":"Numerical experiments"}],"recommendation":"minor_revision","confidential_remarks":"The manuscript fits comfortably within the scope of a control-theory or mean-field-games journal; the citation list appears balanced and does not raise red flags about self-referential loops."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the careful and accurate summary of our work, as well as the positive assessment of its significance. The recommendation for minor revision is appreciated. We note that the major comments section of the report is empty, so we have no specific points requiring point-by-point rebuttal at this stage. We are happy to implement any minor clarifications or improvements the editor or referee may suggest in the next version.","responses":[],"tokens_in":1385,"tokens_out":102,"duration_ms":24562,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing to know is that after you solve the linear stationary HJB for the value function under a fixed randomized policy, the gradient with respect to policy parameters drops out from a direct Gâteaux derivative expressed via the instantaneous advantage function. The cylindrical-function representation for the population law is what makes the HJB tractable and lets them close the derivation without further equations for the mean-field coupling effects.","headline":"This paper gives an explicit Gâteaux formula for the policy gradient in continuous-time entropy-regularized mean-field control that comes directly from the value function and instantaneous advantage without an extra auxiliary equation.","tokens_in":2425,"tokens_out":162,"would_cite":false,"duration_ms":23898,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[{"relation":"unclear","rs_module":"IndisputableMonolith/Cost/FunctionalEquation.lean","rs_theorem":"washburn_uniqueness_aczel","paper_passage":"after computing the value function under a fixed randomized feedback policy, the policy gradient can be obtained directly via an explicit Gâteaux formula in terms of the value function and the instantaneous advantage function, without solving an additional equation"},{"relation":"unclear","rs_module":"IndisputableMonolith/Foundation/ArithmeticFromLogic.lean","rs_theorem":"LogicNat ≃ Nat recovery","paper_passage":"linear stationary Hamilton–Jacobi–Bellman equation ... using cylindrical functions"}],"headline":"Mean-field policy gradient via explicit Gâteaux advantage formula; no RS cost or ratio structure","alignment":"orthogonal","rationale":"The paper derives an explicit policy-gradient formula (Thm 4.6, 4.9) for entropy-regularized MFC using instantaneous representative/population advantage functions q_rep^π and q_pop^π built directly from the value function V^π and the infinitesimal generator L^π on R^d × P_2(R^d). This avoids an auxiliary sensitivity equation. The construction lives entirely in the domain of continuous-time stochastic control and linear stationary HJB equations (Thm 3.3) with cylindrical/Galerkin approximations. No J-cost functional equation, reciprocal symmetry, golden-ratio identities, 8-tick periodicity, or parameter-free constant derivation appears. The RS forcing chain (reality_from_one_distinction, J-uniqueness via Aczél, φ-ladder) is therefore neither confirmed nor contradicted.","tokens_in":62761,"confidence":"moderate","tokens_out":378,"duration_ms":10103,"cache_read_input_tokens":38528,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Policy gradients for continuous-time mean-field control follow directly from an explicit Gâteaux formula in the value function and instantaneous advantage function.","keywords":["mean-field control","policy gradient","Gâteaux derivative","Hamilton-Jacobi-Bellman equation","actor-critic","McKean-Vlasov","entropy regularization","randomized feedback policy"],"falsifier":"A direct numerical check in which a small policy perturbation produces an objective change that fails to match the directional derivative predicted by the Gâteaux formula applied to the computed value function.","tokens_in":2697,"feed_emoji":"📈","tokens_out":670,"duration_ms":36384,"temperature":0.7,"pith_summary":"This paper develops a policy gradient method for entropy-regularized mean-field control in the discounted infinite-horizon setting. It considers randomized feedback policies acting on a representative particle whose state evolves jointly with a population law that satisfies a McKean-Vlasov equation, so the value function is defined on the product space of state and probability measures. After the value function is obtained by solving the linear stationary Hamilton-Jacobi-Bellman equation, the policy gradient is recovered from a first-order Gâteaux variation formula that uses only the value function and an instantaneous advantage function measuring the gain of a chosen action relative to the current policy. The resulting actor-critic procedure updates the policy parameters with this explicit direction, after the critic step represents the population dependence via cylindrical functions.","feed_headline":"Explicit Gâteaux formula yields mean-field policy gradient","feed_subtitle":"Value function and advantage term supply the ascent direction after one HJB solve, without an extra equation.","key_machinery":"The Gâteaux policy-gradient formula expressed through the instantaneous advantage function that quantifies the gain of taking a given action relative to the current randomized policy.","core_discovery":"After computing the value function under a fixed randomized feedback policy, the policy gradient is obtained directly via an explicit Gâteaux formula in terms of the value function and the instantaneous advantage function, without solving an additional equation.","pith_inferences":["The single-PDE-per-iteration structure could reduce the number of coupled solves required in high-dimensional or many-agent control settings compared with methods that recompute a gradient-specific equation.","The cylindrical-function representation of measure dependence may transfer to other control problems where the state law enters the dynamics or cost.","The explicit advantage-based variation formula invites direct comparison with discrete-time mean-field policy gradients obtained by taking continuous-time limits."],"forward_implications":["The method produces a model-based actor-critic scheme in which the critic solves the linear stationary HJB equation and the actor updates via the derived explicit gradient formula.","Well-posedness of the underlying PDE is established in the chosen polynomial-growth function class with cylindrical dependence on the population law.","The approach is illustrated by numerical experiments on an LQR model and a crowd-motion problem."],"fun_headline_variants":["Gâteaux formula gives mean-field policy gradient","Value function yields direct mean-field gradients","Advantage term supplies policy ascent direction","Mean-field control gradient from one HJB solve"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The value function belongs to a suitable polynomial-growth function class in which the linear stationary HJB equation is well-posed when the population law is represented via cylindrical functions.","fun_headline_variants_meta":{"raw":{"variants":["Gâteaux formula gives mean-field policy gradient","Value function yields direct mean-field gradients","Advantage term supplies policy ascent direction","Mean-field control gradient from one HJB solve"]},"model":"grok-4.3","cost_usd":0.008294,"raw_usage":{"total_tokens":3772,"prompt_tokens":694,"num_sources_used":0,"completion_tokens":52,"cost_in_usd_ticks":82937000,"prompt_tokens_details":{"text_tokens":694,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3026,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":694,"tokens_out":52,"duration_ms":28679,"temperature":1.0,"reasoning_tokens":3026,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-21T04:11:54.306972+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A direct numerical check in which a small policy perturbation produces an objective change that fails to match the directional derivative predicted by the Gâteaux formula applied to the computed value function.","supporting_citations":[],"review_version":1}