{"id":"df9da6ba-7caf-45e2-a025-905efc63ca20","arxiv_id":"2607.11005","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"Model-free deterministic policy gradients and a continuous-time deep actor-critic algorithm solve extended mean-field control problems whose dynamics and rewards depend on the joint state-control law.","lead":"The paper gives a model-free actor-critic method for continuous-time extended mean-field control that uses deterministic policies so the joint state-action law is just a push-forward of the state law. The resulting continuous-time deep DPG algorithm is shown to be stable and efficient on consensus and crowded-liquidation problems that depend on the control distribution.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the already-flagged classical regularity postulate.","rationale":"The central claim is that a model-free deterministic policy gradient for continuous-time extended MFC can be expressed via a local advantage-rate function on the joint state–action law and turned into a practical actor-critic algorithm. The only place this claim is not fully secured by the given arguments is the a-priori C^{1,2} regularity of the lifted value function (Assumptions 2.2 and 3.1). The reader already isolates this assumption accurately; the remainder of the derivation (performance difference, invariance along McKean–Vlasov flows, local decomposition, martingale TD loss) is standard once regularity is granted. The numerical evidence, while lacking error bars and public code, is consistent with the theory on both LQ and non-LQ instances and does not introduce a new correctness risk. Consequently no adjustment to the ACCEPT verdict is warranted.","tokens_in":27782,"tokens_out":527,"duration_ms":6235,"concrete_test":"Verify that the linear-quadratic Cucker–Smale instance (γ=0) satisfies the C^{1,2} regularity of Assumption 2.2 by direct solution of the associated Riccati equation; if the explicit value function is C^{1,2} and the policy-gradient formula (3.22) recovers the known closed-form optimum (5.4), the regularity postulate holds in the regime where the algorithm is claimed to be exact.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest assumption (Assumption 2.2/3.1: V(·,·,θ) ∈ C^{1,2}([0,T]×P_2) for every θ, and continuous differentiability of A[w] in θ) is correctly identified and is the only load-bearing gap. Once that classical-solution regularity is granted, Theorems 2.1–2.3 and 3.1–3.2 follow by standard Itô calculus on the Wasserstein space, the performance-difference identity (Proposition 6.2), and the chain rule for the push-forward measure (3.18). The martingale characterization (3.20) that drives Algorithm 1 is then a direct consequence, and the numerical experiments on the Cucker–Smale and liquidation problems are consistent with the claimed gradient representation. No additional internal inconsistency, hidden structural restriction, or unsupported leap in the deterministic-policy argument was found.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper develops a model-free continuous-time actor-critic framework for extended mean-field control (MFC), in which both dynamics and rewards may depend on the joint law of states and controls. Deterministic feedback policies are used so that the state-action measure is the push-forward of the state law, avoiding optimization over stochastic kernels. A model-free sensitivity formula for parameterized McKean-Vlasov dynamics (Theorems 2.1-2.3) yields a deterministic policy-gradient identity on the Wasserstein space (Theorem 3.1). This is refined via local value and advantage-rate functions of (state, action, joint law), producing a gradient that contains both ordinary action derivatives and L-derivatives with respect to the control marginal (Theorem 3.2). The local objects are characterized by a martingale condition that is turned into a continuous-time deep DPG algorithm (CT-DDPG, Algorithm 1) with particle approximations, measure-dependent networks, TD learning, and action- or parameter-space exploration. Numerical experiments on stochastic Cucker-Smale consensus and optimal liquidation with trade crowding illustrate efficiency and robustness, including settings with explicit control-distribution dependence.","tokens_in":28106,"tokens_out":1018,"duration_ms":7592,"significance":"If the regularity assumptions hold, the work supplies a clean, first-principles policy-gradient theory for continuous-time extended MFC that removes the separable-structure and known-control-dependence restrictions of earlier exploratory-policy methods. The local martingale characterization (3.20) and the resulting CT-DDPG algorithm are practically useful and are supported by consistent numerical evidence on both LQ and non-LQ problems. The deterministic-policy route is a genuine conceptual contribution relative to the stochastic-policy literature, and the paper is careful to state the classical-solution hypotheses under which the identities are derived.","major_comments":[{"comment":"Assumption 2.2 (and the induced Assumption 3.1) postulates that the lifted value V(·,·,θ) already belongs to C^{1,2}([0,T]×P_2(R^n)) for every policy parameter θ and that the advantage-rate map A[w] is continuously differentiable in θ. This classical regularity is used both for the sensitivity formula (Theorem 2.1) and for the martingale characterization that drives learning (Theorem 3.2 / (3.20)). The paper does not derive it from the coefficients; a short discussion of sufficient conditions (or a pointer to existing viscosity/regularity results for McKean-Vlasov HJB equations) would make the scope of the claims clearer.","section":null},{"comment":"Section 3.3 constructs candidate local functions V_D^θ and q_D via a decoupled dynamics and the integrated Hamiltonian (3.27). While this shows existence under extra smoothness, the uniqueness claim in Theorem 3.2 is only for the integrated objects V̂ and q̂. The algorithm learns the local networks V^φ and q^ψ; a brief remark on whether different local representatives can produce the same integrated gradient (and therefore the same policy update) would strengthen the link between theory and practice.","section":null}],"minor_comments":[{"comment":"In (3.18) and (3.22) the independent copy is written eξ / eX; a single consistent notation (e.g., ξ̃) would improve readability.","section":null},{"comment":"Figure 1 caption states that AC and q-Learning exploit the LQ structure while CT-DDPG does not; the main text already makes this clear, but the caption could briefly note that the comparison is therefore not fully model-agnostic.","section":null},{"comment":"The terminal-penalty weight w=0.002 and soft-update τ=0.1 appear only in the experimental section; a short sensitivity remark (or a default recommendation) would help reproducibility.","section":null},{"comment":"A few typographical inconsistencies remain (e.g., “T echnical” in the section heading of 6.1, occasional missing spaces after commas in displayed equations).","section":null}],"recommendation":"minor_revision","confidential_remarks":"The classical-regularity gap is real but standard in continuous-time MFC; once granted, the derivations are solid. The deterministic-policy contribution is genuine and the numerics are convincing. Minor revision is appropriate; I would not insist on a full regularity theory before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The real advance here is a clean model-free deterministic policy gradient for continuous-time extended MFC that works when both dynamics and reward depend on the joint state-control law. By sticking to deterministic feedback, the state-action measure is just the push-forward of the state law, so they avoid optimizing over stochastic kernels and the structural restrictions that plague the exploratory-policy literature (separable coefficients, known control dependence, high-frequency sampling of relaxed controls).\n\nWhat is new is the model-free sensitivity formula for parametric McKean-Vlasov dynamics (Theorems 2.1–2.3), obtained from a performance-difference lemma plus Itô on Wasserstein space, and the refined local advantage-rate representation (Theorem 3.2) that produces both ordinary action derivatives and L-derivatives with respect to the control marginal. That local object is what lets them write a martingale characterization and turn it into a practical continuous-time deep DPG algorithm with particle approximations and measure-dependent nets. The numerics on Cucker-Smale (LQ and non-LQ) and liquidation with trade crowding look consistent and show the method is stable even when the control law appears explicitly.\n\nThe only load-bearing soft spot is the classical-solution assumption (V(·,·,θ) already in C^{1,2} for every θ, and continuous differentiability of A[w] in θ). It is postulated rather than derived from the coefficients. Once you grant it, the rest of the calculus is standard and clean. No code or error bars, free parameters for learning rate and exploration noise, but that is normal for this literature and does not undercut the identities. Citation pattern is honest; they engage the right prior work on both the continuous-time and discrete-time sides.\n\nThis is for people who actually need to learn continuous-time mean-field control problems with control-distribution dependence. It deserves a serious referee. I would accept it for peer review and would cite the gradient formulas and the algorithm if I were working in the area.","headline":"Solid, usable model-free DPG for continuous-time extended MFC via deterministic policies; the classical C^{1,2} regularity is the only real soft spot and is typical of the subfield.","tokens_in":28631,"tokens_out":511,"would_cite":true,"duration_ms":5248,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["49N80","93E20","68T05","60H10"],"pacs":[],"model":"grok-4.5","headline":"Deterministic policies make continuous-time extended mean-field control learnable without stochastic kernels or known action dependence.","keywords":["extended mean field control","deterministic policy gradient","McKean-Vlasov dynamics","advantage-rate function","continuous-time reinforcement learning","actor-critic","Wasserstein space"],"falsifier":"Run the CT-DDPG algorithm on the linear-quadratic Cucker–Smale or liquidation problem with known closed-form optimum; if the learned return systematically fails to approach the analytic optimum as particle number and episode count increase, the claimed gradient and martingale characterizations are false.","tokens_in":28674,"feed_emoji":"🎯","tokens_out":622,"duration_ms":5195,"temperature":0.7,"pith_summary":"The paper shows that continuous-time extended mean-field control—where both dynamics and rewards can depend on the joint law of states and controls—can be solved by model-free reinforcement learning once one restricts to deterministic feedback policies. Under that restriction the joint state–action law is simply the push-forward of the state law, so policy search reduces to ordinary parameter optimization rather than optimization over stochastic kernels. From a model-free sensitivity formula for parameterized McKean–Vlasov dynamics the authors derive a deterministic policy gradient expressed through an advantage-rate function on the Wasserstein space; they then refine it into local value and advantage-rate functions of state, action and joint law. The resulting martingale characterization is turned into a continuous-time deep deterministic policy-gradient algorithm that uses particle approximations, measure-dependent networks, temporal-difference updates and either action- or parameter-space exploration. Experiments on Cucker–Smale consensus and crowded liquidation confirm that the method converges stably even when the control distribution appears explicitly in the coefficients.","feed_headline":"Deterministic policies unlock model-free learning for extended mean-field control","feed_subtitle":"Push-forward of the state law replaces stochastic kernels; a local advantage-rate martingale drives a continuous-time deep DPG algorithm.","key_machinery":"The deterministic policy-gradient formula (Theorems 3.1–3.2): the policy gradient is expressed via both the ordinary action derivative and an L-derivative with respect to the control marginal of a local advantage-rate function that is identified by a martingale condition along observed trajectories.","core_discovery":"Under natural regularity, the gradient of the value with respect to a deterministic-policy parameter equals the integral of the derivative of a local advantage-rate function that depends on state, action and the joint state–action law; this identity supplies a martingale characterization that can be learned model-free and yields a practical continuous-time deep DPG algorithm for extended mean-field control.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Deterministic policies enable model-free gradients for extended mean-field control","Local advantage-rate yields policy gradient for continuous-time mean-field control","Push-forward of state law replaces kernels in extended mean-field actor-critic","Martingale characterization drives continuous-time deep DPG for mean-field control","Model-free learning of deterministic policies for joint state-action mean-field problems"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The lifted value function is assumed already to be a classical C1,2 solution on the Wasserstein space for every policy parameter, rather than being proved from the coefficients.","fun_headline_variants_meta":{"raw":{"variants":["Deterministic policies enable model-free gradients for extended mean-field control","Local advantage-rate yields policy gradient for continuous-time mean-field control","Push-forward of state law replaces kernels in extended mean-field actor-critic","Martingale characterization drives continuous-time deep DPG for mean-field control","Model-free learning of deterministic policies for joint state-action mean-field problems"]},"model":"grok-4.5","effort":"low","cost_usd":0.0048,"raw_usage":{"total_tokens":1396,"prompt_tokens":798,"num_sources_used":0,"completion_tokens":104,"cost_in_usd_ticks":48000000,"prompt_tokens_details":{"text_tokens":798,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":494,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":798,"tokens_out":104,"duration_ms":4800,"temperature":1.0,"reasoning_tokens":494,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T07:42:00.399139+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Run the CT-DDPG algorithm on the linear-quadratic Cucker–Smale or liquidation problem with known closed-form optimum; if the learned return systematically fails to approach the analytic optimum as particle number and episode count increase, the claimed gradient and martingale characterizations are false.","supporting_citations":[],"review_version":1}