{"id":"086f996b-2194-4198-996e-ba9f23cb6377","arxiv_id":"2501.18201","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A SAC controller with a DeepONet pretrained on backstepping stabilizes a first-order hyperbolic PDE with spatially-varying delay faster and with less steady-state error than plain SAC.","lead":"This preprint combines two established control techniques, neural operators and soft actor-critic reinforcement learning, to stabilize an unstable delayed PDE. It reports that adding a DeepONet trained on backstepping control as a feature extractor speeds up learning and removes a restriction on the delay function.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claim that the delay assumption is eliminated rests on an untested distribution shift: DeepONet is trained only on delays satisfying |tau'(x)|<1, then deployed outside that class, with no OOD validation or ablation.","rationale":"The reader's weakest assumption exactly identifies the same load-bearing concern: the DeepONet is pretrained on the admissible class D but is expected to work for delays outside D, and the paper provides no training distribution, no out-of-distribution validation, and no ablation isolating the contribution of the pretrained features. The central empirical claim is plausible and the simulation setup is concrete, so the appropriate verdict remains CONDITIONAL rather than REJECT: the paper should either supply the missing OOD evidence and ablations, or soften the claim to a demonstration on a selected delay. My reading adds no new independent objection; it confirms the reader's assessment and keeps the verdict unchanged.","tokens_in":7482,"tokens_out":3404,"duration_ms":34556,"concrete_test":"Run the following ablation: (1) Specify the tau distribution used to pretrain the DeepONet; (2) evaluate the DeepONet's relative L2 prediction error on the test delay tau(x)=0.7+0.3 cos(4 arccos x) against a numerically computed backstepping gain if it exists; (3) run NO-SAC with the pretrained DeepONet replaced by (a) randomly initialized weights and (b) a DeepONet trained on a distribution that includes the test delay, with at least 10 random seeds per configuration and reporting mean and variance of reward and L2 state norm. If configuration (a) performs as well as the reported NO-SAC, or if the OOD prediction error is large, the central claim that the backstepping prior eliminates the delay assumption is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline claim ('eliminates the assumption on the delay function required for the backstepping design') hinges on the DeepONet feature extractor remaining informative for delays outside the class D defined in Eq. (11). Section 2 defines the learned operator U only on D x C^1 x C^1, and the manuscript never states the distribution of tau used to generate the DeepONet training data. The single out-of-assumption test in Section 4 uses tau(x)=0.7+0.3 cos(4 arccos x), which violates |tau'(x)|<1, yet no OOD prediction error, no comparison with a DeepONet trained on a distribution covering this delay, and no ablation replacing the pretrained features with random or frozen features are reported. Because the actor and critic networks fine-tune the DeepONet weights during SAC training (Section 3.2, Algorithm 1), observed success could come from RL adapting to the plant rather than from backstepping prior knowledge. Without an ablation or OOD evaluation, the claim that the assumption is eliminated is not established. This is an evidential gap, not an internal inconsistency: the approach may work, but the current single experiment plus missing training details does not support the strong generalization claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes NO-SAC, a soft actor-critic (SAC) controller for a first-order hyperbolic PIDE with spatially varying state delay. A DeepONet is first trained to approximate a backstepping boundary controller designed for delays satisfying the slow-variation assumption, and the trained DeepONet is embedded as a feature extractor in the SAC actor and critic networks. The authors claim that this architecture removes the assumption on the delay function required by backstepping, and report simulations in which NO-SAC converges faster than vanilla SAC and exhibits better transient behavior than the analytical backstepping controller.","tokens_in":7680,"tokens_out":5464,"duration_ms":54922,"significance":"If substantiated, the idea of using a neural operator pretrained on an analytic backstepping controller as a feature extractor inside SAC is a useful and timely contribution to learning-based PDE control. The paper makes the right comparisons in principle: it evaluates against both a no-prior RL baseline and the analytic backstepping controller, and the experimental configuration is described in reasonable detail. However, the central generalization claim is currently supported only by a single out-of-assumption simulation, with no training distribution for the DeepONet, no out-of-distribution validation, no ablations, and no repeated trials. The paper does not provide proofs, code, or error bars, so the strengths lie in the clarity of the problem formulation and the plausibility of the approach rather than in the strength of the evidence.","major_comments":[{"comment":"The controller operator U is defined only on D × C^1 × C^1 with D given by Eq. (11), yet the abstract's central claim is that NO-SAC eliminates the delay assumption by working for delays outside D. The manuscript never states the distribution of τ used to pretrain the DeepONet, reports no out-of-distribution prediction error, and evaluates only one delay, τ(x)=0.7+0.3 cos(4 arccos x), that violates the assumption. Because the DeepONet is trained on examples generated by the backstepping controller for delays in D, this single experiment does not establish that the learned operator remains informative outside D; this is a load-bearing evidential gap in the claim that the delay assumption is eliminated.","section":"§2, Definition 1, and §4.2"},{"comment":"During SAC training, the DeepONet weights φ_N and ϑ_N are updated by backpropagation together with the fully connected layers (Algorithm 1, lines 13-14). Without an ablation that freezes the DeepONet features or replaces them with random features, the observed improvement over baseline SAC cannot be attributed to the backstepping prior; it could instead result from the RL agent adapting directly to the plant dynamics. The paper should report such an ablation to support the stated role of the DeepONet as the mechanism transferring prior knowledge.","section":"§3.2 and Algorithm 1"},{"comment":"All conclusions are drawn from single training runs. The reward curves in Fig. 4 and the state evolutions in Figs. 5-7 contain no error bars, no multiple random seeds, and no sensitivity analysis with respect to the reward weights Γ, σ, ζ or the SAC hyperparameters. Consequently, the statement in §4.2 that NO-SAC 'consistently outperforms' the baseline is unsupported; the evidence shows one favorable trajectory, not a statistically reliable comparison.","section":"§4.1 and §4.2"},{"comment":"The comparison with the backstepping controller may be confounded by implementation differences. The RL control is updated every 100 steps with zero-order hold and is bounded by [-30,30], but no implementation details are given for the backstepping controller used in Figs. 6-7; if the backstepping controller is evaluated as a continuous-time signal, the claim of smaller overshoot and shorter settling time is not an apples-to-apples comparison. In addition, the conclusion's statement that NO-SAC 'stabilizes' the PDE is stronger than what a 5-second finite-horizon simulation can establish, since no asymptotic or quantitative convergence criterion is reported.","section":"§4.1, §4.2, and §5"}],"minor_comments":[{"comment":"The delay function is written inconsistently: Fig. 3 defines τ(x)=0.7+0.3 cos(4 arccos(x)), while the Fig. 5 caption says τ(x)=0.7+0.3 cos(arccos(x)). Please correct this, as the exact delay affects reproducibility.","section":"§4.1 and §4.2 captions"},{"comment":"The reward r_mid uses s_{t-1} without defining s_{-1} at the start of an episode; please specify the initial previous state.","section":"§3.1, Eq. (15)"},{"comment":"There are several typos: 'bacsktepping' in Section 2, 'actot-critic' in Section 3.2, and 'whitout' in the Fig. 5 caption. The manuscript should be proofread.","section":"Throughout"},{"comment":"The symbol U is overloaded: it denotes the control input in Eq. (2), the controller operator in Definition 1, and the bound of the action space in Section 3.1. Consider using distinct symbols for these quantities.","section":"§2 and §3.1"},{"comment":"The text says the DeepONet inputs consist of τ, x, and u, while Definition 1 defines the operator as U(τ,v,u). The figure and text should clarify whether the state v is also an input to the DeepONet.","section":"§2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a promising empirical contribution, but its headline claim of eliminating the delay assumption needs substantially more evidence: an explicit training distribution for the DeepONet, out-of-distribution validation, ablations isolating the prior, and repeated trials. The authors should also clarify the relationship with their own prior work, arXiv:2412.08219, which appears to contain the underlying DeepONet backstepping controller. Given the IFAC scope, a revision with these additions could be suitable; the current version is not yet convincing on its central claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a plausible architecture paper with a real new combination—DeepONet as a pretrained feature extractor inside SAC—but the headline claim that the delay assumption is eliminated is not supported by the evidence. The single out-of-assumption simulation is a demonstration, not a proof, and the missing DeepONet training distribution matters.\n\nWhat's genuinely new: no prior work cited here combines a neural operator as a feature extractor in SAC for PDE control. The warm-start idea is sensible, and the comparisons against plain SAC and the analytical backstepping controller are appropriate. The paper also tests both a delay satisfying |τ'(x)|<1 and one that violates it, which is a reasonable first pass.\n\nWhere it's soft: (1) The DeepONet is defined on D (delays satisfying the assumption) and presumably trained there, yet it is deployed on a τ outside D. The paper never states the training distribution or measures out-of-distribution prediction error. So the 'eliminates the assumption' claim rests on an untested distribution shift. (2) The actor and critic fine-tune the DeepONet weights during SAC training (Section 3.2, Algorithm 1). Observed success could therefore come from RL adapting to the plant rather than from backstepping prior knowledge. Without an ablation replacing the pretrained features with random or frozen ones, you can't attribute the gain to the prior. (3) No error bars, no multiple seeds, no sensitivity analysis; every conclusion rests on one run per condition. (4) Stability and convergence are only empirical, which is fine for a demo but not for the strong wording. (5) Minor: Fig. 5 caption writes cos(arccos x) while Fig. 3 uses cos(4 arccos x); likely a typo.\n\nThe math and citation pattern look solid. The self-citation to Qi et al. 2024a is legitimate since that is the backstepping DeepONet prior, and the final comparison is against external baselines, so the work is not circular.\n\nWho this is for: people working on data-driven control of delayed PDEs. A serious referee should engage with it. The verdict should be conditional: add multiple seeds and error bars, specify the DeepONet training distribution, add an OOD evaluation and a feature-ablation, and soften 'eliminates' to 'demonstrates on selected delays' unless the new experiments support the stronger claim.","headline":"A genuinely new architecture—DeepONet as a feature extractor inside SAC—but the headline claim about eliminating the delay assumption outruns the evidence; needs seeds, training details, and an ablation.","tokens_in":8238,"tokens_out":3865,"would_cite":false,"duration_ms":33836,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["93C20","68T07"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proposes NO-SAC, a soft actor-critic controller whose actor and critic networks read features from a DeepONet pretrained on backstepping controllers, and argues this removes the slow-variation delay assumption while stabilizing…","keywords":["first-order hyperbolic PDE","spatially-varying delay","backstepping control","DeepONet","neural operator","soft actor-critic","reinforcement learning","boundary control"],"falsifier":"Train the DeepONet only on delays satisfying $|\\tau'(x)| < 1$, then run NO-SAC on a delay with $|\\tau'(x)|$ significantly greater than 1, for example $\\tau(x) = 0.1 + 1.9x$, and record the closed-loop $L^2$ norm; if the state fails to converge or the NO-SAC reward collapses while SAC still learns, the claim that the delay assumption is eliminated would be contradicted.","tokens_in":7223,"feed_emoji":"⚙️","tokens_out":7651,"duration_ms":67456,"temperature":0.7,"pith_summary":"The paper tries to show that a reinforcement-learning controller can stabilize an unstable first-order hyperbolic PDE whose state is delayed by a spatially-varying function $\\tau(x)$, without the slow-variation assumption $|\\tau'(x)| < 1$ that the analytic backstepping design requires. The proposed method, NO-SAC, trains a DeepONet to approximate the backstepping controller operator and then embeds copies of that network as feature extractors in the actor and critic networks of a soft actor-critic agent. In simulations on the delay $\\tau(x) = 0.7 + 0.3\\cos(4\\arccos x)$, which violates the assumption, NO-SAC converges faster and eliminates steady-state error compared to plain SAC. When the delay satisfies the assumption, both RL controllers give smaller overshoot and shorter settling times than the analytic backstepping controller. The significance would be a general recipe: analytical PDE control knowledge, converted into an operator approximator, can warm-start model-free RL and widen the class of delay functions that can be controlled.","feed_headline":"DeepONet boosts RL control of delayed PDEs, beating backstepping","feed_subtitle":"Soft actor-critic with a DeepONet feature extractor converges faster and removes steady-state error in simulations.","key_machinery":"The mechanism is the composition of a DeepONet and an SAC agent. The DeepONet is a branch-and-trunk neural operator trained to approximate the backstepping boundary-control map $U(\\tau, v, u)$; its branch network encodes the three functions sampled on a $21 \\times 21$ grid and its trunk network encodes coordinates on the same grid, with the Cartesian product of the two outputs giving a 441-dimensional feature vector. Five copies of this trained operator are inserted into the policy network and the two action-value networks, so the RL agent's decisions are conditioned on features extracted from the analytical backstepping law. To turn the non-Markovian delayed evolution into an MDP, the state is augmented with the transport-delay coordinate $u(x,r,t)$ that carries the delayed information, with the reward split into a running term and a terminal term.","core_discovery":"The central claim is that the delay assumption needed for backstepping can be removed by giving the RL agent a neural-operator representation of the backstepping solution rather than requiring the delay to belong to the admissible class $\\mathcal{D} = \\{\\tau \\in C^2[0,1] : \\tau(x) > 0 \\text{ for all } x \\text{, and if } \\tau(x) < x \\text{ then } |\\tau'(x)| < 1\\}$. The paper constructs a DeepONet that learns the controller operator $U(\\tau, v, u)$, mapping the delay function, the current state $v(x,t)$, and the delayed state $u(x,r,t)$ to the boundary input, and then uses copies of the trained DeepONet as feature-extraction layers in the SAC actor and critic. The resulting NO-SAC policy is evaluated on a delay that violates the assumption and is reported to stabilize the closed loop faster than baseline SAC, and on an admissible delay where it compares favorably with the analytic controller in transient performance.","pith_inferences":["Editorial inference: The strongest test of the no-assumption claim is a systematic sweep over delays far outside $\\mathcal{D}$ with $|\\tau'(x)| > 1$; the paper reports only a single violating delay, so the claim's breadth remains untested.","Editorial inference: The same operator-pretraining scheme should transfer to other backstepping designs, such as actuator or sensor delay compensation, because the DeepONet only needs to approximate the controller map rather than the PDE coefficients.","Editorial inference: A natural testable extension is to vary the $21 \\times 21$ spatial discretization and the zero-order-hold update rate, which would reveal how much of the observed gain comes from the operator prior and how much from RL exploration."],"forward_implications":["Controllers for this class of delayed PDEs can be obtained without checking $|\\tau'(x)| < 1$, provided the operator network generalizes beyond its training set.","Analytic backstepping laws can be packaged as pretrained feature extractors, giving RL a warm start that reduces steady-state error and training time.","In the admissible-delay regime, the learned policy can match or beat the analytic controller in transient performance, suggesting RL can refine rather than only replace analytical designs.","The augmented-state MDP formulation makes delayed boundary-control problems amenable to standard off-policy RL algorithms."],"supporting_citations":[{"why":"Defines the backstepping controller and the admissible delay class $\\mathcal{D}$ that the paper aims to relax, and supplies the operator $U$ and baseline analytic controller.","marker":"Zhang and Qi (2021, 2024)"},{"why":"Provides the earlier DeepONet-based backstepping controller for the same PIDE, whose architecture is reused for feature extraction.","marker":"Qi et al. (2024a)"},{"why":"Demonstrates neural-operator backstepping control for first-order hyperbolic PIDEs, supporting the feasibility of offline operator learning for this plant.","marker":"Qi et al. (2024b)"},{"why":"Supplies the SAC algorithm and the baseline agent that NO-SAC is compared against.","marker":"Haarnoja et al. (2018)"},{"why":"Gives the DeepONet branch/trunk architecture and the operator-approximation rationale.","marker":"Lu et al. (2021)"},{"why":"The PDE control Gym environment template is used to build the simulation environment for equations (1)-(4).","marker":"Bhan et al. (2024)"},{"why":"Supports the non-Markovian characterization of delayed dynamics that motivates the augmented-state MDP.","marker":"Bouteiller et al. (2020)"},{"why":"Frames backstepping boundary control for first-order hyperbolic PDEs, the analytical basis being approximated by the DeepONet.","marker":"Krstic and Smyshlyaev (2008)"}],"fun_headline_variants":["DeepONet removes delay assumption, RL beats backstepping","Neural operator RL stabilizes delayed PDEs without delay constraints","SAC with DeepONet features outdoes analytic controller for delayed PDEs","Delayed PDE control via DeepONet-enhanced SAC: no delay class needed","Backstepping knowledge distilled into RL via DeepONet for PDE delays"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the DeepONet, pretrained on backstepping controllers generated for delays in the admissible class $\\mathcal{D}$, generalizes to delay functions outside that class; the paper does not report the training distribution of $\\tau$ or any out-of-distribution test beyond the single violating delay used in simulation.","fun_headline_variants_meta":{"raw":{"variants":["DeepONet removes delay assumption, RL beats backstepping","Neural operator RL stabilizes delayed PDEs without delay constraints","SAC with DeepONet features outdoes analytic controller for delayed PDEs","Delayed PDE control via DeepONet-enhanced SAC: no delay class needed","Backstepping knowledge distilled into RL via DeepONet for PDE delays"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000179,"raw_usage":{"total_tokens":1281,"prompt_tokens":904,"completion_tokens":377,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":520,"completion_tokens_details":{"reasoning_tokens":283}},"tokens_in":520,"tokens_out":377,"duration_ms":4039,"temperature":1.0,"reasoning_tokens":283,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T00:18:53.716523+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the DeepONet only on delays satisfying $|\\tau'(x)| < 1$, then run NO-SAC on a delay with $|\\tau'(x)|$ significantly greater than 1, for example $\\tau(x) = 0.1 + 1.9x$, and record the closed-loop $L^2$ norm; if the state fails to converge or the NO-SAC reward collapses while SAC still learns, the claim that the delay assumption is eliminated would be contradicted.","supporting_citations":[{"cited_title":"and Qi, J","cited_arxiv_id":null,"evidence_quote":"Defines the backstepping controller and the admissible delay class $\\mathcal{D}$ that the paper aims to relax, and supplies the operator $U$ and baseline analytic controller."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the SAC algorithm and the baseline agent that NO-SAC is compared against."},{"cited_title":"PDE Control Gym: A Benchmark for Data-Driven Boundary Control of Partial Differential Equations","cited_arxiv_id":"2405.11401","evidence_quote":"The PDE control Gym environment template is used to build the simulation environment for equations (1)-(4)."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the non-Markovian characterization of delayed dynamics that motivates the augmented-state MDP."},{"cited_title":"and Smyshlyaev, A","cited_arxiv_id":null,"evidence_quote":"Frames backstepping boundary control for first-order hyperbolic PDEs, the analytical basis being approximated by the DeepONet."}],"review_version":1}