{"id":"e17fbd91-e105-4718-bed6-a095a3f54038","arxiv_id":"2505.07244","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Neural DDEs are universal approximators only when the product of Lipschitz constant and delay is large enough; small memory capacity makes them behave like non-universal neural ODEs.","lead":"Neural networks with long-range connections can be modeled as delay differential equations. This paper shows their ability to approximate arbitrary functions is governed by a memory capacity threshold, explaining why well-connected architectures like DenseNets are more expressive.","discovery_kind":"first_principles","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No-go theorem is proved only for globally defined vector fields; the general NDDE architecture with proper open state-space domain is not covered, and the K in the positive/negative regimes is not the same object.","rationale":"The reader's weakest assumption identifies the decisive gap: the negative theorem is proved for a strictly smaller class than the general neural DDEs defined in Definition 2.7. The manuscript itself flags this in Section 4.5, where globally defined weakly nonlinear vector fields are introduced as necessary for the small-delay machinery, and Theorem 4.14 handles only vector fields already defined on Omega_t x C. I add one compounding detail: the K in the positive embedding theorem is a Lipschitz constant on the set of constant initial data, whereas the K in the no-go theorem is a global Lipschitz constant on C, so the two parameter regions should not be drawn on the same K axis without an extension argument. This does not invalidate the proofs within their stated class, and the positive constructions are explicit and useful, but it makes the abstract-level conclusion conditional. No machine-checked verification or reproducible code is supplied, so the conditional verdict is appropriate.","tokens_in":64743,"tokens_out":19322,"duration_ms":209478,"concrete_test":"Settle the extension step used in Step 3 of Theorem 5.4. Take C = C([-1,0],R), Omega_y = {u in C : ||u||_inf < 2}, and define F(t,u) = phi(u(0)) with phi(x)=integral_0^x sin(1/(2-s)) ds on (-2,2); this F is C^{0,1}_b on [0,T] x Omega_y, globally 1-Lipschitz in the second variable, with F(t,0)=0. Prove or disprove that F admits a global extension F~ in C^{0,1}_b(R x C,R) with global Lipschitz constant 1. If no such extension exists, Theorem 3.13 does not apply to this admissible Phi in NDDE^1_{tau,N,K}, and the small-delay obstruction remains unproved for the general architecture; optionally, check whether this local vector field with K tau < 1/e can still approximate a non-degenerate extremum.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Theorem 3.13 and Corollary 5.5 restrict to vector fields F: Omega_t x C -> R^m that are globally K-Lipschitz on Omega_t x C with sup_t ||F(t,0)|| <= A. The architecture in Definition 2.7 only requires F on an open Omega = Omega_t x Omega_y, and Theorem 4.14 extends only vector fields already defined on Omega_t x C. Step 3 of the proof of Theorem 5.4 relies on that extension, so the exponential-attraction/inertial-manifold obstruction is not shown for local vector fields. The gap is not merely cosmetic: the positive construction in Theorem 3.9 controls Lipschitz constants only on Omega_0 = {c_{lambda(x)}: x in X}, not on all of C, so the memory capacity K in Figure 3.2 has different meanings in the two regimes. Without an extension preserving C^{0,k}_b regularity and the same K,A--or a separate argument for local vector fields--the abstract's unconditional small-capacity no-go statement overstates what is proved.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies universal approximation and universal embedding for neural delay differential equations (neural DDEs), interpreted as infinite-depth limits of DenseResNets, and investigates how the product Kτ of the vector-field Lipschitz constant and the delay controls expressivity. The main positive result (Theorem 3.9) is an explicit construction embedding any globally Lipschitz map into a non-augmented neural DDE when Kτ is sufficiently large, with a similar augmented construction (Theorem 3.14) when the DDE dimension is at least n+q. The main negative result (Theorem 3.13, Corollary 5.5, Theorem 5.4) states that when Kτ is sufficiently small, non-augmented neural DDEs cannot have the universal approximation property, using the exponential attraction of DDE solutions to finite-dimensional special solutions and a topological separation argument involving local extreme points. The paper concludes that the infinite-dimensional phase space alone does not give universal approximation and that a memory threshold Kτ is needed.","tokens_in":64875,"tokens_out":15118,"duration_ms":158795,"significance":"If the claims held in the stated generality, this would be a valuable contribution: the positive construction in Theorem 3.9 is explicit, and the negative proof in Section 5 carefully tracks constants while combining small-delay theory, exponential attraction, Morse theory, and Jordan-Brouwer separation. The paper also connects the continuous-time results to DenseResNet architectures through the Euler-discretization discussion in Section 2. However, both sides of the claimed phase transition require attention before the advertised conclusion is justified: the positive construction has an unverified domain-of-definition issue, and the negative theorems are proved only for vector fields that are globally defined on all of C, not for the general architecture of Definition 2.7.","major_comments":[{"comment":"The no-go theorem is proved only for vector fields F: Ω_t×C→R^m that are globally defined in the second variable. Theorem 4.14, which is invoked in Step 3 of the proof of Theorem 5.4, extends exactly this class: it assumes Ω_y=C and does not apply to the general architecture of Definition 2.7, where F may be defined only on an open Ω=Ω_t×Ω_y. Consequently, Corollary 5.5 and Theorem 3.13 inherit the global-domain restriction, while the abstract and Theorem 3.13 present the result for non-augmented neural DDEs without this restriction. A separate argument, or a Lipschitz/C^{0,k}_b extension theorem for local vector fields that preserves K and A, is needed; without it the small-capacity no-go claim overstates what is proved.","section":"Section 5.5, Step 3; Theorem 4.14"},{"comment":"The Lipschitz constant K in the positive and negative regimes is not the same object. In Theorem 3.9 the constructed vector field is only shown to be Lipschitz on R×Ω_0, the set of constant initial data, and the proof's estimate is performed for y_t,z_t∈Ω_0; no Lipschitz bound on all of C or on Ω_y is established. In Theorem 3.13 and Theorem 5.4, K is a global Lipschitz constant on Ω_t×C. Thus the regions AUE and AnUA in Figure 3.2 are defined under different hypotheses, and the claimed Kτ-transition is not a phase diagram for one fixed function class. The comparison requires either strengthening the positive theorem to a globally Lipschitz vector field on all of C or relaxing the negative theorem's assumption.","section":"Section 3.3; Figure 3.2"},{"comment":"The proof does not show that the solution segments of the constructed DDE remain in the domain Ω_y, so the vector field may not be well-defined on the whole interval [0,T]. The solution y(t) interpolates between λ(x) and (1/\\tilde w)Id_{m,q}Ψ(x); for a general open, nonconvex X, the first n components of y(t) need not remain in wX for t∈(0,τ], so y_t may leave Ω_y before the vector field becomes zero on (τ,∞). The proof asserts unique solvability without verifying y_t∈Ω_y. This affects the central positive embedding claim. A fix is to enlarge the domain (for example, require the condition only at the delayed evaluation point u(-τ) for t∈[0,τ] and define F=0 on all of C for t≥τ, or extend Ψ to all of R^n with the same Lipschitz constant).","section":"Section 3.3, proof of Theorem 3.9"}],"minor_comments":[{"comment":"Corollary 5.5 states τ∈[0,τ0(K)], while Theorem 5.4 is proved only for τ∈[0,τ0(K)); the proof's final contradiction uses strict inequalities, so either the closed interval should be justified or the statement should use τ<τ0(K).","section":"Corollary 5.5; Theorem 5.4"},{"comment":"The notation NDDE^k_{τ,N,K} is used before the subscript K is formally defined in the class notation; please state explicitly that K denotes the Lipschitz constant and specify on which domain it is measured in each theorem.","section":"Theorems 3.9 and 3.13"},{"comment":"Calling the property 'globally Lipschitz continuous on Ω_t×Ω_0' is potentially confusing when Ω_0 is a subset; consider using 'Lipschitz continuous on Ω_t×Ω_0' to avoid suggesting global Lipschitz continuity on all of C.","section":"Definition 3.8"},{"comment":"In the definition of τ3, the condition r_0^2-2ε-δ*>0 is guaranteed by the assumptions but should be stated explicitly before dividing by ln(2C_2/(r_0^2-2ε-δ*)).","section":"Lemma 5.15"}],"recommendation":"major_revision","confidential_remarks":"The paper is technically competent and the central idea is interesting, but the two load-bearing gaps—the global-domain restriction of the no-go theorem and the domain issue in the positive construction—prevent the advertised phase-transition claim from being established as stated. I see no circularity: the paper relies on appropriate external results and detailed explicit constructions. I expect the issues can be addressed in revision, either by extending the results to the general architecture or by carefully restricting the statements."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this paper gives a real expressivity threshold for a natural class of neural DDEs: small Kτ blocks universal approximation, large Kτ gives universal embedding. Second, the class is narrower than the abstract suggests, and the K in the no-go region and the K in the embedding region are not the same object. The no-go theorem (Thm 3.13 / 5.4) only covers vector fields globally defined on all of C with a global Lipschitz bound and sup ||F(t,0)|| ≤ A; the general architecture in Def 2.7 allows open-domain vector fields, and no extension theorem preserving the same K,A is proved. The positive embedding (Thm 3.9) controls the Lipschitz constant only on the set of constant initial histories Ω0, not on all of C. So the clean phase diagram in Fig 3.2 mixes two different Lipschitz constants and overstates what is proven.\n\nCredit where earned. The construction in Thm 3.9 is explicit and genuinely goes beyond the τ=T special case of Zhu et al. The negative proof is a serious piece of work: small-delay reduction to an ODE inertial manifold, exponential-attraction estimates from Driver and Jarnik-Kurzweil, then topological separation via Morse functions and Jordan–Brouwer to rule out approximating local extrema. The error bookkeeping is careful, and the authors do note the global-definition requirement in the body. The overstatement is in the headline, not hidden in the proofs.\n\nSoft spots, in proportion. The domain gap above is the main one and it is load-bearing for the 'small capacity implies no universal approximation' narrative. A second, minor overstatement: the abstract says universal approximation for continuous functions, but the positive theorem embeds only Lipschitz (or C^k_b) maps. The transition region Kτ in [1/e, roughly 2] is untouched, which the authors honestly concede.\n\nWho this is for: researchers working on expressivity of neural differential equations and on the connection between DenseNets and continuous-depth models. The paper deserves a serious referee. It should go to peer review with the expectation that the authors tighten the abstract and either extend the no-go theorem to open-domain vector fields or explicitly scope the claim.","headline":"A solid threshold result for globally defined neural DDEs, but the abstract sells a cleaner and more general picture than the theorems actually deliver.","tokens_in":65485,"tokens_out":3160,"would_cite":true,"duration_ms":32509,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["34K07","34K19","58K05","58K45","68T07"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that the expressivity of neural delay differential equations is controlled by the memory capacity Kτ — the product of the vector field's Lipschitz constant and the delay — so that universal approximation fails for small…","keywords":["neural networks","neural DDEs","DDEs with small delay","Morse functions","universal approximation","universal embedding","memory capacity","DenseResNets"],"falsifier":"Decisive test: try to approximate the scalar target x ↦ x² to accuracy ε < r₀²/2 on a ball using any non-augmented neural DDE with Kτe < 1 whose vector field is defined only on a proper open subset of the history space; success would show the small-memory obstruction does not cover the most general architecture, while failure would confirm that the obstruction is tied to the memory threshold itself.","tokens_in":64471,"feed_emoji":"🧠","tokens_out":11037,"duration_ms":97225,"temperature":0.7,"pith_summary":"This paper asks whether adding memory to residual networks — in the continuous limit, letting the state derivative depend on a delayed segment of its own history, which is what a neural delay differential equation does — can remove the known expressivity ceiling of non-augmented neural ODEs and feed-forward networks. Its answer is that memory helps only past a threshold: the decisive quantity is Kτ, the product of the vector field's Lipschitz constant and the delay, which the paper calls the memory capacity. Below the threshold, every solution is exponentially attracted to a finite-dimensional family of special solutions governed by an ordinary differential equation, so the DDE inherits the topological obstruction that prevents non-augmented ODEs from approximating smooth functions with local extrema, and universal approximation fails. Above the threshold, an explicit construction embeds every globally Lipschitz continuous map exactly, and widening the DDE beyond the sum of input and output dimensions relaxes the requirement further. If correct, this gives a quantitative rule for when long shortcut connections in densely connected networks buy genuine expressive power rather than just more parameters.","feed_headline":"Memory capacity, not delay, decides what neural DDEs can learn","feed_subtitle":"Below the Lipschitz-times-delay threshold these networks hit an ODE-style wall; above it, every Lipschitz map fits.","key_machinery":"The load-bearing object is the memory capacity Kτ: the delay τ is the length of the longest shortcut connection in the DenseResNet whose infinite-depth limit is the neural DDE, and K is the global Lipschitz constant of the vector field, the continuous analogue of how strongly the activation functions amplify differences between layers. The negative side rests on the small-delay theory of functional differential equations. When Kτe < 1, every solution of the DDE is exponentially attracted to a finite-dimensional inertial manifold whose points are 'special solutions', and those special solutions are exactly the solutions of a single ordinary differential equation, so the infinite-dimensional system behaves like an ODE with only a mild Lipschitz inflation. The positive side rests on a direct construction: the vector field is built so that on [0, τ] the delayed argument equals the constant input, the target value Ψ(x) is reached by explicit integration at time τ, and the field then switches off, yielding exact embedding for every globally Lipschitz Ψ. The negative proof enforces the obstruction by choosing a Morse function — smooth with non-degenerate critical points — as the target and using the Jordan–Brouwer separation theorem to show that the output map's level sets cannot separate the interior from the boundary of a ball around the extremum.","core_discovery":"The paper's central claim is that the universal approximation property of neural DDEs is governed by the memory capacity Kτ — the product of the global Lipschitz constant K of the vector field and the delay τ — not by the mere fact that the phase space is infinite-dimensional. For globally defined, weakly nonlinear vector fields with Kτe < 1, no non-augmented neural DDE with bounded weights can approximate certain smooth targets: the paper constructs a smooth function with a non-degenerate local extreme point and proves, via exponential attraction toward ODE-governed special solutions and a level-set separation argument, that every such network misses it by a fixed positive error (Theorems 5.4 and 3.13). For Kτ ≥ 2(1 + KΨ/(w w̃)), in contrast, any globally Lipschitz continuous map can be embedded exactly as the time-T map of a non-augmented neural DDE with Lipschitz constant K and delay τ — the universal embedding property (Theorem 3.9) — and with dimension m ≥ n + q the embedding works for every delay, including zero (Theorem 3.14). Read together, the theorems map out a transition: increasing Kτ carries the architecture from ODE-like failure, through an uncharacterized intermediate window, to exact representation of all Lipschitz targets.","pith_inferences":["The paper's parameter regions leave an uncharacterized window between Kτe < 1 (no approximation) and Kτ ≥ 2 (universal embedding); a natural conjecture, not tested here, is that the true critical curve lies inside this window and is set by the spectral gap of the linearized small-delay operator rather than by the two explicit thresholds.","Transferred back through the Euler discretization of Section 2.1, the threshold predicts a critical shortcut length for discrete DenseResNets: networks whose longest inter-layer connection spans fewer than roughly 1/(Kδ) layers should exhibit the same approximation obstructions as plain ResNets — a quantitative architectural prediction that could be tested numerically without training full continu","Because the positive construction writes the target Ψ directly into the vector field, its practical counterpart is a parameterized family rich enough to approximate every weakly nonlinear field on [0, T] × C; Theorem 3.4(b) makes universality for concrete parameterized architectures conditional on exactly that richness, so checking it becomes the real engineering question.","The level-set separation mechanism suggests the obstruction is topological: local extrema force an interior-versus-boundary separation that Lipschitz-small DDE flows cannot achieve, whereas saddle points are explicitly expected to be approximable — implying that classification-type output functions with well-separated level sets may sit right at the boundary of what small-memory networks can learn"],"forward_implications":["Non-augmented neural DDEs and their parameterized versions inherit the expressivity limits of non-augmented neural ODEs whenever Kτ is below roughly 1/e: no amount of weight tuning or parameterization choice can give them the universal approximation property (Theorems 3.13 and 3.4(a)).","Once Kτ ≥ 2(1 + KΨ/(w w̃)), any globally Lipschitz target is exactly representable, so the memory threshold functions as a design prescription: to enlarge what a DenseResNet-style architecture can express, increase the longest shortcut length or the amplification of the vector field.","For augmented architectures with m ≥ n + q, universal embedding holds even at τ = 0, so memory is not what unlocks expressivity when the state dimension already exceeds input plus output; the benefit of Kτ is specific to the non-augmented case.","The obstruction is generic rather than pathological: Morse functions with local extrema are dense in the spaces of smooth functions, so the failing target is not a specially crafted exception."],"supporting_citations":[{"why":"Supplies the basic neural DDE architecture, the DenseResNet modeling link, and the τ = T universal embedding result that the paper extends to intermediate delays.","marker":"[59]"},{"why":"Supplies the inertial and slow manifold framework for delay equations with small delay, the source of the special-solution machinery used in Section 4.","marker":"[4]"},{"why":"Supplies the existence of special solutions and the ordinary differential equation that generates them under Kτe < 1 (Theorems 4.5 and 4.11).","marker":"[12]"},{"why":"Supplies the exponential attraction theorem (Theorem 4.7) showing every solution converges to a special solution.","marker":"[27]"},{"why":"Supplies the universal embedding definition, the augmented neural ODE construction mimicked in Theorem 3.14, and the non-universal approximation result for non-augmented neural ODEs.","marker":"[32]"},{"why":"Supplies the local existence, uniqueness, and differentiability theory for functional differential equations underlying Lemma 2.5 and Theorem 4.1.","marker":"[17]"},{"why":"Supplies the Jordan-Brouwer separation theorem used in Lemma 5.13 and in the proof of Theorem 5.4.","marker":"[35]"},{"why":"Supplies the Morse-function viewpoint and the density of Morse functions that makes the obstruction generic.","marker":"[33]"},{"why":"Supplies the Morse-Palais lemma used to normalize the target function to quadratic form near the extreme point.","marker":"[40]"}],"fun_headline_variants":["Neural DDEs need a memory threshold for universal approximation","Memory capacity, not delay, sets neural DDE approximation limits","Kτ threshold: from ODE-like limits to full representation","Infinite dimension isn't enough—memory capacity decides neural DDE power","Small memory wall: neural DDEs miss smooth targets until Kτ grows"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The negative results assume the vector field is defined on the whole space of continuous histories, globally Lipschitz, and bounded at zero, and the paper does not prove that vector fields living only on proper open subsets can be extended with the same Lipschitz constant and bound.","fun_headline_variants_meta":{"raw":{"variants":["Neural DDEs need a memory threshold for universal approximation","Memory capacity, not delay, sets neural DDE approximation limits","Kτ threshold: from ODE-like limits to full representation","Infinite dimension isn't enough—memory capacity decides neural DDE power","Small memory wall: neural DDEs miss smooth targets until Kτ grows"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000278,"raw_usage":{"total_tokens":1767,"prompt_tokens":1170,"completion_tokens":597,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":786,"completion_tokens_details":{"reasoning_tokens":506}},"tokens_in":786,"tokens_out":597,"duration_ms":5617,"temperature":1.0,"reasoning_tokens":506,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:21:29.428160+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Decisive test: try to approximate the scalar target x ↦ x² to accuracy ε < r₀²/2 on a ball using any non-augmented neural DDE with Kτe < 1 whose vector field is defined only on a proper open subset of the history space; success would show the small-memory obstruction does not cover the most general architecture, while failure would confirm that the obstruction is tied to the memory threshold itself.","supporting_citations":[{"cited_title":"Neural Delay Differential Equations","cited_arxiv_id":"2102.10801","evidence_quote":"Supplies the basic neural DDE architecture, the DenseResNet modeling link, and the τ = T universal embedding result that the paper extends to intermediate delays."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the inertial and slow manifold framework for delay equations with small delay, the source of the special-solution machinery used in Section 4."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the existence of special solutions and the ordinary differential equation that generates them under Kτe < 1 (Theorems 4.5 and 4.11)."},{"cited_title":"Jarnik and J","cited_arxiv_id":null,"evidence_quote":"Supplies the exponential attraction theorem (Theorem 4.7) showing every solution converges to a special solution."},{"cited_title":"Kuehn and S.-V","cited_arxiv_id":null,"evidence_quote":"Supplies the Morse-function viewpoint and the density of Morse functions that makes the obstruction generic."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Morse-Palais lemma used to normalize the target function to quadratic form near the extreme point."}],"review_version":1}