{"id":"89b88c97-1ec7-469d-b06b-a97d708ba386","arxiv_id":"2608.10738","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Two training strategies, data-reset MPC windows and a Floquet-regularized periodic encoding, convert short-horizon accuracy of semi-autonomous neural ODEs into long-horizon guarantees, linear in elapsed periods for stable limit cycles.","lead":"This paper shows how to keep a neural-network model of a dynamical system accurate over long time horizons by resetting it in two ways: restarting it from fresh data on short windows, or exploiting the natural contraction of a stable limit cycle. It proves error bounds that grow only linearly with elapsed periods for the limit-cycle strategy, and runs experiments that measure the hypotheses behind each guarantee.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Assumption 4.4 is the uninstantiated pivot: no trained model in the paper is shown to satisfy one-period tube closeness at Theorem 4.13's smallness scale, so the linear-in-period claim remains conditional.","rationale":"The reader's weakest_assumption already identifies Assumption 4.4, and I agree. The paper is an honest conditional analysis and the theorems are not obviously wrong, so no change to CONDITIONAL. I would sharpen the reason: this assumption is not just 'hard to verify' - the paper gives no route from its training objectives to it, and the measurements show the deployed models far outside it. The grid-verification gap in Assumption 3.4 is a second, real limitation, but the Floquet theorem is the strongest claim and its premise is the less secure. A targeted numerical check on the warm-started models would settle whether any trained instance enters the theorem's regime; absent that, the abstract's linear bound should be read as conditional on an unshown hypothesis, not as an achieved behavior of the method.","tokens_in":52711,"tokens_out":9993,"duration_ms":101297,"concrete_test":"Choose the warm-started autonomous models of Table 5.2 (the only models in the presumed regime of Theorem 4.13) and compute r1 from Lemma 4.7 for each benchmark. Then evaluate E_tube := sup_{x in N_{2r1}, t in [0,T_hat+2]} ||Phi_Theta(t;x)-Phi_f(t;x)|| on a dense grid or Monte Carlo sample of the tube, using high-accuracy reference integration. Compare E_tube with the smallness threshold epsilon0 of Theorem 4.13 (or, if epsilon0 is not made explicit, with the explicit bounds delta0/4 and the confinement/closeness requirements in the proof). If E_tube exceeds the admissible range for every trained model, then Assumption 4.4 is not instantiated and the linear-in-period claim (4.6) is not demonstrated for the trained models; only the orbital claim remains experimentally supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The internal logic of Theorem 4.13 is coherent; the load-bearing problem is that the proof's key premise is never established for any trained model. Assumption 4.4 demands one-period flow closeness epsilon on the whole tube N_{2r1}, and the paper itself states (Section 4.1) that neither the base-orbit trajectory loss nor the Floquet loss establishes it. No theorem in Section 4 derives it from training: Corollary 4.18 certifies only the return-map multiplier, and Proposition 4.16's C1 route controls field closeness, not flow closeness over the tube. The experiments measure the gap: Section 5.4 reports median tube errors 0.88 (Stuart-Landau) and 4.84 (van der Pol), i.e., 1124x and 328x the orbit Hausdorff distance epsilon_Gamma, with C1 oscillation eta missing Lemma 4.15's admissible regime by about ten orders of magnitude. Even the warm-started autonomous arm, the one place the autonomous theorem could apply, reports one-period tube errors 0.13 and 0.61-0.78 and is explicitly labeled 'not a certified instance' (smallness conditions and constants not evaluated). For the deployed periodic architecture the difficulty is structural: Proposition 4.20 shows an exactly T_hat-periodic field cannot have both a q-contraction of its stroboscopic map on a set containing Gamma and one-period closeness below (1/2)(1-q)diam(Gamma). Thus the abstract's 'dynamical reset' linear trajectory bound is a theorem whose main hypothesis the paper never shows its training can meet; the experimentally grounded guarantee is the orbital bound (4.15), not (4.6).","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies how the certified error of a learned SA-NODE flow grows with the horizon and proposes two state-reset strategies that avoid the double-exponential barrier of the monolithic certificate (Theorem 2.1; constant (2.1)). The model-predictive strategy partitions the horizon and restarts each window from the true state: under per-window tolerance and a bounded, time-uniformly regular reachable tube (Assumptions 3.1–3.4), Theorem 3.5 yields uniform error ≤ ε with total width O(N C^2 ε^{-2}), linear in T for the uniform partition, while predicted-IC deployment obeys only the exponential bound of Proposition 3.7. The Floquet strategy targets autonomous systems with a hyperbolic stable limit cycle: with a certified contraction of the learned return map (Assumption 4.6) and one-period flow closeness on a tube (Assumption 4.4), Theorem 4.13 gives ||Φ_Θ(t;x0) − Φ_f(t;x0)|| ≤ Cε(1 + t/T̂); Corollary 4.18 links the Floquet loss to the planar multiplier; Remark 4.19 shows the surrogate degenerates to det M for the deployed time-periodic encoding; Propositions 4.20–4.21 prove a quantitative incompatibility of one-period closeness with stroboscopic contraction; Theorem 4.22 supplies the orbital bound (4.15) with measured hypotheses. Experiments: forced Duffing and pendulum for MPC; Stuart–Landau and van der Pol for the Floquet ablation, plus a warm-started autonomous arm.","tokens_in":52907,"tokens_out":31654,"duration_ms":262890,"significance":"The conditional theorems appear sound: Proposition 4.20's three-line proof is correct, the adapted-norm contraction machinery of Lemmas 4.9–4.12 is standard and coherently assembled, Corollary 4.18 constructs the learned fixed point without assuming Assumption 4.6 (a genuine non-circular step), and the MPC budget of Theorem 3.5 follows from its stated assumptions. The paper ships code, models, configs, and figure scripts; reports an independent adaptive-integrator recomputation reproducing the reported multipliers to 10^{-3}; and makes falsifiable quantitative predictions (tube-error lower bounds ε ≥ 0.30 and ε ≥ 1.93; transient decay governed by the measured ρ(M); window count linear in T) that its own measurements confirm. Its candor is exemplary: the demanding character of Assumption 4.4, the measured-not-certified status of Theorem 4.22's inputs, and the grid-wise verification of Assumption 3.4 are all disclosed in place. The significance is nonetheless capped by a wide gap between headline and instantiation: the abstract's 'certified contraction … confines the error to linear growth in the number of elapsed periods' is realized by no trained model in the paper.","major_comments":[{"comment":"Assumption 4.4 is the load-bearing premise of Theorem 4.13 and Corollary 4.18, and no trained model in the paper is shown to satisfy it. The paper states (Section 4.1, paragraph following Assumption 4.4) that the assumption 'asks for one-period flow closeness at every initial state of the tube N_{2r1}', that a small trajectory loss on the base orbit does not establish it, that a small Floquet loss does not establish it, and that the two certification routes 'presuppose it rather than prove it'. Section 5.4 then quantifies the gap: median tube errors are 0.88 (Stuart–Landau) and 4.84 (van der Pol), i.e. 1124× and 328× the orbit Hausdorff distance ε_Γ, and the measured C^1 oscillation η misses Lemma 4.15's admissible regime by about ten orders of magnitude. The warm-started autonomous arm, the only setting in which the autonomous Theorem 4.13 could apply, has tube errors 0.13 and 0.61–0.78 and is labeled 'not a certified instance: the smallness conditions and constants are not evaluated'. The honest conclusion the paper itself draws in Section 5.4 — 'one-period closeness on the tube, and with it Theorem 4.13 and Lemma 4.15, is not observed at any tuning' — should therefore be reflected in the abstract's Floquet claim, which currently presents the linear bound as the strategy's outcome rather than as an idealized conditional statement.","section":"§4.1 (Assumption 4.4); §5.4"},{"comment":"The obstruction results make the gap structural for the deployed architecture, so the framing issue is not merely empirical. Proposition 4.20 shows that an exactly T̂-periodic learned field with q-contraction of the stroboscopic map on a set containing Γ forces one-period error ε ≥ (1/2)(1−q) diam(Γ); the consistency check in Section 5.4 measures grid Jacobian norms 0.052 (Stuart–Landau) and 0.186 (van der Pol), giving lower bounds ε ≥ 0.30 and ε ≥ 1.93, and the measured tube errors (0.88, 4.84) lie above both. Proposition 4.21 extends the incompatibility to the locally measured ρ(M), and Remark 4.19 shows the scalar surrogate controls det M rather than the spectral radius. Taken together these imply that the linear trajectory bound (4.6) cannot hold for the periodic-encoding models of Section 5.4, regardless of tuning, and that the training target ρ̃_T(Θ) ≤ ρ* is not a stability certificate in that class. The abstract's sentence 'a certified contraction of the learned return map confines the error to linear growth in the number of elapsed periods' promises exactly the statement that the paper's own theory rules out for its deployed architecture. The Floquet contributions should be led, in the abstract and Section 1.2, by the orbital guarantee (Theorem 4.22, (4.15)) — the one actually instantiated — with Theorem 4.13 presented explicitly as a conditional result for the idealized autonomous case.","section":"§4.4 (Propositions 4.20–4.21); abstract"},{"comment":"Even the instantiated guarantee is delivered at the level of measurement rather than certificate, which sits uneasily with the paper's 'certified' vocabulary. Theorem 4.22's discussion concedes that the implementation returns 'floating-point eigenvalues of a numerically integrated monodromy and a distance between finite samples, without residual or quadrature bounds, so the reported values are measurements of the hypotheses and not certificates of them', and that the radius r_0 'is likewise not evaluated for the initial conditions used in Section 5.4'. In the warm-started arm the fitted stroboscopic growth 0.09–0.16 per period is described as 'consistent with the envelope Cε(1+k)' of Theorem 4.13(iii), but since C is left unevaluated this is a plausibility check rather than a test of the bound. The paper itself names the remedy (validated integration with interval bounds on both the monodromy eigenvalues and the sampled distances); completing it for at least one representative model per benchmark would convert one instance of each guarantee from 'measured' to 'certified', and the abstract and Section 6.1 should be adjusted accordingly even if that upgrade is not undertaken.","section":"§4.4 (Theorem 4.22); §5.4 (warm-started arm)"}],"minor_comments":[{"comment":"Five pairwise distinct spectral quantities (ρ_T(f), ρ*, ρ_Θ, ρ̃_T, ρ(M)) are introduced in a short span; the paragraph distinguishing them is helpful, but a one-line table or footnote at first use would substantially reduce the reader's verification burden.","section":"§4.1"},{"comment":"The phrase 'the operator norm of the Jacobian of the stroboscopic map over the 0.2-tube... is at most 0.052' should read 'the sampled maximum is', since the values are grid evaluations and the next sentence correctly notes they are not a certificate; the current wording could mislead a skimming reader.","section":"§5.4"},{"comment":"The surrogate integrates the divergence over the nominal period T̂ of a curve whose actual period is T_Θ; the proof of Corollary 4.18 Step 4 handles the resulting O(ε) gap, but the definition itself should state that γ̂_Θ denotes a length-T̂ segment of the relaxed periodic orbit rather than 'the curve... sampling one nominal period'.","section":"Definition 4.17"},{"comment":"The first window of the Duffing run has length 10.4, exceeding the training horizon H = 10; the text explains that the stopping rule (3.4) is evaluated on the full remaining interval while training covers [τ_k, min(τ_k + H, T)], but a parenthetical at first mention in the experiment would prevent a misreading.","section":"§5.2"},{"comment":"The phrase 'it makes the warm start Θ_k ← Θ_{k-1} shape-preserving' is used without definition; one clause explaining that local time keeps the time-input bias in [0, τ_max] so that copied weights remain on comparable scales would remove the ambiguity.","section":"§3.1"},{"comment":"The monolithic error is reported to grow as e^{0.39t} with correlation r = 0.88 'over the growth phase', but the growth phase is not precisely delimited; a definition of the fitting window would make the reported rate reproducible.","section":"§5.3 (Figure 5.2(a))"}],"recommendation":"major_revision","confidential_remarks":"This assessment agrees with the stress-test concern: Assumption 4.4 is the uninstantiated pivot, and the paper says so itself. The mathematical content is, on my spot-checks, sound; the revision I am asking for is primarily about matching the headline claims (abstract, Section 1.2, Section 6.1) to what is actually established for trained models, and optionally upgrading one or two representative models to genuine certificates. The heaviest unscrutinized part is Appendix A.6 (Lemma 4.15); I would welcome a second referee pass there. Note for the editor: the paper is effectively a numerical-analysis contribution posted in cs.LG; the relevant readership should be kept in mind when selecting reviewers."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here is my take. The paper is a careful, conditional analysis of two reset strategies for long-horizon SA-NODE approximation, and it earns most of its claims. What is actually new: the MPC data-restarted composite bound with explicit summed-width budget (Theorem 3.5), the Floquet propagation law converting a certified contraction of a learned return map into a linear-in-period trajectory bound (Theorem 4.13), the orbital guarantee for the periodic encoding (Theorem 4.22), and the clean obstruction results (Propositions 4.20–4.21). The appendix proofs look standard but thorough, and the authors ship code and data. Credit where due: they measure the hypotheses of every guarantee and report the failures plainly.\n\nSoft spots: the headline Floquet result is load-bearing conditional. Theorem 4.13 requires Assumption 4.4, one-period flow closeness on the whole tube N_{2r1}. No trained model in the paper is shown to satisfy it. Section 5.4 reports tube errors 0.88 (Stuart–Landau) and 4.84 (van der Pol), 1124x and 328x the orbit Hausdorff distance, and the oscillation parameter misses the Lemma 4.15 regime by about ten orders of magnitude. The warm-started autonomous arm does better but is labeled \"not a certified instance.\" The authors openly state this, and they pivot to Theorem 4.22, the orbital guarantee. That is honest, but it means the abstract's \"dynamical reset\" linear trajectory bound is a theorem whose main hypothesis is never instantiated. That should be fixed in revision, either by a certified instance or by reframing.\n\nThe MPC half is the cleaner part. Theorem 3.5 is solid; its Assumption 3.4 is verified on grids rather than continuously, which the authors disclose. The predicted-IC bound Proposition 3.7 is honest about exponential growth when resets are withheld, and the width budget statement is substantive. The obstruction Proposition 4.20 is a clean three-line proof and looks correct.\n\nVerdict: the central logic holds up as a conditional analysis. The gap is in the evidence, not the mathematics. I would send this to serious peer review; a good referee should press on Assumption 4.4 instantiation and ask for a calibrated presentation. The paper is useful for people working on neural ODE long-horizon guarantees, and the MPC theorem plus obstruction are likely to be cited.","headline":"Honest conditional analysis of two reset strategies; the MPC bound is solid, the Floquet linear-in-period theorem is coherent but its key tube-closeness hypothesis is never met by any deployed model.","tokens_in":53619,"tokens_out":2732,"would_cite":true,"duration_ms":28447,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["68T07","93B45","34C25","65L05","41A30"],"pacs":[],"model":"deepseek-v4-flash","headline":"Two reset strategies — restarting from data or exploiting a limit cycle's contraction — turn one-period accuracy into long-horizon guarantees with error linear in elapsed periods, not double-exponential in the horizon.","keywords":["Neural ODEs","long-time approximation","model predictive control","Floquet theory","limit cycles","SA-NODE","Poincaré return map","time horizon barrier"],"falsifier":"Train any sequence of SA-NODEs on an autonomous limit-cycle target and measure, for each, the one-period flow error $\\varepsilon$ over a tube of radius $r$ around the cycle and the spectral radius $\\rho$ of the learned first-return map at its fixed point. A single model with $\\rho \\le 0.5$, tube-level $\\varepsilon \\le 0.01$, and stroboscopic error after 100 periods above $10\\varepsilon$ would refute the linear envelope $C\\varepsilon(1 + t/\\hat T)$ of Theorem 4.13; conversely, Proposition 4.20 predicts that for an exactly periodic learned field, stroboscopic contraction $q$ forces one-period error on the cycle to be at least $(1-q)\\operatorname{diam}(\\Gamma)/2$, so a systematic sweep over trained periodic models would show whether that obstruction is tight.","tokens_in":52354,"feed_emoji":"🔁","tokens_out":19904,"duration_ms":154578,"temperature":0.7,"pith_summary":"The paper takes on a specific failure of learned dynamics: the certified error bound of a single neural ODE trained over a long horizon grows double exponentially in the horizon length, which would make long-time prediction look hopeless. It argues this barrier is an artifact of training one monolithic network, not an intrinsic limit of the model class, and develops two training strategies that circumvent it, each built on a 'reset' of the state. The model predictive strategy partitions the horizon into short windows, restarts each window from observed data, and proves the composite model holds a user-chosen tolerance uniformly in time with a total network width that is linear in the horizon. The Floquet strategy targets autonomous systems with a stable limit cycle, trains a single periodic network whose first-return map contracts, and proves that one-period accuracy on a tube around the cycle yields trajectory error bounded by $C\\varepsilon(1 + t/\\hat T)$ — linear in elapsed periods, with no data needed at deployment. What makes the guarantees usable is that each rests on hypotheses that can be measured on the trained model; the paper measures them, and its own experiments show where the deployed architecture satisfies them and where it does not.","feed_headline":"Two resets cut neural-ODE error from double-exponential to linear","feed_subtitle":"Long-horizon prediction gains provable error bounds that grow only linearly in elapsed periods.","key_machinery":"The argument is carried by two reset mechanisms and the quantities that certify them. The data reset is a model predictive partition: one SA-NODE per window with local time, warm-started parameter copies, and the true state injected at each switch time, which confines the Grönwall factor $e^{L\\tau_{\\max}}$ to a single window; the load-bearing certificate is the uniform window constant $C_{\\tau_{\\max},K_\\infty,f}$ of Assumption 3.3, finite whenever the reachable tube is bounded with uniformly regular data. The dynamical reset is the Poincaré first-return map $P_\\Theta$ of the learned flow: a certified contraction $\\rho(DP_\\Theta(p)) \\le \\rho_* < 1$ contracts transverse errors geometrically at every return, while the one quantity contraction cannot control — the period mismatch between the learned and true clocks — accumulates linearly and produces the $1 + t/\\hat T$ factor. The Floquet loss trains the surrogate $\\tilde\\rho_T(\\Theta) = \\exp\\big(\\int_0^{\\hat T} \\operatorname{div} f_\\Theta(\\hat\\gamma_\\Theta(t))\\,dt\\big)$, which by the Liouville–Abel identity is exactly the scalar return-map multiplier in the autonomous plane, so a vanishing Floquet loss implies contraction and hence the linear bound. For the deployed periodic encoding the same surrogate degenerates to the determinant of the monodromy, so the paper computes the full stroboscopic spectrum instead and proves the orbital guarantee of Theorem 4.22, whose two hypotheses are exactly the quantities its training loop already returns.","core_discovery":"On the paper's own terms, the central claim is that long-horizon approximation error for semi-autonomous neural ODEs is governed by the reset mechanism, not by network size. For the data-restarted composite, if training meets a prescribed tolerance $\\varepsilon$ on every window and the target's reachable tube is bounded with uniformly regular data, then $\\sup_{t,x_0}\\|\\Phi(t;x_0) - \\hat\\Phi(t;x_0)\\| \\le \\varepsilon$ with a total width $\\lesssim N C^2_{\\tau_{\\max},K_\\infty,f}\\,\\varepsilon^{-2}$, which is linear in the horizon $T$ for a uniform partition (Theorem 3.5). For an autonomous target with a hyperbolic stable limit cycle, if the learned first-return map is certified to contract, $\\rho(DP_\\Theta(p)) \\le \\rho_* < 1$, and one-period flow closeness $\\varepsilon$ holds on a tube around the cycle, then every trajectory starting near the cycle satisfies $\\|\\Phi_\\Theta(t;x_0) - \\Phi_f(t;x_0)\\| \\le C\\,\\varepsilon\\,(1 + t/\\hat T)$ for all $t \\ge 0$: transverse errors are contracted at every return, and only the phase drift between the two clocks accumulates (Theorem 4.13). The paper also proves an obstruction for its deployed time-periodic architecture — an exactly periodic learned field cannot simultaneously have small one-period error and a contracting stroboscopic map — so the deployed models are covered instead by a uniform-in-time orbital bound whose two hypotheses (monodromy spectral radius and orbit distance to the cycle) are directly measurable (Theorem 4.22).","pith_inferences":["Implicit in the paper but not pursued: if the learned period $T_\\Theta$ were explicitly fit to $\\hat T$ during training, the dominant linear term of (4.6) would nearly vanish and the error envelope would flatten far below $C\\varepsilon(1+t/\\hat T)$ for many periods; the polar attainment example of Remark 4.14 makes this a directly measurable prediction — stroboscopic error should track $k\\,|T_\\The","The gap the authors report between tube-level and orbit-level one-period error (median factors of 1124 and 328 in their deployed models) points to the next bottleneck: a training objective that penalizes one-period error uniformly over the tube, rather than only on the base orbit, is the natural route to bring deployed models inside the regime of Theorem 4.13.","The obstruction for exactly periodic fields plausibly extends to any entraining learned system: a contracting stroboscopic map forgets the initial phase, so any application that needs phase timing — circadian, cardiac, power-grid models — should track the phase observable explicitly and be content with an orbital guarantee for amplitude.","Whether the horizon barrier is intrinsic is left open; if a super-linear width lower bound is ever proved, the MPC composite is optimal in $T$ up to constants, and a longer horizon sweep at fixed tolerance than the four points reported would give the first empirical evidence separating linear from mildly super-linear window growth."],"forward_implications":["A monolithic model trained once on the whole horizon cannot be certified beyond a few multiples of $1/L$: both strategies replace the double-exponential certified width of (2.1) with budgets that grow at most linearly in $T$ (data reset) or that keep a single network with error linear in elapsed periods (dynamical reset).","In the data-assisted regime the uniform guarantee is carried by the resets: the same trained windows chained from their own predictions obey only the exponential bound $\\varepsilon\\,(e^{N\\bar L\\tau_{\\max}}-1)/(e^{\\bar L\\tau_{\\max}}-1)$, so deployment on novel initial conditions needs either observed states at the switch times or dynamical stability.","The linear envelope of Theorem 4.13 is attained up to constants by a pure period mismatch: two dynamics with identical radial contraction and angular speeds differing by $O(\\varepsilon)$ achieve error at least $c\\varepsilon k$ at the $k$-th return until the phase wraps.","For the deployed exactly-periodic architecture, trajectory-wise accuracy and stroboscopic contraction are quantitatively incompatible, so the available guarantee is orbital — distance to the target cycle bounded by $\\varepsilon_\\Gamma$ plus a geometrically decaying transient — and launch-phase fidelity is lost by design.","For near-autonomous learned fields the linear bound survives with $\\varepsilon$ replaced by $\\varepsilon + C\\eta$, where $\\eta$ is the $C^1$ oscillation of the learned field about its time average; the paper's measured oscillations place its deployed models roughly ten orders of magnitude outside that regime."],"supporting_citations":[{"why":"Supplies the SA-NODE architecture, the universal approximation theorem (Theorem 2.1), and the explicit double-exponential constant (2.1) that defines the time horizon barrier the paper removes.","marker":"[16]"},{"why":"Introduces neural ODEs as continuous-time flows, the model class whose long-horizon error growth this paper studies.","marker":"[7]"},{"why":"Classical multiple shooting for parameter identification, whose node-pinning mechanics the model predictive strategy reuses with node values fixed to data.","marker":"[2]"},{"why":"Error-growth theory for one-step integration of attracting periodic orbits, the structural template for the linear-in-period law of Theorem 4.13.","marker":"[6]"},{"why":"Floquet theory source supplying the adapted-norm contraction construction, spectral-radius continuity, and the scalar multiplier identity used in certification.","marker":"[8]"},{"why":"Matrix perturbation bound used to turn C¹ closeness of the learned field into a quantitative spectral-radius certificate in higher dimension.","marker":"[1]"},{"why":"Shows activation saturation obstructs transverse contraction in standard neural ODEs, motivating the periodic time encoding the deployed architecture uses.","marker":"[20]"}],"fun_headline_variants":["Resets cut neural ODE error from double-exponential to linear","Model-predictive resets give uniform error for long horizons","Floquet strategy confines neural ODE error to linear growth","Resets turn double-exponential neural ODE error linear"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the trained model's one-period flow error stays below a small $\\varepsilon$ uniformly over a whole tube of initial conditions around the target cycle, not just on the single orbit used for training — and the paper's own measurements put its deployed models outside that regime, with tube-level errors hundreds of times the orbit-level error.","fun_headline_variants_meta":{"raw":{"variants":["Resets cut neural ODE error from double-exponential to linear","Model-predictive resets give uniform error for long horizons","Floquet strategy confines neural ODE error to linear growth","Resets turn double-exponential neural ODE error linear"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000715,"raw_usage":{"total_tokens":3302,"prompt_tokens":1122,"completion_tokens":2180,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":738,"completion_tokens_details":{"reasoning_tokens":2109}},"tokens_in":738,"tokens_out":2180,"duration_ms":17478,"temperature":1.0,"reasoning_tokens":2109,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T18:21:43.523478+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train any sequence of SA-NODEs on an autonomous limit-cycle target and measure, for each, the one-period flow error $\\varepsilon$ over a tube of radius $r$ around the cycle and the spectral radius $\\rho$ of the learned first-return map at its fixed point. A single model with $\\rho \\le 0.5$, tube-level $\\varepsilon \\le 0.01$, and stroboscopic error after 100 periods above $10\\varepsilon$ would refute the linear envelope $C\\varepsilon(1 + t/\\hat T)$ of Theorem 4.13; conversely, Proposition 4.20 predicts that for an exactly periodic learned field, stroboscopic contraction $q$ forces one-period error on the cycle to be at least $(1-q)\\operatorname{diam}(\\Gamma)/2$, so a systematic sweep over trained periodic models would show whether that obstruction is tight.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the SA-NODE architecture, the universal approximation theorem (Theorem 2.1), and the explicit double-exponential constant (2.1) that defines the time horizon barrier the paper removes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces neural ODEs as continuous-time flows, the model class whose long-horizon error growth this paper studies."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Classical multiple shooting for parameter identification, whose node-pinning mechanics the model predictive strategy reuses with node values fixed to data."},{"cited_title":"Cano and J","cited_arxiv_id":null,"evidence_quote":"Error-growth theory for one-step integration of attracting periodic orbits, the structural template for the linear-in-period law of Theorem 4.13."},{"cited_title":"Chicone , Ordinary differential equations with applications , vol","cited_arxiv_id":null,"evidence_quote":"Floquet theory source supplying the adapted-norm contraction construction, spectral-radius continuity, and the scalar multiplier identity used in certification."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Matrix perturbation bound used to turn C¹ closeness of the learned field into a quantitative spectral-radius certificate in higher dimension."}],"review_version":1}