{"id":"e01d0e58-782d-4952-ad0a-e65c9050fd37","arxiv_id":"2507.12563","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"On a Berger plate benchmark, state-of-the-art neural surrogates fail in long autoregressive rollouts, and time-domain error metrics miss the resulting spectral errors.","lead":"This paper tests several neural network models as fast stand-ins for simulating a nonlinear drum membrane, comparing them against a traditional physics-based solver. It finds that the models look accurate on short clips but fail over long audio sequences, and the usual time-domain metrics hide the failures.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The §4 spectral floor is not computed and the stated Gaussian-width argument is directionally wrong: the minimum σ yields the broadest spectrum, with only ~4 dB roll-off at 50 m⁻¹, so 'excess' high-wavenumber energy may be mislabeled.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the spectral diagnosis depends on an uncomputed expectation that the ground truth has negligible energy above roughly 50 m⁻¹. My analysis strengthens this concern by showing the stated reasoning is not merely unverified but directionally incorrect: the minimum Gaussian standard deviation, 0.02 m, produces the broadest spatial spectrum, with only about 4 dB roll-off in power at 50 m⁻¹ and 8.5 dB at 70 m⁻¹. Therefore, unless the actual FTM ground-truth spectra happen to be much steeper than the initial-condition spectrum (e.g., due to strong damping over the 0.25 s averaging window), the label 'excess' high-wavenumber energy is not justified. This is the single most load-bearing concern because the paper's second main claim — that time-domain metrics are insufficient and spectral diagnostics reveal model failures — rests on this floor. Other issues, such as the missing nonlinear coefficient and dataset/code availability, are reproducibility gaps rather than correctness threats to the central argument, and the mismatch between the 0.25 s quantitative rollout and the full-second qualitative spectrograms is secondary because the models already fail within 0.25 s. If the concrete test shows the ground truth does have non-negligible energy in the 50–100 m⁻¹ band, the paper's spectral evidence weakens, but the high MAE in Fig. 1 still supports the negative conclusion that the models are not suitable in their current state for long-sequence plate synthesis. The reader's CONDITIONAL verdict is therefore appropriate, and this stress-test does not move it.","tokens_in":10809,"tokens_out":9006,"duration_ms":100437,"concrete_test":"Compute the radial spatial power spectrum of the ground-truth test trajectories used in Fig. 3 (or regenerate them with the stated FTM parameters, including the missing nonlinear coefficient C_NL) and compare the average power in the 50–100 m⁻¹ band against the peak. Also compute the analytic spectrum of the Gaussian initial velocity with σ = 0.02 m. If the ground-truth power in that band is within 10 dB of the peak, the 'excess energy' labeling is mis-calibrated; if it is more than 20 dB down, the paper's floor assumption is validated. Re-plot Fig. 3 with the actual ground-truth floor indicated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Section 4 (Fig. 3), the paper labels model outputs as having an 'excess of energy in the higher wavenumbers, specially beyond 50 m−1, where due to the minimum standard deviation of the gaussian-shaped initial conditions (see Appendix B) we would not expect to see much energy.' This expected-energy floor is asserted qualitatively and never computed from the actual FTM-generated ground-truth spectra. The inference is also mathematically backwards for a Gaussian: an initial velocity profile with spatial std σ has a wavenumber power spectrum proportional to exp(−σ²k²). The minimum σ in the dataset (0.02 m, Appendix B) therefore yields the *broadest* spectrum and the *most* high-wavenumber energy, not the least. Quantitatively, at k = 50 m⁻¹ the power is only exp(−1) ≈ 0.37 (−4.3 dB) of the peak, and at k = 70 m⁻¹ it is exp(−1.96) ≈ 0.14 (−8.5 dB). Thus the statement that 'we would not expect to see much energy' beyond 50 m⁻¹ is unsupported by the stated parameters. If the actual ground-truth trajectories contain comparable energy in the 50–100 m⁻¹ band, the central diagnostic claim that models 'add' excess high-wavenumber energy is overstated, and the paper's key motivation for spectral evaluation is weakened. The core negative result (high rollout MAE) may still hold, but this load-bearing spectral evidence needs verification against the true ground-truth spectrum.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates seven neural surrogates for the Berger nonlinear plate model used in physical modelling synthesis: a Fourier Neural Operator, three Koopman/LTI variants, S5, and LRU with and without the pushforward trick. Each model is trained on short temporal blocks and then evaluated both on single-block prediction and on autoregressive long-sequence rollout. The main empirical claims are that single-block prediction errors do not predict rollout quality, that LRU is the strongest single-block model but still degrades substantially in rollout, and that spectral analysis reveals an excess of high-wavenumber spatial energy and a decoupling between spatial and temporal spectral behavior in the model outputs. The paper concludes that current neural surrogates are not suitable for long-sequence audio synthesis and that spectral diagnostics are needed in evaluation.","tokens_in":11175,"tokens_out":4883,"duration_ms":57672,"significance":"If the findings hold, the paper provides a useful negative benchmark for neural PDE surrogates in the audio physical-modelling setting. The study is honest in reporting failure rather than overclaiming success, and it addresses a real gap in evaluation practice by going beyond scalar time-domain errors and examining spectral content. The normalized metrics, the comparison across several model families, and the three-seed reporting are strengths, as is the availability of the architecture code. The main weakness is that the spectral 'excess energy' conclusion rests on an unsupported ground-truth spectral floor, and the strong universal negative claim is based on a small test set. The central rollout-error finding is credible, but the spectral evidence needs to be verified against the actual FTM ground-truth spectra before the paper's main evaluation recommendation is fully established.","major_comments":[{"comment":"The claim that 'we would not expect to see much energy' beyond wavenumber 50 m^-1 is not supported by the stated initial-condition parameters. For a Gaussian velocity profile with spatial standard deviation sigma, the wavenumber power spectrum decays as exp(-sigma^2 k^2); the minimum sigma of 0.02 m yields the broadest spectrum, with only -4.3 dB at k = 50 m^-1 and -8.5 dB at k = 70 m^-1 relative to the peak. The paper should compute the actual radial power spectrum of the FTM ground truth, quantify the model excess in specific wavenumber bands, and revise the 'models add high-wavenumber energy' claim accordingly. This is load-bearing because the spectral diagnosis is a central motivation for the paper's main recommendation that spectral evaluation is necessary.","section":"Section 4, Fig. 3"},{"comment":"The conclusion that 'none of the models are able to capture the dynamics of the plate in the long sequence prediction task' is a strong universal negative statement supported by a test set of only 10 trajectories and three seeds. The paper should report the distribution of rollout errors across the test trajectories and seeds, show confidence intervals, and frame the negative conclusion as applying to the evaluated benchmark rather than to the entire class of neural surrogates. This does not require new experiments, but the current wording overgeneralizes from a small sample.","section":"Section 4 and Appendix C.4"},{"comment":"The paper describes the long-sequence task as generating up to 4000 steps, which at a 16 kHz sampling rate is only 250 ms and is far shorter than the 'hundreds of thousands of timesteps' mentioned in the introduction. The text also refers to spectrograms 'for the full second (16000 samples)', while Figure 3 is described as averaged over the first 4000 samples. Please clarify exactly which rollout lengths are used in each figure and, if possible, report how the errors evolve over the full 1 s trajectory. Without this clarification, the link between the reported rollout failure and audio-scale synthesis remains unclear.","section":"Section 4, Fig. 2 and Fig. 3"}],"minor_comments":[{"comment":"There are several typographical errors: 'This is is particularly relevant' in Section 1, 'neccesarily' and 'woud' in Section 5, and 'as see in Fig. 2b' in Section 4. These should be fixed in a revision.","section":"Section 1 and Section 5"},{"comment":"The caption says 'output legth 49'; this should be 'output length 49'.","section":"Figure 1 caption"},{"comment":"The sentence 'the last step of the output is used as input for the next prediction of the next block' is redundant; 'the last step of the output is used as input for the next block' would be clearer.","section":"Section 4"},{"comment":"The legend and line styles for the ground-truth curve should be made more visually distinct, for example with a thicker black line, to make the comparison easier for readers with color vision deficiencies.","section":"Figure 3"},{"comment":"The symbols S(u) and S0 are used in the nonlinear tension term but S0 is not explicitly defined in the main text or appendix; please define it at first use.","section":"Appendix A, Eq. (8)"}],"recommendation":"major_revision","confidential_remarks":"This is an honest and potentially useful negative benchmark. The rollout-error result is credible and the architecture comparisons are useful, but the spectral-floor argument in Section 4 is currently incorrect as written and is load-bearing for one of the paper's two main claims. The small test set and the ambiguous rollout-length reporting also need attention. I would be willing to accept after the spectral evidence is recomputed against the actual ground-truth spectra and the claims are appropriately calibrated."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a worthwhile empirical comparison of FNO, Koopman-style LTI autoencoders, LRU, and S5 on the Berger plate at audio rates, and the core time-domain negative result holds up. The paper's most interesting contribution—the spectral discrepancy analysis—is also its weakest link, because the expected-energy floor is asserted from a mathematically backwards reading of Gaussian initial conditions.\n\nWhat's new and good: this is the first head-to-head of these model families on a nonlinear plate model for synthesis. Tables 1 and Figure 1 cleanly show that single-block error does not predict long-rollout error, that LRU is the best single-block model yet still degrades, and that the pushforward trick does not meaningfully help. The observation that FNO's spatial and temporal frequency content diverge during rollouts is a genuinely useful diagnostic for the learned-PDE-surrogate community. Credit where due: the negative result is stated plainly and the training details are mostly transparent.\n\nSoft spots, in proportion:\n\n1. The Section 4 spectral floor is not computed, and the argument as written is wrong. A Gaussian velocity profile with smaller standard deviation gives a broader wavenumber spectrum, not a narrower one. With the dataset's minimum sigma of 0.02 m, the power at 50 m−1 is only about 4 dB below the peak—hardly 'not much energy.' So the label of 'excess' high-wavenumber energy is not established. The authors should compute the actual ground-truth power spectra and compare against them. This is load-bearing for the spectral diagnostic, though the time-domain negative result survives independently.\n\n2. Reproducibility gaps: no dataset or data-generation code, the Berger nonlinear coefficient C_NL is never given, and the test set is ten trajectories. Ten is thin for a benchmark meant to guide evaluation practice.\n\n3. The quantitative evaluation covers 250 ms while the spectrograms show a full second; the text should reconcile this timing gap.\n\nWho this is for: people working on neural surrogates for physical modeling synthesis, and anyone evaluating learned PDE solvers on audio. It deserves a serious referee and likely publication after revision, with the spectral analysis recomputed against ground truth and the missing artifacts supplied. I would cite it for the negative result, with a caveat on the spectral floor.","headline":"Useful negative-result benchmark for neural plate surrogates, but the spectral 'excess' diagnosis rests on a backwards Gaussian-width argument and needs recomputation against ground truth.","tokens_in":11676,"tokens_out":3224,"would_cite":true,"duration_ms":33828,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper finds that neural surrogates trained on short sequences fail to capture the long-sequence dynamics of nonlinear elastic plates, and shows why time-domain error alone is a misleading evaluation metric.","keywords":["physical modelling synthesis","nonlinear elastic plates","Berger plate model","neural surrogates","long sequence prediction","spectral diagnostics","Koopman operator","state space models"],"falsifier":"Compute the time-averaged radial spatial power spectrum of the ground-truth FTM trajectories themselves for the stated 0.02 to 0.1 m Gaussian initial widths, and compare the energy above 50 $m^{-1}$ in ground truth and model outputs; if the ground truth carries non-negligible energy in that band, the excess-energy diagnosis collapses. A simpler check is whether any model, when rolled out for a full second, conserves total modal energy or grows without bound within a few hundred steps.","tokens_in":10617,"feed_emoji":"🥁","tokens_out":6141,"duration_ms":63921,"temperature":0.7,"pith_summary":"The paper tests whether neural networks can replace numerical solvers for synthesizing drum-like plate sounds, by training six architectures on short blocks and rolling them out autoregressively. It finds that although some models look good on single-block prediction, none captures the plate dynamics over long sequences. The key diagnostic is spectral: models inject excess energy at high spatial wavenumbers, and the Fourier Neural Operator decouples spatial from temporal frequency behaviour. If the result holds, current neural surrogates are not suitable for real-time physical-modelling audio synthesis, and evaluation practice must include spectral diagnostics.","feed_headline":"Neural drum-plate simulators fall apart on long sequences","feed_subtitle":"Rolled out for a second, even the best model adds spurious high-frequency energy that time-domain error hides.","key_machinery":"The load-bearing machinery is the Berger plate model, a nonlinear PDE for thin elastic plates, projected onto the plate's linear eigenfunctions so that the nonlinearity appears as a global tension modulation of otherwise linear modal oscillators. Ground-truth trajectories are generated with the Functional Transformation Method (FTM), and models are trained with MSE loss using short output blocks, temporal bundling, and the pushforward trick. The evaluation compares normalized time-domain errors but also uses spectrograms and time-averaged radial spatial power spectra to reveal where energy accumulates in space and time. The Koopman-inspired LTI models use autoencoders with diagonal linear latent dynamics, whose eigenvalues are clipped to remain stable.","core_discovery":"The paper's central claim is stated in its conclusions: \"none of the models are able to capture the dynamics of the plate in the long sequence prediction task, making them not suitable in the current state.\" Even the best time-domain model, the Linear Recurrent Unit, degrades badly in autoregressive rollout, and training with the pushforward trick does not meaningfully help. The paper shows that looking only at prediction error in the time domain is insufficient, because the models tend to add energy higher in the spectrum, especially beyond roughly 50 $m^{-1}$ where the Gaussian initial conditions imply little energy should exist. In the FNO case, the spatial spectrum looks similar across sequence lengths while the temporal spectrum becomes unstable, suggesting the model is not learning coherent spatiotemporal modes. This decorrelation is linked to a conceptual issue: Koopman-inspired autoencoder models compress into a low-dimensional latent space rather than expanding into the high-dimensional space that Koopman theory would suggest.","pith_inferences":["A direct test of a spectral-loss training variant, such as multi-resolution spectrogram loss, could reveal whether the high-wavenumber excess is an artifact of MSE training or an architectural limit.","Because the ground truth is generated from Gaussian velocity profiles, percussive or impulse-like excitations with broadband spatial content may change the failure pattern; evaluating on such inputs is a natural next experiment.","The plate's dispersion relation, linking spatial and temporal frequencies, could serve as a physics-based metric: a surrogate that has truly learned the modes should trace the same dispersion curve as the FTM solution.","The diagnostic recipe implied by the paper, reporting energy in physical bands above the excitation floor, could be applied to other learned physical-modelling synthesis methods."],"forward_implications":["None of the evaluated models, trained and rolled out this way, is suitable as a stand-in for numerical solvers in nonlinear plate audio synthesis.","Single-block prediction quality does not transfer to autoregressive rollout, so reported time-domain error on short blocks should not be used as the sole success metric.","The excess energy concentrated above roughly 50 m^-1 means neural surrogates colour the sound with spurious high-frequency content even when overall error looks low.","Pushforward training and temporal bundling, as applied here, do not meaningfully stabilise long rollouts.","Spatial and temporal frequency content must be evaluated together, because spatial agreement can coexist with temporal instability, as the FNO results show."],"supporting_citations":[{"why":"The nonlinear plate model under study: large-deflection plate dynamics with homogeneous tension modulation.","marker":"Berger (1954)"},{"why":"Projection of the Berger PDE onto linear eigenfunctions and the modal nonlinearity used to build the dataset.","marker":"Avanzini et al. (2012)"},{"why":"The Functional Transformation Method that generates the ground-truth trajectories.","marker":"Trautmann & Rabenstein (2003)"},{"why":"Defines the autoregressive rollout and distribution-shift problem, and the temporal bundling and pushforward tricks used in training.","marker":"Brandstetter et al. (2022b)"},{"why":"The Fourier Neural Operator architecture evaluated in the comparison.","marker":"Kovachki et al. (2023)"},{"why":"The autoencoder-based Koopman architecture underlying the LTI models.","marker":"Lusch et al. (2018)"},{"why":"The Linear Recurrent Unit architecture that gives the best single-block and long-sequence time-domain results.","marker":"Orvieto et al. (2023)"},{"why":"The S5 state-space model evaluated as a competitor.","marker":"Smith et al. (2023)"},{"why":"Prior evidence that neural operators degrade beyond roughly ten times the training length, motivating the long-rollout evaluation.","marker":"Michałowska et al. (2024)"},{"why":"Source of the physical parameters of the mylar membrane used to generate the dataset.","marker":"Fletcher & Rossing (1991)"}],"fun_headline_variants":["Neural drum sims fail long sequences, add spurious highs","Best neural drum model still degrades on long audio","Time-domain error hides spectral blowup in plate nets","Koopman autoencoders miss high-dimensional plate dynamics","No neural plate model stable for long audio synthesis"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that models add excess high-wavenumber energy rests on the assumption that the true plate response has almost no energy above roughly 50 $m^{-1}$, inferred only from the minimum 0.02 m width of the Gaussian initial conditions, and this expected floor is never computed from the actual ground-truth power spectra.","fun_headline_variants_meta":{"raw":{"variants":["Neural drum sims fail long sequences, add spurious highs","Best neural drum model still degrades on long audio","Time-domain error hides spectral blowup in plate nets","Koopman autoencoders miss high-dimensional plate dynamics","No neural plate model stable for long audio synthesis"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000745,"raw_usage":{"total_tokens":3276,"prompt_tokens":852,"completion_tokens":2424,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":468,"completion_tokens_details":{"reasoning_tokens":2345}},"tokens_in":468,"tokens_out":2424,"duration_ms":22611,"temperature":1.0,"reasoning_tokens":2345,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:44:31.188107+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the time-averaged radial spatial power spectrum of the ground-truth FTM trajectories themselves for the stated 0.02 to 0.1 m Gaussian initial widths, and compare the energy above 50 $m^{-1}$ in ground truth and model outputs; if the ground truth carries non-negligible energy in that band, the excess-energy diagnosis collapses. A simpler check is whether any model, when rolled out for a full second, conserves total modal energy or grows without bound within a few hundred steps.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The nonlinear plate model under study: large-deflection plate dynamics with homogeneous tension modulation."},{"cited_title":"Efficient synthesis of tension modulation in strings and membranes based on energy estimation","cited_arxiv_id":null,"evidence_quote":"Projection of the Berger PDE onto linear eigenfunctions and the modal nonlinearity used to build the dataset."},{"cited_title":"and Rabenstein, R","cited_arxiv_id":null,"evidence_quote":"The Functional Transformation Method that generates the ground-truth trajectories."},{"cited_title":"Neural Operator : Learning Maps Between Function Spaces With Applications to PDEs","cited_arxiv_id":null,"evidence_quote":"The Fourier Neural Operator architecture evaluated in the comparison."},{"cited_title":"L., Gu, A., Fernando, A., Gulcehre, C., Pascanu, R., and De, S","cited_arxiv_id":null,"evidence_quote":"The Linear Recurrent Unit architecture that gives the best single-block and long-sequence time-domain results."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The S5 state-space model evaluated as a competitor."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source of the physical parameters of the mylar membrane used to generate the dataset."}],"review_version":1}