{"id":"b9f71721-2919-4253-82e2-aa18543b80c7","arxiv_id":"2608.04222","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"TIDE is a DNS-verified, physically diverse 3D turbulence benchmark with independent ensembles that shows current neural operators barely beat persistence and that low pointwise error does not guarantee physical fidelity.","lead":"TIDE is a new high-resolution 3D turbulence dataset with 15 physical configurations, independent simulation runs, and equation-level checks, built to test machine learning models against real turbulent dynamics. A benchmark on it shows today's learned models barely beat trivial persistence and remain about twice as inaccurate as a classical spectral solver, and that lower error can hide badly distorted small-scale physics.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Equation-level verification reuses the generating solver and, for the momentum equation, is run on separate frozen-forcing triplets rather than on the released frames; an independent-solver cross-check is needed to sustain the 'DNS-verified' claim.","rationale":"The reader's weakest assumption already names the self-consistency problem; my pass sharpens it by pointing out that D2 is not even computed on the released frames, only on auxiliary frozen-forcing triplets, so the released fields' momentum balance is certified only indirectly. This matters because the benchmark's external value—'DNS-verified' and the transfer/conditioning conclusions—depends on the data representing the Navier-Stokes dynamics rather than the quirks of one implementation. I do not see fraud or sloppiness; the authors disclose single-solver scope and provide extensive gates. But the central claim as worded exceeds what the verification demonstrates. An independent-solver regeneration of one configuration would settle the question. If the cross-check passes, the current ACCEPT stands with no change; if it fails or shows systematic offsets, the dataset should be described as one-solver-verified, and the learned-model conclusions should be re-scoped accordingly. Therefore I would adjust the reader's verdict to CONDITIONAL rather than unconditional ACCEPT: acceptance is justified, but the DNS-verified label should be tied to a cross-solver check or explicitly qualified.","tokens_in":32104,"tokens_out":7815,"duration_ms":75878,"concrete_test":"Regenerate the flagship Re_lambda=86 configuration with an independently written DNS code (different dealiasing, time integrator, and forcing implementation, or a high-order finite-volume code) using the same box, grid, Reynolds number, forcing band, and OU forcing statistics, and compare the independent ensemble to the released fields on energy spectra, structure functions, derivative skewness, and the A-gate/D-gate envelopes. If the independent ensemble overlaps TIDE within seed spread and passes the same gates, the self-consistency concern is retired; if systematic offsets appear, the dataset and benchmark findings must be re-scoped to the original solver family.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim that TIDE is 'DNS-verified' rests on the Section 4/Appendix B D-group. D1 (divergence) and D4 (pressure-Poisson) are computed on released frames, but D2 (momentum residual) is not: it uses dedicated frame triplets exported with the Eswaran-Pope OU forcing state frozen between frames, so the centered time derivative can be compared with P(u×ω)+ν∇²u+f without stochastic contamination. D3 then confirms second-order time truncation on those same triplets. This validates the solver's temporal accuracy on auxiliary runs, not the momentum balance of any released frame. Moreover every check reuses the same pseudo-spectral operators, dealiasing, projection, and forcing implementation that produced the data, so a systematic error in those operators appears in both generator and verifier and passes D-group. The statistical gates are either not verdict-bearing at this Reynolds number (Kolmogorov constant, C_epsilon in Table 6) or are single-point quantities the solver family is known to reproduce. Thus 'DNS' is currently certified only as self-consistency of one solver family, and the benchmark conclusions about learned models and the missing forcing-conditioning variable may be artifacts of that family. The authors disclose the one-solver scope in Appendix K, but the verification procedure itself does not remove the circularity.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces TIDE, a 256^3 fp64 DNS corpus and benchmark for 3D incompressible turbulence, comprising 15 configurations across eight controlled physics axes, 8–16 independent realizations per configuration, pressure channels, and an acceptance protocol combining statistical gates with equation-level residuals. The companion benchmark defines five tasks (forecasting, super-resolution, sparse reconstruction, pressure recovery, subgrid-stress closure) with five audited learned baselines and trivial/informed references, plus generalization splits along the controlled axes. Headline findings are that learned models barely outperform persistence and remain about twice as far from the truth as an equations-informed spectral integrator; that lower pointwise error can coincide with badly distorted small-scale dynamics; and that forced-to-decay transfer exposes a missing conditioning input, since the stochastic drive is not part of the model input.","tokens_in":32387,"tokens_out":3892,"duration_ms":32218,"significance":"If the claims hold, TIDE is a valuable shared resource: it is, to my knowledge, the first 3D incompressible turbulence corpus that combines physically diverse configurations with per-configuration ensembles, a certified pressure channel, and a fixed acceptance standard. The paper ships machine-checked acceptance scripts, audited baseline variants with tabulated deviations, reproducible code (per-run costs and seeds), and a measured replicate-seed study, which are concrete strengths. The three-axis evaluation (pointwise error, enstrophy ratio, high-band spectral error) is a genuine methodological contribution, and the forced-to-decay conditioning finding is an interesting, falsifiable open problem. However, the strength of the 'DNS-verified' claim depends on the verification protocol, and that protocol has a circularity that the paper acknowledges but does not resolve; this is the main point requiring revision.","major_comments":[{"comment":"The momentum residual check D2 is computed on dedicated frame triplets exported with the OU forcing state frozen, not on the released frames; D3 then confirms second-order time truncation on those same triplets. As a result, the released fields' momentum balance is not directly verified—D2 validates the solver's temporal accuracy on auxiliary runs. The abstract and Section 4 currently state that the corpus ships 'equation-level verification,' which overstates what D2 provides. Please either compute momentum residuals on released frames (e.g., by storing forcing realizations or using the frozen-forcing protocol on the released trajectory), or qualify the claim to indicate that D1 (divergence) and D4 (pressure–Poisson) are checked on released frames while D2/D3 certify the solver's time integration on auxiliary triplets.","section":"Section 4 / Appendix B (D2, D3)"},{"comment":"Every D-group check reuses the same pseudo-spectral operators, dealiasing, projection, and Eswaran–Pope forcing implementation that produced the data. A systematic implementation error would pass both the generator and the verifier, so the verification certifies self-consistency of one solver family rather than DNS fidelity in an absolute sense. The authors disclose the one-solver scope in Appendix K, but the central 'DNS-verified' claim (abstract, Section 1, Table 1) needs either an independent-solver spot check on a subset of released frames or an explicit restatement as 'solver-verified with equation-level residuals computed within the generating discretization.' This is load-bearing because the benchmark's conclusions about learned models are framed as turbulence conclusions, not solver-family conclusions.","section":"Appendix B / Section 4"},{"comment":"The statistical gates carry little independent certification weight at this Reynolds number: the Kolmogorov constant lacks a verdict because the plateau lies in the bottleneck region, and C_epsilon is compared against a low-Re trend. This is disclosed, but it means the certification burden falls almost entirely on the equation-level checks, which strengthens the need for the independent cross-check raised in the previous comment.","section":"Section 4 / Table 6"}],"minor_comments":[{"comment":"The phrase 'DNS-verified' appears in the abstract and Section 1; consider a precise definition in Section 4 of what 'verified' means relative to the D-group scope, so readers do not infer a stronger guarantee than D1–D4 provide.","section":"Abstract / Introduction"},{"comment":"The symbols (✓, partial, ✗) are not defined in the caption; consider adding a legend for readability.","section":"Table 1"},{"comment":"The statement 'at roughly one Kolmogorov time per frame' is useful, but the exact relation between Δt=0.05T_L and τ_η is not given; please add the measured value for the flagship configuration.","section":"Section 3.3"},{"comment":"The footnote about median and divergence count is clear, but the table would benefit from explicit divergence counts for all cells, not only the median cells.","section":"Table 4 caption"}],"recommendation":"major_revision","confidential_remarks":"The paper is well organized and the resource appears genuinely useful. The main risk is overclaiming 'DNS-verified' when the verification is single-solver and D2 is not run on released frames. If the authors add an independent-solver spot check or carefully re-scope the language, I would support publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: TIDE is worth having and worth reviewing. It is, to my knowledge, the first 3D incompressible turbulence dataset organized as a benchmark with independent ensembles per configuration, controlled physics axes, a forced/decay contrast, pressure fields, and an explicit acceptance protocol. That alone fills a real gap — everything else in the ML-for-PDE space is 2D, single-realization, or a different governing system. The paper also reports benchmark results that are actually informative: learned models barely beat persistence, pointwise accuracy does not track physical fidelity, and forced-to-decay transfer exposes a missing conditioning input. Those are useful findings, not just dataset plumbing.\n\nWhat the paper does well: the acceptance gates are documented with measured envelopes, the verification scripts are validated on known-answer fields including deliberately violating ones, the flagship configuration is independently reproduced, and the metric layer is audited against real DNS fields with mutation tests. That is more rigor than most dataset papers. The baseline audit is honest about deviations from published architectures, and the replicate-seed study is a real strength — single-seed leaderboards are a known fragility and they show it.\n\nThe main soft spot is exactly what the stress-test note says: the equation-level verification is self-referential. D1 and D4 run on released frames, but the momentum residual D2 runs on auxiliary frame triplets with the forcing frozen, not on the released frames. And every check uses the same pseudo-spectral operators that generated the data, so a systematic error in the solver, forcing, dealiasing, or projection would pass the D-group. The paper discloses the one-solver scope in Appendix K, so it is not hiding anything, but the phrase \"DNS-verified\" overstates what the gates certify. It is closer to \"self-consistent with one solver family.\" An independent-solver cross-check would tighten this, and I would suggest that as a condition for strongest claims, not as a rejection. Other limitations — moderate Reynolds number, Neff≈9, dissipation-peak under-sampling, single seed for the main table — are disclosed and are minor in context.\n\nWho it is for: anyone developing or evaluating learned surrogates for 3D turbulence, and anyone building benchmarks for physics-ML. It deserves a serious referee; the dataset itself is the artifact and the code is public, so a good reviewer should spot-check the acceptance scripts and ideally run one configuration through a different solver. I would cite it if I worked in this area.","headline":"TIDE is a genuinely useful, carefully documented 3D turbulence benchmark that fills a real gap, though its 'DNS-verified' label is slightly stronger than the evidence warrants because the verification reuses the generating solver.","tokens_in":32909,"tokens_out":2089,"would_cite":true,"duration_ms":21361,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"TIDE introduces a DNS-verified, physically diverse 3D turbulence corpus showing that learned emulators barely beat persistence and stay about twice the error of a spectral solver given the true equations.","keywords":["3D turbulence","direct numerical simulation","scientific machine learning","benchmark dataset","neural operators","physical fidelity","generalization","Navier-Stokes"],"falsifier":"Take a released forced-isotropic frame triplet and evolve it forward with an independent DNS implementation that differs in dealiasing, projection, and forcing scheme, then compare the ensemble statistics and short-horizon trajectories against TIDE's; if the fields diverge beyond sampling uncertainty while TIDE's own residual gates still pass, the acceptance standard certified self-consistency of one solver rather than fidelity to the incompressible Navier-Stokes equations.","tokens_in":31908,"feed_emoji":"🌀","tokens_out":8730,"duration_ms":72173,"temperature":0.7,"pith_summary":"TIDE is a corpus and benchmark purpose-built to let machine learning be tested on three-dimensional incompressible turbulence, the setting where vortex stretching actually lives. The paper's central claim is that current learned forecasting models, trained and evaluated on this corpus under a fixed protocol, barely outperform simply copying the last frame forward, and remain about twice as far from the truth as a classical spectral solver that is handed the Navier-Stokes equations. The paper also claims that a lower pointwise error does not imply a physically faithful rollout: models can rank best by nRMSE while inflating the small-scale enstrophy by an order of magnitude. What makes the claim measurable is the corpus design: 15 configurations on eight controlled physical axes, each with 8-16 independent realizations, verified fields, and a forced-to-decay contrast that isolates whether an operator is conditioned on the external drive. If correct, TIDE turns 'did the model learn the dynamics?' from a slogan into a quantitative comparison.","feed_headline":"3D turbulence benchmark: AI emulators barely beat persistence","feed_subtitle":"TIDE's 15 verified configurations with independent ensembles separate accuracy, physical fidelity, and conditioning.","key_machinery":"The central object is the configuration axis: one incompressible Navier-Stokes system, one pseudo-spectral solver on a $256^{3}$ periodic box, one acceptance standard, with only the physics varied across 15 configurations grouped into forced isotropic, extended physics (rotation, stratification, passive scalar), and free decay. The mechanism that carries the argument is the combination of independent realizations, each seed having its own initial condition and its own forcing sequence, with equation-level verification: divergence residuals at solver precision, momentum residuals measured on frame triplets with the stochastic forcing frozen, a step-halving check confirming second-order time truncation, and a pressure-Poisson consistency check. The forced-to-decay axis is what isolates conditioning: removing the drive changes no term of the governing equations, yet no single snapshot reveals whether a flow is driven or freshly decaying, so a forced-trained operator's persistence-level error on decay is evidence of a missing input rather than of how well the physics was learned.","core_discovery":"On its own terms, the paper's discovery is a negative result made precise: after standardizing training and evaluation on a $256^{3}$ direct-numerical-simulation corpus with independent ensembles, the benchmarked learned operators improve on persistence by at most about 8% in mean forecasting skill and their rollout error is roughly twice that of an equations-informed spectral integrator (nRMSE around 0.426 versus 0.843 persistence mean, with the best learned models near 0.78-0.82). No single architecture leads on all tasks: the attention-based operator is the most consistently positive on forecasting while the spectral convolution operator wins super-resolution on every configuration, and the pressure-recovery task is dominated by the exact Poisson solve that no learned model approaches. The paper further establishes that pointwise accuracy and physical fidelity separate: a model with lower nRMSE than another can have an order-of-magnitude larger enstrophy ratio and a high-band spectral error tens to hundreds of times larger. Finally, forced-to-decay transfer shows a missing conditioning variable: because the stochastic drive cannot be read off a single frame, operators trained under forcing continue to predict driven evolution when the drive is removed, at persistence-level error with in-distribution single-step error.","pith_inferences":["A testable extension follows directly: condition an operator on the last two or three frames, or on an estimated injection rate, and measure whether the forced-to-decay error drops from persistence level toward the near-exact decay performance of the equations-informed integrator; TIDE's released frame sequences make this experiment immediate.","The separation of accuracy from physical fidelity suggests that small nRMSE improvements, on the order of a few hundredths, are within seed noise on this corpus, so future studies should compare error distributions or multi-seed intervals rather than point estimates.","The same missing-conditioning mechanism may explain failures in learned emulators of other stochastic PDEs where the noise realization is invisible in a single observation; the forced/decay contrast is a template for detecting such omitted state variables.","The corpus's controlled axes make it possible to train a joint model across regimes with the regime as an input, the defining ingredient of a foundation model for 3D turbulence, giving the field a concrete pretraining target rather than an abstract one."],"forward_implications":["Any credible learned surrogate for 3D turbulence should be reported against persistence and against an equations-informed spectral solver; the interval between these two references is the unclaimed predictability, at least half of the total on these configurations.","Pointwise error alone cannot rank models: a leaderboard should report at least one small-scale physical metric, because the model with lowest nRMSE can have an enstrophy ratio of 51 against a target of 1.0.","The forced-to-decay result implies that single-frame surrogate formulations lack a conditioning input for the stochastic drive; adding a short history or an explicit forcing input is the direct remedy the benchmark points to.","Because rollout stability is configuration- and seed-dependent, single-seed leaderboard cells are not decisive; replicate seeds set the resolution of model comparisons.","Regime shifts whose signature is visible in the input field are mostly coverage problems, while the drive is not visible; generalization tests should be labeled by which category of shift they exercise."],"supporting_citations":[{"why":"supplies the stochastic Ornstein-Uhlenbeck forcing used for all statistically steady configurations, the protocol that makes deterministic forcing metastability a non-issue.","marker":"[14]"},{"why":"the general-purpose PDE benchmark against which TIDE's 3D-native, DNS-verified structure is positioned.","marker":"[46]"},{"why":"the landmark 3D DNS database documented with one realization per configuration, the resource TIDE's independent ensembles are designed to fix.","marker":"[29]"},{"why":"a large-scale physics simulation collection whose few 3D incompressible turbulence entries motivate TIDE's dedicated corpus.","marker":"[38]"},{"why":"the Fourier spectral operator family that serves as one of the five capacity-matched learned baselines.","marker":"[30]"},{"why":"the attention-based operator that is the most consistently positive forecasting baseline and the one that transfers losslessly onto rotation.","marker":"[51]"},{"why":"the learned-correction hybrid cited as the only prior demonstration of forcing transfer, and the closest precedent for the missing-conditioning finding.","marker":"[24]"},{"why":"the dynamic eddy-viscosity closure used as the trivial reference for subgrid-stress prediction, which returns zero backscatter by construction.","marker":"[17]"},{"why":"the global branch-operator baseline whose encoding failure on pressure recovery supports the paper's architecture-dependence argument.","marker":"[34]"}],"fun_headline_variants":["3D turbulence AI: barely beats persistence, misses physics","TIDE: AI turbulence models fail physical fidelity","3D turbulence: pointwise accuracy ≠ physical fidelity","AI turbulence operators fail when forcing stops","New 3D turbulence benchmark exposes AI limits"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's value rests on the assumption that the acceptance procedure, including equation-level residual checks computed with the same pseudo-spectral discretization that generated the data, certifies that the released fields are faithful DNS of the incompressible Navier-Stokes equations rather than merely self-consistent outputs of one solver family.","fun_headline_variants_meta":{"raw":{"variants":["3D turbulence AI: barely beats persistence, misses physics","TIDE: AI turbulence models fail physical fidelity","3D turbulence: pointwise accuracy ≠ physical fidelity","AI turbulence operators fail when forcing stops","New 3D turbulence benchmark exposes AI limits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001213,"raw_usage":{"total_tokens":5043,"prompt_tokens":1047,"completion_tokens":3996,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":663,"completion_tokens_details":{"reasoning_tokens":3924}},"tokens_in":663,"tokens_out":3996,"duration_ms":25958,"temperature":1.0,"reasoning_tokens":3924,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T00:12:27.910858+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a released forced-isotropic frame triplet and evolve it forward with an independent DNS implementation that differs in dealiasing, projection, and forcing scheme, then compare the ensemble statistics and short-horizon trajectories against TIDE's; if the fields diverge beyond sampling uncertainty while TIDE's own residual gates still pass, the acceptance standard certified self-consistency of one solver rather than fidelity to the incompressible Navier-Stokes equations.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the stochastic Ornstein-Uhlenbeck forcing used for all statistically steady configurations, the protocol that makes deterministic forcing metastability a non-issue."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"the general-purpose PDE benchmark against which TIDE's 3D-native, DNS-verified structure is positioned."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"the landmark 3D DNS database documented with one realization per configuration, the resource TIDE's independent ensembles are designed to fix."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"a large-scale physics simulation collection whose few 3D incompressible turbulence entries motivate TIDE's dedicated corpus."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"the dynamic eddy-viscosity closure used as the trivial reference for subgrid-stress prediction, which returns zero backscatter by construction."}],"review_version":1}