{"id":"bf0a6934-859c-45b8-8193-c7e07b5825b2","arxiv_id":"2607.23667","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":7,"one_line_summary":"No single flow-surrogate architecture transfers from a boundary-driven Stokes film to a self-sustained Kármán wake; time treatment decides the winner and pointwise RMSE ranks the wrong models.","lead":"Eight flow surrogates were tested on a quiet slurry film and a self-sustained vortex wake; no architecture won both. The result warns that a model validated on a simple startup flow can fail once the physics gets richer, so surrogate choice should follow the flow’s dynamics, not a single error score.","discovery_kind":"new_application","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The film-side win may be an encoder/wall-shear bottleneck rather than dynamics: POD k=8 (98.8% energy) can miss low-energy near-wall modes that set Eq. (3) cum-WSS, so the >4× cwL2 gap needs an encode–decode floor/control before the causal time-treatment reading.","rationale":"I partially agree with the reader: the named weakest assumption is the right family—encoder–regime confounding—but the stress-test identifies a sharper film-side mechanism: global-energy POD truncation can be exactly wrong for a boundary-shear process target. This does not overturn the reported measurements. The paper gives real support for the narrower claims: multi-seed tables, frozen shared encoders within regimes, GRU/S4D one-shot-vs-AR pairs in both regimes, candid §8 limitations, and metric inversions that are physically interpretable. The wake claim that autoregressive feedback preserves a limit cycle is comparatively strong because representation is held fixed and P/RMSEw/ESt move together. The film claim that one-shot beats AR is also supported within the same POD encoder. What remains conditional is the stronger practical inference for CMP programmes: that a richer CMP flow would necessarily demand the wake-style latent AR architecture, or that the direct operator’s film advantage is due to dynamics rather than wall-resolving representation. The proposed encode–decode floor and wall-weighted encoder sweep would settle that cheaply before any new solver campaigns. I would keep the verdict CONDITIONAL rather than downgrade: the concern targets generality and causal attribution, not the internal validity of the benchmark, and the authors already disclose the broader confound and data-availability limits.","tokens_in":27456,"tokens_out":3284,"duration_ms":74345,"concrete_test":"Compute the film encoder reconstruction floor before comparing dynamics: encode each reference test field and immediately decode it, then evaluate cwL2 from Eq. (3). Compare POD k=8 against k=31/99%, a wall/shear-weighted POD basis, and a boundary-weighted conv-AE, all frozen across the same GRU/S4D one-shot and AR dynamics. If k=8 decode-only truth already has cwL2 ≳0.1 while wall-preserving encoders drop the floor to ≲0.05 and latent one-shot then approaches the field DeepONet, the direct win is encoder-driven. If decode-only floor is ≪0.032 yet latent cwL2 stays ≥4× behind at matched encoder, the dynamical claim survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The least secure condition is not whether the tables are internally consistent—they look careful—but whether the film result licenses the causal slogan that dynamical character/time treatment selects the architecture. The CMP target S(x)=∫|τw|dt with τw=μ(u_top−u_bot)/h (§3.1, Eq. 3) is a boundary-shear functional over a 40 µm gap. It can be dominated by near-wall, high-gradient structure that contributes little global variance. The latent film pipeline deliberately truncates POD at k=8 because it holds 98.8% energy (Fig. 5, §4.2), an energy criterion not a shear-functional criterion. The missing 1.2%, if localized at moving walls or the swept high-shear ridge, is divided by h and integrated in time, so it can produce large cwL2 while barely changing nRMSE3. The direct field DeepONet predicts wall planes at full resolution and bypasses this bottleneck; the latent models cannot exceed their encoder’s reconstruction floor even with perfect dynamics. Thus the headline 0.032 vs 0.144 cwL2 gap could measure representation/encoder loss rather than “boundary-driven film rewards direct map.” The within-film GRU one-shot vs GRU-DON comparison (0.095 vs 0.157, same POD encoder) does support one-shot over AR on film, and the wake GRU/S4D pairs support AR over one-shot there. But the cross-regime no-free-lunch claim and the CMP-program caution depend on the direct operator’s film lead being dynamically necessary. §8 admits encoder–regime confounding generally; this is the concrete mechanism by which it could invert the practical conclusion.","agreement_with_reader":"partial"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The manuscript compares eight surrogate architectures — spanning direct vs latent representation and one-shot vs autoregressive time traversal — on two transient flows driven by piecewise-linear boundary-condition ramps: a quasi-static 3D Stokes slurry film motivated by chemical-mechanical planarisation (CMP), and the 2D Kármán vortex street behind a cylinder. The central finding is a no-free-lunch result: a one-shot full-field DeepONet wins the film on the process target (cumulative wall shear stress, cwL2≈0.032) while latent autoregressive DeepONets (S4D/Mamba/GRU) win the wake (retaining up to 96% of shedding power where direct and one-shot models damp it to ~0). The authors attribute the reversal to the treatment of time: the self-sustained limit cycle needs autoregressive phase memory, the boundary-driven film rewards a direct map. A second contribution is a five-aspect failure-mode metric suite (field, structure, invented motion, amplitude, timing), with the demonstration that pointwise RMSE inverts the physically relevant ranking in both regimes. A third is a break-even cost analysis showing the surrogate pays off only beyond ~70 (film) and ~647 (wake) queries. Models share splits, frozen encoders within regime, fixed epoch budgets, three seeds, and metrics fixed before ranking.","tokens_in":27981,"tokens_out":4158,"duration_ms":38151,"significance":"If the results hold, this is a useful and unusually careful empirical contribution to operator learning for transient PDEs. The concrete strengths: (i) a controlled within-regime ablation isolating time treatment (frozen shared encoder, same branch family, one-shot vs autoregressive), which supports the central causal reading far better than the cross-regime contrast alone; (ii) a failure-mode-resolved metric suite fixed before model ranking, with a demonstrated, quantified inversion of model ranking under pointwise RMSE in both regimes (Table 2: nRMSE3 ranks GRU-DON ahead of the field DeepONet while cwL2 separates them 4× in the opposite direction; Table 3: the damped model wins RMSE at P=0.004 vs 0.96); (iii) the mirror-branch diagnostic (§7), which cleanly explains the S4D RMSE deficit as an unpinned symmetry choice (1.10→0.68 under mirror-rescoring) — a falsifiable, mechanistic check; and (iv) honest cost accounting with break-even query counts Q*≈70/647 (Eq. 27), directly answering the McGreivy–Hakim reporting critique. The CMP motivation is real and the \"validation does not transfer across regimes\" caution is practically relevant to surrogate programmes. The claims are approp","major_comments":[{"comment":"§4.2 / Table 2 / Fig. 7: The film-side headline gap (direct DeepONet cwL2=0.032 vs best latent-AR 0.144) is interpretable as a time-treatment effect only if the POD k=8 encoder does not itself set a cwL2 floor well above 0.032. The truncation is chosen by an energy criterion (98.8%, Fig. 5), but the target S(x)=∫|τw|dt (Eq. 3) divides the velocity difference by h=40 µm and integrates in time, so it can be dominated by the discarded 1.2% if that energy sits in near-wall, high-gradient structure. The latent models cannot beat their encoder's reconstruction floor regardless of dynamics quality, so part of the >4× gap may measure representation, not the direct map. There is a cheap, decisive control that requires no retraining: project the reference test fields through the frozen POD k=8 basis and report cwL2, corr, and ΔWIWNU of the encode–decode reconstruction. If the floor is ≪0.1, the ca","section":"§4.2, Table 2"},{"comment":"§7, Table 1: The claim that 'the treatment of time decides who wins and the representation influences the margin' is, on the wake, established only within the latent family: the direct×autoregressive cell is empty because neither the autoregressive field DeepONet nor the FNO variant could be trained stably (§8). A direct-AR model that held the shedding would break the clean attribution (it would suggest representation/time interact rather than separate). The failed-training disclosure is commendable, but the conclusion should be scoped accordingly: on the wake, autoregressive feedback is shown necessary for latent models and sufficient given the frozen conv-AE, while for full-field models only the one-shot failure is demonstrated. Please qualify the abstract and §7 mechanism paragraph (e.g., 'no representation rescues the one-shot map' is supported; the converse necessity for direct mode","section":"§7, Table 1"},{"comment":"§5.1–5.2: Several suite instruments carry free thresholds — the hallucination velocity threshold τ=0.1×peak (Eq. 16), the onset detector's 0.1-of-developed-fluctuation bar, and the DTW alignment window in RMSEw (Eq. 15) — and the text asserts the detector parameters 'affect the model comparison only marginally' without evidence. The central RMSE-inversion claim rests mainly on P (Eq. 21) and cwL2 (Eq. 20), which are threshold-free, so the risk is contained; still, since the suite is itself a contribution and is recommended for adoption by CMP surrogate programmes, a small sensitivity table (τ and onset threshold varied over a plausible range, e.g. 0.05–0.2, reporting H and Δt∗ rankings for the leading models) would substantiate the invariance claim rather than assert it.","section":"§5"}],"minor_comments":[{"comment":"§3.1: The difference-quotient shear τw=μ(u_top−u_bot)/h is said to agree with the resolved wafer-side wall shear 'to within a few percent across the pad land' — please give the actual number or a figure reference, since Eq. (20) scores surrogates against this convention.","section":"§3.1"},{"comment":"§5.3/Table 2: Film metrics are means over only 15 test cases per split, and Fig. 8 shows cwL2 ranging from 0.011 to 0.139 across cases. Please report the per-case distribution (median/IQR) for the headline cwL2 numbers, not only the seed spread, so the reader can separate case variance from seed variance.","section":"Table 2"},{"comment":"Fig. 5 vs §4.2: The conv-AE latent dimension 32 for KVS coincides almost exactly with the k=31 at which POD reaches 99%; one sentence on whether this was chosen by that coincidence or independently would help.","section":"§4.2"},{"comment":"§2: The claim that a flow driven by a boundary condition varying during the transient 'remains to be addressed' should be tempered slightly — e.g., [43] and time-dependent-input DeepONet variants [59] are adjacent; the novelty is specifically the ramp-BC startup setting with the two-regime contrast.","section":"§2"},{"comment":"Typos/typesetting: abstract 'pay offfrom' (missing space); inconsistent spacing around numbers ('0 .77 h', '103 to 104' for 10^3–10^4 in abstract); 'Reh' renders with subscript spacing issues; 'Sec.' vs 'Section' usage varies.","section":"Throughout"},{"comment":"Table 3: for DMDc and S4D one-shot, Δt∗ is reported as '—' with a footnote ('sheds on fewer than half of the test cases'); please state in the caption what fraction of cases those models do shed, since 'undefined onset' is itself informative about the damping failure mode.","section":"Table 3"},{"comment":"§6.3/Table 4: The CPU-vs-GPU speedup S is honestly flagged as a deployment comparison; please also state the GPU inference batch size used for tinfer, since per-case latency at batch 1 vs full batch can differ by an order of magnitude and Q∗ depends on it weakly through Eq. (27).","section":"§6.3"}],"recommendation":"major_revision","confidential_remarks":"Data availability is stated prospectively (\"will be made available,\" raw data \"on reasonable request\"). Given the paper's own alignment with the McGreivy–Hakim reporting recommendations, I suggest the editor require a persistent public deposit of configs, evaluation harness, and at least the evaluation subset at acceptance. The generative-AI declaration is transparent and appropriate. The CMP framing is well motivated and the work fits the journal's scope at the NA/operator-learning interface; the single substantive risk is the unquantified encoder floor on the film-side process target, which is cheap to resolve."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is simple and well backed: among the eight models they actually ran, nothing wins both a boundary-driven Stokes film and a self-sustained KVS wake under the same ramp-BC setup. One-shot full-field DeepONet hits ~3% cum-WSS on the film and kills the wake; latent AR DeepONets (especially S4D) keep ~96% shedding power and lose badly on the film target. The within-regime one-shot vs AR controls with a frozen encoder are the cleanest part—they isolate time treatment without hand-waving.\n\nWhat is new is not the operators (DeepONet, S4D, Mamba, DMDc are stock) but the paired transfer test aimed at CMP startup, the failure-mode metric suite (field, structure, invented motion, amplitude, timing), and the honest break-even accounting that pins cost on N_train rather than GPU training. Losses are plain MSE on field/latent; metrics were fixed before ranking; three seeds and shared splits. That is careful empirical work, not a circular bake-off.\n\nSoft spots, in proportion. The stress-test on POD k=8 missing near-wall shear modes has teeth for the *direct vs latent* film gap: cum-WSS is a thin-gap functional, energy truncation is not a shear criterion, and the direct model sees full wall resolution. So the >4× cwL2 lead is not pure “dynamics demand a direct map.” But the same-encoder film GRU one-shot vs AR still favors one-shot, and the wake AR vs one-shot flip is solid—so the time-treatment slogan mostly holds; the cross-regime causal slogan is weaker, which §8 already admits. Other limits are real but labeled: open direct-AR cell, small CMP test n, FE error unquantified, data on request. Citation pattern is fine; no load-bearing self-deal.\n\nThis is for people building fluid/CMP surrogates or writing operator-learning papers who still report only RMSE and transfer wins. Math and tables look internally consistent. I would send it to referees; it deserves a serious read, not a desk reject. Engage if you care about validation practice in scientific ML; skip if you only want new architectures.","headline":"Solid two-regime bake-off with a real practice punchline: time treatment flips the winner, and RMSE lies in both regimes—worth engaging, with the encoder confound already half-owned by the authors.","tokens_in":28877,"tokens_out":581,"would_cite":true,"duration_ms":16861,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["65M60","76D05","68T07"],"pacs":[],"model":"grok-4.5","headline":"No single flow-surrogate architecture wins both a boundary-driven film and a self-sustained wake; how the model treats time decides the winner, and pointwise error ranks the wrong model in both regimes.","keywords":["surrogate modeling","reduced-order modeling","operator learning","DeepONet","Kármán vortex street","chemical-mechanical planarisation","time-varying boundary conditions","failure-mode metrics"],"falsifier":"Hold the encoder, dimension, and ramp family fixed and swap only whether the target flow is boundary-driven or self-sustained; if the one-shot versus autoregressive ranking no longer flips with that dynamical character, the central claim fails.","tokens_in":28542,"feed_emoji":"🌊","tokens_out":1173,"duration_ms":26289,"temperature":0.7,"pith_summary":"Engineers often treat a surrogate that works on a simplified flow as proof it will work on a richer one. This paper tests that habit on two startup-driven flows: a quasi-static three-dimensional slurry film from chemical-mechanical planarisation, and a two-dimensional Kármán vortex street that sheds on its own. Eight models that either map the full field or a latent state, and either predict the whole trajectory at once or step by step, are scored on one shared pipeline. No architecture wins both. A one-shot full-field map reconstructs the film’s process target—cumulative wall shear—to about three percent relative error, while latent autoregressive models keep nearly all of the wake’s shedding power that direct and one-shot models damp away. The deciding choice is the treatment of time: the self-sustained wake needs the phase memory of feedback, while the boundary-driven film rewards a direct map. Aggregate root-mean-square error inverts the physically relevant ranking in both regimes, so the paper scores field accuracy, structure, invented motion, amplitude, and timing instead. The practical message is that surrogate choice should follow the flow’s dynamical character, validation should resolve failure modes, and the speedup only pays off after enough queries to amortise the training simulations.","feed_headline":"No free lunch: flow surrogates flip winners across regimes","feed_subtitle":"Time treatment decides who wins a film versus a wake, and plain RMSE picks the wrong model both times.","key_machinery":"A two-axis design matrix of representation (full-field versus frozen latent encoder) against time treatment (one-shot map versus autoregressive rollout), evaluated by a fixed suite of five physical questions—field accuracy, spatial structure, invented motion, fluctuation amplitude, and event timing—rather than by aggregate RMSE alone.","core_discovery":"Among eight surrogates compared on a shared pipeline, no single architecture wins both regimes. On the CMP film a one-shot full-field DeepONet reaches 3.2% relative error on cumulative wall shear stress; on the Kármán wake a latent autoregressive DeepONet retains about 96% of the shedding power that direct and one-shot models collapse to nearly zero. The axis that flips the winner is the treatment of time—autoregressive feedback for the self-sustained limit cycle, a direct map for the boundary-driven Stokes film—while representation only changes the margin. Pointwise RMSE ranks the wrong model in both regimes, so failure-mode-resolved metrics are required; neither the winning architecture no","pith_inferences":["The same time-treatment split is likely to reappear in other manufacturing flows that mix a forced ramp with possible spontaneous oscillation, such as coating, filling, or stirred reactors.","Once a process model adds particle transport or free-surface dynamics on top of the film, the winning cell may flip mid-programme, so architecture should be re-chosen when the physics enrichment crosses into self-sustained behaviour.","Active learning or smaller training sets could move the break-even count down, but only if accuracy saturates early enough that the offline simulation bill shrinks without losing the physical metrics.","Unpinning of mirror limit-cycle branches is a silent deployment risk for any wake-like surrogate scored only by pointwise error."],"forward_implications":["A CMP surrogate programme cannot treat success on a simplified Stokes film as evidence that the same architecture will hold once richer dynamics appear.","Self-sustained oscillatory flows should be assigned autoregressive latent models; boundary-driven quasi-static flows should be assigned one-shot direct maps.","Model selection by pointwise RMSE alone will deploy the damped wake predictor and a film model several times worse on the process target.","Surrogates pay off only as many-query instruments: break-even is set almost entirely by training-set size (about 70 queries for the film, about 647 for the wake), not by training compute.","Validation reports should include failure-mode-resolved scores for structure, hallucination, amplitude, and timing alongside any aggregate error."],"fun_headline_variants":["No free lunch: surrogates flip winners on film vs wake","Time treatment flips surrogate winners across two flow regimes","One-shot wins film, autoregressive wins wake: no free lunch","RMSE picks wrong surrogate in both CMP film and Kármán wake","Flow surrogates: architecture win depends on time treatment"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That contrasting these two flows is enough to pin the architecture flip on dynamical character alone, even though the regimes also differ in encoder, dimension, governing equations, and how the ramp family is built.","fun_headline_variants_meta":{"raw":{"variants":["No free lunch: surrogates flip winners on film vs wake","Time treatment flips surrogate winners across two flow regimes","One-shot wins film, autoregressive wins wake: no free lunch","RMSE picks wrong surrogate in both CMP film and Kármán wake","Flow surrogates: architecture win depends on time treatment"]},"model":"grok-4.5","effort":"low","cost_usd":0.003538,"raw_usage":{"total_tokens":1264,"prompt_tokens":957,"num_sources_used":0,"completion_tokens":73,"cost_in_usd_ticks":35384000,"prompt_tokens_details":{"text_tokens":957,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":234,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":957,"tokens_out":73,"duration_ms":5776,"temperature":1.0,"reasoning_tokens":234,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T16:20:29.599024+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Hold the encoder, dimension, and ramp family fixed and swap only whether the target flow is boundary-driven or self-sustained; if the one-shot versus autoregressive ranking no longer flips with that dynamical character, the central claim fails.","supporting_citations":[],"review_version":1}