{"id":"be6c04d2-5df4-48d9-90dc-e0ba930de890","arxiv_id":"2607.09763","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A knowledge-to-DFFD design-space pipeline plus MoE neural-operator surrogate with Mahalanobis uncertainty gates yields CFD-validated 4–10% vehicle drag reductions and ~1.16% MAPE on heterogeneous aero data.","lead":"The paper builds a vehicle-shape optimizer that turns design rules into editable free-form deformation boxes and uses a mixture-of-experts neural operator to predict drag, calling CFD only when uncertainty is high. It reports CFD-validated drag cuts of about 4–10% on MPV, SUV, and sedan cases while improving surrogate accuracy on mixed vehicle families.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Constraint fidelity is asserted via RAG+VLM→DFFD embedding of H, but never independently validated against engineer-admissible designs.","rationale":"The reader correctly isolates the softest load-bearing assumption: that RAG+VLM grounding plus DFFD embedding of H produces a design-valid search space. Empirical support for MoE-NO (MAPE 1.16%, Acc_tre 94.34%) and for CFD-validated drag cuts is present and not internally contradicted; the issue is whether those cuts are achieved inside a faithfully knowledge-constrained space. Because H is never checked as independent inequalities and DFFD can move non-editable nodes, the optimization results alone do not close the loop on “engineering-aware” validity. That keeps the verdict CONDITIONAL rather than ACCEPT or REJECT: the contribution is accept-shaped if constraint fidelity (and related reproducibility) is stress-tested; it is not rejectable on the CFD numbers as written. No stronger internal inconsistency (e.g., in MoE routing or Mahalanobis gating math) is needed to explain the moderate confidence.","tokens_in":25652,"tokens_out":613,"duration_ms":7338,"concrete_test":"Have independent vehicle aerodynamicists (blind to the agent outputs) mark editable regions and deformation bounds on the same multi-view baselines used in Fig. 7 / Table 4; compute IoU of 3D boxes and fraction of optimized geometries that violate their H (packaging, clearance, silhouette, manufacturability). If IoU is low or a non-trivial fraction of the CFD-validated optima violate expert H, the knowledge-constraint claim does not support the reported design reductions.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim couples two parts: (i) MoE-NO accuracy/trend gains and (ii) CFD-validated Cd reductions of ~4–10% under knowledge-constrained DFFD search. Part (i) is supported by Table 5 and baselines. Part (ii) depends on Sec. 2.3: RAG+VLM produce S, Ωm, and H, which are fused into C and Ω(C) and then “embedded in C, Ω, and D” rather than enforced as independent analytic inequalities (explicitly stated after Eq. 11 and in Sec. 2.1). DFFD (Eqs. 23–29) further allows small induced displacements outside editable boxes. There is no quantitative check that the resulting (C,Ω,H) match engineer-specified admissible regions, packaging/safety/manufacturability bounds, or design-identity constraints—only qualitative multi-view figures and Table 4 boxes. If the grounded space is systematically too loose or biased, the reported CFD Cd drops can be real physics improvements that still fail industrial admissibility, so the “knowledge-constrained / high-confidence design” framing is not established by the optimization numbers alone.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes a knowledge-constrained, surrogate-assisted shape-optimization framework for vehicle aerodynamics. Domain knowledge and multi-view geometry are converted via RAG, LLM, and VLM into editable DFFD control boxes C, admissible deformation space Ω(C), and preservation constraints H; a Mixture-of-Experts Neural Operator (MoE-NO) predicts Cd on heterogeneous MPV/SUV/Sedan data; and dual physics-solver gates use encoder-space OOD detection and Mahalanobis-percentile uncertainty to trigger local CFD enrichment and online refinement. On an in-house 450-sample ID database, MoE-NO reports test MAPE 1.16% and trend accuracy 94.34% versus best baselines 1.52% and ~90.34%. CFD-validated optimizations yield Cd reductions of roughly 4–10% (MPV 10.52%, Sedan 4.17%, refined SUV 7.96%, adapted OOD Sedantest 7.43%), with an OOD adaptation ablation improving R² from −18.42 to 0.9525 on a held-out sedan family.","tokens_in":26057,"tokens_out":1869,"duration_ms":27423,"significance":"If the results hold under broader validation, the work is a useful systems-level contribution at the intersection of geometry parameterization, neural operators, and industrial aerodynamic design. Strengths include CFD-validated optima rather than surrogate-only claims, explicit dual-gate solver-in-the-loop design (distribution-level OOD adaptation and optimization-level Mahalanobis refinement), a multi-family train/val/test protocol with Transolver and DragSolver baselines (Table 5), grid-convergence for labeling (Table 1), an online-refinement ablation on SUV (Table 7), and a public geometry dataset. The MoE-NO accuracy/trend gains and selective CFD enrichment are practically relevant for expensive external-flow optimization. The knowledge-grounding pipeline is ambitious and timely given multi-agent design work, but its industrial significance depends on whether grounded (C, Ω, H) truly encode engineer-admissible packaging, safety, manufacturability, and design-identity constraints—an aspect the manuscript asserts more than it measures.","major_comments":[{"comment":"Sec. 2.1 (after Eq. 11) and Sec. 2.3 state that preservation constraints H are not imposed as independent analytic inequalities but are “embedded” in C, Ω(C), and D. DFFD (Eqs. 23–29) further allows induced displacements outside editable boxes. The central “knowledge-constrained / high-confidence design” claim therefore rests on the fidelity of RAG+VLM grounding (Fig. 7, Table 4). There is no quantitative check against engineer-specified admissible regions, packaging/safety/manufacturability bounds, or design-identity criteria—only qualitative multi-view figures and reconstructed boxes. CFD Cd reductions (Tables 7, 9) can be real physics improvements while still violating industrial admissibility. Please add an independent validation protocol (e.g., expert review scores, constraint-violation rates, or comparison to hand-specified engineer bounds) or narrow the claim to “knowledge-informe","section":"Sec. 2.1, 2.3; Eqs. 11, 23–29; Fig. 7; Table 4"},{"comment":"ID and OOD quantitative claims rest on very small held-out sets: 15 test samples per ID family (Table 2) and 15 OOD Sedantest test samples (Tables 5, 8). MAPE 1.16%, Acc_tre 94.34%, and the OOD jump from R²=−18.42 / MAPE 14.88% to R²=0.9525 / MAPE 1.84% are directionally convincing but statistically fragile; a few outliers can dominate. Please report confidence intervals or bootstrap variability for MAPE/Acc_tre, clarify whether family-wise metrics differ, and temper absolute claims in the abstract/conclusions accordingly. Larger held-out or cross-family leave-one-family-out tests would substantially strengthen the MoE-NO generalization argument.","section":"Tables 2, 5, 8; Sec. 3.3, 4.2, 4.5"},{"comment":"The dual-gate mechanism depends on several free thresholds (η_σ=0.6, η_CFD=0.025, α_ex=0.1, N_KNN=5, p_cal, DFFD β, expert-addition schedule) that are stated without sensitivity analysis (Secs. 2.4.2–2.4.4). For SUV, online refinement is activated by σ>η_σ and e_phys>η_CFD and improves CFD-validated reduction from 4.36% to 7.96% (Table 7)—a useful ablation, but it is unclear how often refinement would fire under alternate thresholds or whether MPV/Sedan acceptance is robust. A short sensitivity study on η_σ and η_CFD (and reporting how many CFD calls each gate consumed) is needed to support the “high-confidence” and cost-efficiency narrative.","section":"Secs. 2.4.2–2.4.4; Fig. 12; Table 7"}],"minor_comments":[{"comment":"Abstract and Sec. 4.2 cite best baseline trend accuracy as 90.34%, while Table 5 lists Transolver test Acc_tre=0.9051 and DragSolver 0.8495; reconcile the rounded figure with the table.","section":"Abstract; Table 5"},{"comment":"Fig. 10 joint density of σ vs δ is informative but would benefit from a calibration curve (e.g., error rate vs σ bins) and explicit false-negative rate for high-error cases, since the text argues low upper-left mass.","section":"Fig. 10; Sec. 4.2"},{"comment":"Clarify MoE training details: number of experts N_exp for the ID model, how the new expert E_{N_exp+1} is initialized, and whether the encoder is frozen during OOD fine-tuning (Sec. 2.4.3).","section":"Secs. 2.4.1, 2.4.3"},{"comment":"Table 5 DragSolver train MAPE 0.98% with 57.20% of samples having MAPE>1% is possible but surprising; a brief note on error distribution would help readers interpret the metric.","section":"Table 5"},{"comment":"Several references are arXiv preprints dated 2025–2026; ensure citation completeness and that concurrent related multi-agent aero-design works are fairly positioned in the introduction.","section":"Sec. 1; References"},{"comment":"Notation: J_phys, Ć_d, and Ć_ϕ are used interchangeably for drag; a short symbol table would reduce friction.","section":"Secs. 2.1, 2.4"}],"recommendation":"major_revision","confidential_remarks":"The MoE-NO + dual-gate CFD loop is the more solid technical contribution; the knowledge-constraint story is currently the weaker half of the title and is the main reason I recommend major rather than minor revision. If the authors can add even a modest expert-admissibility study or reframe claims, the paper would be a good fit for a methods-oriented computational design / ML-for-fluids venue. Scope is more engineering design + neural operators than pure cs.CV, but the multi-view grounding component justifies the category. No integrity concerns; dataset release is a plus."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is practical, not foundational: they wire RAG/VLM grounding into DFFD control boxes, put a mixture-of-experts neural operator on heterogeneous vehicle families, and only call CFD when OOD or Mahalanobis uncertainty says so. That package is what is new. The pieces (DFFD, DragSolver-style encoders, MoE, surrogate-in-the-loop) are all in the literature they cite; the integration and the dual-gate enrichment loop are the contribution.\n\nWhat they do well is the empirical side of the surrogate and the optimization loop. On the 450-sample MPV/SUV/Sedan set, MoE-NO beats Transolver and DragSolver on MAPE (1.16% test) and trend accuracy (94.34%). Grid convergence for the LBM labels is reported. Optimization optima are CFD-validated, not accepted from the surrogate alone; the SUV online-refinement ablation actually moves the needle (discrepancy 2.78% → 1.25%, Cd cut 4.36% → 7.96%), and the OOD Sedantest adaptation is honest (R² from −18 to 0.95). Dataset is public on GitHub. That is real evidence, not circular self-scoring.\n\nThe soft spot the stress-test flags is real but proportional: H is “embedded” in C, Ω, and D rather than checked as independent inequalities, DFFD can induce small displacements outside the boxes, and there is no quantitative engineer-admissibility audit—only multi-view figures and Table 4. So the “knowledge-constrained / high-confidence design” framing rests more on process description than on a fidelity metric. Gate thresholds (ησ=0.6, ηCFD=0.025, etc.) are hand-set; held-out sets are 15 per family. None of that overturns the CFD Cd numbers; it just means those numbers do not by themselves prove industrial constraint fidelity.\n\nMath and citation pattern look standard and clean for this genre. Who it is for: people doing surrogate-assisted vehicle aero or MDO who care about fewer blind CFD calls and less manual design-space setup. I would send it to peer review; a referee can demand constraint-fidelity checks and threshold sensitivity without the paper collapsing. Worth engaging if that is your lane.","headline":"Solid industrial systems paper: MoE-NO + dual CFD gates deliver real accuracy and validated Cd cuts; the knowledge-constraint story is the softest link, not the surrogate math.","tokens_in":26696,"tokens_out":579,"would_cite":true,"duration_ms":6511,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A knowledge-constrained framework with a mixture-of-experts neural operator turns engineering rules into editable shape variables and uses selective CFD feedback to cut vehicle drag by about 4–10%.","keywords":["Shape optimization","Knowledge-constrained design","Surrogate-assisted optimization","Mixture-of-Experts Neural Operator","Aerodynamic drag","DFFD","Out-of-distribution detection","Vehicle design"],"falsifier":"Run the full pipeline on a new production vehicle family and check whether the optimized shapes either violate packaging, safety, or manufacturability constraints a human engineer would reject, or lose their reported drag reductions when re-meshed and re-solved under an independent high-fidelity CFD configuration.","tokens_in":26527,"feed_emoji":"🚗","tokens_out":944,"duration_ms":21722,"temperature":0.7,"pith_summary":"Industrial shape optimization still depends on experts to hand-define where a geometry may change, and on surrogates that often fail when databases mix design families or when search leaves the training distribution. This paper claims both problems can be closed in one loop: design knowledge and multi-view images are converted into Direct Free-Form Deformation control boxes and bounds, while a Mixture-of-Experts Neural Operator predicts drag and uses its latent space to flag out-of-distribution shapes and uncertain optima so that CFD is called only for local enrichment. On in-house MPV, SUV, and Sedan data the surrogate reaches 1.16% test MAPE and 94.34% trend accuracy, beating strong baselines, and CFD-validated optimizations reduce drag by roughly four to ten percent while preserving overall design intent. The result is a practical path to automated yet engineering-aware aerodynamic design that spends expensive simulation only where the model is unreliable.","feed_headline":"Neural operator cuts car drag 4–10% with selective CFD","feed_subtitle":"Design rules become shape variables; uncertainty gates call the solver only when the model is unsure.","key_machinery":"Mixture-of-Experts Neural Operator (MoE-NO): a geometry encoder with a gating network that routes each shape to multiple expert predictors; the same latent features drive Mahalanobis-distance OOD detection and uncertainty-gated CFD enrichment inside a knowledge-derived DFFD design space.","core_discovery":"The paper establishes that translating domain knowledge and user intent into DFFD editable control boxes, admissible deformation spaces, and preservation constraints, then guiding search with a Mixture-of-Experts Neural Operator and Mahalanobis-percentile uncertainty gates, yields high-confidence knowledge-constrained shape optimization: MoE-NO improves drag MAPE to 1.16% and trend accuracy to 94.34% on heterogeneous vehicle data, and selective physics-solver feedback for OOD and uncertain candidates produces CFD-validated Cd reductions of about 4% to 10%.","pith_inferences":["The same knowledge-to-DFFD grounding could carry packaging and manufacturability rules into multi-objective vehicle design without rewriting the search loop.","Selective Mahalanobis-gated enrichment is portable to other expensive physics loops such as thermal or structural shape optimization.","Richer internal design-rule corpora become a direct competitive advantage because they define a tighter admissible space before any CFD is run.","Trend accuracy near 94% suggests the surrogate can serve as a ranking prior inside evolutionary or Bayesian optimizers even when absolute error is imperfect."],"forward_implications":["Engineering constraints can live inside the optimizer as editable regions and bounds rather than as post-hoc filters.","Heterogeneous vehicle databases benefit from multi-expert routing instead of a single global neural operator for reliable drag ranking.","High-fidelity CFD can be reserved for distribution-level and uncertainty gates instead of every candidate design.","An unseen vehicle family can be brought into reliable optimization by local resampling and adding one new expert branch.","Latent-space uncertainty can decide when a surrogate optimum is safe to accept versus when it must be refined with CFD."],"fun_headline_variants":["MoE neural operator cuts vehicle drag 4–10% under knowledge constraints","Knowledge rules become DFFD controls for 4–10% CFD-validated Cd cuts","Mixture-of-experts operator hits 1.16% MAPE and 4–10% drag reduction","Uncertainty-gated MoE-NO enables high-confidence car shape optimization","Expert mixture neural operator yields 4–10% lower vehicle drag"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The method assumes that language and vision models correctly turn design rules into editable 3D boxes and deformation bounds that match what engineers would actually allow.","fun_headline_variants_meta":{"raw":{"variants":["MoE neural operator cuts vehicle drag 4–10% under knowledge constraints","Knowledge rules become DFFD controls for 4–10% CFD-validated Cd cuts","Mixture-of-experts operator hits 1.16% MAPE and 4–10% drag reduction","Uncertainty-gated MoE-NO enables high-confidence car shape optimization","Expert mixture neural operator yields 4–10% lower vehicle drag"]},"model":"grok-4.5","effort":"low","cost_usd":0.006746,"raw_usage":{"total_tokens":1749,"prompt_tokens":841,"num_sources_used":0,"completion_tokens":110,"cost_in_usd_ticks":67460000,"prompt_tokens_details":{"text_tokens":841,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":798,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":841,"tokens_out":110,"duration_ms":7945,"temperature":1.0,"reasoning_tokens":798,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T16:10:35.946981+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Run the full pipeline on a new production vehicle family and check whether the optimized shapes either violate packaging, safety, or manufacturability constraints a human engineer would reject, or lose their reported drag reductions when re-meshed and re-solved under an independent high-fidelity CFD configuration.","supporting_citations":[],"review_version":1}