{"id":"328c729a-b011-45fd-b491-356f739d9c18","arxiv_id":"2509.05305","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"A physics-transfer graph network trained on simulated growth of simple elastic shells predicts curvature and short-term shape change on a fetal brain atlas, but validation is limited to a single population-averaged atlas.","lead":"This preprint trains a graph neural network on computer simulations of growing spheres and ellipsoids, then applies it without retraining to human brain surfaces from fetal MRI atlases to estimate curvature and short-term shape changes. It is a candidate route to 'digital twin' brain models that forecast development from sparse imaging data.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"PT-vs-SL comparison is confounded: SL omits normal vectors, though curvature is locally determined by positions and normals; the claimed physics-transfer advantage may be an input-feature artifact.","rationale":"The reader's REJECT verdict is well supported, but I identify a more fundamental weakness than the one highlighted as weakest_assumption. The reader focused on the population-averaged atlas invalidating the claim of individual morphogenesis prediction. That concern is real: a cross-sectional average cannot validate individual trajectories. However, even if longitudinal individual data were used, the main quantitative claim that physics-transfer outperforms statistical learning would still be undercut by the asymmetric input features. Curvature is a local differential invariant of a surface; providing positions and normals to a GNN is nearly providing the answer in a form a sufficiently expressive network can decode. The SL control, by withholding normals, is not a baseline for 'statistical learning' — it is a baseline for missing information. Therefore the PT-vs-SL gap of 0.5 vs 7.01 mm^-1 does not demonstrate that the learned nonlinear elasticity from simple geometries transfers; it demonstrates that normals are useful features for estimating curvature. The paper does not report ablations, error bars, code, or data, so this confound is not resolvable from the manuscript. My proposed test directly isolates the role of normals and would settle whether the physics-transfer mechanism contributes beyond a purely geometric operation. If the fair baseline matches PT, the central claim of physics embedding collapses; if it does not, the atlas concern still prevents the stronger 'individual digital twin' claim. Either way, the current evidence does not support the abstract's assertion that 'the physics of nonlinear elasticity from simple geometries is embedded into a neural network and applied to brain models.' The verdict should remain REJECT/UNCHANGED, with the path to acceptance requiring a fair baseline and longitudinal individual validation.","tokens_in":14496,"tokens_out":5027,"duration_ms":66075,"concrete_test":"Retrain the SL control with the identical GNN architecture, training data, and loss as PT, but with normal vectors included as node features (i.e., SL+normals); evaluate on the same brain-atlas curvature task of Fig. 3d. Also compute a classical discrete curvature estimate (e.g., cotangent Laplacian or vertex-normal shape operator) directly from the brain mesh as a no-training baseline. If SL+normals or the direct estimator reaches MAE <0.5 mm^-1, the PT advantage over SL is explained by input information, not by physics transfer; if not, the confound is less severe, though the atlas/individual-trajectory concern remains.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative evidence for physics transfer is the curvature result (Figs. 3c,d): PT MAE <0.5 mm^-1 vs 7.01 mm^-1 for SL. But in Methods, the PT model's node features are 'spatial coordinates and normal vectors', while the SL control 'excluded normal vectors from the inputs'. For a triangular surface, curvature is a local differential quantity: mean curvature is half the divergence of the normal field and Gaussian curvature follows from the shape operator, so positions plus normals largely determine the target. A GNN given these inputs can learn a generic discrete geometry operator from any training set; no nonlinear-elasticity/growth physics is required. The SL baseline is therefore not a fair control for isolating 'physics transfer' — it is a control for whether normals are informative. The observed gap may be entirely due to this input asymmetry. A secondary but independent problem is that validation uses a population-averaged fetal atlas (Methods, Collection of medical data), so one-week-ahead MAE tests mean-atlas progression, not the claimed individual morphogenesis. Both issues attack the central claim, but the baseline confound is more fundamental because it undermines the main evidence that physics, rather than geometric features, transfers.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a physics-transfer (PT) learning framework for brain morphogenesis. A graph neural network is pretrained on finite-element simulations of growth-driven nonlinear elastic instabilities on spheres and ellipsoids, then applied zero-shot to human fetal brain surface meshes. The reported results are a curvature-prediction task, where PT achieves MAE below 0.5 mm^-1 versus 7.01 mm^-1 for a statistical-learning (SL) control, and a one-week-ahead morphogenesis prediction task with MAE below 0.01. The authors further analyze network weights and activations to argue that the PT model encodes transferable physics of cortical folding. The core claim is that nonlinear-elasticity physics learned on simple geometries transfers to complex brain geometries despite data scarcity.","tokens_in":14887,"tokens_out":6369,"duration_ms":74479,"significance":"If substantiated, the framework would be valuable: it addresses a real data-scarcity problem in fetal brain MRI, offers a computationally feasible route from mechanical simulations to patient-specific morphology prediction, and includes interpretability analyses that go beyond black-box accuracy. The digital-library idea and the explicit FEA-plus-GNN pipeline are strengths, and the paper is clearly written in its high-level structure. However, the current experimental design does not isolate physics transfer from geometric feature engineering, and the validation data cannot support the individual-level prediction claims. The significance of the contribution is therefore not established by the evidence presented.","major_comments":[{"comment":"The head-to-head comparison between PT and SL (Figs. 3c,d) is confounded by an input-feature asymmetry. The PT node features are 'spatial coordinates and normal vectors', while the SL control 'excluded normal vectors from the inputs' (Methods, Machine learning models). For a smooth surface, mean curvature is half the divergence of the normal field and the full curvature tensor is determined by position and normals; hence a model supplied with normals can learn a generic discrete differential-geometry operator from any training corpus and does not require elasticity/growth physics. The reported MAE of <0.5 mm^-1 vs 7.01 mm^-1 therefore does not establish that nonlinear-elasticity physics transfers; it may simply reflect the availability of normals. A valid control would give SL the same input features or compare against a standard cotangent-Laplacian curvature estimator, and would train b","section":"Results, 'Predictions of curvature maps and 3D morphology'; Methods, 'Machine learning models'"},{"comment":"The morphogenesis validation uses a population-averaged spatio-temporal atlas of the fetal brain spanning 21–36 weeks of gestation, with 32,492 vertices per hemisphere, not longitudinal scans of individual subjects. Consequently, the one-week-ahead MAE <0.01 measures how well the model tracks the population-mean developmental trajectory, not how well it predicts an individual brain's evolution. The Introduction and Discussion repeatedly claim individualized trajectory prediction and patient-specific digital twins; these claims are not supported by this dataset. Either individual longitudinal data are needed, or the claims must be explicitly restricted to population-mean morphogenesis.","section":"Methods, 'Collection of medical data'; Fig. 3f"},{"comment":"Equation (8) asserts p(θ|D'_L) ≈ p(θ|D'_H), but the supporting experiment reports only the mean layer-wise weight µ (Eq. 11). Equality of means is not equality of distributions; the identical mean could conceal very different variances or higher-order structure. Moreover, no quantitative distributional distance or statistical test is given, and the comparison appears to be based on a single training run. The claim of distributional similarity between PT models trained on spheres and ellipsoids is therefore unsupported. A proper analysis would compute, e.g., a Wasserstein distance between weight distributions across multiple random seeds and compare it with the same metric for SL.","section":"Model interpretability; Eq. (8) and Fig. 4d-e"}],"minor_comments":[{"comment":"The units for the one-week-ahead MAE (<0.01) are not specified. If the coordinates are normalized or scaled, please state this explicitly; otherwise the reader cannot interpret the magnitude.","section":"Results, Fig. 3f"},{"comment":"The composite loss includes a 'global gyrification index' but does not define it. Please specify how it is computed from the surface and how it enters the loss weighting.","section":"Methods, 'Machine learning models'"},{"comment":"The digital FEA library is not described in enough detail for reproducibility: no number of simulations, parameter sampling distributions, train/test splits, or mesh statistics are given. Please add these details or a reference to a public dataset.","section":"Methods, 'Digital libraries'"},{"comment":"The information-bottleneck comparison reports 34% versus 31% compression. Without error bars or significance testing, this small difference does not support the conclusion that deformation curvature 'plays a more critical role' in morphogenesis.","section":"Supplementary Note S2, Fig. S7d"},{"comment":"The notation p(θ|D), D' ⊂ D, and p(θ|D_L') ≈ p(θ|D_H') is informal and not standard in machine learning. Clarify the intended probabilistic model, or replace these statements with precise definitions of the data domains and model parameter distributions.","section":"Equations (5)–(8)"},{"comment":"No data or code availability statement is provided. Given the paper's reliance on FEA-generated data and GNN training, a public release or clear availability statement is essential for reproducibility.","section":"General"}],"recommendation":"reject","confidential_remarks":"The paper has potentially useful components, but the main claim is not supported by the current experiments. The SL confound is readily fixable and should be addressed in any resubmission: the control must receive the same input features or the PT model must be compared against a purely geometric baseline. The individual-morphogenesis claim requires longitudinal individual data or a substantial reframing to population-mean prediction. If the authors can provide a normals-matched baseline and show the PT advantage persists, a new submission could be viable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe headline: this is a plausible and timely idea—pretrain a GNN on FEA simulations of sphere/ellipsoid growth, then apply zero-shot to brain morphology—but the validation as it stands does not support the central claim of predicting individual brain morphogenesis. The paper is not a waste of time; it just needs a much stricter evaluation design.\n\nWhat's new: the specific pipeline of transferring nonlinear-elasticity physics from simple geometries to brain surfaces via a graph network is not something I've seen spelled out for brains. The digital library of finite-element simulations and the latent-space distance metric for estimating OOD error are sensible pieces. The architecture is borrowed from Pfaff et al., and the transfer idea from the authors' own prior work, but the application is new and could be useful.\n\nThe main problem is the PT-versus-SL comparison in Figure 3. The SL control is deprived of normal vectors, while PT gets positions plus normals. For a triangular mesh, curvature is a local function of positions and normals (mean curvature is divergence of the normal field). So the dramatic gap on curvature maps is exactly what you'd expect from an input-feature asymmetry; it says nothing about whether physics transferred. To make the claim, the control needs the same features, or a control that gets normals but not physics. As written, the 'physics-transfer advantage' is not established.\n\nSecond, the brain validation uses a single population-averaged atlas (32,492 vertices/hemisphere, 21–36 weeks). That's a cross-sectional average, not an individual trajectory. The one-week-ahead MAE below 0.01 is a fit to mean-atlas progression; it doesn't demonstrate the individual forecasting promised in the introduction. No error bars, no code, no data, which makes it hard to check.\n\nMinor: the interpretability section (weights distributions, activation similarities) is suggestive but rests on informal comparisons; the IB claims are thin. The limitations paragraph is honest about rare-event gaps but doesn't touch the validation issues above.\n\nWho this is for: researchers building surrogate models for biomechanical simulations, and anyone interested in domain transfer for scarce medical data. The paper deserves a serious referee, because the core idea is worth testing properly, but as submitted the central numerical evidence is confounded. I'd ask the authors for individual longitudinal data, a feature-matched SL baseline, and open code before making the digital-twin claim.","headline":"Plausible transfer-learning pipeline, but the headline results are confounded by an unfair baseline and an averaged atlas.","tokens_in":15257,"tokens_out":2157,"would_cite":false,"duration_ms":26117,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A network trained on the growth physics of spheres and ellipsoids can map fetal brain surfaces to curvature maps and forecast one-week shape changes without brain-specific training data.","keywords":["physics-transfer learning","cortical folding","morphogenesis prediction","graph neural networks","tangential growth model","fetal brain atlas","nonlinear elasticity","digital twin"],"falsifier":"Take a cohort with repeated fetal MRI scans of the same individuals, run the PT model from a baseline scan, and compare one-week-ahead predictions to the actual follow-up surface. If predictions are no closer to the follow-up than a no-growth template or a statistical model trained on the same atlas, the transfer claim is falsified. A second check: remove normal-vector inputs from the PT model; if curvature error jumps to the statistical-learning level, the advantage is geometric preprocessing rather than learned elasticity.","tokens_in":14435,"feed_emoji":"🧠","tokens_out":4935,"duration_ms":53613,"temperature":0.7,"pith_summary":"This paper claims that the mechanics of cortical folding can be learned once, from simple geometries, and then transferred directly to the human brain. The authors build a digital library of finite-element growth simulations on spheres and ellipsoids, train a graph neural network on those simulations, and apply its fixed weights to brain surface meshes. On curvature characterization the transferred model achieves mean absolute errors below 0.5 mm^-1, compared with 7.01 mm^-1 for a morphology-only statistical baseline, and on one-week-ahead morphogenesis the error stays below 0.01. If correct, the result matters because it offers a route around the scarcity of longitudinal brain MRI data: synthetic growth physics from elementary shapes stands in for missing individual developmental data.","feed_headline":"Sphere-growth physics forecasts fetal brain folding","feed_subtitle":"A network trained on simple elastic shells maps brain surfaces to curvature and one-week-ahead shape change.","key_machinery":"The load-bearing mechanism is physics-transfer learning: a graph neural network pretrained on a dense digital library of tangential-growth simulations of spherical and ellipsoidal core-shell structures, with weights frozen when applied to human brain meshes. The network is an encoder-decoder graph architecture whose nodes carry spatial coordinates and normal vectors, whose decoder produces local curvature, and whose rollout module integrates nodal accelerations via Newton's second law to step morphology forward in time. The tangential growth model—where outer gray matter grows faster than inner white matter and the mismatch drives buckling instabilities—supplies the physical content that is","core_discovery":"The paper's central claim is that the nonlinear-elasticity physics driving cortical pattern formation is geometry-transferable. Using a core-shell model with tangential cortical growth, the authors simulate morphogenesis on spheres and ellipsoids, then transfer the learned representations to human fetal brain surfaces. The transferred network predicts local curvature (sum of absolute mean and Gaussian curvatures) and short-horizon shape evolution with errors far below a statistical-learning control trained only on morphology. The authors further argue that the transferred model's internal representations—weight distributions and neuron activations—stay consistent across simple and brain geom","pith_inferences":["A direct test of the clinical headline would require per-subject longitudinal MRI, not just a population-averaged atlas; until then, 'individual trajectory prediction' remains a transfer claim awaiting individual-level validation.","The same recipe—a synthetic FEA library on simple shells, a frozen-weight graph network, zero-shot application to an organ mesh—should port to other growth-instability morphologies such as intestinal villi, tumor spheroids, or swelling gels, since the paper's own logic makes background geometry secondary.","The normal-vector input features carry local orientation information, so the PT-versus-SL gap may partly reflect a geometric descriptor rather than learned elasticity; an ablation withholding normals from the PT model would isolate the physics contribution.","The latent-distance/KDE metric could be deployed as a clinical early-warning system for when the digital twin is untrustworthy, but its calibration on real individual trajectories is not established in the paper."],"forward_implications":["Curvature characterization: the transferred model maps brain surface geometry to curvature maps with MAE below 0.5 mm^-1, versus 7.01 mm^-1 for a morphology-only statistical model.","Short-horizon forecasting: one-week-ahead brain morphology is predicted with MAE below 0.01, and multi-step autoregressive rollouts preserve key structural features.","Uncertainty estimation: the distance between test brain data and training sphere data in latent space correlates with prediction error, providing an a priori warning when the model is likely to fail.","Interpretability: PT-trained models show similar weight distributions and neuron activations on sphere and brain data, whereas statistical models diverge across domains.","Scale separation: information-bottleneck analysis indicates that localized deformation curvature is more informative for morphogenesis than large-scale background curvature."],"supporting_citations":[{"why":"Supplies the tangential-growth finite-element simulation parameterization used to build the digital library.","marker":"[19]"},{"why":"Establishes the constrained core-shell expansion model of gyrification that the simulations extend.","marker":"[18]"},{"why":"Provides cortical-thickness and stiffness ranges and the folding-pattern simulation approach used to sample the library.","marker":"[22]"},{"why":"Frames the continuum-mechanics and growth-tensor relationship between tissue growth and cortical folding instabilities.","marker":"[17]"},{"why":"Supplies the mesh-graph network architecture with learned accelerations and rollout used for spatiotemporal prediction.","marker":"[43]"},{"why":"Introduces the physics-transfer learning formulation this framework builds on.","marker":"[28]"},{"why":"Documents the scarcity of fetal brain MRI atlases and datasets, motivating the synthetic-library strategy.","marker":"[32]"},{"why":"Provides the population-averaged fetal cortical surface atlas used as ground truth for brain validation.","marker":"[60]"},{"why":"Documents the scarcity of longitudinal MRI data and normative brain charts that frame the data-scarcity problem.","marker":"[9]"}],"fun_headline_variants":["Physics-transfer model predicts brain folding from simple shapes","Brain growth predicted via sphere-based physics transfer","Simple-geometry physics maps fetal brain curvature","Transferring shell physics to forecast cortex folding","Brain morphogenesis learned from spheres and ellipses"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The evaluation treats a population-averaged fetal brain atlas spanning 21 to 36 weeks of gestation as valid ground truth for testing prediction of individual brain morphogenesis; cross-sectional averages may not represent any individual's developmental trajectory.","fun_headline_variants_meta":{"raw":{"variants":["Physics-transfer model predicts brain folding from simple shapes","Brain growth predicted via sphere-based physics transfer","Simple-geometry physics maps fetal brain curvature","Transferring shell physics to forecast cortex folding","Brain morphogenesis learned from spheres and ellipses"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000341,"raw_usage":{"total_tokens":1679,"prompt_tokens":671,"completion_tokens":1008,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":415,"completion_tokens_details":{"reasoning_tokens":940}},"tokens_in":415,"tokens_out":1008,"duration_ms":8181,"temperature":1.0,"reasoning_tokens":940,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T17:24:38.065434+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a cohort with repeated fetal MRI scans of the same individuals, run the PT model from a baseline scan, and compare one-week-ahead predictions to the actual follow-up surface. If predictions are no closer to the follow-up than a no-growth template or a statistical model trained on the same atlas, the transfer claim is falsified. A second check: remove normal-vector inputs from the PT model; if curvature error jumps to the statistical-learning level, the advantage is geometric preprocessing rather than learned elasticity.","supporting_citations":[{"cited_title":"On the growth and form of cortical convolutions.Nat","cited_arxiv_id":null,"evidence_quote":"Supplies the tangential-growth finite-element simulation parameterization used to build the digital library."},{"cited_title":"Gyrification from constrained cortical expansion","cited_arxiv_id":null,"evidence_quote":"Establishes the constrained core-shell expansion model of gyrification that the simulations extend."},{"cited_title":"The influence of biophysical parameters in a biomechanical model of cortical folding patterns","cited_arxiv_id":null,"evidence_quote":"Provides cortical-thickness and stiffness ranges and the folding-pattern simulation approach used to sample the library."},{"cited_title":"Computational models of cortical folding: A review of common approaches.J","cited_arxiv_id":null,"evidence_quote":"Frames the continuum-mechanics and growth-tensor relationship between tissue growth and cortical folding instabilities."},{"cited_title":"Learning mesh-based simula- tion with graph networks","cited_arxiv_id":null,"evidence_quote":"Supplies the mesh-graph network architecture with learned accelerations and rollout used for spatiotemporal prediction."},{"cited_title":"Discovering High-Strength Alloys via Physics-Transfer Learning","cited_arxiv_id":"2403.07526","evidence_quote":"Introduces the physics-transfer learning formulation this framework builds on."},{"cited_title":"Fetal brain MRI atlases and datasets: A review.Neu- roImage, 292:120603, 2024","cited_arxiv_id":null,"evidence_quote":"Documents the scarcity of fetal brain MRI atlases and datasets, motivating the synthetic-library strategy."},{"cited_title":"Developing Human Connectome Project spatio-temporal surface atlas of the fetal brain","cited_arxiv_id":null,"evidence_quote":"Provides the population-averaged fetal cortical surface atlas used as ground truth for brain validation."},{"cited_title":"Brain charts for the human lifespan","cited_arxiv_id":null,"evidence_quote":"Documents the scarcity of longitudinal MRI data and normative brain charts that frame the data-scarcity problem."}],"review_version":1}