{"id":"570eaade-db6a-4c00-9011-0ae58d1ba8ef","arxiv_id":"2505.12228","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A synthetic-data-trained network (recon-any) enables cortical surface reconstruction and morphometry from portable low-field MRI, achieving strong agreement with high-field MRI for 3 mm T2 scans.","lead":"A new deep learning pipeline, recon-any, reconstructs brain cortical surfaces from low-field portable MRI scans after training only on synthetic data. On 15 paired scans, a 3 mm T2 low-field scan matches high-field MRI for surface area (r=0.96) and parcellation (Dice=0.98), though thickness agreement is weaker.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Single-scanner validation and uncalibrated synthetic domain randomization are untested against other LF systems; 'out-of-the-box' robustness claim is underdetermined.","rationale":"The paper's core value is the claim that a 3 mm isotropic T2 LF-MRI scan can be processed out of the box with morphometric accuracy close to HF-MRI. For this to hold generally, the U-Net must have learned representations of LF-MRI appearance from synthetic training data that cover the real distribution across scanners. The reader's weakest assumption identifies exactly this risk, and I agree. The evidence in the paper is internally consistent: a synthetic-trained model performs well on 15 Hyperfine subjects and on qualitative postmortem cases. However, the four domain-randomization knobs are not derived from measured LF-MRI physics, and there is no quantitative comparison between synthetic and real image statistics, so we cannot tell whether the augmentation ranges bracket the true distribution or simply happen to cover the single evaluated scanner. This is a correctness risk for the generalization clause, not an internal inconsistency. A cross-scanner or cross-protocol evaluation would directly test whether the transfer holds. Given the reader's CONDITIONAL verdict, my stress test does not change it; it sharpens the condition that must be met to support the broad 'robust across LF-MRI' claim.","tokens_in":18345,"tokens_out":6084,"duration_ms":63704,"concrete_test":"Assemble paired high-field/low-field scans from a second portable LF scanner (or a different Hyperfine site with different sequence parameters) and run the frozen recon-any model without retraining; compute AAD, Dice, and Pearson correlations for 3 mm isotropic T2 as in Table 1. If volume or surface-area correlations fall outside the reported 95% confidence intervals (e.g., r<0.85) or lobe-level Dice drops below 0.90, the synthetic domain randomization does not cover real LF-MRI variability.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that recon-any works 'out of the box' across low-field MRI depends on the unverified premise that the domain-randomized synthetic training distribution (Methods 4.3) reproduces real portable LF-MRI contrast, noise, and artifacts. The four hand-set augmentation ranges (in-plane resolution to 2 mm, isotropic to 4 mm, 2x bias field, 2x deformation and rotation) are relative to the original SynthSeg model, not derived from measured scanner physics, and the underlying Gaussian mixture model is not calibrated to 64 mT relaxation parameters. All in vivo validation uses 15 healthy adults on a single Hyperfine 64 mT scanner, with no cross-scanner or cross-protocol test; postmortem evidence is only qualitative expert QC. The paper itself states that all evaluations were on Hyperfine and that generalizability to other systems is future work (Section 3). Therefore the reported r=0.96/0.93 and Dice=0.98 may not transfer, and the 'robust across LF-MRI' claim is underdetermined.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents recon-any, a deep-learning pipeline for cortical surface reconstruction, parcellation, and morphometry from portable low-field MRI (LF-MRI). The method is a 3D U-Net trained on domain-randomized synthetic LF-MRI-like images to predict signed distance functions for the white and pial surfaces, followed by geometric post-processing. The authors validate on paired 3T and 64 mT Hyperfine scans from 15 healthy adults across multiple resolutions (1.6x1.6x5 mm axial, 2, 3, 4 mm isotropic) and contrasts (T1, T2), reporting surface error metrics, Dice parcellation overlap, and Pearson correlations for gray matter volume, surface area, and cortical thickness. They additionally report qualitative postmortem validation on fresh and cadaveric brains. The headline results are that 3 mm isotropic T2 LF-MRI yields surface area correlation r=0.96, gray matter volume r=0.93, and cortical parcellation Dice=0.98 relative to FreeSurfer on high-field MRI.","tokens_in":18590,"tokens_out":6937,"duration_ms":64981,"significance":"If the reported results hold, the method would substantially lower the barrier to cortical morphometry by enabling surface-based analysis on portable, inexpensive LF-MRI systems, with potential applications in acute care, resource-limited settings, and postmortem neuropathology. The paper's strengths include paired high-field/low-field data from the same subjects, evaluation across multiple resolutions and contrasts, surface-based error metrics, confidence intervals on correlations, and public release of the tool within FreeSurfer. The domain-randomized synthetic training approach is a sensible strategy for a modality with scarce annotated data. However, the central 'robust across LF-MRI' claim is currently supported by data from a single scanner model (Hyperfine 64 mT), and the synthetic augmentation ranges are not calibrated to measured scanner physics, so the cross-scanner generality remains unproven.","major_comments":[{"comment":"The manuscript's headline robustness claim is underdetermined by the evidence: Section 3 explicitly states that all evaluations used Hyperfine scanners and that generalization to other LF systems is future work, while Section 4.3 describes the four hand-set augmentation ranges (in-plane resolution to 2 mm, isotropic to 4 mm, 2x bias field, 2x deformation and rotation) as multipliers relative to the original SynthSeg model rather than values derived from measured 64 mT scanner physics. Because the 'out of the box' generalization claim rests on this uncalibrated synthetic distribution, the paper should either provide a quantitative synthetic-to-real domain-shift analysis (e.g., noise statistics, intensity distributions, or artifact profiles) or evaluate on at least one additional low-field scanner; otherwise the claims should be narrowed to Hyperfine-specific performance.","section":"Section 3, Section 4.3"},{"comment":"No baseline comparison against existing LF-capable analysis pipelines is provided. The Introduction and Discussion assert that tools such as FreeSurfer 'struggle' with LF-MRI and that recon-any is 'the only viable method currently available,' but the manuscript only compares recon-any to FreeSurfer on high-field reference surfaces. Adding a comparison against LF-SynthSR followed by recon-all, or recon-all-clinical applied directly to the same LF scans, would quantify the claimed advantage; if such a comparison is not feasible, the 'only viable method' claim should be tempered.","section":"Section 2, Section 3"},{"comment":"There is a direct contradiction between the text and Table 1 regarding cortical thickness correlations. The text states that T1 LF-MRI thickness correlations (r~0.3) 'remain statistically significant,' but Table 1 marks every T1 thickness entry with a dagger, indicating 95% confidence intervals that include zero (e.g., 0.30 [-0.25, 0.70] for 3 mm T1). This is load-bearing for the claim that thickness estimation is feasible across contrasts; the text should be corrected and the interpretation revised to reflect that T1 thickness correlations are not statistically distinguishable from zero in this sample.","section":"Section 2.3, Table 1"},{"comment":"The reported scan counts are internally inconsistent. Section 2.1 describes 15 subjects each having 1.6x1.6x5 mm axial, 2 mm, 3 mm, and 4 mm scans for both T1 and T2, which implies 120 LF-MRI scans, yet the text states '60 scans total.' Similarly, Section 2.4 reports 31 postmortem LF-MRI scans that passed QC, while Section 4.2 reports 32 postmortem brains (21 fresh + 11 cadaveric). These counts should be reconciled, and the number of subjects contributing to each condition should be stated explicitly.","section":"Section 2.1, Section 2.4, Section 4.2"}],"minor_comments":[{"comment":"The abstract states gray matter volume correlates at r=0.93 for 3 mm isotropic T2, but Table 1 reports 0.95 for that condition and 0.92 for 4 mm T2; please clarify which acquisition or pooling the abstract value refers to.","section":"Abstract, Table 1"},{"comment":"The text says the parcellation results 'are based on 3 mm isotropic T1 and T2 LF-MRI scans' but also that 'the mean Dice coefficient was computed for each parcel across all resolutions'; Figure 4 caption does not state the resolution. Please disambiguate which resolutions are shown in the main figure.","section":"Section 2.2"},{"comment":"The L2 loss equation contains a typographical error: the notation [$SDF_i$ and $SDF_i$ is unclear, and the predicted SDF should be denoted with a hat or similar. Please correct the formula.","section":"Section 4.3"},{"comment":"The demographic description states 10 females and 5 males who were 'either White (n=12) or Asian (n=2)' and 'non-Hispanic (n=11) or unknown (n=3)'; these numbers sum to 14, not 15. Please reconcile the missing participant or clarify that one participant's demographics were unreported.","section":"Section 4.2"},{"comment":"The caption says the 1 mm HF-MRI reference values come from scans 'used during training,' but Section 2.1 describes these as 15 held-out high-resolution test scans. Please reword to avoid implying the reference surfaces came from the training set.","section":"Figure 3 caption, Section 2.1"}],"recommendation":"major_revision","confidential_remarks":"This is a well-executed engineering contribution with a valuable paired HF/LF dataset and transparent confidence intervals. The main risk is overclaiming cross-scanner generality from a single-scanner validation set; the editor should require either additional validation or a narrowed claim. The declared financial interest of an author in Hyperfine is disclosed and appears appropriately managed. The statistical contradiction in the thickness results needs correction before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a useful, honest engineering contribution. It delivers the first quantitative demonstration I know of that cortical surfaces can be reconstructed from portable 64 mT MRI with agreement close to high-field FreeSurfer for volume and area, and it releases the code. The practical standout is the 3 mm isotropic T2 scan, under 4 minutes, which gives surface area r=0.96, gray matter volume r=0.93, and parcellation Dice near 0.98. They also report confidence intervals on the correlations and are upfront that cortical thickness is hard, especially for T1, where correlations are nonsignificant. That transparency is a real strength.\n\nThe main soft spot is the scope of the robustness claim. All real data come from a single Hyperfine scanner, and the synthetic domain-randomization ranges are hand-set rather than derived from measured scanner physics. The paper itself acknowledges cross-scanner generalizability is future work, so the title and abstract's \"robust across LF-MRI\" oversells what is actually shown. I would not call this a fatal flaw; the method genuinely handles multiple contrasts, resolutions, and orientations on real out-of-distribution inputs. But the \"out of the box\" claim is underdetermined until a second scanner or a measured forward model is tested.\n\nThe other issues are smaller. There is no baseline comparison against LF-SynthSR plus FreeSurfer, or against any other existing tool, which weakens the assertion that recon-any is \"the only viable method.\" A simple baseline would make the contribution concrete rather than rhetorical. Postmortem validation is purely qualitative expert QC; reasonable given the lack of an HF reference, but it doesn't quantify anything. There are also minor inconsistencies: 31 postmortem brains in the results section versus 32 in methods, and \"60 scans total\" does not obviously match 15 subjects times eight reported sequences. These are fixable.\n\nOn circularity: the model is trained on FreeSurfer-generated SDFs and evaluated against FreeSurfer on HF-MRI. That is a shared convention in the field, and the LF inputs are genuinely out-of-distribution, so I do not treat it as a serious problem.\n\nMy recommendation: give this a proper peer review. The central single-scanner claim holds up, the method is a useful step for bedside and autopsy settings, and the paper is honestly written. I would ask the authors for a baseline comparison, a tempering of the generalization language, and a cleanup of the small inconsistencies.","headline":"Solid, honest engineering paper that gives low-field MRI a cortical surface pipeline with real paired validation, but the 'out of the box' claim is only demonstrated on one scanner and would benefit from a baseline comparison.","tokens_in":19149,"tokens_out":3020,"would_cite":true,"duration_ms":31930,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 3D U-Net trained on synthetic low-field images reconstructs cortical surfaces from portable low-field MRI with high-field-level surface area and gray-matter volume correlations.","keywords":["low-field MRI","portable MRI","cortical surface reconstruction","signed distance functions","domain randomization","synthetic training data","cortical morphometry","postmortem brain imaging"],"falsifier":"Run recon-any without retraining on paired high-field and low-field scans from a second portable low-field system with different coil geometry, and compare isotropic T2 reconstructions to high-field surfaces; if mean surface error exceeds about 2 mm or surface-area correlation drops below about 0.9, the synthetic generator has not captured the new hardware. A cheaper probe is to measure the actual noise and bias-field statistics of real low-field scans and check whether they fall outside the simulated ranges, namely resolutions above 4 mm, bias fields stronger than twice the base model, or deformations stronger than twice the base model; the out-of-box claim should fail exactly where the simulator stops.","tokens_in":18198,"feed_emoji":"🧠","tokens_out":11473,"duration_ms":101125,"temperature":0.7,"pith_summary":"This paper tries to establish that cortical surface reconstruction and morphometry—normally reserved for high-field 1 mm T1 MRI—can be done reliably from portable low-field MRI. The proposed pipeline, recon-any, is a 3D U-Net that predicts signed distance functions to the white-matter and pial surfaces from scans of arbitrary contrast and resolution, trained only on synthetic low-field-like images produced by domain randomization. On paired 64 mT low-field and 3 T high-field scans of the same subjects, a 3 mm isotropic T2 scan acquired in under four minutes yields surface area correlation $r=0.96$, gray-matter volume $r=0.93$, and parcellation Dice of 0.98; cortical thickness is the weak point, with best correlation $r=0.70$. If the claim holds, cheap bedside scanners could support quantitative cortical analysis in clinics, field settings, and postmortem neuropathology rather than just emergency screening.","feed_headline":"Low-field scan approaches high-field cortical surface maps","feed_subtitle":"Trained only on synthetic scans, it reaches surface-area r=0.96, volume r=0.93, and parcellation Dice=0.98 vs high-field.","key_machinery":"The load-bearing object is the signed distance function (SDF): a volumetric field that records, at every voxel, the signed distance to the white-matter or pial surface. A 3D U-Net with four encoder and four decoder levels is trained with an L2 loss on SDFs clipped to $\\pm 5$ mm, using synthetic low-field images generated by a Bayesian-style generative model under domain randomization—voxel intensities sampled from tissue-class Gaussian mixtures, with added noise, bias fields, and nonlinear deformations, resolutions up to 4 mm, doubled bias-field and deformation strength, and simulations of ex vivo brains. At test time, marching cubes extracts an initial surface from the predicted SDF, and a geometric optimization step enforces smoothness, correct topology, and freedom from self-intersections. This SDF representation is what lets the method bypass contrast-specific segmentation and voxel-resolution limits.","core_discovery":"The central discovery is that direct regression of cortical signed distance functions from synthetic low-field MRI transfers to real portable scans without retraining, provided the synthetic generator spans the relevant resolution, contrast, noise, bias-field, and deformation ranges. The paper shows that reconstruction accuracy is acquisition-dependent: isotropic T2 contrast outperforms T1 at low field, and 3 mm isotropic is a practical sweet spot, since 2 mm buys little accuracy at about three times the acquisition time while the default axial $1.6\\times1.6\\times5$ mm sequences are clearly suboptimal. Against high-field reference surfaces, mean surface placement errors stay around 1–2 mm, below the voxel size, and lobe-level gray-matter volume errors are typically 3–6%. Cortical thickness is the exception: with 3 mm voxels and low contrast, sub-millimeter thickness estimates reach only $r\\approx0.70$ on T2 and $r\\approx0.30$ on T1, so thickness is reported as feasible but not high-field-equivalent. The same pipeline also reconstructs postmortem brains, including fresh, deformed tissue, with all 31 cases passing expert quality control.","pith_inferences":["Beyond the paper: the domain-randomization recipe is testable on other low-field hardware; if a second scanner's noise, bias-field, or deformation statistics fall inside the simulated ranges, the same out-of-box model should transfer, but the paper does not demonstrate this.","Beyond the paper: because the evaluation uses one scanner model, the reported correlations are best read as an upper bound for portable low-field systems generally; a scanner with different coil geometry or stronger field inhomogeneity could exceed the simulated artifact envelope.","Beyond the paper: the observation that T2 outperforms T1 at low field reverses the high-field convention and suggests that portable scanner vendors should prioritize isotropic T2 sequences for cortical morphometry, a design implication the authors raise but do not pursue quantitatively.","Beyond the paper: error maps concentrate in deep sulci and small parcels, so an uncertainty map per surface vertex—flagging low-confidence regions—would be a natural clinical extension; the paper does not provide one."],"forward_implications":["A 3 mm isotropic T2 low-field protocol of under four minutes becomes a practical acquisition target for cortical morphometry, while the default axial $1.6\\times1.6\\times5$ mm sequences should be avoided for surface analysis.","Surface area and gray-matter volume from low-field scans can support group-level and longitudinal studies, with correlations above 0.9 at every tested resolution from axial to 4 mm and across both T1 and T2 contrasts.","Cortical thickness from low-field scans should be interpreted cautiously; even the best T2 protocol leaves $r=0.70$, and anisotropic T1 drops to $r\\approx0.30$.","Lobe-level parcellation is reliable, with Dice above 0.85 everywhere and above 0.90 for most regions, enabling region-based analyses without high-field access.","Because the model works out of the box, a new low-field site needs only a scan inside the simulated augmentation ranges—no retraining or local training data are required."],"supporting_citations":[{"why":"It supplies the signed-distance-function pipeline and the domain-randomized training scheme that recon-any extends to low-field and postmortem imaging.","marker":"[28]"},{"why":"It provides the synthetic-data generation and surface-refinement machinery, including the deformations and bias-field model that the low-field modifications build on.","marker":"[32]"},{"why":"It supplies the Bayesian Gaussian-mixture image synthesis model from which the low-field contrast, noise, and bias-field simulations are drawn.","marker":"[33]"},{"why":"It generates the volumetric segmentations used to build the SDF training targets from high-resolution datasets.","marker":"[31]"},{"why":"It defines the high-field reference surfaces and parcellations against which low-field reconstructions are compared.","marker":"[8]"},{"why":"It supplies the super-resolution and T1 synthesis approach used for registration and downstream analysis, representing the prior super-resolution route.","marker":"[25]"},{"why":"It is the low-field-specific super-resolution baseline whose cortical-surface limitations motivate direct SDF prediction.","marker":"[26]"},{"why":"It reports subcortical segmentation results on the same portable scanner and the T2-over-T1 acquisition preference that the surface results corroborate.","marker":"[18]"},{"why":"It defines the cortical parcellation atlas used to compute region-level Dice coefficients.","marker":"[27]"}],"fun_headline_variants":["Low-field MRI cortical surfaces match high-field: r=0.96, Dice=0.98","AI maps cortex from low-field MRI, matching high-field quality","Quick low-field MRI plus AI yields high-field-grade cortical maps","Synthetic-trained network reconstructs cortex from portable low-field MRI","Low-field MRI with AI: cortical maps near high-field accuracy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the synthetic low-field images produced by the domain-randomized generator faithfully reproduce the contrast, resolution, noise, bias-field, and deformation properties of real portable low-field scanners, so that a network trained only on synthetic data transfers to real scans; the evaluation tests this on a single 64 mT scanner model.","fun_headline_variants_meta":{"raw":{"variants":["Low-field MRI cortical surfaces match high-field: r=0.96, Dice=0.98","AI maps cortex from low-field MRI, matching high-field quality","Quick low-field MRI plus AI yields high-field-grade cortical maps","Synthetic-trained network reconstructs cortex from portable low-field MRI","Low-field MRI with AI: cortical maps near high-field accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000278,"raw_usage":{"total_tokens":1748,"prompt_tokens":1132,"completion_tokens":616,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":748,"completion_tokens_details":{"reasoning_tokens":522}},"tokens_in":748,"tokens_out":616,"duration_ms":6201,"temperature":1.0,"reasoning_tokens":522,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:38:06.122363+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run recon-any without retraining on paired high-field and low-field scans from a second portable low-field system with different coil geometry, and compare isotropic T2 reconstructions to high-field surfaces; if mean surface error exceeds about 2 mm or surface-area correlation drops below about 0.9, the synthetic generator has not captured the new hardware. A cheaper probe is to measure the actual noise and bias-field statistics of real low-field scans and check whether they fall outside the simulated ranges, namely resolutions above 4 mm, bias fields stronger than twice the base model, or deformations stronger than twice the base model; the out-of-box claim should fail exactly where the simulator stops.","supporting_citations":[{"cited_title":": Synthetic data in generalizable, learning-based neuroimaging","cited_arxiv_id":null,"evidence_quote":"It provides the synthetic-data generation and surface-refinement machinery, including the deformations and bias-field model that the low-field modifications build on."},{"cited_title":": Synthseg: Segmentation of brain mri scans of any contrast and resolution without retraining","cited_arxiv_id":null,"evidence_quote":"It supplies the Bayesian Gaussian-mixture image synthesis model from which the low-field contrast, noise, and bias-field simulations are drawn."},{"cited_title":"NeuroImage 143, 235– 249 (2016)","cited_arxiv_id":null,"evidence_quote":"It generates the volumetric segmentations used to build the SDF training targets from high-resolution datasets."},{"cited_title":"Neuroimage 9(2), 195–207 (1999)","cited_arxiv_id":null,"evidence_quote":"It defines the high-field reference surfaces and parcellations against which low-field reconstructions are compared."},{"cited_title":"Science advances 9(5), 3607 (2023)","cited_arxiv_id":null,"evidence_quote":"It supplies the super-resolution and T1 synthesis approach used for registration and downstream analysis, representing the prior super-resolution route."},{"cited_title":"Radiology 306(3), 220522 (2022)","cited_arxiv_id":null,"evidence_quote":"It is the low-field-specific super-resolution baseline whose cortical-surface limitations motivate direct SDF prediction."},{"cited_title":"Nature Communications 15(1), 1–12 (2024)","cited_arxiv_id":null,"evidence_quote":"It reports subcortical segmentation results on the same portable scanner and the T2-over-T1 acquisition preference that the surface results corroborate."},{"cited_title":"NeuroImage 31(3), 968–980 (2006)","cited_arxiv_id":null,"evidence_quote":"It defines the cortical parcellation atlas used to compute region-level Dice coefficients."}],"review_version":1}