{"id":"7f159266-1b77-45dc-b283-9aa22f8a91b8","arxiv_id":"2411.14630","paper_version":1,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"ACE-Net estimates B0 and eddy-current field imperfections from autofocus metrics and spiral diffusion images, enabling correction without per-scan calibration.","lead":"ACE-Net uses a neural network plus autofocus image-sharpness metrics to estimate magnetic field errors (B0 and eddy currents) from spiral diffusion MRI scans, without separate calibration scans at scan time. It was tested on two volunteers and produced clearer diffusion images after correction.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unvalidated basis-order and cross-b-value generalization leave the 'accurate estimation' claim contingent; a quantitative residual test against the b=3000 Skope ground truth is needed.","rationale":"The reader's conditional verdict is well-founded. For the central claim to hold, ACE-Net must both represent the true b=3000 fields within its compact basis and generalize from training data that were not acquired at b=3000. The paper states the basis order was determined from field-probe data at lower b-values and then applied at b=3000, which is a genuine extrapolation. The test acquisitions did include a Skope ground truth, yet no quantitative field-error or residual-blur metrics are reported, so the key condition is asserted rather than demonstrated. My proposed test is feasible immediately from the reported acquisitions and would either confirm the basis and network at the target b-value or expose the missing condition. I do not see an internal inconsistency or a reason to reject; the paper needs the quantitative residual analysis to support the headline claim. Since the reader already recommends CONDITIONAL, no verdict change is needed.","tokens_in":3001,"tokens_out":4642,"duration_ms":47698,"concrete_test":"Reuse the existing b=3000 two-volunteer Skope ground truth. (1) Fit the measured spatiotemporal field with the 3rd-order SH/polynomial basis and with higher-order alternatives (e.g., spatial order 5 and temporal order 5); compute out-of-basis residual RMS Hz in the brain over the diffusion-encoding window. If the residual energy exceeds ~10% of the total field energy, the basis is insufficient at the target b-value. (2) Compute pointwise error maps between ACE-Net predictions (CNN and unrolled) and the same Skope ground truth: report mean and 95th-percentile Hz, and compare with the baseline 3rd-order fit. If image sharpness improves while field error remains large, hallucination is not excluded. This uses data already in hand and does not require new scans.","verdict_should_be":"UNCHANGED","load_bearing_attack":"ACE-Net's central claim requires that the spatiotemporal field at the target b=3000 be accurately representable by 3rd-order spatial spherical harmonics times 3rd-order temporal polynomials, and that the CNN generalizes from training data acquired at b=1000/2000 (different mixing times/TEs) plus synthesized directions. The basis order is justified only from field-probe data at lower b-values, not from a reported analysis of the b=3000 ground truth acquired on the two test volunteers. If real b=3000 eddy currents contain higher-order spatial modes or temporal components beyond a cubic polynomial, the network's output space cannot contain the true field, and every downstream correction—autofocus or unrolled—will be biased. The paper presents no quantitative field-error metric (e.g., RMS/percentile Hz against Skope) and no quantitative blur-reduction metric; the visual figures are insufficient to rule out that the network finds a field that improves sharpness but is not the true field. This is the load-bearing unverified condition for 'accurate estimation without external calibrations.'","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes ACE-Net, a deep-learning method for estimating spatiotemporal field imperfections (B0 inhomogeneity and diffusion-encoding-induced eddy currents) in high-b-value spiral diffusion MRI. The method combines autofocus blur metrics with a CNN and uses a compact basis of 3rd-order spherical harmonics in space and 3rd-order polynomials in time to represent the field. A static B0 map is estimated from a b=0 image, and a subsequent network estimates per-diffusion-direction spatiotemporal eddy/dB0 fields, with an optional unrolled data-consistency refinement. Training data are synthesized from previously acquired B0 maps, MRF data, BUDA-EPI DWIs, and Skope measurements at b=1000/2000, while testing is performed on two volunteers scanned with b=3000 spiral diffusion with multi-echo GRE and Skope measurements as ground truth. The results are presented qualitatively through figures showing estimated fields and corrected images.","tokens_in":3215,"tokens_out":2500,"duration_ms":26712,"significance":"If the claimed accuracy were quantitatively established, ACE-Net would be a practically valuable contribution: it could eliminate scan-time calibration for eddy-current and B0 correction in high-b-value spiral diffusion MRI, a setting where external field probes and lengthy calibration scans are burdensome. The paper has some notable strengths: the evaluation uses external ground truth (multi-echo GRE and Skope), the training and test data are separate, and the unrolled variant is a sensible attempt to reduce hallucination risk. However, the evidence as presented is not sufficient to support the abstract's 'accurate estimation' claim, and the central modeling assumption about the basis order is not validated on the target b=3000 protocol. The work is therefore best viewed as a promising proof-of-concept that requires stronger quantitative validation before its main claims can be accepted.","major_comments":[{"comment":"The evaluation is entirely qualitative. No quantitative field-error metric (e.g., RMS or percentile error in Hz against the Skope ground truth), no image-quality metric (e.g., sharpness, SSIM, or blur reduction), and no error bars or per-volunteer breakdown are reported. With only two test volunteers, visual inspection of selected slices is insufficient to support the abstract's claim of 'accurate estimation' and 'high quality image reconstruction'. The authors should report quantitative residuals between estimated and ground-truth fields (at least on the two test volunteers, ideally across slices and diffusion directions) and quantitative measures of image-quality improvement.","section":"Results, Figs. 3-5"},{"comment":"The compact basis representation (3rd-order spatial spherical harmonics times 3rd-order temporal polynomials) is justified only by 'field-probe data' without specifying which protocols or b-values those data came from. Because this basis defines the output space of the network, any field component outside this subspace at b=3000 cannot be represented, causing a systematic bias regardless of training quality. The authors should provide a basis-truncation analysis on the b=3000 Skope ground truth (e.g., residual field energy or Hz error as a function of expansion order) to demonstrate that the chosen order is adequate for the target protocol.","section":"Methods, 'Spatiotemporal eddy and dB0 fields'"},{"comment":"The spatiotemporal training data are derived from Skope measurements at b=1000 and b=2000 with different mixing times and TEs than the b=3000 test protocol, supplemented by synthesized non-principal diffusion directions. No evidence is given that eddy-current fields at b=3000 lie in the same distribution after the 30% coefficient augmentation. The cross-b-value generalization is a load-bearing assumption and should be explicitly tested, for example by comparing the network's b=3000 field estimates against the Skope ground truth with quantitative metrics, as requested above.","section":"Methods, 'Data'"},{"comment":"The text states that the unrolled ACE-Net 'outperformed the CNN' in field-parameter estimation, but no numerical results or statistical comparison are provided. If the unrolled variant is claimed to be superior, the authors should report quantitative error metrics for both variants, ideally across multiple test cases, to substantiate this comparison.","section":"Results, Fig. 4"}],"minor_comments":[{"comment":"The autofocus metric definition contains typographical artifacts in the equation (e.g., misplaced parentheses and superscript notation); please rewrite it clearly with standard LaTeX.","section":"Methods, 'Static B0 inhomogeneity'"},{"comment":"The choice of autofocus search range (±120 Hz at 10 Hz increments) and the reduction to a valid 40 Hz range are described but not justified; a brief explanation of how these parameters affect performance would help.","section":"Methods, 'Static B0 inhomogeneity'"},{"comment":"The training data synthesis is described only briefly; please specify the forward-model details (e.g., how spiral k-space data are generated from the B0 maps and synthesized T2-w images) and the exact augmentation procedure for the basis coefficients.","section":"Methods, 'Data'"},{"comment":"The figures are referenced as 'Fig.3' and 'Fig.4' inconsistently in the Results section; please renumber and reference all figures consistently and ensure each figure is called out in order.","section":"General"},{"comment":"No information is provided about network training details (loss functions, optimizer, learning rate, number of epochs, or computational cost); adding a short training-details section would aid reproducibility.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The manuscript appears to be a short conference-style report, and the figures suggest the method works in examples. However, the absence of any quantitative evaluation is a serious gap for a journal submission, and the basis-order assumption is not validated for the target b=3000 protocol. The reader's low confidence in the central claim is warranted. I recommend major revision with a request for quantitative field-error and image-quality metrics, a basis-truncation analysis on the b=3000 ground truth, and clarification of the generalization from lower-b training data. If the authors can provide these, the paper could be acceptable; if the quantitative results show large errors at b=3000, the central claim would need to be substantially weakened."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You asked about ACE-Net. Quick take: this is a plausible integration of known pieces, but the abstract-level evidence doesn't support the \"accurate estimation\" claim as stated.\n\nWhat's genuinely new: combining autofocus metrics with a physics-informed CNN and a compact spatiotemporal basis, plus the unrolled refinement for robustness. Each ingredient exists in the literature (Noll, Anderson, ResoNet, SAFE, Valsamis), but the combination for high-b spiral diffusion without external calibration is a reasonable workflow improvement. The authors also deserve credit for acquiring Skope and multi-echo GRE ground truth on the two test subjects and for training on separate data—so the evaluation is not circular in the naive sense.\n\nThe soft spots are real and load-bearing. First, no quantitative field-error metric anywhere. The figures show visual similarity, but you cannot rule out that the network finds a field that improves sharpness while being systematically wrong in ways that matter for DTI. Second, the spatiotemporal basis order (3rd-order spherical harmonics × 3rd-order polynomials) is justified from field-probe data at lower b-values, and the paper never reports a residual analysis against the b=3000 Skope ground truth it already has. If real eddy fields at b=3000 contain higher spatial modes or temporal components beyond cubic, the network's output space cannot contain the true field, and every downstream correction is biased. This is the weakest joint, and the stress-test note is right to flag it. Third, two test volunteers is thin for a claim of \"accurate\" estimation; that is a feasibility demonstration, not validation. Fourth, the training data come from different protocols and scanner sessions (b1000/2000 with different TEs, synthesized DWIs from BUDA-EPI), which is fine for an abstract but leaves cross-b-value generalization unexamined.\n\nI want to be fair: this is an extended abstract, so we shouldn't expect full reproducibility or error bars. But the authors' own conclusion goes beyond what they show. The one missing experiment that would transform this is a quantitative residual of estimated vs. Skope field at b=3000 (RMS or percentile Hz, plus a blur metric before/after). They have the ground truth; they just didn't report the numbers.\n\nRecommendation: I would not send this to referees as-is. It reads as a work-in-progress abstract, and the evidence is too thin for a serious journal. But the method has merit, and if the authors add the residual analysis, more subjects, and code/data, it becomes a refereed paper. For now: desk reject with a clear signal to strengthen and resubmit.","headline":"Promising integration of autofocus + CNN for field estimation, but the 'accurate' claim rests on unquantified visuals and an unvalidated basis order; needs a residual analysis against existing Skope data.","tokens_in":3734,"tokens_out":2086,"would_cite":false,"duration_ms":23308,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A convolutional network, ACE-Net, estimates the spatiotemporal B0 and eddy-current fields that corrupt spiral diffusion MRI directly from the blurred images themselves, using autofocus metrics and a compact basis, and reconstructs…","keywords":["spiral diffusion MRI","field imperfection estimation","autofocus metric","eddy currents","B0 inhomogeneity","deep learning","unrolled network","spherical harmonics"],"falsifier":"Acquire a spiral diffusion scan at b=3000 while deliberately inducing an eddy field with a fourth-order spherical-harmonic spatial component or a temporal variation faster than a third-order polynomial, measure the true field with a field-probe system, and run ACE-Net; if the estimated field misses the injected component and residual blur remains after correction, the compact-basis assumption is falsified.","tokens_in":2818,"feed_emoji":"🧲","tokens_out":14378,"duration_ms":115458,"temperature":0.7,"pith_summary":"This paper seeks to establish that a deep convolutional network can recover the magnetic-field imperfections that blur and distort high-b-value spiral diffusion MRI, directly from the images being corrected. The method, ACE-Net, feeds autofocus blur metrics and slice information into a network that outputs coefficients of a compact physical basis for the field: spatial spherical harmonics up to third order multiplied by temporal polynomials up to third order. The payoff is that scans no longer need lengthy calibration acquisitions or external field-probe measurements, which are time-consuming and can be invalidated by scanner heating or subject motion. On two volunteer scans at b=3000, the estimated fields closely matched ground-truth measurements, and the corrected DTI showed white-matter tracts that were not discernible before correction.","feed_headline":"Neural net fixes spiral MRI field errors without calibration scans","feed_subtitle":"ACE-Net combines autofocus blur metrics with a learned compact field model, so high-b diffusion scans correct themselves.","key_machinery":"The load-bearing mechanism is the autofocus blur metric combined with a compact basis representation. The autofocus metric measures how blurred an image is at a given off-resonance frequency offset; minimizing it locates the field error, but the metric is noisy and prone to local minima, so the network supplies a learned prior and the basis constrains the solution. The compact basis, $\\phi(\\mathbf{r},t) = \\sum_n \\phi_n(\\mathbf{r})\\delta_n(t)$, with third-order spatial spherical harmonics and third-order temporal polynomials, reduces the unknown field to a small set of coefficients, making the estimation tractable. The unrolled architecture adds a data-consistency step that re-encodes the estimated field into the forward model and updates the blurred images and autofocus metrics, counteracting hallucination.","core_discovery":"The paper's central claim is that ACE-Net can estimate spatiotemporal field imperfections from the corrupted spiral diffusion data alone, with no external calibration. Static B0 inhomogeneity is estimated first from the b=0 image, using autofocus metrics across a ±120 Hz range, narrowed to ±40 Hz by a rough CNN B0 estimate, to produce a refined B0 map. Spatiotemporal eddy-current and dynamic B0 fields are then estimated per diffusion direction as $\\phi(\\mathbf{r},t) = \\sum_n \\phi_n(\\mathbf{r})\\delta_n(t)$, with $\\phi_n$ spatial spherical harmonics and $\\delta_n$ temporal polynomials truncated at third order based on field-probe data. The network outputs the basis coefficients, and an unrolled version with a data-consistency block re-encodes the estimated field into the forward model to prevent hallucination. In the demonstrated application, the estimated fields matched the ground-truth field-probe and multi-echo GRE measurements, and incorporating them into reconstruction removed blur and made the white-matter tracts clearly visible.","pith_inferences":["Editorial inference: the third-order spatial/temporal basis is the most fragile component; fields with sharp structure near metal or the sinuses, or fast transient eddy behavior, would fall outside what the network can represent, so a learned or adaptive basis is the natural next step.","Editorial inference: because the autofocus metric is a generic sharpness signal, the same architecture could be applied to other phase-error sources that lack external references, such as motion-induced phase or chemical shift, provided a physical basis for those errors is available.","Editorial inference: a cross-platform generalization test would reveal whether ACE-Net learned scanner-specific eddy dynamics or a universal physical constraint; if estimates degrade on another scanner, the training distribution would need to cover more hardware.","Editorial inference: with only two volunteer test scans, a larger cohort study is needed to confirm that the estimates stay accurate when B0 drifts during long protocols and when subject motion changes the field during the acquisition."],"forward_implications":["High-b-value spiral diffusion protocols can drop dedicated calibration scans and field-probe measurements for B0 and eddy correction, since ACE-Net estimates the fields from the data already acquired.","Corrected reconstructions make white-matter fiber tracts visible in 22-direction DTI, so quantitative diffusion metrics in affected regions should become more reliable.","The unrolled data-consistency variant provides a check against network hallucination by requiring the estimated field to reproduce the measured k-space data.","Because the field is represented by a small coefficient set over physical bases, the method also produces a compact characterization of the scanner's eddy-current state during a scan.","The same autofocus-plus-network estimation principle is intended by the authors to extend to EPI and motion-robust reconstruction, where the same class of spatiotemporal field errors degrades image quality."],"supporting_citations":[{"why":"Supplies the expanded encoding model and field-probe-based correction for single-shot spiral diffusion that ACE-Net aims to replace.","marker":"(1)"},{"why":"Introduces the autofocus deblurring concept for non-2DFT MRI that underlies the blur metric.","marker":"(2)"},{"why":"Refines off-resonance maps via autofocusing for spiral imaging, providing the metric ACE-Net uses.","marker":"(3)"},{"why":"Establishes the physics-informed deep-learning approach for off-resonance correction that ACE-Net integrates with autofocus.","marker":"(4)"},{"why":"Characterizes time-varying eddy currents and supplies the compact spatiotemporal basis representation.","marker":"(5)"},{"why":"Provides the rough B0 estimate that restricts the autofocus search window on a voxel-wise basis.","marker":"(6)"},{"why":"Supplies clean diffusion datasets used to synthesize corrupted spiral training data for the spatiotemporal network.","marker":"(7)"}],"fun_headline_variants":["ACE-Net fixes spiral MRI artifacts without calibration scans","Autofocus deep learning estimates MRI field errors without calibration","Spiral MRI self-corrects via autofocus neural net","Field imperfections from data alone fix high-b spiral diffusion","No calibration needed: ACE-Net corrects spiral MRI field errors"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes that every field error it needs to fix can be written as a few smooth spatial patterns (spherical harmonics up to third order) changing along a few smooth time curves (polynomials up to third order); if a real scan contains sharper or faster variations, the network cannot represent them, so the correction would fail no matter how well the network was trained.","fun_headline_variants_meta":{"raw":{"variants":["ACE-Net fixes spiral MRI artifacts without calibration scans","Autofocus deep learning estimates MRI field errors without calibration","Spiral MRI self-corrects via autofocus neural net","Field imperfections from data alone fix high-b spiral diffusion","No calibration needed: ACE-Net corrects spiral MRI field errors"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000739,"raw_usage":{"total_tokens":3265,"prompt_tokens":875,"completion_tokens":2390,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":491,"completion_tokens_details":{"reasoning_tokens":2307}},"tokens_in":491,"tokens_out":2390,"duration_ms":16973,"temperature":1.0,"reasoning_tokens":2307,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:04:19.162850+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Acquire a spiral diffusion scan at b=3000 while deliberately inducing an eddy field with a fourth-order spherical-harmonic spatial component or a temporal variation faster than a third-order polynomial, measure the true field with a field-probe system, and run ACE-Net; if the estimated field misses the injected component and residual blur remains after correction, the compact-basis assumption is falsified.","supporting_citations":[],"review_version":1}