{"id":"6d0584c5-d877-4f01-b0c6-0a620c8e9618","arxiv_id":"2506.13293","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"SUSEP-Net is a simulation-trained dual-branch U-net with contrastive learning that separates QSM into paramagnetic and diamagnetic source maps, outperforming three existing methods in reported experiments.","lead":"This paper trains a dual-branch neural network, SUSEP-Net, to split MRI susceptibility maps into paramagnetic (iron-like) and diamagnetic (myelin-like) components using simulated training data and a contrastive learning loss. If the method holds up, it gives clinicians and researchers a faster, more consistent way to measure iron and myelin signals in brain disease, with public code that others can test.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Simulated benchmarks reuse the same APART-QSM labels and Eq. (1) forward model used to train SUSEP-Net, and the synthetic test lesions lie inside the training distribution; the claimed numerical advantage may measure self-consistency rather than independent accuracy.","rationale":"The reader's weakest assumption centers on the accuracy of APART-QSM labels and Eq. (1) as the true physical model. My concern is closely related but focuses on the evaluation protocol: the simulated test data are not independent of the training data because both are generated from the same APART-QSM reconstructions and the same forward model. Additionally, the synthetic lesions in the test set are drawn from the same ranges as the training lesions, making the test in-distribution for SUSEP-Net but out-of-distribution for the non-trained baselines. These issues do not invalidate the paper, but they substantially weaken the quantitative superiority claim; the phantom experiment and qualitative in vivo results are insufficient to fully compensate. The appropriate verdict remains CONDITIONAL, requiring independent evaluation against a ground truth not derived from APART-QSM/Eq. (1). I partially agree with the reader because we identify the same root cause (dependence on APART-QSM and Eq. (1)), but I emphasize benchmark fairness rather than label accuracy per se.","tokens_in":16152,"tokens_out":5841,"duration_ms":62716,"concrete_test":"Regenerate the simulated healthy and pathological test sets from an independent source of ground truth—e.g., a digital brain phantom with known χpos/χneg from a different model, or a Bloch-equation simulation using a constant A kernel rather than APART-QSM-derived A maps—and re-run SUSEP-Net, APART-QSM, χ-separation, χ-sepnet, and a χ-sepnet retrained on the SUSEP-Net simulation data. If SUSEP-Net's NRMSE/HFEN/XSIM advantage is not reproduced on this independent test set, the claimed superiority is not established. A minimal version: include the retrained χ-sepnet in the Fig. 3 simulated pathological-brain comparison; if it approaches SUSEP-Net's metrics, the reported gains are largely train/test distribution effects.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim (best NRMSE/HFEN/XSIM on simulated healthy and pathological brains, Sections 4.1-4.2) rests on test data generated by the exact pipeline of Fig. 2: ground-truth χpos/χneg and A(r) maps come from APART-QSM reconstructions of 32 subjects (Section 2.3), and inputs R2', local field, and QSM are synthesized via Eq. (1). SUSEP-Net is trained to invert this same generative process, so its simulation performance largely measures self-consistency. The comparison is further biased because the synthetic hemorrhage (1 ppm) and calcification (-0.2 ppm) in the test brain lie within the uniform ranges [0.4, 1.2] and [-0.3, -0.1] used to generate training lesions (Section 2.3), giving SUSEP-Net an in-distribution advantage over APART-QSM, χ-separation, and the original χ-sepnet, which were not trained on these synthetic lesions. The agarose phantom (Section 4.3) provides independent quantitative support, but only for homogeneous CaCO3/Fe3O4 solutions, not realistic brain tissue mixtures; the in vivo patient results (Section 4.4.2) are qualitative. If the simulation benchmarks are removed, the 'consistently improved results' claim has much weaker quantitative support.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes SUSEP-Net, a dual-branch U-net with two auxiliary encoders and a contrastive learning loss, trained on simulated data generated by Eq. (1) from APART-QSM-derived χpos, χneg, and A(r) maps, with synthetic hemorrhage and calcification lesions added. The method is evaluated against APART-QSM, χ-separation, and χ-sepnet on simulated healthy and pathological brains, an agarose gel phantom, and in vivo healthy and pathological subjects, and the authors report improved NRMSE/HFEN/XSIM metrics, better lesion contrast, and reduced artifacts.","tokens_in":16342,"tokens_out":4514,"duration_ms":50519,"significance":"If the reported simulation accuracy reflected independent ground truth, SUSEP-Net would be a meaningful advance for susceptibility source separation, and the contrastive-learning design is a plausible contribution. The paper has concrete strengths: the code and trained networks are released, an ablation study supports the contrastive loss, and the phantom experiment provides an independent test of generalization. However, the central quantitative claim is weakened because the simulated test data are generated by the same APART-QSM-based forward model used for training, so the numerical superiority over untrained methods partly measures self-consistency rather than accuracy against independent truth; the in vivo evidence is qualitative.","major_comments":[{"comment":"The simulation benchmark is self-referential: the test brains in §4.2 are synthesized with Eq. (1) from APART-QSM-derived χpos, χneg, and A(r) maps, which is exactly the training-data generation pipeline of Fig. 2. Consequently, SUSEP-Net's large NRMSE/HFEN advantages over APART-QSM, χ-separation, and χ-sepnet partly measure consistency with the teacher model and the generator, not independent accuracy. The concern is compounded by the test lesions: the hemorrhage value of 1 ppm and calcification of −0.2 ppm in §3.1(2) lie inside the training ranges [0.4, 1.2] and [−0.3, −0.1] given in §2.3, so the test lesions are in-distribution for SUSEP-Net. The paper itself concedes in Section 5 that training labels still rely on APART-QSM. Please add a validation setup that does not reuse the training generator, or substantially qualify the simulation-based superiority claims.","section":"§2.3, §4.2, Fig. 2"},{"comment":"The agarose gel phantom is a genuinely independent test, but it uses only homogeneous CaCO3 and Fe3O4 solutions in simple cylinders. It does not exercise realistic sub-voxel coexistence of paramagnetic and diamagnetic sources in brain tissue, so it cannot by itself support the claim of improved accuracy in pathological brains. The paper should state this limitation explicitly and should not present the phantom results as sufficient evidence for the general in vivo claim.","section":"§4.3"},{"comment":"The in vivo pathological evaluations are qualitative and have no independent ground truth. The abstract's claim that SUSEP-Net shows \"improved high-intensity hemorrhage and calcification lesion contrasts, and reduced artifacts\" is thus stronger than the evidence provided; the ROI comparisons in §4.4.1 show similarity to APART-QSM rather than superiority, and the patient results rely on visual inspection. Please temper the wording or add blinded reader scoring, lesion delineation metrics, or another quantitative outcome.","section":"§4.4.2 and Abstract"}],"minor_comments":[{"comment":"The text refers to \"Table 3\" for the 10 simulated healthy brain results, but the table shown is numbered Table 2.","section":"§4.2"},{"comment":"In Eq. (6), the notation \"χ&#'∗ and χ&#'∗\" uses the same subscript for both reconstructions; the second should be χneg∗ (i.e., χ()*∗).","section":"Eq. (6)"},{"comment":"The last term in Eq. (8) is written as R1!, which appears to be a typo for R2′; please correct this to match the notation used throughout the paper.","section":"Eq. (8)"},{"comment":"The sentence referring to \"AFTER-QSM\" appears to mean \"APART-QSM\" based on the context and the methods compared in Fig. 5.","section":"§4.3"},{"comment":"There is a typo in the sentence introducing the contrastive loss: \"coxntrastive\" should be \"contrastive.\"","section":"§2.2.2"},{"comment":"The figure caption and text describe the calcification lesion as 0.2 ppm, but §3.1 specifies −0.2 ppm; this sign inconsistency should be fixed.","section":"Fig. 3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within scope and the method is potentially useful, but the quantitative evaluation relies heavily on a simulation pipeline shared with the training data. I would not reject the paper; the phantom and in vivo results show promise, but the authors should either add independent validation or clearly downgrade the simulation-based claims before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a credible engineering contribution to susceptibility source separation, but the headline simulation numbers partly reflect self-consistency with the APART-QSM teacher. The more convincing evidence is the agarose phantom and the retrained χ-sepnet baseline.\n\nWhat is new: a dual-branch U-net with two QSM-derived guidance encoders and a branch-wise contrastive loss, trained with a simulation-supervised recipe that synthesizes pathological lesions. The code and trained network are public. The authors also retrain χ-sepnet on their own synthetic data, which is the right control for isolating the training strategy from the architecture. The ablation shows the contrastive loss is worth about 3% NRMSE on simulated brains, and the phantom linearity (R² = 0.96 for CaCO3, 0.97 for Fe3O4) is solid.\n\nWhere it is soft: the simulated evaluation reuses the same forward model (Eq. 1) and APART-QSM labels that generated the training data, and the simulated test lesions (1 ppm, -0.2 ppm) fall inside the training ranges ([0.4,1.2] and [-0.3,-0.1]). So the large NRMSE gains over APART-QSM and χ-separation on simulated brains partly measure how well the network inverts its own training generator, not independent accuracy. The Discussion honestly concedes the APART-QSM dependence, but the abstract's 'consistently improved results' overreaches what the simulation can support. The in vivo pathological comparisons are qualitative, and there are mechanical errors: a reference to a nonexistent Table 3 in Section 4.2 (the numbers are in Table 2), a copied typo in Eq. (6), and at least one typo in the definition of the χpos/χneg reconstructions.\n\nThe phantom experiment partially compensates. SUSEP-Net's single-source vs mixed-source regression is nearly as good as APART-QSM, and retrained χ-sepnet also improves, which indicates the synthetic training data, not just the architecture, is doing useful work.\n\nBottom line: the method is real, the evaluation has a self-consistency flaw that is acknowledged but not fully resolved by the current benchmarks, and the paper needs revision before publication. The right next step is a serious referee: ask for test lesions outside the training range, for a decomposition of the simulation metric into self-consistency versus true generalization, and for fixing the reporting errors.\n\nFor whom: QSM researchers and anyone building a baseline for paramagnetic/diamagnetic separation. I'd cite it with the caveat noted.\n\nRecommendation: send to peer review, not desk reject.","headline":"Solid engineering contribution to susceptibility source separation, but the headline simulation numbers partly reflect self-consistency with the APART-QSM teacher; the phantom and retrained-baseline comparisons are the more convincing evidence.","tokens_in":17014,"tokens_out":3401,"would_cite":true,"duration_ms":32980,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A dual-branch U-net trained purely on simulated forward-model data separates paramagnetic and diamagnetic brain susceptibility maps with lower error than APART-QSM, χ-separation, and χ-sepnet.","keywords":["Susceptibility source separation","Quantitative susceptibility mapping (QSM)","Paramagnetic and diamagnetic separation","Simulation-supervised training","Contrastive learning","Dual-branch U-net","R2' relaxometry","Sub-voxel QSM"],"falsifier":"Acquire an agarose phantom with known concentrations of mixed paramagnetic and diamagnetic compounds (e.g., Fe3O4 and CaCO3) and compare SUSEP-Net's separated χpos and χneg values against gravimetrically known single-source and mixed-source values; systematic deviations in the mixture cylinders that APART-QSM does not show would indicate label bias. Alternatively, measure R2' and local field on a phantom using an A(r) map different from the APART-QSM-derived one and check whether the network's predictions follow the physical model.","tokens_in":15847,"feed_emoji":"🧠","tokens_out":6530,"duration_ms":64386,"temperature":0.7,"pith_summary":"The paper sets out to solve susceptibility source separation: in a single brain voxel, paramagnetic iron and diamagnetic myelin can cancel in standard quantitative susceptibility maps, hiding the contribution of each. SUSEP-Net is a dual-branch deep network that takes R2', local field, and QSM as inputs and outputs separate paramagnetic (χpos) and diamagnetic (χneg) maps, trained entirely on simulated data generated from the forward model rather than on large multi-orientation in vivo acquisitions. The authors report that it beats APART-QSM, χ-separation, and χ-sepnet on simulated brains, on an agarose phantom, and on patient brains with CO poisoning and MOGAD, with lower NRMSE and sharper lesion contrast. A contrastive loss keeps the two output branches disentangled by aligning each branch's latent features with a branch-specific guidance feature derived from QSM. If the results hold, the method would make iron and myelin mapping more practical and more reliable for clinical MRI, without needing special training data.","feed_headline":"Simulation-trained network separates brain iron and myelin better","feed_subtitle":"SUSEP-Net cuts diamagnetic-map error to 9.66 percent on simulated brains and sharpens lesion contrast in patients.","key_machinery":"The load-bearing object is the complex forward model of Eq. (1): R2'(r) + i ΔB_local(r) = A(r)·(χpos(r) − χneg(r)) + i D(r) ⊗ (χpos(r) + χneg(r)), where A(r) is a voxel-specific magnitude decay kernel relating R2' to absolute susceptibility and D(r) is the unit dipole kernel. This equation both defines what the network must invert and generates the simulation-supervised training data. The network itself is a dual-branch U-net: one shared encoder fed by the concatenated inputs, two decoders producing χpos and χneg, and two additional guidance encoders fed by QSM. The contrastive loss ties Guide_pos to F_pos and Guide_neg to F_neg while repelling cross-branch pairs. The claim is that this architecture, with pure-synthetic high-intensity lesions added to the simulated patches, is what lets the network recover sub-voxel paramagnetic and diamagnetic content without in vivo ground truth.","core_discovery":"SUSEP-Net's central claim is that a simulation-supervised dual-branch U-net with contrastive feature constraints can separate susceptibility sources more accurately than the methods used to generate its own labels. The network learns the mapping from R2', local field, and QSM to χpos and χneg by training on 3024 patches synthesized from APART-QSM reconstructions of 32 brains plus purely synthetic hemorrhage and calcification lesions. On the simulated pathological brain it reports χneg NRMSE 9.66% versus 24.72% (APART-QSM), 31.97% (χ-separation), and 15.85% (χ-sepnet), and the closest lesion values (hemorrhage 0.97 ± 0.018 ppm, calcification 0.19 ± 0.017 ppm). On 10 simulated healthy brains it reports average χpos NRMSE 5.08% against 10.27–16.41% for the competitors. The authors interpret this as evidence that the network has learned the underlying physics of Eq. (1) rather than simply imitating a traditional algorithm.","pith_inferences":["Because the labels are generated by APART-QSM, SUSEP-Net's superior phantom and patient contrast likely reflects a combination of learned physics and regularization or denoising, not proof that it is closer to true sub-voxel concentrations wherever APART-QSM itself is biased.","The same simulation-supervision recipe should transfer to other ill-posed two-channel inverse problems with opposite-sign sources, such as susceptibility and R2* mapping or fat-water separation, wherever a forward model like Eq. (1) is available.","A direct test of whether SUSEP-Net truly learns physics would be to retrain it with labels from a different algorithm, such as χ-separation instead of APART-QSM: if the outputs change substantially, the network is partly imitating the label generator rather than the underlying forward model.","The method may also be adapted to estimate A(r) jointly instead of taking it as a fixed input, which would remove one of the main sources of label dependence."],"forward_implications":["On simulated pathological brains, SUSEP-Net cuts χneg NRMSE to 9.66% versus 15.85% for the best prior deep network and recovers hemorrhage and calcification values close to ground truth (0.97 ppm and 0.19 ppm against inserted 1 ppm and 0.2 ppm).","Because training patches can be synthesized from a handful of subjects plus geometric lesions, the method avoids the expensive multi-orientation acquisitions that χ-sepnet requires.","The agarose phantom results (R2 = 0.96 for CaCO3 and 0.97 for Fe3O4 linearity) suggest the network generalizes to data with different acquisition hardware and contrast sources than its training set.","On CO poisoning and MOGAD patients, SUSEP-Net shows fewer reconstruction artifacts than APART-QSM and χ-separation and renders lesions more visible, which could improve clinical delineation of iron-related and myelin-related damage.","If the method is adopted, standard single-orientation 3T mGRE scans could yield sub-voxel iron/myelin separation, making the technique feasible for routine clinical protocols."],"supporting_citations":[{"why":"Supplies the APART-QSM reconstruction pipeline that generates the training labels (χpos, χneg, A maps) and serves as the main comparison method.","marker":"(Li et al., 2023)"},{"why":"Defines the χ-separation formulation and Eq. (1) that the simulation-supervised training is based on, and is one of the compared methods.","marker":"(Shin et al., 2021)"},{"why":"Provides the deep-learning baseline χ-sepnet, including the U-net-style network the dual-branch design extends, and the in vivo label approach the paper argues against.","marker":"(Kim et al., 2025)"},{"why":"Produces the iQFM local-field reconstruction used to compute in vivo inputs and the iQSM-style simulation protocol that generates synthetic training patches.","marker":"(Gao et al., 2022)"},{"why":"Supplies iQSM+ for QSM reconstruction of the in vivo evaluation data and the latent-feature editing framework behind the fusion module.","marker":"(Gao et al., 2024)"},{"why":"Contributes the U-net backbone that the dual-branch encoder-decoder architecture is built from.","marker":"(Ronneberger et al., 2015)"},{"why":"Provides the contrastive representation learning formulation adapted for the same-branch and cross-branch feature constraint.","marker":"(Chen et al., 2020)"}],"fun_headline_variants":["Simulation-trained dual-branch net reduces QSM separation errors","SUSEP-Net outperforms three established QSM separation tools","Contrastive constraints boost susceptibility source separation accuracy","Simulation-supervised U-net sharpens brain iron and myelin maps","AI net separates magnetic brain sources with 9.66% error"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The training labels (χpos, χneg, and the A(r) decay kernel) come from APART-QSM, and Eq. (1) is assumed to describe exactly how R2' and local field are produced from those labels; if APART-QSM mis-assigns susceptibility or A(r) is wrong, SUSEP-Net inherits the bias and the simulated benchmarks cannot reveal it, since they are generated from the same labels.","fun_headline_variants_meta":{"raw":{"variants":["Simulation-trained dual-branch net reduces QSM separation errors","SUSEP-Net outperforms three established QSM separation tools","Contrastive constraints boost susceptibility source separation accuracy","Simulation-supervised U-net sharpens brain iron and myelin maps","AI net separates magnetic brain sources with 9.66% error"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000209,"raw_usage":{"total_tokens":1451,"prompt_tokens":1032,"completion_tokens":419,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":648,"completion_tokens_details":{"reasoning_tokens":342}},"tokens_in":648,"tokens_out":419,"duration_ms":5063,"temperature":1.0,"reasoning_tokens":342,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:36:32.995269+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Acquire an agarose phantom with known concentrations of mixed paramagnetic and diamagnetic compounds (e.g., Fe3O4 and CaCO3) and compare SUSEP-Net's separated χpos and χneg values against gravimetrically known single-source and mixed-source values; systematic deviations in the mixture cylinders that APART-QSM does not show would indicate label bias. Alternatively, measure R2' and local field on a phantom using an A(r) map different from the APART-QSM-derived one and check whether the network's predictions follow the physical model.","supporting_citations":[],"review_version":1}