Pith. sign in

REVIEW 2 major objections 5 minor 8 references

Deep learning brain conductivity mapping using a patch-based 3D U-net

T0 review · 2 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Simulation-trained networks fail on real brain scans

desk verdict Solid extension of DLEPT with honest limits: the simulation story is clean, but the in-vivo evaluation leans on HHEPT for both labels and reference, so the central artifact-reduction claim is only partly established. read the letter →

arxiv 1908.04118 v1 pith:Q23IR73F submitted 2019-08-12 physics.med-ph eess.IVq-bio.NC

classification physics.med-pheess.IVq-bio.NC
keywords deeplearningelectricalpropertiestomographybrainconductivity3DU-netpatch-basedreconstructiontransceivephasein-vivogeneralizationMRI
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper studies how far a deep-learning network can go in turning MRI radio-frequency phase maps into brain conductivity maps, a technique called deep-learning electrical properties tomography. It reports that a 3D patch-based U-net trained on electromagnetic simulations reconstructs simulated brains with high fidelity, but when the same network is applied to real volunteer and patient scans it produces visible artifacts, especially in cerebrospinal fluid. Adding homogeneous Gaussian noise to simulated training data does not fix the problem. The paper then shows that training the same architecture on real in-vivo phase maps, using conventional Helmholtz-based EPT reconstructions as conductivity labels, sharply reduces those artifacts. If the authors are right, deep-learning EPT is clinically promising but only when the training data closely matches the target anatomy, acquisition noise, and tissue types.

What carries the argument

The load-bearing machinery is two-part. The first part is the Helmholtz-based phase-only conductivity reconstruction, HHEPT, which computes conductivity as $\sigma = \Delta\phi^{+} / (\mu_0 \omega)$ from the Laplacian of the B1 transceive phase under the transceive-phase assumption; the paper uses HHEPT maps as both the in-vivo training labels and the evaluation reference. The second part is a 3D patch-based U-net with two downsampling steps, three convolutions between each step, residual units, batch normalization, and ReLU activations, trained on 24x24x24 phase patches with mean subtraction and overlapping-patch averaging. The network learns a surrogate for the inverse mapping from local B1 phase patches to conductivity, and the key experimental lever is which data supplies its training patches: pure simulations with added Gaussian noise, or real in-vivo phase maps carrying genuine acquisition artifacts.

What would settle it

A concrete check is to run the same simulation-trained network on a physical phantom or numerical phantom with known ground-truth conductivity and realistic in-vivo-like phase artifacts (motion, CSF pulsation). If artifacts still appear despite matching the phantom's anatomy and artifacts, the claimed role of in-vivo training data would be weakened; conversely, simulating those artifacts and seeing artifacts disappear would confirm it.

Watch

Extended reading notes

Core claim

The central claim is that the ability of DLEPT networks to generalize is governed less by network architecture than by how faithfully the training data reproduces the measurement conditions of the target data. Simulation-trained networks achieve correlations up to 0.991 on simulated test data, but applying them to in-vivo data produces artifact-heavy reconstructions, and training them with added Gaussian noise (even at the empirically optimal SNR=200) degrades generalization to geometries not present in training. Networks trained on in-vivo phase data, with conductivity labels provided by HHEPT reconstructions, reconstruct both volunteers and patients with far fewer artifacts and higher correlation with the HHEPT reference. The authors conclude that in-vivo training works because the training phase maps already contain acquisition-related artifacts such as head motion and CSF pulsation, and that simulation-trained networks will need realistic simulated artifacts, not just homogeneous noise, to become clinically transferable.

Load-bearing premise

The load-bearing premise is that the HHEPT conductivity maps used as in-vivo training labels and as the evaluation reference are accurate enough to stand in for true tissue conductivity; if HHEPT is systematically wrong at tissue boundaries or under CSF pulsation, the network learns and is scored by those same errors.

Editorial extensions

If this is right

  • Performance on simulated test data does not predict performance on real scans: correlations above 0.99 drop to visibly artifact-laden reconstructions when the same network is applied in vivo.
  • Training on in-vivo data, with HHEPT maps as labels, suppresses acquisition-related artifacts in the network output because the training phase maps already contain those artifacts.
  • Networks trained on one in-vivo population (volunteers) do not generalize to another (patients), and combining both populations gives the best results, so training sets must span the intended target group.
  • Adding homogeneous Gaussian noise during simulation training does not bridge the simulation-to-clinic gap; realistic artifact simulation is the missing ingredient.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If HHEPT labels carry boundary or pulsation errors, then the reported artifact reduction may partly be the network learning to reproduce HHEPT's own artifacts; replacing labels with more accurate references (segmented literature values, ex-vivo measurements, or forward-model fits) would tell how much of the gain is real conductivity recovery.
  • The stronger dependence on anatomy and artifacts suggests a practical deployment path of fine-tuning or domain adaptation on local scanner data rather than relying on one universal model; this follows from the cross-population performance drop but is not tested in the paper.
  • A testable extension is to simulate CSF pulsation and head motion as phase perturbations in the synthetic training data; if the artifact reduction seen with in-vivo training can be reproduced synthetically, simulation-based DLEPT could be made transferable without acquiring large labeled patient cohorts.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper investigates a 3D patch-based U-net for brain conductivity mapping from B1 transceive phase data (DLEPT). The authors train networks on electromagnetic simulations (Duke/Ella models) with and without added Gaussian noise and test on simulated and in-vivo data (healthy volunteers and patients with brain lesions). They also train networks on in-vivo data using conductivity labels from conventional Helmholtz-based EPT (HHEPT) and evaluate by correlation with HHEPT references. The main findings are: (1) networks trained on simulations reconstruct simulated data well but produce artifacts when applied to in-vivo data, especially in CSF; (2) training on in-vivo data with HHEPT labels reduces these artifacts; and (3) generalization between volunteer and patient datasets is limited, suggesting sensitivity to geometry and acquisition-related artifacts.

Significance. If the central claims hold, the paper would be a useful contribution to EPT by quantifying the domain-shift problem in deep-learning-based conductivity mapping and by showing that training on in-vivo-like data improves robustness to acquisition artifacts. The study has several strengths: it uses realistic electromagnetic simulations with known ground truth, it applies cross-validation and excludes validation data from training, it includes pathology, and it explicitly discusses the limitations of HHEPT labels. The negative result that sim-trained networks do not transfer to in-vivo data is valuable for the community. However, the in-vivo evaluation is partly circular because the same HHEPT estimator provides both the training labels and the evaluation reference, so the reported artifact reduction is not an independent measure of true conductivity accuracy. This limits the support for the paper's central conclusion.

major comments (2)
  1. [Methods/Experimental outline and Results (Tab. 2, Fig. 4, Fig. 5)] The in-vivo networks are trained on conductivity labels from HHEPT, computed from Eq. (1) with the transceive-phase assumption and median-filtered method of reference [4], and are then evaluated by correlation with the same HHEPT reference. This makes the reported artifact reduction partly a measure of agreement with the training-target estimator. Since HHEPT has known systematic errors at tissue boundaries and is affected by CSF pulsation (as the authors note in the Discussion and reference [26]), the central claim that "training with realistic phase data and conductivity labels from conventional EPT allows for severely reducing these artifacts" is not independently established. The Discussion does acknowledge this limitation, but the Abstract and Conclusion still present the artifact reduction as a robust finding. To support the claim, the authors should either (a) evaluate on simulated data that mimic in-vivo phase artifacts and have known ground-truth conductivity, (b) use an independent reference (e.g., tissue-segmentation-based literature conductivity values or ex-vivo measurements), or (c) explicitly reframe the conclusion as "reduces deviation from the HHEPT estimator" rather than "reduces artifacts" in an absolute sense.
  2. [Results (Tab. 2, Tab. 3)] The quantitative evaluation reports only average correlation coefficients without confidence intervals, per-fold ranges, or statistical significance tests. With 14 patients and 18 volunteers, the observed differences between NWvol, NWpat, and NWvol+pat could be within sampling variability. The authors should provide per-subject correlations, standard deviations or bootstrap confidence intervals, and, where appropriate, paired significance tests. This is particularly important for the cross-validation comparisons in Tab. 3, where the claim that accuracy saturates with half the training data rests on average correlations alone.
minor comments (5)
  1. [Abstract and Methods/Experimental outline] The abstract mentions "different levels of homogeneous Gaussian noise introduced in training and testing," but the methods describe only specific SNR values (100 and 200). Please clarify which SNR values were used for each network and why the level was changed from 100 to 200.
  2. [Fig. 1] The network architecture figure would be easier to interpret with a legend or caption listing the patch size, number of channels per layer, and the locations of residual units and batch normalization, as these details are described only in the text.
  3. [Theory, Eq. (1)] The notation Δφ+ is introduced without a formal definition of the Laplacian convention or the transceive-phase assumption. Adding a sentence defining these terms would improve reproducibility.
  4. [Results, Tab. 1] Tab. 1 lists the lesions included in the patient dataset, but the table content is not described in the text. Please refer to it explicitly and state how lesion presence might affect the reconstruction evaluation.
  5. [Discussion] The comparison between DSpat and DSvol uses sdR as a proxy for brain-shape differences, but the volunteer and patient data were acquired at different sites and likely with different populations. The observed generalization gap could be confounded by acquisition-site or sequence differences. A sentence acknowledging this confound and suggesting a matched-site study would strengthen the interpretation.

Circularity Check

1 steps flagged · score 6.0 of 10

In-vivo evaluation is partly circular: training labels and the quantitative reference are both HHEPT (Eq. 1), so the reported artifact reduction partly measures regression to the training target rather than true conductivity recovery.

  1. fitted input called prediction [Abstract; Methods (Experimental outline); Eq. (1)]
    "Secondly, to investigate potential robustness towards systematical differences between simulated and measured phase maps, in-vivo data with conductivity labels from conventional EPT is used for training. ... Results are evaluated quantitatively by calculating the correlation of the reconstructed conductivity map with the HHEPT reference."

    The same HHEPT estimator (Eq. 1, with the transceive-phase assumption and the median-filtered solution of [4]) supplies both the supervised training labels for the in-vivo networks and the quantitative evaluation reference. A network trained to regress HHEPT maps is therefore scored by correlation with its own training target; high correlation on held-out scans mostly confirms that the phase-to-HHEPT mapping was learned, not that true conductivity was recovered. The paper concedes this limitation: 'This is unfortunately not guaranteed for in-vivo reconstructions using networks trained with HHEPT, given the intrinsic inaccuracies of this technique.' The simulation experiments with known ground truth remain independent, which is why the circularity is partial rather than total.

full rationale

The simulation arm of the paper is self-contained: networks trained on simulated B1+ maps with known ground-truth conductivity are tested on simulated data with known conductivity, so those correlations (e.g., 0.991/0.947 for Duke0) are genuine and do not reduce to the training target. The circularity is confined to the in-vivo arm. The in-vivo networks are trained on HHEPT conductivity labels (Eq. 1 with transceive-phase assumption and the median-filtered numerical method of ref. [4]) and then evaluated by correlation of their reconstructions with HHEPT references generated by the same estimator. Since the training target and the evaluation metric are the same estimator, the reported 'severe reduction of artifacts' for NWvol, NWpat, and NWvol+pat partly reflects the network learning to reproduce the HHEPT transform, including its smoothing and boundary inaccuracies, rather than independently recovering true conductivity. The authors themselves acknowledge that accurate in-vivo labels are not guaranteed: 'This is unfortunately not guaranteed for in-vivo reconstructions using networks trained with HHEPT, given the intrinsic inaccuracies of this technique.' Visual inspection of artifact reduction provides some qualitative independent evidence, and the simulation results show DLEPT can recover known conductivity from simulated phase; therefore the circularity is partial, not total.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The paper contributes a trained surrogate, not a physical derivation. The load-bearing inputs are: (1) HHEPT maps from prior work [4] used as both training labels and evaluation reference for in-vivo data; (2) FDTD simulations of Duke and Ella models as proxies for real in-vivo phase; and (3) empirically selected network hyperparameters. No new physical entity is introduced, but the absence of an independent conductivity ground truth means the in-vivo results depend on the accuracy of HHEPT.

free parameters (4)
  • Training noise level (SNR) = 200
    Chosen by empirical investigation rather than from a physical model; it affects whether simulation-trained networks transfer to in-vivo data. Appears in Methods, Experimental outline.
  • Patch size = 24x24x24
    Empirically selected; patch size controls the local context available to the network and affects generalization. Methods, Network Architecture and Preprocessing.
  • Number of downsampling steps = 2
    Empirical choice for U-net depth; controls receptive field and model capacity. Methods, Network Architecture.
  • Number of convolutions between down- and upsampling = 3
    Inspired by V-net but empirically selected; impacts network capacity and training behavior. Methods, Network Architecture.
assumptions (4)
  • domain assumption Phase-based Helmholtz approximation sigma = Delta phi / (mu0 omega) with the transceive-phase assumption
    Used to define HHEPT reference labels and the evaluation metric, Eq. (1) and Theory section; known to be inaccurate at boundaries and under CSF pulsation.
  • domain assumption HHEPT reconstructions, including the median filter of [4], provide adequate in-vivo conductivity labels
    All in-vivo training and evaluation depend on HHEPT labels from [4]; the paper acknowledges intrinsic inaccuracies of HHEPT in the Discussion.
  • domain assumption Simulated FDTD phase maps of Duke and Ella brain models capture in-vivo phase distributions up to additive Gaussian noise
    Central to transfer from simulation-trained networks; the paper's results show this assumption fails for in-vivo data, so it is load-bearing and partially contradicted.
  • ad hoc to paper A supervised U-net trained on mean-subtracted patches can learn the inverse mapping from phase to conductivity
    No proof or unsupervised alternative is given; the method is a surrogate model fitted to training data, stated in the Theory section.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Deep learning brain conductivity mapping using a patch-based 3D U-net." pith.science (2026). https://pith.science/paper/Q23IR73F

@misc{pith2026190804118,
  author       = {Pith},
  title        = {Pith review of: Deep learning brain conductivity mapping using a patch-based 3D U-net},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Q23IR73F}},
  note         = {Machine review of arXiv:1908.04118}
}
read the original abstract

Purpose: To investigate deep learning electrical properties tomography (EPT) for application on different simulated and in-vivo datasets including pathologies for obtaining quantitative brain conductivity maps. Methods: 3D patch-based convolutional neural networks were trained to predict conductivity maps from B1 transceive phase data. To compare the performance of DLEPT networks on different datasets, three datasets were used throughout this work, one from simulations and two from in-vivo measurements from healthy volunteers and cancer patients, respectively. At first, networks trained on simulations are tested on all datasets with different levels of homogeneous Gaussian noise introduced in training and testing. Secondly, to investigate potential robustness towards systematical differences between simulated and measured phase maps, in-vivo data with conductivity labels from conventional EPT is used for training. Results: High quality of reconstructions from networks trained on simulations with and without noise confirms the potential of deep learning for EPT. However, artifact encumbered results in this work uncover challenges in application of DLEPT to in-vivo data. Training DLEPT networks on conductivity labels from conventional EPT improves quality of results. This is argued to be caused by robustness to artifacts from image acquisition. Conclusions: Networks trained on simulations with added homogeneous Gaussian noise yield reconstruction artifacts when applied to in-vivo data. Training with realistic phase data and conductivity labels from conventional EPT allows for severely reducing these artifacts.

Figures

Figures reproduced from arXiv: 1908.04118 by the authors.

Figure 1
Figure 1. Network architecture employed in this work [PITH_FULL_IMAGE:figures/full_fig_p012_1.png] view at source ↗
Figure 2
Figure 2. Example slices of conductivity reconstructions of Duke0 and Ella0 from [PITH_FULL_IMAGE:figures/full_fig_p013_2.png] view at source ↗
Figure 3
Figure 3. Example slices of reconstructions of volunteer A and patient A by N [PITH_FULL_IMAGE:figures/full_fig_p014_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Example slices of conductivity reconstructions of volunteer A and patient B from [PITH_FULL_IMAGE:figures/full_fig_p014_4.png]
Figure 5
Figure 5. Figure 5: Example slices of conductivity reconstruction of patient A by N [PITH_FULL_IMAGE:figures/full_fig_p015_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

8 extracted references · 8 canonical work pages

  1. [4]

    University Medical Centre Utrecht, Imaging Division, Department of Radiotherapy Heidelberglaan 100 Utrecht, NL 3584CX

  2. [1]

    Philips Research Hamburg Roentgenstrasse 24 Hamburg, Hamburg, DE 22335

  3. [2]

    University Medical Center Utrecht, Division of Imaging & Oncology, Department of Radiotherapy Heidelberglaan 100 Utrecht, NL 3582JX

  4. [3]

    University Medical Centre Utrecht, Centre of Image Sciences, Computational Imaging Group Heidelberglaan 100 Utrecht, NL 3584CX +31 88 75 506 66

  5. [5]

    Hokkaido University Hospital, Department of Diagnostic and Interventional Radiology Kita14, Nishi5, Kita-Ku Sapporo, Hokkaido, JP 060-8648

  6. [6]

    Methods: 3D patch-based convolutional neural networks were trained to predict conductivity maps from B1 transceive phase data

    Hokkaido University, Global Station for Quantum Medical Science and Engineering, Global Institution for Collaborative Research and Education 5 Chome Kita 8 Jonishi, Kita Ward Sapporo, Hokkaido, JP 060-0808 Abstract: Purpose: To investigate deep learning electrical properties tomography (EPT) for application on different simulated and in-vivo datasets incl...

  7. [7]

    Deep learning for undersampled MRI reconstruction

    were scanned at Hokkaido University (Japan) yielding DSpat, and 18 healthy volunteers (mean age 44 +/- 7 yrs) were scanned at Philips Research Hamburg (Germany) yielding DSvol. For all subjects, a bSSFP sequence (TR/TE=3.4/1.7ms, voxel size=1x1x1mm, flip angle=25°, 2 averages, scan duration 3:40 minutes) was acquired as suggested by [ 25]. HHEPT reconstru...

  8. [18]

    Opening a new window on MR-based Electrical Properties Tomography with deep learning

    Mandija S, Meliadò E, Huttinga N, Luijten P, van den Berg CAT, Opening a new window on MR-based Electrical Properties Tomography with deep learning, arXiv:1804.00016. 2018 Mar. [19]: Hampe N, Herrmann M, Amthor T, Findeklee C, Doneva M, Katscher U, Dictionary -based Electric Properties Tomography, Magn Reson Med. 2019;81:342-349. [20]: Isola P, Zhu J-Y, Z...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.