REVIEW 1 major objections 1 minor 40 references
Dc-EEMF: Pushing depth-of-field limit of photoacoustic microscopy via decision-level constrained learning
T0 review · 1 major / 1 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read By fusing two single-focus photoacoustic microscopy scans with a lightweight learned network, Dc-EEMF reports extending the imaging depth of field to 560 µm while preserving lateral resolution.
desk verdict A credible engineering paper whose two central quantitative claims are not actually measured; send to review with requests for direct resolution and DoF measurements. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the artifact-resistant channel-wise spatial frequency fusion rule inside a lightweight siamese network. For every channel of the paired feature maps, the rule computes row and column gradient magnitudes in an 11×11 window, forms a decision tensor by comparing the two source frequencies, and selects the higher-activity feature at each position, which suppresses misalignment and boundary artifacts. Complementary to this, the U-Net-based decision-level focus property perceptual (dFPP) loss maps the fused image and each source to focus property maps and penalizes mismatch with ground-truth focus regions, giving the network the focus-boundary accuracy of decision-map methods while keeping the information-preservation behavior of end-to-end methods.
What would settle it
Image a phantom containing a USAF-style resolution target extending over 560 µm in depth, record source images at Z=0 µm and Z=500 µm, fuse them with Dc-EEMF, and determine the maximum resolvable spatial frequency at both focal planes and at intermediate depths. If the fused image cannot resolve at Z=0 and Z=500 µm the features each source individually resolves at its own focus, or if vessel-density measurements on a deliberately misaligned in vivo pair show double edges, the central claim fails.
Extended reading notes
Core claim
The central discovery claimed is that a decision-level constrained end-to-end network can fuse two OR-PAM source images, one focused at Z=0 µm and one at Z=500 µm, into a single image with a 560 µm depth of field and acceptable transverse resolution. The network extracts paired features with inverted residual dual-attention modules, fuses them by comparing channel-wise spatial frequencies computed over 11×11 windows, and reconstructs the fused image with a two-layer head. A U-Net-based decision-level focus property perceptual loss is trained to supply focus/defocus region constraints, and the total loss combines this with perceptual, SSIM, and focal frequency terms. Across simulations, leaf and liver-tissue experiments, and in vivo mouse cerebral vasculature, the paper reports top-ranked fusion quality metrics and statistically significant increases in vessel and junction density in fused versus single-focus images.
Load-bearing premise
The load-bearing premise is that the two source images are well enough co-registered that the 11×11 windowed channel-wise spatial frequency rule can absorb respiration- and heartbeat-induced misalignment, since the method has no explicit registration step.
Editorial extensions
If this is right
- An OR-PAM system can extend its DoF to 560 µm by post-processing two consecutive raster scans, with no change to the optical or acoustic hardware.
- Vascular parameters such as vessel density and junction density become quantifiable across a much thicker tissue slab than a single Gaussian focus allows.
- Acquisition stays fast: only two or three focal layers are scanned, compatible with MHz-rate lasers, and no axial z-stack is needed.
- The same trained model transfers to other microscopy modalities, as demonstrated on optical and reflectance confocal microscopy images.
- At 0.24 million parameters and about 29.7 ms per 128×128 pair, the fusion step is light enough for near-real-time clinical use.
Reading between the lines
- A natural untested extension is to add an explicit registration step before fusion; the authors note that respiration and cardiac motion cause incomplete registration, so deformable alignment could extend Dc-EEMF to freer-moving animals.
- The 11×11 window size effectively sets a tolerance for residual motion; a future study could measure this tolerance directly and adapt window size to scan speed or motion level.
- The decision-level focus property constraint is modality-agnostic, so the same loss could be applied to multimodal fusion (e.g., photoacoustic plus optical or pathological slide images), an idea the paper gestures toward but does not test.
- Combining Dc-EEMF with a hardware extended-DoF excitation such as a Bessel beam, rather than only a Gaussian focus, might push the fused depth range beyond 560 µm; this combination is not explored in the paper.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Dc-EEMF, a siamese convolutional network for multi-focus image fusion intended to extend the depth of field (DoF) of optical-resolution photoacoustic microscopy (OR-PAM). The network combines an end-to-end fusion architecture with a U-Net-based decision-level focus property perceptual loss and a channel-wise spatial frequency feature fusion rule. The authors train on simulated PAM data generated by Gaussian blurring and validate on simulation, leaf phantom, in vitro liver tissue, in vivo mouse brain, and additional optical/confocal microscopy datasets. They report favorable quantitative results against thirteen baselines and claim, in the Discussion, that the method achieves a 560 µm DoF by fusing two source images at 0 and 500 µm while preserving acceptable transverse resolution.
Significance. If the central claim is supported, the work would be practically valuable: it offers a purely computational route to DoF extension without hardware modifications, which is relevant for OR-PAM imaging of uneven biological surfaces. The paper has clear strengths: a well-described architecture, a broad comparison against 13 existing methods, ablation studies for the loss and fusion rule, detailed training specifics, and a low parameter count (0.24 million). The simulation results in Table I are strong, and the real-data results show qualitative improvements over the baselines. However, the paper's headline claim about preserving lateral resolution over a 560 µm DoF is not directly measured anywhere in the manuscript. The reported metrics are either similarity to a synthetic Gaussian-blur model or reference-free sharpness/fidelity statistics, neither of which establishes resolution preservation. The in vivo experiment compares two focal planes but does not demonstrate continuous performance at intermediate depths.
major comments (1)
- [II-A.1 and III-C] The Discussion also claims 'accurate vasculature quantifications across the extended DoF' (last paragraph before the Conclusion), but the only quantification in Sec. III-C is vessel/junction density, which is computed without a ground-truth reference. Density can increase from false positives or artifact-enhanced structures. The authors should provide a ground-truth comparison (e.g., a registered all-in-focus reference or a manually annotated vascular map) for at least one in vivo or phantom dataset to support the quantification claim.
minor comments (1)
- [II-B.3] The description of Q_CV says 'smaller absolute value of Q_CV demonstrates the better fused result', but for a human-perception-based metric this is surprising and is not explained. Please provide a citation or a brief explanation of why smaller Q_CV is better, and ensure the signs are consistent in Tables II–IV.
Circularity Check
No significant circularity: the network is trained end-to-end against independently generated simulated ground truth, and the reported metrics are not encoded as training objectives; the self-citation in the introduction is contextual rather than load-bearing.
full rationale
The derivation chain of Dc-EEMF is self-contained. Training uses simulated multi-focus PAM images generated from external PAM datasets [31] by Gaussian blurring with randomly generated focus property maps (Sec. II-B.1), so the ground truth is not derived from the method's outputs. The U-Net focus-property predictor is trained with Eq. 9 against these independent FPM labels, and the full network is optimized with the supervised losses in Eqs. 4-8 against the original PAM image as ground truth; no reported metric (PSNR, SSIM, Q_E, Q_CV, Q_P, SD) is fitted or used as a training loss. The channel-wise spatial-frequency fusion rule (Eqs. 1-3) is a hand-designed decision criterion, not a quantity learned from the evaluation data. The Introduction mentions prior work [14] by the same authors, but this self-citation only motivates the registration problem and is not used to justify the method's correctness or to exclude alternatives. The Discussion's '560 µm DoF' and 'preserving acceptable transverse resolution' claims are under-supported because no direct lateral-resolution or axial-response measurement is reported and the real-data metrics are reference-free; however, this is a missing-evidence or correctness limitation, not a case where a predicted quantity reduces by construction to a fitted parameter or to a self-citation chain. Therefore no circular step meeting the required evidentiary standard was found.
Assumptions & free parameters
free parameters (4)
- Loss weights alpha1, alpha2, alpha3 =
alpha1=0.2, alpha2=1, alpha3=8
- Spatial frequency window size =
11x11
- Training hyperparameters =
LR 1e-3, decay 0.9 every 5 epochs, batch 16, 120 epochs
- Simulation blur parameters =
unspecified
assumptions (4)
- domain assumption Gaussian blurring of sharp images models optical-resolution PAM defocus
- domain assumption Source images from two raster scans are adequately co-registered; residual motion artifacts are small enough to be suppressed by the windowed channel-wise spatial frequency rule
- domain assumption The U-Net focus property map predictor, trained on simulated data, generalizes to real PAM images
- domain assumption No-reference metrics Q_E, Q_CV, Q_P, SD, and the Borda count are valid proxies for real-data fusion quality
Cite this review
Pith. "Pith review of Dc-EEMF: Pushing depth-of-field limit of photoacoustic microscopy via decision-level constrained learning." pith.science (2026). https://pith.science/paper/6LBLQWXK
@misc{pith2026250603181,
author = {Pith},
title = {Pith review of: Dc-EEMF: Pushing depth-of-field limit of photoacoustic microscopy via decision-level constrained learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/6LBLQWXK}},
note = {Machine review of arXiv:2506.03181}
}
read the original abstract
Photoacoustic microscopy holds the potential to measure biomarkers' structural and functional status without labels, which significantly aids in comprehending pathophysiological conditions in biomedical research. However, conventional optical-resolution photoacoustic microscopy (OR-PAM) is hindered by a limited depth-of-field (DoF) due to the narrow depth range focused on a Gaussian beam. Consequently, it fails to resolve sufficient details in the depth direction. Herein, we propose a decision-level constrained end-to-end multi-focus image fusion (Dc-EEMF) to push DoF limit of PAM. The DC-EEMF method is a lightweight siamese network that incorporates an artifact-resistant channel-wise spatial frequency as its feature fusion rule. The meticulously crafted U-Net-based perceptual loss function for decision-level focus properties in end-to-end fusion seamlessly integrates the complementary advantages of spatial domain and transform domain methods within Dc-EEMF. This approach can be trained end-to-end without necessitating post-processing procedures. Experimental results and numerical analyses collectively demonstrate our method's robust performance, achieving an impressive fusion result for PAM images without a substantial sacrifice in lateral resolution. The utilization of Dc-EEMF-powered PAM has the potential to serve as a practical tool in preclinical and clinical studies requiring extended DoF for various applications.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Functional aspects of meningeal lymphatics in ageing and Alzheimer’s disease,
S. Da Mesquitaet al., “Functional aspects of meningeal lymphatics in ageing and Alzheimer’s disease,”Nature, vol. 560, no. 7717, pp. 185– 191, Jul. 2018
work page 2018
-
[2]
CNS lymphatic drainage and neuroinflammation are regulated by meningeal lymphatic vasculature,
A. Louveauet al., “CNS lymphatic drainage and neuroinflammation are regulated by meningeal lymphatic vasculature,”Nature Neuroscience, vol. 21, no. 10, pp. 1380–1391, Sep. 2018
work page 2018
-
[3]
Multiscale photoacoustic microscopy and computed to- mography,
L. V . Wang, “Multiscale photoacoustic microscopy and computed to- mography,”Nature Photonics, vol. 3, no. 9, pp. 503–509, Aug. 2009
work page 2009
-
[4]
M. Xu and L. V . Wang, “Erratum: Universal back-projection algorithm for photoacoustic computed tomography [Phys. Rev. E71, 016706 (2005)],”Physical Review E, vol. 75, no. 5, May 2007
work page 2005
-
[5]
X. Gong, T. Jin, Y . Wang, R. Zhang, W. Qi, and L. Xi, “Photoacoustic microscopy visualizes glioma-induced disruptions of cortical microvas- cular structure and function,”Journal of Neural Engineering, vol. 19, no. 2, p. 026027, Apr. 2022
work page 2022
-
[6]
Y . Hu, Z. Chen, L. Xiang, and D. Xing, “Extended depth-of-field all- optical photoacoustic microscopy with a dual non-diffracting Bessel beam,”Optics Letters, vol. 44, no. 7, p. 1634, Mar. 2019
work page 2019
-
[7]
Deep Learning-Powered Bessel-Beam Multiparametric Photoacoustic Microscopy,
Y . Zhou, N. Sun, and S. Hu, “Deep Learning-Powered Bessel-Beam Multiparametric Photoacoustic Microscopy,”IEEE Transactions on Medical Imaging, vol. 41, no. 12, pp. 3544–3551, Dec. 2022
work page 2022
-
[8]
Optical-resolution photoacoustic microscopy with a needle-shaped beam,
R. Caoet al., “Optical-resolution photoacoustic microscopy with a needle-shaped beam,”Nature Photonics, vol. 17, no. 1, pp. 89+, JAN 2023
work page 2023
Show all 40 references
-
[9]
High-speed dual-layer scanning photoacoustic microscopy using focus tunable lens modulation at resonant frequency,
K. Lee, E. Chung, S. Lee, and T. J. Eom, “High-speed dual-layer scanning photoacoustic microscopy using focus tunable lens modulation at resonant frequency,”Optics Express, vol. 25, no. 22, p. 26427, Oct. 2017
2017
-
[10]
Continuous optical zoom microscope with extended depth of field and 3D reconstruction,
C. Liu, Z. Jiang, X. Wang, Y . Zheng, Y .-W. Zheng, and Q.-H. Wang, “Continuous optical zoom microscope with extended depth of field and 3D reconstruction,”PhotoniX, vol. 3, no. 1, Sep. 2022
2022
-
[11]
High-speed transport-of- intensity phase microscopy with an electrically tunable lens,
C. Zuo, Q. Chen, W. Qu, and A. Asundi, “High-speed transport-of- intensity phase microscopy with an electrically tunable lens,”Optics Express, vol. 21, no. 20, p. 24060, Oct. 2013
2013
-
[12]
High-speed adaptive photoacoustic microscopy,
L. Li, W. Qin, T. Li, J. Zhang, B. Li, and L. Xi, “High-speed adaptive photoacoustic microscopy,”Photonics Research, vol. 11, no. 12, pp. 2084–2092, DEC 1 2023
2023
-
[13]
Miniature probe for optomechanical focus-adjustable optical-resolution photoacoustic endoscopy,
S. Lianget al., “Miniature probe for optomechanical focus-adjustable optical-resolution photoacoustic endoscopy,”IEEE Transactions on Medical Imaging, vol. 42, no. 8, pp. 2400–2413, AUG 2023
2023
-
[14]
Multi-focus image fusion with enhancement filtering for robust vascular quantification using photoacoustic microscopy,
W. Zhouet al., “Multi-focus image fusion with enhancement filtering for robust vascular quantification using photoacoustic microscopy,”Optics Letters, vol. 47, no. 15, p. 3732, Jul. 2022
2022
-
[15]
Drpl: Deep regression pair learning for multi-focus image fusion,
J. Liet al., “Drpl: Deep regression pair learning for multi-focus image fusion,”IEEE Transactions on Image Processing, vol. 29, pp. 4816– 4831, 2020
2020
-
[16]
Multi-focus image fusion with deep residual learning and focus property detection,
Y . Liu, L. Wang, H. Li, and X. Chen, “Multi-focus image fusion with deep residual learning and focus property detection,”Information Fusion, vol. 86–87, pp. 1–16, Oct. 2022
2022
-
[17]
Ifcnn: A general image fusion framework based on convolutional neural network,
Y . Zhang, Y . Liu, P. Sun, H. Yan, X. Zhao, and L. Zhang, “Ifcnn: A general image fusion framework based on convolutional neural network,” Information Fusion, vol. 54, pp. 99–118, Feb. 2020
2020
-
[18]
U2fusion: A unified unsupervised image fusion network,
H. Xu, J. Ma, J. Jiang, X. Guo, and H. Ling, “U2fusion: A unified unsupervised image fusion network,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 1, pp. 502–518, Jan. 2022
2022
-
[19]
Mff-gan: An unsuper- vised generative adversarial network with adaptive and gradient joint constraints for multi-focus image fusion,
H. Zhang, Z. Le, Z. Shao, H. Xu, and J. Ma, “Mff-gan: An unsuper- vised generative adversarial network with adaptive and gradient joint constraints for multi-focus image fusion,”Information Fusion, vol. 66, pp. 40–53, Feb. 2021
2021
-
[20]
Swinfusion: Cross-domain long-range learning for general image fusion via swin transformer,
J. Ma, L. Tang, F. Fan, J. Huang, X. Mei, and Y . Ma, “Swinfusion: Cross-domain long-range learning for general image fusion via swin transformer,”IEEE/CAA Journal of Automatica Sinica, vol. 9, no. 7, pp. 1200–1217, Jul. 2022
2022
-
[21]
Mufusion: A general unsupervised image fusion network based on memory unit,
C. Cheng, T. Xu, and X.-J. Wu, “Mufusion: A general unsupervised image fusion network based on memory unit,”Information Fusion, vol. 92, pp. 80–92, Apr. 2023
2023
-
[22]
Exploit the best of both end-to-end and map-based methods for multi-focus image fusion,
J. Zhang, Q. Liao, H. Ma, J.-H. Xue, W. Yang, and S. Liu, “Exploit the best of both end-to-end and map-based methods for multi-focus image fusion,”IEEE Transactions on Multimedia, vol. 26, pp. 6411–6423, Jan. 2024. XXXXXXXXXXXXXXX 13
2024
-
[23]
Multi-scale weighted gradient-based fusion for multi-focus images,
Z. Zhou, S. Li, and B. Wang, “Multi-scale weighted gradient-based fusion for multi-focus images,”Information Fusion, vol. 20, pp. 60–72, Nov. 2014
2014
-
[24]
Guided filter-based multi-focus image fusion through focus region detection,
X. Qiu, M. Li, L. Zhang, and X. Yuan, “Guided filter-based multi-focus image fusion through focus region detection,”Signal Processing: Image Communication, vol. 72, pp. 35–46, Mar. 2019
2019
-
[25]
Multi-focus image fusion based on multi-scale gradients and image matting,
J. Chen, X. Li, L. Luo, and J. Ma, “Multi-focus image fusion based on multi-scale gradients and image matting,”IEEE Transactions on Multimedia, vol. 24, pp. 655–667, 2022
2022
-
[26]
A general framework for image fusion based on multi-scale transform and sparse representation,
Y . Liu, S. Liu, and Z. Wang, “A general framework for image fusion based on multi-scale transform and sparse representation,”Information Fusion, vol. 24, pp. 147–164, Jul. 2015
2015
-
[27]
Searching for mobilenetv3,
A. Howardet al., “Searching for mobilenetv3,” in2019 IEEE In- ternational Conference on Computer Vision (ICCV 2019), ser. IEEE International Conference on Computer Vision, 2019, pp. 1314–1324
2019
-
[28]
CBAM: Convolutional Block Attention Module,
S. Woo, J. Park, J.-Y . Lee, and I. S. Kweon, “CBAM: Convolutional Block Attention Module,”Computer Vision – ECCV 2018, pp. 3–19, 2018
2018
-
[29]
Perceptual Losses for Real-Time Style Transfer and Super-Resolution,
J. Johnson, A. Alahi, and L. Fei-Fei, “Perceptual Losses for Real-Time Style Transfer and Super-Resolution,”Computer Vision – ECCV 2016, pp. 694–711, 2016
2016
-
[30]
Focal frequency loss for image reconstruction and synthesis,
L. Jiang, B. Dai, W. Wu, and C. C. Loy, “Focal frequency loss for image reconstruction and synthesis,” in2021 IEEE International Conference on Computer Vision (ICCV 2021), 2021, pp. 13 899–13 909
2021
-
[31]
Simultaneous photoacoustic imaging of intravascular and tissue oxygenation,
M. Chenet al., “Simultaneous photoacoustic imaging of intravascular and tissue oxygenation,”Optics Letters, vol. 44, no. 15, p. 3773, Jul. 2019
2019
-
[32]
IFCNN: A general image fusion framework based on convolutional neural network,
Y . Zhang, Y . Liu, P. Sun, H. Yan, X. Zhao, and L. Zhang, “IFCNN: A general image fusion framework based on convolutional neural network,” Information Fusion, vol. 54, pp. 99–118, Feb. 2020
2020
-
[33]
Deep Learning-based Multi-focus Image Fusion: A Survey and A Comparative Study,
X. Zhang, “Deep Learning-based Multi-focus Image Fusion: A Survey and A Comparative Study,”IEEE Transactions on Pattern Analysis and Machine Intelligence, pp. 1–1, 2021
2021
-
[34]
Multi-focus image fusion: A Survey of the state of the art,
Y . Liu, L. Wang, J. Cheng, C. Li, and X. Chen, “Multi-focus image fusion: A Survey of the state of the art,”Information Fusion, vol. 64, pp. 71–91, Dec. 2020
2020
-
[35]
Investigating loss functions for extreme super-resolution,
Y . Jo, S. Yang, and S. J. Kim, “Investigating loss functions for extreme super-resolution,” in2020 IEEE Computer Society Conference on Com- puter Vision and Pattern Recognition Workshops (CVPRW 2020). IEEE; CVF; IEEE Comp Soc, 2020, pp. 1705–1712
2020
-
[36]
Model-based 2.5-d decon- volution for extended depth of field in brightfield microscopy,
F. Aguet, D. V . De Ville, and M. Unser, “Model-based 2.5-d decon- volution for extended depth of field in brightfield microscopy,”IEEE Transactions on Image Processing, vol. 17, pp. 1144–1153, Jul. 2008
2008
-
[37]
Activation of Ras in the Vascular Endothelium Induces Brain Vascular Malfor- mations and Hemorrhagic Stroke,
Q.-f. Li, B. Decker-Rockefeller, A. Bajaj, and K. Pumiglia, “Activation of Ras in the Vascular Endothelium Induces Brain Vascular Malfor- mations and Hemorrhagic Stroke,”Cell Reports, vol. 24, no. 11, pp. 2869–2882, Sep. 2018
2018
-
[38]
Intravital longitudinal imaging of hepatic lipid droplet accumulation in a murine model for nonalcoholic fatty liver disease,
J. Moonet al., “Intravital longitudinal imaging of hepatic lipid droplet accumulation in a murine model for nonalcoholic fatty liver disease,” Biomedical Optics Express, vol. 11, no. 9, p. 5132, Aug. 2020
2020
-
[39]
Label-free intraoperative histology of bone tissue via deep-learning-assisted ultraviolet photoacoustic microscopy,
R. Caoet al., “Label-free intraoperative histology of bone tissue via deep-learning-assisted ultraviolet photoacoustic microscopy,”Nature Biomedical Engineering, vol. 7, no. 2, pp. 124–134, Sep. 2022
2022
-
[40]
Content aware multi-focus image fusion for high- magnification blood film microscopy,
P. Manescuet al., “Content aware multi-focus image fusion for high- magnification blood film microscopy,”Biomedical Optics Express, vol. 13, no. 2, p. 1005, Jan. 2022
2022
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.