Pith. sign in

REVIEW 3 major objections 5 minor 12 references

Digital twins enable full-reference quality assessment of photoacoustic image reconstructions

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read By building numerical digital twins of tissue-mimicking phantoms and the imaging system, this paper supplies the missing reference image needed for full-reference quality assessment of photoacoustic reconstruction algorithms, and shows…

desk verdict A calibrated digital-twin framework for full-reference IQA on experimental photoacoustic data is a genuine step forward, but the p0 reference needs a known-ground-truth check and the 'comparable' claim about the FFT algorithm is a bit generous. read the letter →

arxiv 2505.24514 v1 pith:H3TKE4FE submitted 2025-05-30 physics.med-ph cs.CVeess.SP

classification physics.med-phcs.CVeess.SP MSC 65R3244A1292C55
keywords photoacousticimagingimagereconstructionfull-referencequalityassessmentdigitaltwintissue-mimickingphantomtimereversalFourierlimited-viewtomography
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Photoacoustic imaging reconstructions cannot usually be scored against a true reference image because the initial pressure distribution depends on both the phantom and the scanner's illumination, so the field falls back on no-reference metrics that do not measure accuracy. This paper argues that a numerical digital twin of the phantom and device can provide that reference: with measured optical properties, a Monte Carlo light transport simulation, an acoustic forward model, and a three-part calibration against real time-series data, the simulated initial pressure distribution becomes a surrogate ground truth. Using this framework, the authors compare five reconstruction algorithms on experimental data, finding that an FFT-based circular-geometry algorithm performs close to iterative time reversal while running far faster. If the digital-twin reference is trustworthy, the approach gives the community a reproducible, quantitative way to rank acoustic inversion schemes on real measurements rather than only in simulation.

What carries the argument

The load-bearing mechanism is the digital-twin pipeline: a numerical phantom built from measured absorption and reduced-scattering coefficients, a Monte Carlo simulation of light fluence, multiplication by a constant Grüneisen parameter to give the initial pressure, a k-space pseudospectral acoustic forward model with a 707-point interpolation of each toroidal transducer surface, and a three-part linear calibration against experimental time-series data. A secondary mechanism is the FFT-based circular reconstruction algorithm, which inverts the 2D wave equation exactly for a full circle of detectors and then applies a microlocal correction that replaces the distorted half of the Fourier domain with reflected values from the accurate half, mitigating the missing 90 degrees of detector coverage.

What would settle it

Acquire a phantom whose true initial pressure distribution is known independently, for example by taking a full-view 360-degree scan and reconstructing with an exact algorithm, then check whether the ranking of reconstruction algorithms by full-reference IQA against the digital-twin reference matches the ranking against that independent ground truth; a mismatch would falsify the framework's central claim.

Watch

Extended reading notes

Core claim

The central discovery is that a calibrated digital-twin pipeline can produce a sufficiently accurate simulated initial pressure distribution p0 for a real, piecewise-constant tissue-mimicking phantom, making full-reference image quality assessment possible on experimental photoacoustic data. The paper shows the calibration step—linear amplitude scaling, addition of a measured noise profile, and convolution with the system impulse response—improves the correlation between simulated and experimental time series from 0.75 to 0.81, and then uses the simulated p0 to rank five algorithms. Two findings stand out: the FFT-based method, tested on experimental data for the first time, yields image quality comparable to iterative time reversal but with average runtimes around 1.4 seconds per image versus about one minute, and a fraction of the memory; and moving from full-view to limited-view data degrades time-reversal-based methods and the FFT method similarly while the model-based method is less affected.

Load-bearing premise

The simulated initial pressure distribution is accurate enough to serve as a full-reference ground truth, and that accuracy rests on the calibrated forward model faithfully representing the real phantom and scanner; the paper explicitly notes that the simulated p0 is only a reference, not exact ground truth.

Editorial extensions

If this is right

  • The FFT-based algorithm can be adopted as a fast, memory-lean reconstruction method for circular-geometry scanners, with image quality close to that of iterative time reversal.
  • Full-reference IQA measures can now be applied to experimental photoacoustic data, not just simulations, enabling algorithm comparisons that no-reference metrics cannot provide.
  • The digital-twin framework is portable to other photoacoustic systems: one characterizes a phantom, builds a device model, and calibrates the forward model, yielding a reference for that specific system.
  • The limited-view comparison indicates that algorithms respond differently to missing detector coverage, and that some IQA measures, such as SSIM, are less sensitive to limited-view artifacts than others.
  • The computational cost gap between the FFT method and iterative time reversal encourages using the FFT method as a building block for model-based or learned reconstruction approaches.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same digital-twin reference could be used to train supervised reconstruction networks, since it gives paired experimental-style measurements and a high-quality p0 target, something the field currently lacks.
  • Because the FFT algorithm is CPU-based and very fast, it could serve as a differentiable surrogate operator in optimization loops, potentially accelerating iterative or learned reconstruction for limited-view geometries.
  • If the forward model's linear calibration misses systematic nonlinear effects, the ranking could be biased toward algorithms whose own assumptions mirror those of the simulation; pushing the calibration toward nonlinear corrections would be a natural stress test.
  • The framework's reliance on piecewise-constant phantoms limits its direct extrapolation to in vivo tissue; extending it to heterogeneous, 3D-printed phantoms would test whether the ranking holds under realistic complexity.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The manuscript introduces a digital-twin evaluation framework for full-reference image-quality assessment (IQA) in photoacoustic imaging. Numerical phantoms are built from optically and acoustically characterized tissue-mimicking materials, and device digital twins (MCX light transport, k-Wave acoustics, MSOT InVision-256TF geometry) are used to simulate the initial pressure distribution p0 and the measurement time series. A three-parameter linear calibration with impulse-response convolution and a measured noise term bridges simulation and experiment. Using the simulated p0 as reference, five full-reference IQA measures and FWHM compare DAS, FBP, model-based reconstruction, time reversal, iterative time reversal, and a circular FFT-based algorithm on simulated and experimental data. The central claim is that the FFT algorithm performs comparably to iterative time reversal at much lower computational cost, and that the digital-twin framework enables objective algorithm comparison on experimental data.

Significance. The proposed framework addresses a genuine need: full-reference IQA requires a known reference, which is unavailable in phantom and in vivo experiments. The authors contribute an open pipeline (SIMPA, PATATO, Zenodo data/code), a new experimental evaluation of a circular FFT-based reconstruction method, and a quantitative calibration methodology that demonstrably improves simulated-versus-experimental time-series agreement (R from 0.75 to 0.81, RMSE from 2.75 to 2.27). The idea of using a calibrated digital twin as a middle ground between pure simulation and unknown ground truth is valuable and likely to be reused. However, the strength of the claims is currently limited by indirect validation of the reference and by the absence of statistical characterization of the ranking.

major comments (3)
  1. [Section III.A, Table I] The forward model is validated only at the time-series level (Pearson R=0.81, RMSE=2.27), while the full-reference IQA scores are computed against the simulated p0, which is never directly validated. Because the FFT-versus-ITTR differences are small (R 0.61 vs 0.77; MAE 140.17 vs 110.91; SSIM 0.75 vs 0.78; JSD 0.41 vs 0.37), a spatially varying error in the simulated fluence, Grüneisen parameter, or geometry could change the algorithm ranking. The manuscript should add a sensitivity analysis that perturbs the forward-model parameters (e.g., illumination profile, segmentation, acoustic attenuation, calibration coefficients) and reports whether the relative ranking of FFT and ITTR is stable, or carry out an independent cross-check against known ground truth in a simulation study.
  2. [Table I] The reported IQA values are single numbers with no error bars, confidence intervals, or significance tests across the 15 test phantoms, although Figure 3 suggests paired data are available. The central claim that FFT is comparable to ITTR requires showing that the observed score differences are not within the noise of the evaluation. Please report per-phantom variability and paired significance tests (or equivalent) for the key pairwise comparisons.
  3. [Section II.A.3, Section IV] The segmentation masks of the numerical phantoms are created from delay-and-sum reconstructions of the same experimental data being simulated, and the calibration parameters are optimized on N=15 calibration phantoms. The manuscript should state explicitly whether the experimental IQA in Table I was computed on held-out test phantoms and should quantify how sensitive the IQA scores are to the manual segmentation and to the choice of calibration subset. The Discussion already concedes that systematic nonlinear changes could not be captured; this limitation needs to be reflected in the confidence of the FFT-versus-ITTR comparison.
minor comments (5)
  1. [Section II.A.3] The stated density of 1000 g/cm3 should be 1000 kg/m3 (or 1 g/cm3); as written the density is off by a factor of 1000.
  2. [Figure 3] The caption uses three asterisks and a Wilcoxon signed-rank test without specifying the sample size or what is paired; please add this information.
  3. [Section III.B, Table I] It would help to state how many experimental phantoms were used for Table I and how they were split from the 15 calibration phantoms; currently the reader must infer this from Section II.A.3.
  4. [Section II.C.6] The FWHM is described as a no-reference measure but is included in the full-reference comparison table; please label it accordingly in the table and caption.
  5. [Section III.D] The conclusion states 'up to 100-times faster' but the results section reports average runtimes without identifying which algorithm pair yields the 100x figure; please specify the comparison.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the simulated p0 reference is produced by a forward model calibrated only against time-series measurements, so the reconstruction comparison is not self-referential.

full rationale

The paper's derivation chain is: (i) characterize phantom optical and acoustic properties; (ii) simulate p0 = Γ·μa·φ with MCX and time series g(t,y) with k-Wave; (iii) calibrate the linear parameters a, b, c plus the impulse response and noise model against experimental time series only; (iv) use the simulated p0 as the reference for full-reference IQA of reconstruction algorithms. No IQA score enters the calibration, and no reconstruction algorithm is used to define p0. The calibration equation in Section II.A.3 fits g(t,y)exp = a + (b·g(t,y)sim)*IRF(t,y) + c·noise(t,y), which is independent of the reconstruction methods compared in Table I. The FFT-based algorithm is a prior result by a co-author (Section II.B.1), but it is tested here on experimental data for the first time; the paper does not invoke the citation as proof of the algorithm ranking. Self-citations to SIMPA and PATATO are tool references, not load-bearing mathematical premises. The Discussion explicitly cautions that 'the simulated p0 is only a reference and not an exact ground truth,' which is a modeling limitation rather than a circular reduction. The only unusual wording is 'careful calibration of the simulations with the reconstructions' in the Discussion; the Methods calibration equation shows the fit is against time-series measurements, so no circular step can be exhibited. No predicted quantity reduces by construction to a fitted input.

Assumptions & free parameters 6 free parameters · 6 assumptions · 0 invented entities

The central evaluation rests on the digital twin simulation being a faithful proxy for the real scanner and phantom. Three linear calibration parameters and the manually optimized illumination and detector geometry are fitted to the same experimental data used for evaluation, which introduces a fitting step but not circularity: the reference p0 is an intermediate simulation output, not derived from the reconstruction algorithms. The assumptions of constant acoustics, constant Grueneisen parameter, and piecewise-constant materials limit fidelity to the experimental system.

free parameters (6)
  • Calibration offset a = 4.5
    Fitted by least squares on N=15 calibration phantoms to match simulated time series to experimental data (Section II.A3).
  • Calibration scaling b = 0.068
    Fitted by least squares on N=15 calibration phantoms to match simulated time series to experimental data (Section II.A3).
  • Noise coefficient c = 0.89
    Fitted by least squares on N=15 calibration phantoms to match simulated time series to experimental data (Section II.A3).
  • Illumination profile parameters = not reported
    Optimized to match experimental radiant exposure; the paper states a Gaussian profile was used but does not report its parameters (Section II.A3).
  • Detector design parameters = not reported
    Optimized along with illumination to match experiments; numbers not given (Section II.A3).
  • Naive scaling factor = 10
    Empirically determined for the comparison baseline in Section III.A.
assumptions (6)
  • domain assumption Wave equation model (ptt = c^2 Delta p) with instantaneous absorption, open-space propagation, no reflections, and negligible attenuation.
    Eq. (1) in Section I; this is the simplified physics assumed by all reconstruction methods and the forward model.
  • domain assumption Constant sound speed and density in phantom and coupling medium.
    Section II.A3 sets sound speeds 1468 and 1489 m/s and density 1000 g/cm^3.
  • domain assumption Constant Grueneisen parameter for p0 = Gamma mu_a phi.
    Section II.A3.
  • domain assumption Piecewise-constant phantom composition with manually segmented masks.
    Section II.A3.
  • standard math Monte Carlo light transport (MCX) and k-space pseudospectral acoustics (k-Wave) are accurate numerical solvers.
    Used as the forward simulation backbone; treated as validated external tools.
  • ad hoc to paper Linear calibration model g_exp = a + (b*g_sim)*IRF + c*noise adequately bridges simulation and experiment.
    Section II.A3; the paper notes nonlinear models were tested without significant improvement.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Digital twins enable full-reference quality assessment of photoacoustic image reconstructions." pith.science (2026). https://pith.science/paper/H3TKE4FE

@misc{pith2026250524514,
  author       = {Pith},
  title        = {Pith review of: Digital twins enable full-reference quality assessment of photoacoustic image reconstructions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/H3TKE4FE}},
  note         = {Machine review of arXiv:2505.24514}
}
read the original abstract

Quantitative comparison of the quality of photoacoustic image reconstruction algorithms remains a major challenge. No-reference image quality measures are often inadequate, but full-reference measures require access to an ideal reference image. While the ground truth is known in simulations, it is unknown in vivo, or in phantom studies, as the reference depends on both the phantom properties and the imaging system. We tackle this problem by using numerical digital twins of tissue-mimicking phantoms and the imaging system to perform a quantitative calibration to reduce the simulation gap. The contributions of this paper are two-fold: First, we use this digital-twin framework to compare multiple state-of-the-art reconstruction algorithms. Second, among these is a Fourier transform-based reconstruction algorithm for circular detection geometries, which we test on experimental data for the first time. Our results demonstrate the usefulness of digital phantom twins by enabling assessment of the accuracy of the numerical forward model and enabling comparison of image reconstruction schemes with full-reference image quality assessment. We show that the Fourier transform-based algorithm yields results comparable to those of iterative time reversal, but at a lower computational cost. All data and code are publicly available on Zenodo: https://doi.org/10.5281/zenodo.15388429.

Figures

Figures reproduced from arXiv: 2505.24514 by the authors.

Figure 1
Figure 1. FIG. 1. Overview of the proposed evaluation strategy for [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. FIG. 2 [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. FIG. 3 [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: A, C, E) and that there is a significant level of noise in the experimental images. Furthermore, the qualitative comparison suggests that the circular FFT-based method produces results that are similar to the computationally expensive iterative time reversal method. Fo…
Figure 5
Figure 5. Figure 5: FIG. 5 [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

12 extracted references · 11 canonical work pages

  1. [1]

    Zero-pad datag(t, y)intby a factor of 2 or more, by adding zeros after the actually measured values

  2. [2]

    Using the FFT, compute the Fourier transform of the zero-padded datag(ct/R, y(θ))intand the Fourier series inθ, wherey(θ) =R(cosθ,sinθ) : ˆgk(ρ) = c 2πR Z R   2πZ 0 g(ct/R,ˆy(θ))e−ikθdθ   eitρdt. The computed values are defined on a computational grid inρ∈[−ρ Nyq, ρNyq], k∈[−N d/2, Nd/2],where ρNyq is the Nyquist frequency of discretization inˆtand Nd...

  3. [3]

    whereH (0) m is the Hankels function of orderm

    For eachkand each positive value ofρin the grid compute coefficientsb k(ρ)by the formula bk(ρ) = 4i|k| ρH (1) |k| (ρ) ˆgk(ρ). whereH (0) m is the Hankels function of orderm

  4. [4]

    For each value ofρin the grid, using the FFT, sum the Fourier series thus obtaining functionˆB(ρ, φ): ˆB(ρ, φ) = Nd/2X k=−Nd/2 bk(ρ)eikφ. Function ˆB(ρ, φ)represents a theoretically exact po- lar grid representation to the Fourier transformˆu(0, ξ) of the functionu(0,ˆx)we seek, assuming thatξ= |ξ|(cosφ,sinφ),ξ∈(0, ρ max),φ∈[0,2π)

  5. [5]

    Interpolate ˆB(ρ, φ)from the polar grid to a Cartesian grid inξ,producing an approximation toˆu(0, ξ).Here, we utilize bilinear interpolation

  6. [6]

    Additional correction: the key observation of37 is that, if a segment of detectors is absent, roughly, a half of the approximation toˆu(0, ξ)will be severely distorted, while the other half will be reconstructed quite ac- curately. Sinceu(0,ˆx)is a real function, its Fourier transform has the propertyˆu(0, ξ) =ˆu(0,−ξ).The ad- ditional correction we deplo...

  7. [7]

    Using a 2D FFT we computeu(0,ˆx)fromˆu(0, ξ), and reconstructp 0(x) =u(0, Rˆx). This FFT-based method, in its current imple- mentation, is asymptotically fast, meaning that it requiresO(n 2 logn)floating point operations (flops) for an(n×n)image, assuming that the data con- tainsO(n 2)values. Other fast algorithms are only known for the cases of linear or...

  8. [8]

    Extract a line profile through the image that in- cludes both background and target boundaries

Show all 12 references
  1. [9]

    Compute the absolute value of the gradient of the line profile

  2. [10]

    Identify the peaks in the absolute gradient

  3. [11]

    For each peak, locate the first positions to the left and right where the absolute gradient drops to half of the peak value

  4. [12]

    Mul- tispectral optoacoustic tomography for assessment of crohn’s dis- ease activity,

    Compute the width by subtracting the left position from the right position. D.Computational Footprint Estimation We evaluated the computational footprint, including both time and memory usage, for the specific algorithm implementations. To measure execution time, we used timin...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.