{"id":"a22d38ad-bc72-4754-8404-d8971b311af9","arxiv_id":"2505.24514","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Digital twins provide a simulated reference that enables full-reference quality comparison of photoacoustic image reconstructions on experimental data.","lead":"The authors built digital twins of tissue-mimicking phantoms and a photoacoustic scanner to generate a simulated reference image, then compared six reconstruction algorithms on experimental data. They found that a Fourier transform-based algorithm matches the quality of iterative time reversal while running much faster.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reference p0 is validated only at the time-series level (R=0.81), so a spatially biased digital-twin reference could alter the FFT-vs-ITTR ranking; the comparability claim needs a known-ground-truth cross-check.","rationale":"The reader's weakest_assumption correctly identifies the simulated p0 as the load-bearing element. My independent reading of the full text reinforces this and localizes it further: the calibration pipeline (Section II.A.3) fits linear parameters a, b, c to the acoustic time series after simulation, but the p0 reference used for IQA is exactly the unmodified simulation output of the coupled MCX/k-Wave model. Thus the 0.81 correlation is an upper bound on how well the model predicts the measured data; the p0 reference may contain errors that are not corrected by the calibration and that correlate with the very features (edges, contrast, spatial frequency) that the IQA metrics score. The paper's own acknowledgment that 'systematic non-linear changes could not be captured' (Section IV) means the reference can be systematically wrong in a way that favors one reconstruction class over another. For the specific claim of FFT-ITTR comparability, the differences between these two methods on several metrics (e.g., R 0.61 vs 0.77, SSIM 0.75 vs 0.78) are relatively small, so a reference bias could flip the ordering on some measures and weaken the conclusion. The proposed test directly benchmarks the framework's reliability: comparing the algorithm rankings obtained with known ground truth (simulations) against those obtained with the digital-twin reference (experiments) would reveal whether the reference biases the ranking. If the rankings diverge, the paper would need to either improve the forward model or restrict its claims to the simulation-validated regime. This does not change the overall verdict of conditional acceptance, but it should become an explicit condition for acceptance.","tokens_in":16829,"tokens_out":9398,"duration_ms":110446,"concrete_test":"Use the existing simulated-phantom data (where the exact p0 is known) to compute the same IQA table (R, MAE, SSIM, JSD, HaarPSI, FWHM) for DAS, FBP, MB, TR, ITTR, and FFT, and compare the algorithm ranking against the experimental-data ranking obtained with the digital-twin reference. If FFT ranks above or alongside ITTR in the simulated ground-truth ranking but below in the experimental ranking (or vice versa), the reference bias is demonstrably affecting the conclusion; if the rankings agree, the concern is mitigated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that digital twins provide a valid full-reference IQA basis rests on the simulated p0 being a sufficiently accurate surrogate for the true initial pressure. The paper's own validation of the forward model (Section III.A) reaches only a Pearson correlation of 0.81 between calibrated simulated and experimental time series, and the calibration (Section II.A.3) fits a, b, c to g(t,y) only; the p0 reference itself is never directly validated. Because the IQA scores in Table I are computed against this p0, any spatially varying error in the simulated fluence, Grüneisen parameter, or geometry (e.g., from the manually segmented masks, the Gaussian illumination profile, or the unmodelled acoustic attenuation) will propagate into the scores. The ranking of FFT vs ITTR is particularly sensitive: on experimental data FFT is worse on R (0.61 vs 0.77) and MAE (140 vs 111) but better on FWHM (4.55 vs 4.92) and close on SSIM (0.75 vs 0.78). A reference that, say, underestimates high-frequency content would systematically favor algorithms that de-emphasize high frequencies, potentially inverting the comparability conclusion. The paper explicitly acknowledges 'the simulated p0 is only a reference and not an exact ground truth' (Section IV), but does not quantify how sensitive the algorithm ranking is to reference errors. Without an independent check of the reference against known ground truth, the claim that FFT yields results 'comparable' to ITTR is not yet secured.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces a digital-twin evaluation framework for full-reference image-quality assessment (IQA) in photoacoustic imaging. Numerical phantoms are built from optically and acoustically characterized tissue-mimicking materials, and device digital twins (MCX light transport, k-Wave acoustics, MSOT InVision-256TF geometry) are used to simulate the initial pressure distribution p0 and the measurement time series. A three-parameter linear calibration with impulse-response convolution and a measured noise term bridges simulation and experiment. Using the simulated p0 as reference, five full-reference IQA measures and FWHM compare DAS, FBP, model-based reconstruction, time reversal, iterative time reversal, and a circular FFT-based algorithm on simulated and experimental data. The central claim is that the FFT algorithm performs comparably to iterative time reversal at much lower computational cost, and that the digital-twin framework enables objective algorithm comparison on experimental data.","tokens_in":17204,"tokens_out":4470,"duration_ms":51341,"significance":"The proposed framework addresses a genuine need: full-reference IQA requires a known reference, which is unavailable in phantom and in vivo experiments. The authors contribute an open pipeline (SIMPA, PATATO, Zenodo data/code), a new experimental evaluation of a circular FFT-based reconstruction method, and a quantitative calibration methodology that demonstrably improves simulated-versus-experimental time-series agreement (R from 0.75 to 0.81, RMSE from 2.75 to 2.27). The idea of using a calibrated digital twin as a middle ground between pure simulation and unknown ground truth is valuable and likely to be reused. However, the strength of the claims is currently limited by indirect validation of the reference and by the absence of statistical characterization of the ranking.","major_comments":[{"comment":"The forward model is validated only at the time-series level (Pearson R=0.81, RMSE=2.27), while the full-reference IQA scores are computed against the simulated p0, which is never directly validated. Because the FFT-versus-ITTR differences are small (R 0.61 vs 0.77; MAE 140.17 vs 110.91; SSIM 0.75 vs 0.78; JSD 0.41 vs 0.37), a spatially varying error in the simulated fluence, Grüneisen parameter, or geometry could change the algorithm ranking. The manuscript should add a sensitivity analysis that perturbs the forward-model parameters (e.g., illumination profile, segmentation, acoustic attenuation, calibration coefficients) and reports whether the relative ranking of FFT and ITTR is stable, or carry out an independent cross-check against known ground truth in a simulation study.","section":"Section III.A, Table I"},{"comment":"The reported IQA values are single numbers with no error bars, confidence intervals, or significance tests across the 15 test phantoms, although Figure 3 suggests paired data are available. The central claim that FFT is comparable to ITTR requires showing that the observed score differences are not within the noise of the evaluation. Please report per-phantom variability and paired significance tests (or equivalent) for the key pairwise comparisons.","section":"Table I"},{"comment":"The segmentation masks of the numerical phantoms are created from delay-and-sum reconstructions of the same experimental data being simulated, and the calibration parameters are optimized on N=15 calibration phantoms. The manuscript should state explicitly whether the experimental IQA in Table I was computed on held-out test phantoms and should quantify how sensitive the IQA scores are to the manual segmentation and to the choice of calibration subset. The Discussion already concedes that systematic nonlinear changes could not be captured; this limitation needs to be reflected in the confidence of the FFT-versus-ITTR comparison.","section":"Section II.A.3, Section IV"}],"minor_comments":[{"comment":"The stated density of 1000 g/cm3 should be 1000 kg/m3 (or 1 g/cm3); as written the density is off by a factor of 1000.","section":"Section II.A.3"},{"comment":"The caption uses three asterisks and a Wilcoxon signed-rank test without specifying the sample size or what is paired; please add this information.","section":"Figure 3"},{"comment":"It would help to state how many experimental phantoms were used for Table I and how they were split from the 15 calibration phantoms; currently the reader must infer this from Section II.A.3.","section":"Section III.B, Table I"},{"comment":"The FWHM is described as a no-reference measure but is included in the full-reference comparison table; please label it accordingly in the table and caption.","section":"Section II.C.6"},{"comment":"The conclusion states 'up to 100-times faster' but the results section reports average runtimes without identifying which algorithm pair yields the 100x figure; please specify the comparison.","section":"Section III.D"}],"recommendation":"major_revision","confidential_remarks":"The reviewer concern about p0 validity is legitimate and should be addressed before publication. I would encourage revision rather than rejection: the open pipeline and first experimental evaluation of circular FFT are useful contributions, and the missing sensitivity analysis is feasible within the manuscript's scope. The main weaknesses are the indirect validation of the reference and the lack of statistical characterization of the ranking."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is the first paper I've seen that uses a calibrated digital twin of a tissue-mimicking phantom and a real scanner to produce a p0 reference for full-reference IQA on experimental photoacoustic data. That is a genuinely useful move: it gives the field a way to compare reconstruction algorithms on real measurements instead of simulation-only or no-reference metrics. Second, the experimental test of the circular FFT algorithm is new and the speed advantage is real, but the \"comparable to iterative time reversal\" claim is softer than the abstract suggests.\n\nWhat I like: the forward model is built from characterized phantom materials (DIS measurements), MCX light transport, k-Wave acoustics, and a device model with vendor geometry. The calibration to time-series data (offset, scaling, convolution with IRF, noise term) is a sensible way to reduce the sim-to-experiment gap, and the comparison of calibrated vs. naive scaling (R 0.81 vs 0.75, RMSE 2.27 vs 2.75) is honest evidence that the calibration does something. The paper ships all data and code on Zenodo, so the numbers are checkable. The full-view vs. limited-view comparison (Table II) is a nice addition and shows the framework can probe algorithm behavior under realistic acquisition geometry.\n\nSoft spots, in order of severity. (1) The reference p0 is never directly validated. The calibration is fit to g(t,y) and reaches R=0.81 at the time-series level; a spatially varying error in fluence, Grüneisen, or geometry would propagate into the IQA scores and could in principle change the FFT-vs-ITTR ranking. The paper acknowledges this in Section IV, but does not quantify sensitivity. I'd want a known-ground-truth cross-check, e.g., a simulation where the true p0 is known, to see how much ranking error the simulation gap introduces. (2) Table I reports single mean values without error bars or significance tests across the 15 test phantoms. The differences between FFT and ITTR on R and MAE are large enough to look real, but the FWHM advantage (4.55 vs 4.92) could be within noise. (3) The \"comparable\" language is a bit generous: on experimental data FFT is clearly worse on R and MAE than ITTR, better on FWHM. \"Comparable\" should be \"slightly worse but much faster in this implementation.\" (4) The 100x speed-up is implementation-dependent; TR runs in 2-3 seconds, so the real contrast is vs. MB/ITTR, and the FFT implementation is CPU-only. That should be stated more carefully.\n\nOverall: the central idea is sound and the paper is a solid methods contribution. The simulation gap is a real limitation but it's acknowledged and partially quantified. I'd send it to peer review and would likely ask for the sensitivity analysis and error bars before acceptance.","headline":"A calibrated digital-twin framework for full-reference IQA on experimental photoacoustic data is a genuine step forward, but the p0 reference needs a known-ground-truth check and the 'comparable' claim about the FFT algorithm is a bit generous.","tokens_in":17678,"tokens_out":2061,"would_cite":true,"duration_ms":22238,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["65R32","44A12","92C55"],"pacs":[],"model":"deepseek-v4-flash","headline":"By building numerical digital twins of tissue-mimicking phantoms and the imaging system, this paper supplies the missing reference image needed for full-reference quality assessment of photoacoustic reconstruction algorithms, and shows…","keywords":["photoacoustic imaging","image reconstruction","full-reference image quality assessment","digital twin","tissue-mimicking phantom","time reversal","Fourier reconstruction","limited-view tomography"],"falsifier":"Acquire a phantom whose true initial pressure distribution is known independently, for example by taking a full-view 360-degree scan and reconstructing with an exact algorithm, then check whether the ranking of reconstruction algorithms by full-reference IQA against the digital-twin reference matches the ranking against that independent ground truth; a mismatch would falsify the framework's central claim.","tokens_in":16672,"feed_emoji":"🖥️","tokens_out":8410,"duration_ms":86443,"temperature":0.7,"pith_summary":"Photoacoustic imaging reconstructions cannot usually be scored against a true reference image because the initial pressure distribution depends on both the phantom and the scanner's illumination, so the field falls back on no-reference metrics that do not measure accuracy. This paper argues that a numerical digital twin of the phantom and device can provide that reference: with measured optical properties, a Monte Carlo light transport simulation, an acoustic forward model, and a three-part calibration against real time-series data, the simulated initial pressure distribution becomes a surrogate ground truth. Using this framework, the authors compare five reconstruction algorithms on experimental data, finding that an FFT-based circular-geometry algorithm performs close to iterative time reversal while running far faster. If the digital-twin reference is trustworthy, the approach gives the community a reproducible, quantitative way to rank acoustic inversion schemes on real measurements rather than only in simulation.","feed_headline":"Digital twins give photoacoustic reconstruction a reference image","feed_subtitle":"Simulated phantom-and-scanner twins enable fair scoring; the FFT method rivals time reversal at far lower cost.","key_machinery":"The load-bearing mechanism is the digital-twin pipeline: a numerical phantom built from measured absorption and reduced-scattering coefficients, a Monte Carlo simulation of light fluence, multiplication by a constant Grüneisen parameter to give the initial pressure, a k-space pseudospectral acoustic forward model with a 707-point interpolation of each toroidal transducer surface, and a three-part linear calibration against experimental time-series data. A secondary mechanism is the FFT-based circular reconstruction algorithm, which inverts the 2D wave equation exactly for a full circle of detectors and then applies a microlocal correction that replaces the distorted half of the Fourier domain with reflected values from the accurate half, mitigating the missing 90 degrees of detector coverage.","core_discovery":"The central discovery is that a calibrated digital-twin pipeline can produce a sufficiently accurate simulated initial pressure distribution p0 for a real, piecewise-constant tissue-mimicking phantom, making full-reference image quality assessment possible on experimental photoacoustic data. The paper shows the calibration step—linear amplitude scaling, addition of a measured noise profile, and convolution with the system impulse response—improves the correlation between simulated and experimental time series from 0.75 to 0.81, and then uses the simulated p0 to rank five algorithms. Two findings stand out: the FFT-based method, tested on experimental data for the first time, yields image quality comparable to iterative time reversal but with average runtimes around 1.4 seconds per image versus about one minute, and a fraction of the memory; and moving from full-view to limited-view data degrades time-reversal-based methods and the FFT method similarly while the model-based method is less affected.","pith_inferences":["The same digital-twin reference could be used to train supervised reconstruction networks, since it gives paired experimental-style measurements and a high-quality p0 target, something the field currently lacks.","Because the FFT algorithm is CPU-based and very fast, it could serve as a differentiable surrogate operator in optimization loops, potentially accelerating iterative or learned reconstruction for limited-view geometries.","If the forward model's linear calibration misses systematic nonlinear effects, the ranking could be biased toward algorithms whose own assumptions mirror those of the simulation; pushing the calibration toward nonlinear corrections would be a natural stress test.","The framework's reliance on piecewise-constant phantoms limits its direct extrapolation to in vivo tissue; extending it to heterogeneous, 3D-printed phantoms would test whether the ranking holds under realistic complexity."],"forward_implications":["The FFT-based algorithm can be adopted as a fast, memory-lean reconstruction method for circular-geometry scanners, with image quality close to that of iterative time reversal.","Full-reference IQA measures can now be applied to experimental photoacoustic data, not just simulations, enabling algorithm comparisons that no-reference metrics cannot provide.","The digital-twin framework is portable to other photoacoustic systems: one characterizes a phantom, builds a device model, and calibrates the forward model, yielding a reference for that specific system.","The limited-view comparison indicates that algorithms respond differently to missing detector coverage, and that some IQA measures, such as SSIM, are less sensitive to limited-view artifacts than others.","The computational cost gap between the FFT method and iterative time reversal encourages using the FFT method as a building block for model-based or learned reconstruction approaches."],"supporting_citations":[{"why":"supplies the tissue-mimicking phantom material and its measured optical and acoustic parameters used to build the numerical phantom","marker":"[28]"},{"why":"the Monte Carlo light transport simulation that computes the fluence field in the digital twin","marker":"[33]"},{"why":"the k-space pseudospectral acoustic method that simulates the measurement time series from the simulated initial pressure","marker":"[34]"},{"why":"the open-source toolkit that assembles the device digital twin and orchestrates the light and acoustic simulations","marker":"[35]"},{"why":"the fast Fourier-based reconstruction algorithm for circular geometries that is tested on experimental data for the first time","marker":"[16]"},{"why":"the microlocal correction that repairs the Fourier-domain distortion caused by the missing detector segment","marker":"[37]"},{"why":"the time-reversal reconstruction baseline that the FFT method is compared against","marker":"[12]"},{"why":"the model-based reconstruction baseline used as a state-of-the-art comparison","marker":"[21]"}],"fun_headline_variants":["Digital twin yields reference image for photoacoustic quality","Photoacoustic reconstructions now get a full-reference score","Calibrated digital twin scores photoacoustic reconstructions","FFT photoacoustic method rivals time reversal at 1/40 cost","Digital twin enables full-reference photoacoustic quality check"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The simulated initial pressure distribution is accurate enough to serve as a full-reference ground truth, and that accuracy rests on the calibrated forward model faithfully representing the real phantom and scanner; the paper explicitly notes that the simulated p0 is only a reference, not exact ground truth.","fun_headline_variants_meta":{"raw":{"variants":["Digital twin yields reference image for photoacoustic quality","Photoacoustic reconstructions now get a full-reference score","Calibrated digital twin scores photoacoustic reconstructions","FFT photoacoustic method rivals time reversal at 1/40 cost","Digital twin enables full-reference photoacoustic quality check"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000564,"raw_usage":{"total_tokens":2678,"prompt_tokens":951,"completion_tokens":1727,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":567,"completion_tokens_details":{"reasoning_tokens":1644}},"tokens_in":567,"tokens_out":1727,"duration_ms":19337,"temperature":1.0,"reasoning_tokens":1644,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:19:38.931370+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Acquire a phantom whose true initial pressure distribution is known independently, for example by taking a full-view 360-degree scan and reconstructing with an exact algorithm, then check whether the ranking of reconstruction algorithms by full-reference IQA against the digital-twin reference matches the ranking against that independent ground truth; a mismatch would falsify the framework's central claim.","supporting_citations":[],"review_version":1}