{"id":"fad64f31-398a-4872-9e6e-69b63819c208","arxiv_id":"2411.16551","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"A coherence-maximizing, pixel-wise average sound speed map corrects ultrasound focusing delays, with corrected fetal images preferred or rated equal in 72.5% of 172 cases.","lead":"An ultrasound imaging team proposes fixing focus errors by computing, for every image pixel, the sound speed that maximizes signal coherence, then re-beamforming with that speed map. In 172 fetal ultrasound images, expert clinicians preferred the corrected image or rated it equal to the standard image in 72.5% of cases, and sharpness gains tracked sound speed deviation with a correlation of 0.67.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The straight-ray, aperture-averaged sound speed model in §2.1 is the load-bearing premise; the paper never quantifies refraction-induced travel-time error against the 10 m/s candidate-sound-speed step, and in vivo examples G/J show the consequence.","rationale":"The reader's weakest assumption and my load-bearing concern are the same: the straight-ray, no-refraction model in §2.1. The paper is commendably transparent about this assumption and about the mixed clinical outcomes in Examples G and J, so the concern is grounded in the manuscript's own evidence rather than in an external standard. The in-silico validation uses Eq. (1) to define ground truth, so it cannot validate the wave-propagation model; it only validates the estimator against itself. The clinical preference data (72.5% preferred or similar) are useful empirical support but do not isolate the model assumption from other factors such as clustering across 13 examinations or the composite endpoint that counts 'similar' as success. The proposed travel-time comparison is the decisive test: it directly quantifies whether refraction-induced TOF errors can reach the level at which coherence maximization would select a wrong sound speed. Until that is measured, the central methodological claim remains plausible but not fully established. Since the reader's verdict is already CONDITIONAL, my analysis does not change the verdict.","tokens_in":12336,"tokens_out":7614,"duration_ms":83438,"concrete_test":"Simulate a realistic abdominal-wall model in k-Wave (curved 1450 m/s fat layer over 1580 m/s muscle, with the Table 1 aperture and a 1540 m/s focused transmit). For each pixel and receive element, compute travel times with (i) the straight-ray model of Eq. (2) and (ii) an eikonal/ray-tracer through the true heterogeneous map. Compare the maximum TOF discrepancy with the TOF difference caused by changing the candidate sound speed by 10 m/s over the same path. If the refraction-induced discrepancy exceeds that threshold in any clinically relevant region, the straight-ray assumption is measurably violated and the coherence selector is biased. As a second check, re-run the mixed in vivo cases G and J with a per-element delay correction and see whether the defocused regions are resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—focusing delays can be computed directly from a per-pixel average sound speed map without local inversion—rests on Eqs. (1)–(4), where every path is a straight ray and one pixel-wide c_avg(r) is applied to all transmit/receive element paths. Refraction is explicitly treated as second-order (§2.1), but the clinical cohort deliberately includes obese patients, and the paper itself reports Examples G and J where correction sharpens some structures and defocuses others. The mechanism is the aberration ambiguity illustrated in Fig. 2: a lens-like aberration (e.g., a curved fat–muscle interface) produces received signals similar to a global sound-speed error, so coherence maximization will absorb the refraction into a biased c_avg instead of correcting it. Whether this concern actually lands depends on the magnitude of refraction-induced travel-time errors relative to the discrimination step of the estimator (10 m/s over the relevant path). If the TOF error from refraction exceeds the TOF difference between adjacent candidate sound speeds, the coherence maximum can select a wrong sound speed, precisely the mixed in vivo behavior reported. The in-silico validation does not test this assumption because its ground-truth average sound speed is computed with the same straight-ray model, Eq. (1), making the simulation internally consistent rather than physically validating.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a coherence-based sound speed aberration correction method for fetal ultrasound. Instead of solving the ill-posed local sound speed inversion, the method estimates a per-pixel average sound speed by beamforming the same channel data at 21 constant candidate sound speeds, selecting for each pixel the speed that maximizes a coherence factor, and then beamforming with the resulting spatially varying average sound speed map. The method is tested in K-wave simulations and on a CIRS phantom, and clinically evaluated on 172 fetal B-mode images with three expert clinicians. The authors report that the corrected images are preferred or rated equivalent to the uncorrected 1540 m/s images in 72.5% of evaluations, and that a Tenengrad sharpness increase correlates with the estimated sound speed deviation (Pearson r = 0.67).","tokens_in":12606,"tokens_out":3769,"duration_ms":37484,"significance":"If the central claims are sustained, the paper offers a practically relevant contribution: retrospective two-way aberration correction that avoids the ill-posed inversion of local sound speed, applied to obstetric imaging with a 2D matrix array. The use of REFoCUS beamforming, the in vitro demonstration that the estimate is largely independent of the transmit sound speed assumption, and the relatively large clinical reader dataset are genuine strengths. The paper also makes its processing pipeline available through the vbeam library and a project website, which supports reproducibility. However, the physical validity of the straight-ray model is the load-bearing assumption, and the current evidence is partly circular and does not yet fully establish the claimed clinical preference.","major_comments":[{"comment":"The in-silico validation is circular with respect to the central modeling assumption. The ground-truth average sound speed maps are computed from the true local maps using Eq. (1), which is exactly the straight-ray harmonic-mean model assumed by the estimator in the TOF expression of Eq. (4). The reported MAEs (4.78, 4.74, and 4.31 m/s) therefore demonstrate that the estimator recovers what the model assumes, not that the model accurately describes wave propagation in tissue. The authors should either provide a simulation ground truth computed without Eq. (1) (e.g., via full-wave travel-time measurements or a different propagation model) or explicitly restrict the claim to internal consistency of the estimator.","section":"Section 3.1.1, Eqs. (1)-(4)"},{"comment":"The aberration ambiguity and the straight-ray assumption are acknowledged in the text and illustrated in Fig. 2, but the magnitude of the resulting error is never quantified. The estimator uses a 10 m/s candidate grid, yet no estimate is given for the travel-time error caused by refraction or frequency-dependent phase aberration relative to the TOF difference corresponding to a 10 m/s step. This matters because Examples G and J in Section 5 show exactly the failure mode predicted by the ambiguity: correction sharpens some structures and defocuses others, which is consistent with coherence maximization selecting a biased average sound speed. The paper should provide a quantitative bound, a simulation with refraction, or an additional validation criterion to rule out this bias.","section":"Section 2.1 and Section 5"},{"comment":"The headline clinical result of 72.5% is a composite of 'prefer corrected' (50.6%) and 'similar quality' (21.9%), but no confidence interval or formal statistical test is reported. The ratings are also clustered: 172 images are evaluated by three clinicians, and the evaluators use the 'similar' option very differently (6 vs. 77 times), so treating the 516 evaluations as independent overstates precision. The authors should report a confidence interval for the composite proportion, account for clustering by image and evaluator, and separately analyze the forced-choice preference rate between corrected and uncorrected images.","section":"Section 3.2.1, Table 4"},{"comment":"The Tenengrad sharpness comparison may be confounded by the gain normalization described in Section 5, where each corrected image is 'gained to have the same 80th percentile pixel intensity value as the 80th percentile value of the corresponding uncorrected image.' If this normalization is applied before computing the lateral Sobel magnitude in Eq. (10), then F depends on the global gain, and the reported κ could reflect changes in brightness rather than focusing alone. The authors should state the exact order of operations and, if normalization is applied before F, repeat the analysis with and without normalization.","section":"Section 3.2.2 and Section 5"}],"minor_comments":[{"comment":"The abbreviation MAD is used both for 'Mean Abdominal Diameter' in the clinical view list and for 'Mean Absolute Deviation' in the sharpness analysis; this is confusing and should be disambiguated with different symbols.","section":"Section 3.2.2"},{"comment":"The caption states that the 'notch indicate 95% confidence' and that the median of 'similar quality' is 'statistically different by the Wilcoxon rank sum test,' but no p-values or test details are given in the text or caption; please add them.","section":"Figure 8"},{"comment":"The probe name is written inconsistently as 'eM6c' and 'eM6C'; please use a single spelling.","section":"Throughout"},{"comment":"The row 'Corrected preferred' uses values 3, 2, 3, 1, etc., that represent counts of clinicians, but this is not stated in the table caption; please clarify the units of the counts.","section":"Table 5"},{"comment":"The K-wave simulation is described as imaging 'a speckle scene with point targets,' but the exact scatterer configuration is not given; a brief description or figure of the simulated medium would help reproducibility.","section":"Section 3.1.1"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the manuscript's main contribution relative to prior work by Ali et al. is incremental in method but adds a substantial clinical dataset and a clinically relevant application. The in-silico circularity and the unquantified straight-ray/refraction assumption are the main correctness risks. If the authors can supply an independent simulation or a quantitative refraction-error bound, and if the clinical statistics are reported with confidence intervals and clustering, the paper could become acceptable. The clinical preference claim should be framed as exploratory given the relatively small number of patients (12) and the composite endpoint."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nShort version: this is a real, clinically grounded attempt at distributed sound speed aberration correction. The core idea—estimating a per-pixel average sound speed map by coherence maximization over constant sound speeds and using it to compute focusing delays directly, without solving the ill-posed local inversion—is a genuine step beyond prior work like Ali et al. The pipeline is clearly described, and the rank filtering and pixel-grid compensation are sensible fixes for known sidelobe and scan-grid artifacts. The authors are honest about failure cases: Examples G and J are discussed as mixed corrections where clinicians preferred the uncorrected image. Few papers in this area have 172 clinical images from a deliberately high-BMI cohort, so the in vivo data is valuable.\n\nWhat the paper does well: the in vitro phantom experiment showing insensitivity to the transmit sound speed is a meaningful external check, and the public vbeam implementation supports reproducibility. The Tenengrad sharpness metric is simple and clearly defined, and the PCC of 0.67 is a reasonable signal, not a ground-breaking one. The citation pattern is appropriate, with clear positioning against the distributed aberration correction literature and the clinical observation from Chauveau et al.\n\nSoft spots, in proportion. The in-silico validation is partially circular: the ground-truth average sound speed is computed with Eq. (1), the same straight-ray model the estimator assumes, so agreement in MAE 4.3–4.8 m/s mostly confirms internal consistency rather than physical accuracy. The straight-ray assumption itself is the load-bearing premise; the paper never quantifies refraction-induced travel-time errors against the 10 m/s search step used in coherence maximization. In an obese cohort, refraction from curved fat–muscle interfaces is not obviously second-order, and the mixed in vivo results in G and J are the kind of behavior you would expect when the coherence peak lands on a biased sound speed. The clinical evaluation is statistically thin: 172 images from 12 patients are clustered, the three clinicians used the 'similar quality' label very differently (77 vs 6 images), and there are no confidence intervals or mixed-effects modelling. The headline 72.5% is 'preferred or equivalent', which hides that 27.5% were explicitly preferred uncorrected and that 'equivalent' is not a strong endorsement.\n\nDespite these caveats, the central claim is plausible and the paper does not overclaim: the conclusion explicitly says 'feasible' and 'preferred or similar', not 'superior'. I would send this to a serious referee—preferably someone who can push on the refraction assumption—rather than desk-reject. It is not a definitive clinical proof, but it is a useful contribution with honest reporting.\n\nFor a reading group, maybe; I would mention it to students working on aberration correction, but it is not a must-read for the whole field.\n\nBest,\n[Your name]","headline":"A serious, clinically grounded sound-speed aberration correction paper whose headline 72.5% is a weaker composite than it looks; worth reviewing, but the straight-ray assumption needs quantification.","tokens_in":13189,"tokens_out":3752,"would_cite":true,"duration_ms":34422,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Estimating a distributed average sound speed map directly from channel data — without solving the ill-posed local sound speed inversion — corrects focusing delays and improves fetal B-mode image quality in a 172-image clinical evaluation.","keywords":["sound speed aberration correction","coherence factor","distributed average sound speed","fetal ultrasound","REFoCUS beamforming","rank filtering","Tenengrad sharpness","clinical validation"],"falsifier":"On a phantom with a known strongly refracting layer (for example a wedge of material near 1450 m/s over a 1540 m/s background) and a grid of point targets, the assumption would be settled: if coherence maximization over the 1440-1640 m/s candidates yields a speed map whose corrected image is sharp at some targets but visibly defocused at others, or whose global estimate is biased well beyond the roughly 5 m/s in-silico error, then refraction cannot be treated as second order.","tokens_in":12144,"feed_emoji":"🩺","tokens_out":10185,"duration_ms":92663,"temperature":0.7,"pith_summary":"Ultrasound image quality depends on the assumed speed of sound, and the standard 1540 m/s value is often wrong, especially in fetal imaging where maternal fat can lower the actual speed. This paper aims to show that the resulting blur can be corrected without solving the ill-posed problem of reconstructing a local sound speed map: the authors beamform the same stored channel data at 21 candidate sound speeds, measure a coherence factor per pixel (how well the delayed channel signals agree in phase), and keep the sound speed that maximizes coherence for each pixel. That produces a distributed average sound speed map, and the focusing delays are computed directly from that map using a straight-ray travel-time model. The paper validates the map against ground truth in simulations and a phantom, then applies the correction to 172 fetal B-mode images from a commercial matrix-array system; three clinicians preferred the corrected image or rated it equivalent to the uncorrected 1540 m/s image in 72.5% of evaluations, and a lateral-sharpness metric rose with the estimated deviation from 1540 m/s (Pearson r = 0.67). If the claim holds, stored fetal channel data can be re-focused with patient-specific sound speed maps, and the clinical default of 1540 m/s may be too high for many pregnancies.","feed_headline":"Coherence maps sharpen fetal ultrasound in 172-image study","feed_subtitle":"Corrected images beat or matched the 1540 m/s standard in 72.5% of clinician evaluations.","key_machinery":"The load-bearing object is the distributed average sound speed map $c_{\\mathrm{avg}}(\\mathbf{r})$ of Eq. (1), formed from the one-way harmonic averages $c_{\\mathrm{har}}(\\mathbf{r}, e)$ of Eq. (2) along straight rays between each array element and each pixel. Estimation uses the coherence factor $C_R$ from Eq. (7), computed per pixel at 21 candidate sound speeds; a 90th-percentile rank filter expands the main lobe of the point-spread function so pixels near strong scatterers still pick the correct speed, and pixel-grid compensation (Eqs. 5-6) keeps structures stationary while the sound speed changes. REFoCUS retrospective beamforming makes the speed estimate independent of the transmit sound speed and focus, and the final B-mode is formed with the same travel-time expression using the estimated map.","core_discovery":"The paper's central claim is that the two-way travel-time aberration correction can be computed directly from the observed distributed average sound speed, bypassing the ill-posed local sound speed inversion. The map used for focusing is $c_{\\mathrm{avg}}(\\mathbf{r})$, the mean over all transmit and receive elements of the one-way harmonic average sound speed along straight rays (Eqs. 1-2), and it is estimated pixel-by-pixel by beamforming the channel data at 21 constant speeds and selecting the speed that maximizes the coherence factor $C_R$ (Eq. 7). The authors report that the estimate is unbiased against the sound speed assumed on transmission, matches the true average map in simulation (MAE about 4-5 m/s) and in a 1540 m/s phantom (MAE about 1 m/s), and, when used to re-beamform 172 fetal B-mode images, increases measured sharpness and is preferred or rated equivalent to the standard 1540 m/s image by three clinicians in 72.5% of evaluations.","pith_inferences":["In high-BMI pregnancies, where thick fat layers cause stronger refraction and frequency-dependent phase aberration, the straight-ray assumption may bias the estimated average map; a natural next test is to compare this method with a ray-tracing local sound speed estimate on the same channel data.","The same coherence-maximization scheme, without local inversion, should transfer to other matrix-array applications; a testable prediction is that correction gains track the deviation of true tissue sound speed from 1540 m/s.","The 'similar quality' category was used very differently by the three evaluators (77, 30, and 6 images), so a standardized equivalence definition would be needed before multi-site deployment of the method.","Because images with estimated sound speed close to 1540 m/s were mostly rated 'similar', the 72.5% preference/equivalence figure likely understates the benefit in patients with larger sound speed deviations; stratifying by estimated global sound speed would test this."],"forward_implications":["Stored fetal channel data can be re-beamformed with patient-specific sound speed maps, so the correction does not require a new scan or a change in acquisition.","The default fetal ultrasound sound speed may need to be lowered from 1540 m/s toward 1500 m/s, matching the estimated global values in this cohort and earlier clinical findings.","Because the estimator is unbiased against the transmit sound speed, data acquired on scanners that focused at 1540 m/s can still be corrected retrospectively.","A lateral Tenengrad sharpness metric can serve as an objective proxy for focusing improvement, with a Pearson correlation of 0.67 between sharpness gain and estimated sound speed deviation.","Per-pixel sound speed maps can be smoothed with depth-dependent kernels because the averaging in Eq. (2) makes the true map smooth, which stabilizes the per-pixel selection."],"supporting_citations":[{"why":"Supplies the tomographic distributed aberration correction approach that the paper bypasses, and a baseline for comparison.","marker":"[10]"},{"why":"Provides the prior local sound speed estimation method in layered media and the simulation comparison for the estimated average map.","marker":"[12]"},{"why":"REFoCUS beamforming is the technique that makes the sound speed estimate independent of the transmit focus assumption, enabling two-way retrospective correction.","marker":"[16]"},{"why":"Provides the pixel-grid compensation used to remove the bias in sound speed estimation when the time-to-space mapping changes with sound speed.","marker":"[13]"},{"why":"Defines the coherence factor $C_R$ used as the per-pixel focusing quality metric that selects the optimal sound speed.","marker":"[20]"},{"why":"Supplies the rank/percentile filtering method that expands the main lobe in coherence images and suppresses off-axis scatterer artifacts.","marker":"[22]"},{"why":"Clinical study showing fetal images focused at 1480 m/s were graded higher than at 1540 m/s, supporting the paper's lower default sound speed finding and the clinical relevance of the correction.","marker":"[4]"},{"why":"K-Wave simulation toolbox used to generate the in silico channel data and ground-truth sound speed maps for validation.","marker":"[24]"},{"why":"Tenengrad focus measure from digital photography, adapted with a lateral Sobel filter, provides the quantitative sharpness metric for clinical images.","marker":"[26]"}],"fun_headline_variants":["Coherence-based sound speed fix boosts fetal ultrasound clarity","Fetal ultrasound sharpened by coherence-optimized sound speed","Adaptive sound speed from coherence beats standard in fetal scans","Sound speed from coherence: 72.5% of fetal scans improved or equal"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes ultrasound travels in straight lines without bending (refraction), so every focusing error is read as a speed-of-sound error; if ray bending through fat or other tissue is significant, the estimated speed map will be biased even where coherence is maximized.","fun_headline_variants_meta":{"raw":{"variants":["Coherence-based sound speed fix boosts fetal ultrasound clarity","Fetal ultrasound sharpened by coherence-optimized sound speed","Adaptive sound speed from coherence beats standard in fetal scans","Sound speed from coherence: 72.5% of fetal scans improved or equal"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000521,"raw_usage":{"total_tokens":2560,"prompt_tokens":1020,"completion_tokens":1540,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":636,"completion_tokens_details":{"reasoning_tokens":1468}},"tokens_in":636,"tokens_out":1540,"duration_ms":10685,"temperature":1.0,"reasoning_tokens":1468,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:59:07.470839+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On a phantom with a known strongly refracting layer (for example a wedge of material near 1450 m/s over a 1540 m/s background) and a grid of point targets, the assumption would be settled: if coherence maximization over the 1440-1640 m/s candidates yields a speed map whose corrected image is sharp at some targets but visibly defocused at others, or whose global estimate is biased well beyond the roughly 5 m/s in-silico error, then refraction cannot be treated as second order.","supporting_citations":[{"cited_title":"Distributed aberration correction tech- niques based on tomographic sound speed esti- mates,","cited_arxiv_id":null,"evidence_quote":"Supplies the tomographic distributed aberration correction approach that the paper bypasses, and a baseline for comparison."},{"cited_title":"Local sound speed estimation for pulse-echo ultrasound in layered media,","cited_arxiv_id":null,"evidence_quote":"Provides the prior local sound speed estimation method in layered media and the simulation comparison for the estimated average map."},{"cited_title":"Sound speed and vir- tual source correction in synthetic transmit focus- ing,","cited_arxiv_id":null,"evidence_quote":"Provides the pixel-grid compensation used to remove the bias in sound speed estimation when the time-to-space mapping changes with sound speed."},{"cited_title":"Method and apparatus for coher- ence filtering of ultrasound images,","cited_arxiv_id":null,"evidence_quote":"Defines the coherence factor $C_R$ used as the per-pixel focusing quality metric that selects the optimal sound speed."},{"cited_title":"Properties, implementations and appli- cations of rank filters,","cited_arxiv_id":null,"evidence_quote":"Supplies the rank/percentile filtering method that expands the main lobe in coherence images and suppresses off-axis scatterer artifacts."},{"cited_title":"Im- proving image quality of mid-trimester fetal sonog- raphy in obese women: Role of ultrasound propa- gation velocity,","cited_arxiv_id":null,"evidence_quote":"Clinical study showing fetal images focused at 1480 m/s were graded higher than at 1540 m/s, supporting the paper's lower default sound speed finding and the clinical relevance of the correction."},{"cited_title":"A comparison of different focus functions for use in autofocus algorithms,","cited_arxiv_id":null,"evidence_quote":"Tenengrad focus measure from digital photography, adapted with a lateral Sobel filter, provides the quantitative sharpness metric for clinical images."}],"review_version":1}