{"id":"6d649a88-b484-43db-99d2-33e04b5a3781","arxiv_id":"2509.01557","paper_version":2,"verdict":"REJECT","confidence":"LOW","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"An image-domain latent diffusion model suppresses HIFU-induced ultrasound interference at up to 15 frames per second, beating a Notch Filter on an 18,802-pair in vitro/ex vivo/in vivo dataset.","lead":"This paper trains a deep-learning model to erase the acoustic interference that HIFU therapy creates in live ultrasound guidance images, working directly on the final pictures rather than raw radiofrequency signals. It removes the need for synchronized hardware and aims to make continuous ultrasound-based HIFU monitoring practical.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's mHC-Diff claim (26.65 dB, ~20 FPS, 6.8x speedup, teacher-student) appears nowhere in the full text, which describes only HIFU-ILDiff with DDIM K=5/30, lower PSNR, and slower FPS.","rationale":"The reader's verdict was REJECT, and I agree that the paper should not be accepted as-is. However, the reader's designated weakest_assumption was the paired-reference ground truth (Section 2.5), whereas the most load-bearing concern is more fundamental: the abstract and the full text describe different methods with different headline numbers. The central claim as submitted—mHC-Diff's superiority and real-time performance—cannot be evaluated because no derivation, architecture, or experiment for mHC-Diff exists in the manuscript. This is not a disagreement with consensus; it is an internal missing-support failure. The body's HIFU-ILDiff contribution is real and defensible, but it does not support the abstract's specific claims. I set agreement_with_reader to 'partial' because the reader's rationale lists the mismatch prominently, even though their weakest_assumption field focuses on the paired-reference issue. The concrete test would settle the mismatch definitively: a simple string search and timing recalculation. If the search somehow found the mHC-Diff content in the body (it does not, as provided), the verdict would change; absent that, REJECT remains appropriate.","tokens_in":14526,"tokens_out":2511,"duration_ms":26758,"concrete_test":"Perform exact-string searches in the submitted PDF and the public GitHub repository for: 'mHC-Diff', 'teacher', 'student', 'one-step', 'distillation', '26.65', '6.8', and 'HIFU-Diff'. Additionally, recompute FPS from Table 6: 1000/75 = 13.3 FPS and 1000/342 = 2.9 FPS. If none of these strings appear in the body or code, the abstract's claim has no supporting derivation; if the repository contains an mHC-Diff implementation, that does not rescue the paper, which still lacks the corresponding method description and evaluation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The submitted central claim is the abstract's description of mHC-Diff: a teacher-student distilled one-step diffusion model achieving 26.65 dB PSNR, ~20 FPS on an RTX 4090, and a ~6.8x speedup over HIFU-Diff. However, the full text never mentions mHC-Diff, teacher-student distillation, one-step inference, 26.65 dB, or HIFU-Diff. Section 2.3 defines HIFU-ILDiff, a VQ-VAE latent diffusion model with DDIM sampling at K=5 or K=30. Section 3.2 reports best PSNR 24.562 dB (Table 4); Section 3.4 reports 75 ms/frame (K=5) and 342 ms/frame (K=30), corresponding to 13.3 FPS and 2.9 FPS. The abstract's architecture, numbers, and baseline are not derivable from any experiment or derivation in the body. Even the body's own 15 FPS claim is contradicted by Table 6: 75 ms/frame is 13.3 FPS, below the 15 FPS acquisition rate. Thus the central claim as submitted is unsupported by the manuscript's content; the body's plausible HIFU-ILDiff results cannot validate the abstract's mHC-Diff claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The submission's abstract describes a teacher-student distilled diffusion framework, mHC-Diff, for real-time HIFU interference suppression, claiming 26.65 dB PSNR, ~20 FPS on an RTX 4090, and a ~6.8x speedup over iterative diffusion baselines such as HIFU-Diff. The full text, however, describes a different method, HIFU-ILDiff, a VQ-VAE-based latent diffusion model with DDIM sampling at K=5 or K=30. The full text presents a paired dataset of 18,802 ultrasound image pairs from phantoms, ex vivo tissues, and in vivo rabbit, with held-out-subject splits, and reports quantitative gains over a Notch Filter baseline (e.g., SSIM 0.796 vs. 0.443, PSNR 23.780 vs. 14.420 in phantoms). It also reports inference times of 75 ms/frame (K=5) and 342 ms/frame (K=30) and claims real-time processing at 15 FPS. None of the abstract's central elements--mHC-Diff, teacher-student distillation, one-step inference, 26.65 dB, HIFU-Diff, or the 6.8x speedup--appear anywhere in the full text.","tokens_in":1750,"tokens_out":1808,"duration_ms":56560,"significance":"If the full text's HIFU-ILDiff results are valid, the underlying contribution is potentially useful: an image-domain latent diffusion approach that avoids RF data and hardware synchronization, supported by a sizeable multi-subject, multi-modality dataset and public code. The held-out-subject evaluation and the consistent visual and quantitative improvements over the Notch Filter are notable strengths. However, the submission as a whole is built around the abstract's mHC-Diff claim, which is entirely unsupported by the body. The mismatch is not a minor presentational issue; it changes the claimed architecture, training paradigm, and performance numbers. As submitted, the central claim of the paper is not verifiable from the manuscript's content, and the full text's own real-time claim is internally inconsistent. The body might form the basis of a legitimate separate paper, but the current submission does not support its stated headline contribution.","major_comments":[{"comment":"The abstract describes 'mHC-Diff' with 'Manifold-Constrained Hyper-Connections', a two-stage teacher-student distillation, a one-step student, 26.65 dB PSNR, ~20 FPS on an RTX 4090, and a ~6.8x speedup over 'HIFU-Diff'. None of these terms, numbers, or architectural components appear in the full text. The full text defines HIFU-ILDiff in Section 2.3, uses DDIM with K=5 or K=30, and reports best PSNR 24.562 dB (Table 4) and inference times 75/342 ms/frame (Section 3.4, Fig. 9). The abstract's central claim is therefore unsupported by any experiment or derivation in the body. This is a load-bearing inconsistency: the contribution claimed in the abstract cannot be validated by the manuscript's content.","section":"Abstract vs. full text"},{"comment":"The full text claims real-time processing at 15 frames per second (Abstract, Section 4, Section 5), but Section 3.4 reports 75 ms/frame for K=5, which is 13.3 FPS, below the 15 FPS acquisition rate for Diverging Wave and Line Scan imaging (Section 2.5). Even the body's own 15 FPS claim is contradicted by its reported inference time. This undermines the central 'real-time' assertion independent of the abstract mismatch.","section":"Section 3.4 / Fig. 9 / Discussion"},{"comment":"Paired training/evaluation assumes the HIFU-off frame immediately following two HIFU-on frames is the ground truth for those contaminated frames. Over the 200 ms HIFU-off interval, in vivo respiratory/cardiac motion and speckle decorrelation change the underlying scene; the reference is not the true clean version of the contaminated frame. The Discussion concedes the additive-noise degradation model is approximate, but it does not address the temporal pairing issue. Reported SSIM/PSNR values therefore partly measure temporal mismatch rather than pure restoration accuracy. A sensitivity analysis (e.g., evaluating against a different HIFU-off frame, or using synthetic data with known ground truth) is needed to support the claimed restoration quality.","section":"Section 2.5"},{"comment":"The only quantitative baseline is the Notch Filter, and Section 2.8 justifies this by stating that existing deep-learning methods rely on RF data. However, references [10]-[12] describe learnable methods, and Section 4 asserts superiority over 'previous methods' without quantitative comparison. For a claim of state-of-the-art suppression, at least one comparison to an image-domain or diffusion-based method (or a clear scoping of the claim) is necessary. As it stands, the 'significantly outperforms' claim is supported only against a classical signal-processing baseline.","section":"Section 2.8 / Table 4"}],"minor_comments":[{"comment":"The abstract in the full text (and Introduction contributions) states 18,872 image pairs, but Section 2.6 and Table 1 report 18,802 frames. The discrepancy should be resolved.","section":"Section 2.6 vs. Abstract"},{"comment":"The attention equation is garbled: 'SoftmaxU𝑄𝐾4-𝑑5W𝑉' should be formatted properly. Also, 'attention head channels of 32' in Section 2.7 is unclear; specify number of heads or head dimension.","section":"Section 2.4"},{"comment":"The text references 'Fig. 11' for inference time comparison, but the manuscript contains only Fig. 9. Check cross-references.","section":"Section 3.4"},{"comment":"The checkmark/backslash notation is ambiguous, especially for Method 1 (VQ-VAE with backslash in noise-predictor column). Clarify what backslash means; also the reported inference times (4 ms, 11 ms) seem to conflict with Section 3.4 values; explain whether these are VQ-VAE-only timings.","section":"Table 6"},{"comment":"Bootstrap confidence intervals are mentioned, but no implementation details (number of resamples, stratification) are given. The 'variance below 0.1' statement in Section 4 needs a clearer definition, as PSNR/SSIM variances are not reported in Table 4.","section":"Section 3.2 / Table 4"}],"recommendation":"reject","confidential_remarks":"The submission appears to combine two different papers: the abstract describes a teacher-student distilled one-step model (mHC-Diff) with specific performance numbers, while the full text describes a non-distilled latent diffusion model (HIFU-ILDiff) with different numbers and no distillation. This is not a local error; the abstract's central claims are entirely absent from the body. Even the body's own real-time claim is internally inconsistent with its reported inference time. The dataset and held-out-subject protocol in the body are valuable, and a corrected manuscript focused on HIFU-ILDiff with proper baselines and a resolved real-time claim could be resubmitted. The editor may wish to verify the integrity of the submission, as the abstract/body mismatch goes beyond typical revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. First, the paper you'd read in the body is not the paper promised in the abstract. The abstract describes mHC-Diff, a teacher-student distilled one-step diffusion model hitting 26.65 dB PSNR, ~20 FPS, and a 6.8x speedup over HIFU-Diff. The full text describes HIFU-ILDiff, a VQ-VAE latent diffusion with DDIM at K=5 or 30, reporting 23.78–24.56 dB and, at best, 13.3 FPS (75 ms/frame). Teacher-student distillation, one-step inference, and the HIFU-Diff baseline appear nowhere in the body. The stress-test note is accurate; this is not a minor terminology slip.\n\nSecond, the body itself is a credible applied contribution. The dataset is the real asset: 18,802 image pairs across in vitro, ex vivo, and in vivo subjects, three imaging modes, four HIFU power levels, with held-out-subject splits. The core result — a latent diffusion model suppressing HIFU interference in the image domain on same-pipeline data — is defensible from the reported experiments. The improvement over the Notch Filter is large and consistent, and the ablations on diffusion steps and self-attention placement are informative.\n\nSoft spots in proportion. The only baseline is a Notch Filter. The paper cites FUS-Net, a 1D U-Net, and its own earlier pixel-space diffusion (ref 12) but compares against none of them; the stated reason — that DL methods rely on raw RF data and are complicated to implement — is weak, especially for its own prior work. The real-time claim is internally inconsistent: Table 6 gives 342 ms at K=30 and 75 ms at K=5, which is 2.9 and 13.3 FPS, below the claimed 15 FPS. The paired ground-truth assumption (HIFU-off frame right after HIFU-on frames) is not stress-tested for respiratory or cardiac motion, and the paper admits the additive-noise degradation model is approximate. Dataset screening/filtering is mentioned but not detailed. The downstream WUE validation is qualitative and uses the group's own method.\n\nNet: the body's science is worth a referee's time, but the submission as it stands is not. The abstract and body need to be reconciled, and the baselines and timing claims need fixing before this can be fairly reviewed. If it's a versioning accident, resubmission with the correct abstract is the obvious path. Don't desk-reject the underlying work, but don't trust any of the headline numbers.","headline":"The body is a plausible applied diffusion paper with a real dataset, but the abstract describes a different system whose numbers never appear in the text — as submitted, the headline claim is unsupported, though the underlying work deserves referee time.","tokens_in":15378,"tokens_out":2285,"would_cite":false,"duration_ms":23647,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A diffusion model that cleans HIFU interference directly from ultrasound B-mode images can run at real-time frame rates on a standard GPU, without RF data or hardware synchronization.","keywords":["HIFU interference suppression","latent diffusion model","ultrasound image restoration","real-time inference","teacher-student distillation","VQ-VAE","B-mode ultrasound","image-domain denoising"],"falsifier":"On a stationary phantom with a fixed geometry, record HIFU-on frames and HIFU-off references at several delays after switch-off, for example 0 ms, 50 ms, and 200 ms. If model outputs compared against the 200-ms settled reference show substantially lower SSIM/PSNR than against the immediate next frame, then part of the reported gain is an artifact of the paired-frame protocol rather than genuine restoration of anatomy.","tokens_in":14206,"feed_emoji":"🩺","tokens_out":6975,"duration_ms":73913,"temperature":0.7,"pith_summary":"This paper is trying to establish that a diffusion model can remove HIFU-induced acoustic interference from ultrasound images in real time, using only the B-mode images themselves—no proprietary RF data, no extra hardware synchronization. If true, it would let clinicians monitor HIFU treatment continuously while the therapy transducer is firing, instead of pausing therapy to take clean images. The central claim is that an image-domain latent diffusion model, and its distilled one-step successor, can both restore image quality far beyond the standard notch filter and run fast enough for live guidance. The paper builds a large paired dataset across imaging modes, power levels, and phantom, ex vivo, and in vivo tissues, and validates restoration quality with SSIM/PSNR, a radiologist-graded scale, and a downstream entropy-imaging check.","feed_headline":"Image diffusion clears HIFU interference at 20 frames per second","feed_subtitle":"No RF data or hardware sync needed; a distilled diffusion model restores B-mode images on a standard GPU.","key_machinery":"The load-bearing element is the latent space: a VQ-VAE encoder maps contaminated and clean ultrasound frames into a compact discrete latent representation, where an attention-based U-Net noise predictor runs iterative denoising conditioned on the contaminated frame's latent. DDIM sampling lets the same model trade speed for fidelity by choosing a small or large number of reverse steps. In the teacher-student formulation, a multi-step UNet teacher supplies high-fidelity prior knowledge that is distilled into a one-step student with manifold-constrained hyper-connections, which is what converts the iterative pipeline into a real-time single forward pass.","core_discovery":"The paper claims that HIFU interference—the acoustic contamination that appears on ultrasound guidance images while the therapy transducer is firing—can be suppressed by a diffusion model that works directly on B-mode images, with no need for raw RF data or hardware synchronization. In the full text, the proposed HIFU-ILDiff encodes the contaminated frame into a VQ-VAE latent space, runs a conditional latent diffusion denoising, and decodes a clean frame, reaching 15 frames per second while lifting in-vitro SSIM/PSNR from 0.443/14.42 (notch filter) to 0.796/23.78. The abstract goes further with mHC-Diff, a teacher-student distilled version that runs at roughly 20 FPS on an RTX 4090 and repor","pith_inferences":["Because the paired ground truth is the frame captured right after HIFU is switched off, the model may also be learning to undo motion and speckle decorrelation between those frames; a motion-controlled phantom study would separate true interference suppression from temporal smoothing.","The teacher-student distillation result suggests that for this class of restoration, a one-step student can match or exceed the teacher's reported PSNR, implying the bottleneck may be data and pairing quality rather than sampling budget.","The same image-domain architecture could be applied to other therapeutic-ultrasound interference sources, such as cavitation or heating artifacts, or to other real-time image restoration tasks where raw RF data is unavailable.","If real-time suppression holds across human clinical datasets, the practical workflow of HIFU ablation could shift from intermittent imaging with therapy pauses to continuous ultrasound guidance, pending regulatory validation."],"forward_implications":["Ultrasound-guided HIFU can be monitored continuously during sonication, instead of pausing therapy to acquire clean frames.","The method works on standard B-mode images, so it can be used with existing ultrasound systems without proprietary RF access or extra hardware synchronization.","The sampling-step knob lets clinicians trade a little image quality for higher frame rate, adapting to different latency requirements.","The denoised frames retain enough acoustic information for downstream quantitative imaging, such as weighted ultrasound entropy imaging, to delineate the treatment zone.","The distilled student achieves a substantial speedup over iterative diffusion baselines, making near-real-time deployment on a single clinical GPU plausible."],"supporting_citations":[{"why":"The authors' earlier diffusion-based HIFU interference suppression model that this work extends into latent space and distillation.","marker":"[12]"},{"why":"Provides the latent diffusion model framework that enables efficient high-resolution image generation in a compressed space.","marker":"[17]"},{"why":"Supplies the VQ-VAE encoder/decoder architecture used to map ultrasound images into a discrete latent space.","marker":"[21]"},{"why":"Multi-head attention mechanism incorporated into the noise predictor to capture long-range dependencies in HIFU-affected regions.","marker":"[22]"},{"why":"Diffusion-based image super-resolution approach that motivates treating HIFU suppression as a super-resolution problem.","marker":"[15]"},{"why":"The earlier RF-based FUS-Net deep learning baseline that motivates the shift away from RF data to image-domain processing.","marker":"[10]"},{"why":"Weighted ultrasound entropy imaging used to validate that the denoised images retain acoustic properties for downstream monitoring.","marker":"[29]"},{"why":"SSIM metric used for quantitative evaluation of restoration quality against the HIFU-off reference frames.","marker":"[27]"}],"fun_headline_variants":["Teacher-student diffusion zaps HIFU noise at 20 FPS","No RF, no sync: diffusion clears HIFU interference in real time","Image-only diffusion cleans HIFU artifacts at 20 FPS","Distilled diffusion gives real-time HIFU de-noising on GPU","One-step diffusion denoises HIFU at 20 FPS, no RF needed"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The only clean reference for a contaminated frame is the frame captured right after HIFU is switched off; if tissue motion or speckle changes the scene during that gap, the reported SSIM/PSNR partly measure temporal consistency rather than true restoration, and the model's additive-noise view of interference is an approximation.","fun_headline_variants_meta":{"raw":{"variants":["Teacher-student diffusion zaps HIFU noise at 20 FPS","No RF, no sync: diffusion clears HIFU interference in real time","Image-only diffusion cleans HIFU artifacts at 20 FPS","Distilled diffusion gives real-time HIFU de-noising on GPU","One-step diffusion denoises HIFU at 20 FPS, no RF needed"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000697,"raw_usage":{"total_tokens":3023,"prompt_tokens":815,"completion_tokens":2208,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":559,"completion_tokens_details":{"reasoning_tokens":2113}},"tokens_in":559,"tokens_out":2208,"duration_ms":18089,"temperature":1.0,"reasoning_tokens":2113,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T12:26:31.082031+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On a stationary phantom with a fixed geometry, record HIFU-on frames and HIFU-off references at several delays after switch-off, for example 0 ms, 50 ms, and 200 ms. If model outputs compared against the 200-ms settled reference show substantially lower SSIM/PSNR than against the immediate next frame, then part of the reported gain is an artifact of the paired-frame protocol rather than genuine restoration of anatomy.","supporting_citations":[{"cited_title":"J Pak Med Assoc 68(9), 1378--80 (2018)","cited_arxiv_id":null,"evidence_quote":"The authors' earlier diffusion-based HIFU interference suppression model that this work extends into latent space and distillation."},{"cited_title":"IEEE Transactions on Medical Imaging 43(4), 1594--1604 (2023)","cited_arxiv_id":null,"evidence_quote":"Provides the latent diffusion model framework that enables efficient high-resolution image generation in a compressed space."},{"cited_title":"IEEE Transactions on Ultrasonics, Ferroelectrics, and Frequency Control 61(9), 1580--1587 (2014)","cited_arxiv_id":null,"evidence_quote":"Supplies the VQ-VAE encoder/decoder architecture used to map ultrasound images into a discrete latent space."},{"cited_title":"Frontiers in Physiology 16, 1602866 (2025)","cited_arxiv_id":null,"evidence_quote":"Diffusion-based image super-resolution approach that motivates treating HIFU suppression as a super-resolution problem."},{"cited_title":"Advances in neural information processing systems 33, 6840--6851 (2020)","cited_arxiv_id":null,"evidence_quote":"The earlier RF-based FUS-Net deep learning baseline that motivates the shift away from RF data to image-domain processing."},{"cited_title":"Frontiers in Physiology , volume=","cited_arxiv_id":null,"evidence_quote":"Weighted ultrasound entropy imaging used to validate that the denoised images retain acoustic properties for downstream monitoring."},{"cited_title":"International Journal of Hyperthermia , volume=","cited_arxiv_id":null,"evidence_quote":"SSIM metric used for quantitative evaluation of restoration quality against the HIFU-off reference frames."}],"review_version":1}