{"id":"e2dcb387-a0f8-408c-a05c-15870b83023b","arxiv_id":"2501.01157","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A Fourier neural operator trained on simulated and tiny real data reconstructs lung aeration maps from raw ultrasound RF signals, with 9.4% mean percent aeration error on ex vivo swine lungs.","lead":"A new AI model, Luna, reads raw ultrasound signals from pig lungs and turns them into a map showing where air and tissue are. If it works in people, it could make lung ultrasound reading more objective and reproducible.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 9.4% ex vivo error does not yet support the central claim: the ground-truth aeration is measured on a different lobe (right lung / left caudal) than the left cranial lobe that Luna scans, so the scanned region's true aeration is never measured.","rationale":"I read the paper as claiming that Luna can reconstruct lung aeration maps from RF data and estimate percent aeration in real ex vivo scans with 9.4% error. For that claim to hold, the reference aeration must describe the region being imaged. The paper's own methods section shows it does not: the bulk aeration used as ground truth comes from a different lobe than the scanned lobe, and the reconstructed map covers only about 0.8 mm depth while the reference spans the whole lobe. The reader's weakest_assumption identifies exactly this gap, and I agree it is the most load-bearing concern. The in-silico validation has matched ground truth but does not address domain transfer; the B-mode visual alignments in Fig. 3c are qualitative; and the calibration and split issues, while real, are secondary. This is not a claim that the model is wrong, only that the reported 9.4% is ambiguous as evidence for the central claim. A conditional verdict requiring lobe-matched ground truth, or explicit acknowledgment of the surrogate target, is appropriate. My read does not change the reader's conditional verdict, so the verdict is unchanged.","tokens_in":20196,"tokens_out":3399,"duration_ms":34520,"concrete_test":"Redo the ex vivo evaluation with ground-truth aeration measured on the exact scanned lobe: after scanning the left cranial lobe, measure its bulk aeration by the same fluid-displacement/de-aeration protocol instead of using the right lung or left caudal lobe; if the scan-level errors remain at or below 10%, the central claim survives. If that measurement is impossible, perform CT densitometry on the left cranial lobe and compare the predicted aeration against CT-derived aeration in the same ~0.8 mm subpleural band. In either case, also report animal-level train/fine-tune/evaluation splits to rule out leakage.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.2 states that 'Units obtained out of the right lung and caudal lobe of the left lung were employed in the calculation of its bulk aeration, while the left cranial lobe was used for scanning.' The reported 9.4% mean error on 103 scans is computed against a bulk-aeration value from a different lobe of the same animal, not from the left cranial lobe being scanned. For the central claim—that raw RF data contain enough information to estimate aeration in the scanned region—the target quantity is never measured. This assumes, without evidence, that aeration is homogeneous across lobes and that the 2.6-wavelength (~0.8 mm) reconstructed slice equals whole-lobe bulk aeration. The ARDS-like instillation protocol could easily produce regional differences, and aeration is known to vary with both lobe and depth. The in-silico error of 5.2% does not rescue this: simulation provides matched ground truth, but the ex-vivo validation is the only evidence for real-data generalizability, and it compares against an off-target reference. The lack of animal-level split disclosure compounds the ambiguity, but the lobe mismatch is the more fundamental gap: unless lobe correspondence and depth integration are validated, the headline error could reflect correlation with something other than the scanned region's aeration.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces LUNA, a Fourier neural operator pipeline that takes delayed ultrasound radio-frequency (RF) data as input and outputs a two-dimensional lung aeration map, from which a percent-aeration metric is computed. The model is first trained on 10,150 simulated RF/aeration-map pairs generated with the Fullwave-2 full-wave solver, then fine-tuned on 18 ex vivo swine lung scans, and finally evaluated on 103 further ex vivo scans. The authors report in silico map-reconstruction metrics (PSNR 20.10 ± 3.51 dB, SSIM 0.605 ± 0.128) and percent-aeration errors of 5.2% ± 4.6% in silico and 9.4% ± 5.4% ex vivo, and claim that this is the first direct reconstruction of lung aeration maps from RF data, bypassing beamforming and manual B-mode interpretation. The paper also claims superiority over the clinical LUS scoring system.","tokens_in":20420,"tokens_out":3976,"duration_ms":40538,"significance":"If the central claim were fully supported, this would be a valuable contribution to quantitative lung ultrasound: the use of raw RF data rather than B-mode images is a plausible way to preserve diagnostically relevant information and reduce reader dependence, and the proposed two-stage simulation-plus-fine-tuning strategy is pragmatic given the difficulty of collecting paired real RF and aeration data. The paper is transparent in several respects: it reports the exact 9.4% error with standard deviation, provides an ablation study of the network components, uses an experimentally validated full-wave simulator, and explicitly acknowledges in the Discussion that aligned ex vivo aeration maps are unavailable, so full verification of the 2D reconstructions is lacking. However, the ex vivo evaluation is undermined by a mismatch between the region whose aeration is measured and the region that is scanned, which means the headline 9.4% error is not a direct measurement of the quantity the paper's central claim concerns. The in silico validation is internally sound but cannot alone establish real-data generalizability for map reconstruction.","major_comments":[{"comment":"The ex vivo ground-truth aeration is computed from units of the right lung and the caudal lobe of the left lung, whereas the scans are of the left cranial lobe. The reported 9.4% mean error therefore compares LUNA's prediction for the scanned left cranial lobe against a bulk-aeration measurement of a different lobe of the same animal. The target quantity for the central claim, i.e., the true aeration of the scanned region, is never measured. The assumption that aeration is homogeneous across lobes is not justified, especially because the experimental protocol instills fluid into the airways to create ARDS-like regional deaeration. This is a load-bearing gap: the headline result does not, as stated, demonstrate that RF data contain sufficient information to estimate aeration in the scanned region.","section":"Section 4.2, 'Aeration Calculation' and Section 2, 'Real Ex-Vivo Data'"},{"comment":"The reconstructed map is limited to a depth of 2.6 ultrasound wavelengths (about 0.8 mm), while the ground-truth aeration is the whole-lobe bulk value obtained by fluid displacement. Even if the scanned lobe matched the measured lobe, the comparison would equate a thin superficial slice with a bulk average over the entire lobe. The paper provides no evidence that the reconstructed slice is representative of the whole lobe, and given known depth-dependent aeration gradients in injured lungs, this assumption is not safe. The in silico experiments cannot resolve this issue because they use matched 2D ground-truth maps rather than a slice-to-bulk comparison.","section":"Section 2, 'Real Ex-Vivo Data' and Section 4.2, 'Aeration Calculation'"},{"comment":"The paper does not state whether the 18 fine-tuning scans and the 103 evaluation scans came from disjoint animals, nor does it describe an animal-level split. If scans from the same animal appear in both the fine-tuning and evaluation sets, the evaluation could be optimistically biased through animal-specific memorization. In addition, the counts are inconsistent: Section 4.2 states that 124 scans were performed, with 118 used for assessment and 6 for calibration, while Table 1 lists 18 fine-tuning plus 103 evaluation, summing to 121. The relationship between these numbers should be clarified and an animal-level split should be reported.","section":"Table 1 and Section 4.2, 'Ex-Vivo Data Acquisition'"},{"comment":"The claim that LUNA 'outperforms the current lung ultrasound scoring system' is not supported by a quantitative comparison. The LUS score is an ordinal 0-3 scale assigned to B-mode images by human readers, not a percent-aeration measurement, and the paper presents no head-to-head comparison of LUNA against LUS scores on the same ex vivo data, nor any interobserver variability data for the LUS scoring. The statement in the Results that 9.4% error 'well outperforms the sensitivity of the current scoring system' is therefore an unsupported comparison. Either provide a direct comparison or temper the claim.","section":"Section 2, 'Luna has strong generalizability to real ex-vivo data' and Section 3"},{"comment":"The paper's own limitation statement says that the lack of aligned ex vivo aeration maps prevents full verification of the 2D reconstructions, and that calibration is not available on real data because true 2D maps are unavailable. Despite this, the abstract and introduction describe the contribution as the first reconstruction of lung aeration maps from RF data and present the ex vivo results as demonstrating robust performance. The central claim is therefore stronger than the evidence: the ex vivo results support percent-aeration estimation (under the lobe-match caveat above), but not spatially resolved 2D map accuracy on real tissue. The manuscript should either present an ex vivo validation that can confirm map reconstruction (e.g., via a surrogate such as coregistered CT or locally measured aeration) or explicitly restrict the claim to percent-aeration estimation.","section":"Discussion, Limitations, and Section 4.3, 'Model Calibration'"}],"minor_comments":[{"comment":"The text contains unresolved placeholders: 'focal depth varied from X to Y cm' and, in Section 4.1, 'Both flattened (X cases) and naturally curved (Y cases)'.","section":"Section 4.2, 'Ex-Vivo Data Acquisition'"},{"comment":"The phrase 'well-calibrated in silicon' is a typo for 'in silico'.","section":"Section 4.3, 'Model Calibration'"},{"comment":"The text in Section 2 refers to panels 4b, 4c, and 4d with somewhat inconsistent descriptions (e.g., the scatter plot versus box plot); please check that each citation points to the intended panel.","section":"Figure 4 and its caption"},{"comment":"The table lists aeration for the ex-vivo fine-tuning set as 33.4 without a standard deviation, while other entries include 'average ± SD'; please provide the missing standard deviation or explain its absence.","section":"Table 1"},{"comment":"The sentence 'It allowed axial column translation for flattening of the lung and its deformation to conform with various parietal pleura curvature' would be clearer as 'These segmentations allowed axial column translation...'.","section":"Section 4.1, 'Simulated Data Generation'"},{"comment":"The paper would benefit from a statement on ethics or institutional approval for the use of animal tissue; the Tissue Sharing Program is mentioned, but the relevant oversight or approval is not explicitly stated.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"This is a technically interesting paper with an honest reporting style, but the ex vivo validation mismatch between the measured lobe and the scanned lobe is a substantive issue that reviewers will likely focus on. The current claims are stronger than the evidence; a revision that either supplies lobe-matched validation or carefully restates the claims to what the data actually support would be needed. I also recommend asking the authors to clarify the animal-level split and the discrepancies in scan counts. These issues are fixable in principle, so I do not recommend rejection at this stage."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is the first paper I know that tries to reconstruct lung aeration maps directly from delayed RF data, and that target is worth taking seriously. The FNO-based per-pixel aeration inversion is a real step beyond B-line detectors and COVID classifiers. The authors also did the unglamorous work: 14.5k simulated pairs, 121 real scans, an ablation showing that FNO and augmentations matter, and a 0.11 s runtime. I believe the 9.4% ex vivo error is honestly reported; I just don't think it means what the abstract says.\n\nThe soft spot is structural. Section 4.2 says ground-truth bulk aeration is computed from the right lung and left caudal lobe, while the scanned region is the left cranial lobe of the same animal. So the true aeration of the region Luna actually sees is never measured. The comparison assumes aeration is homogeneous across lobes and that a 2.6-wavelength (~0.8 mm) slice represents whole-lobe bulk. With an ARDS-like instillation protocol, regional differences are likely. That makes the headline error an upper bound on something, but not a clean estimate of the scanned region's aeration. The in silico 5.2% error doesn't rescue this, because simulation provides matched ground truth by construction.\n\nOther soft spots, in order: no animal-level split is described, so I can't tell if the 103 evaluation scans are effectively clustered by 12 animals; the simulator is from the same group and was calibrated with 6 real scans from the same setup, so the sim-to-real claim is less independent than it appears; no code or data; and the histology brightness threshold is arbitrary, though the ablation and consistent errors across aeration levels make me think that is a minor issue.\n\nThat said, I don't think any of this sinks the paper. The authors explicitly list the missing ex vivo ground-truth map as a limitation, and the method is clearly described enough that a referee can ask for the right fixes. What the paper needs before the 9.4% becomes a clinical claim: lobe-matched ground truth or a defense of the homogeneity assumption, animal-level evaluation, simulator calibration disclosure, and artifact release.\n\nWho it's for: anyone working on ultrasound inverse problems, quantitative lung imaging, or sim-to-real for medical ultrasound. I'd send it to review and expect a revision. It's not ready as-is for the claim it makes, but it deserves referee time and a fair shot.","headline":"Luna is a genuinely new RF-to-aeration inverse map with real ex vivo work, but the 9.4% headline is measured against a different lobe than the one scanned, so the central validation is weaker than it looks.","tokens_in":21047,"tokens_out":1607,"would_cite":true,"duration_ms":16167,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that Luna, a Fourier neural operator, reconstructs lung aeration maps directly from delayed ultrasound RF data, reaching a mean percent aeration error of 9.4% on 103 ex vivo swine lung scans and outperforming the…","keywords":["lung ultrasound","lung aeration map","radiofrequency (RF) data","Fourier neural operator","neural operator","full-wave simulation","ex vivo swine lung","deep learning"],"falsifier":"A decisive test is to image and measure the same tissue: scan a lobe, then freeze, section, or micro-CT the exact scanned region (or the very lobe that is later weighed) to obtain a true local aeration value. If the region-matched mean absolute error is clearly above the reported 9.4%, say above 15%, the central claim would be contradicted; if it stays near 9.4%, the RF signal genuinely carries the aeration information the paper says it does.","tokens_in":119,"feed_emoji":"🫁","tokens_out":9066,"duration_ms":124313,"temperature":0.7,"pith_summary":"Lung ultrasound is widely available, but its interpretation depends on years of training because clinicians read indirect B-mode artifacts rather than the lung itself. This paper claims that a deep learning model called Luna can skip that step entirely: given the raw radiofrequency (RF) data received by the transducer, Luna reconstructs a two-dimensional lung aeration map, a direct air-versus-tissue image, and from it computes the percent aeration, a key clinical quantity. Trained mostly on full-wave physics simulations and fine-tuned on only 18 real ex vivo swine scans, Luna reports a mean percent aeration estimation error of 9.4% (SD 5.4%) on 103 ex vivo scans, which the authors state outperforms the current clinical LUS scoring system. If this holds, raw ultrasound signals carry enough information to quantify lung air content without beamforming or human artifact interpretation, which would make consistent, quantitative lung monitoring more accessible.","feed_headline":"9% error: AI reads lung aeration from raw ultrasound","feed_subtitle":"Luna skips B-mode beamforming and beats clinical scoring on 103 ex vivo swine lung scans.","key_machinery":"The load-bearing component is a Fourier neural operator applied along the temporal dimension of the RF data: because the Fourier magnitude of a delayed signal is invariant to that delay, the temporal FNO learns to map the time history of received echoes onto the depth axis of the aeration map without being sensitive to chest-wall depth. A spatial convolutional network handles the lateral transducer-element and event dimensions, and a UNet supplies the chest-wall segmentation that is concatenated with the RF data. The training signal comes from Fullwave-2, a nonlinear full-wave solver that simulates ultrasound propagation through maps built from human chest-wall anatomy and swine lung histology; Luna is trained on 10,150 simulated pairs, fine-tuned on 18 real ex vivo samples, and regularized by temporal and lateral masking augmentation plus a direct percent-aeration loss.","core_discovery":"The central claim is that the inverse problem of lung ultrasound, recovering the spatial distribution of air and tissue from the pressure field measured at the body surface, can be solved directly by a neural operator, and that the result is clinically meaningful. Luna takes delayed RF data $p$, computes a chest-wall segmentation $S(p)$ as an auxiliary task, and maps $(p, S(p))$ to a per-pixel aeration map $\\hat{\\rho}_A$; the percent aeration $\\gamma$ is then obtained by averaging the map. On synthetic data the reconstruction reaches an average PSNR of 20.10 dB and SSIM of 0.605, with a percent aeration error of 5.2% (SD 4.6%); on real ex vivo swine lungs, after fine-tuning on 18 samples, the mean percent aeration error is 9.4% (SD 5.4%), with a worst case of 20.3%. The authors present this as the first demonstration that lung aeration maps can be reconstructed from ultrasound RF data, and as evidence that training on full-wave physics simulation followed by a small amount of real fine-tuning transfers to real tissue, offering a quantitative alternative to the semi-quantitative LUS score.","pith_inferences":["If aeration is regionally heterogeneous, as in ARDS, the lobe-mismatched ground truth would probably inflate the reported error; imaging and measuring the same lobe could tighten the number, or reveal that a thin slice cannot represent a whole lobe.","The delay-invariance mechanism, learning on Fourier magnitudes so the output ignores when echoes arrive, is not lung-specific; the same temporal-FNO design could be applied to other ultrasound inverse problems where the depth of the target structure varies between patients.","A testable extension is multi-frequency training: if the simulator creates RF data at several center frequencies, one could check whether the RF-to-aeration map generalizes across transducers, which is the main barrier to device-independent clinical use."],"forward_implications":["Clinicians could read pixel-level aeration directly from a scan, replacing the coarse 0-to-3 LUS score with a continuous quantitative estimate per examination.","Because the input is RF data rather than beamformed B-mode images, the result avoids the variability that time-gain compensation and device-specific display settings introduce.","Large real paired datasets may not be required: after simulation training, fine-tuning on 18 real samples was enough to reach 9.4% error on 103 ex vivo scans.","The 0.11-second per-case runtime (about 9 Hz) puts the reconstruction in reach of the real-time, interactive operation that lung ultrasound demands.","The small gap between in silico error (5.2%) and ex vivo error (9.4%) suggests the learned wave-propagation mapping transfers to real tissue, supporting the path toward in vivo validation."],"supporting_citations":[{"why":"Supplies the Fourier neural operator architecture that Luna uses to learn the RF-to-aeration map in Fourier space.","marker":"[27]"},{"why":"Fullwave-2, the full-wave nonlinear solver that generated the 10,150 simulated RF/aeration training pairs.","marker":"[28]"},{"why":"The experimentally validated simulation approach pairing human chest-wall anatomy with swine lung histology that the synthetic data generation follows.","marker":"[31]"},{"why":"The clinical LUS scoring system, the semi-quantitative baseline whose sensitivity Luna's 9.4% error is compared against.","marker":"[17]"},{"why":"Defines the nonlinear full-wave equation with attenuation that Fullwave-2 solves, grounding the forward model.","marker":"[30]"}],"fun_headline_variants":["LUNA maps lung aeration straight from raw ultrasound RF","AI reads lung aeration from raw ultrasound, no beamforming","Neural operator turns ultrasound RF into lung aeration maps","Lung aeration from echo: 9% error on real swine lungs","Bypass beamforming: LUNA reconstructs aeration maps directly"],"cache_read_input_tokens":23040,"weakest_assumption_plain":"The load-bearing premise is that the lobe actually scanned has the same air content as the different lobes whose fluid displacement is measured for ground truth, the right lung and left caudal lobe are weighed while Luna images the left cranial lobe, and that the air content of the thin reconstructed slice represents the whole lobe's bulk aeration.","fun_headline_variants_meta":{"raw":{"variants":["LUNA maps lung aeration straight from raw ultrasound RF","AI reads lung aeration from raw ultrasound, no beamforming","Neural operator turns ultrasound RF into lung aeration maps","Lung aeration from echo: 9% error on real swine lungs","Bypass beamforming: LUNA reconstructs aeration maps directly"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00068,"raw_usage":{"total_tokens":3165,"prompt_tokens":1100,"completion_tokens":2065,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":716,"completion_tokens_details":{"reasoning_tokens":1976}},"tokens_in":716,"tokens_out":2065,"duration_ms":11218,"temperature":1.0,"reasoning_tokens":1976,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:34:41.448846+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive test is to image and measure the same tissue: scan a lobe, then freeze, section, or micro-CT the exact scanned region (or the very lobe that is later weighed) to obtain a true local aeration value. If the region-matched mean absolute error is clearly above the reported 9.4%, say above 15%, the central claim would be contradicted; if it stays near 9.4%, the RF signal genuinely carries the aeration information the paper says it does.","supporting_citations":[{"cited_title":"& Pinton, G","cited_arxiv_id":null,"evidence_quote":"The experimentally validated simulation approach pairing human chest-wall anatomy with swine lung histology that the synthetic data generation follows."},{"cited_title":"International evidence-based recommendations for point-of-care lung ultrasound","cited_arxiv_id":null,"evidence_quote":"The clinical LUS scoring system, the semi-quantitative baseline whose sensitivity Luna's 9.4% error is compared against."},{"cited_title":"F., Dahl, J., Rosenzweig, S","cited_arxiv_id":null,"evidence_quote":"Defines the nonlinear full-wave equation with attenuation that Fullwave-2 solves, grounding the forward model."}],"review_version":1}