REVIEW 2 major objections 4 minor 92 references
Self-supervised pretraining on unlabeled complex IQ ultrasound data cuts the labels needed for sound-speed estimation by roughly threefold at 10,000 labels, and more than fourfold at 1,000 labels.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 05:01 UTC pith:PXHOHBV6
load-bearing objection A careful, honestly scoped in-silico study showing JEPA pretraining on raw IQ helps label-efficient sound-speed estimation, but the headline 3-4x gain is partly confounded by unequal optimizer budgets. the 2 major comments →
IQ-JEPA: A Joint-Embedding Predictive Architecture with a Hermitian Vision Transformer for Sound Speed and Attenuation Estimation from Ultrasound IQ Data
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that sound-speed maps can be estimated from raw pulse-echo IQ channel data with far fewer labeled simulations if the encoder is first pretrained to predict masked latent representations from visible context, in a joint-embedding predictive architecture (JEPA). The key design is a Hermitian Vision Transformer that operates on the complex IQ signal directly, with attention equivariant to the global demodulation phase and a conjugate-product feed-forward invariant to it, so it naturally reads the phase differences in which sound-speed information lives. On a held-out test set of Fullwave 2.5 simulations at 2.5 MHz, the method reaches 15.60 m/s MAE at 10,000 labels v
What carries the argument
The Hermitian Vision Transformer, built from a complex patch embedding, U(1)-equivariant Hermitian attention, and a conjugate-product feed-forward that is invariant to global phase, paired with a JEPA pretraining objective that predicts the latent representation of masked IQ patches from visible context using an exponential-moving-average target encoder.
Load-bearing premise
The reported gains are measured on held-out simulations drawn from the same simulator, frequency, and transmit sequence as the training data, so they quantify interpolation inside one in-silico distribution rather than performance on real tissue.
What would settle it
Fine-tune the pretrained encoder on real ex vivo or in vivo IQ acquisitions with independently measured sound-speed maps (e.g., from ultrasound tomography or calibrated phantoms) at 2.5 MHz and at other frequencies; if the label-efficiency gain over supervised training shrinks below 2x or disappears, the central claim is not transferable to real data.
If this is right
- Self-supervised pretraining on unlabeled IQ data can replace a large fraction of expensive labeled wave simulations, cutting label requirements by 3x at 10,000 labels and more than 4x at 1,000 labels.
- Sound-speed error keeps decreasing with more unlabeled data, so scaling the unlabeled corpus is the most direct route to improved accuracy.
- The same pretrained encoder can be fine-tuned to a distinct property, attenuation, with a consistent though smaller label-efficiency gain, supporting multi-property reuse.
- Pretraining on one phantom-family distribution transfers to another with a penalty under 2% of the error, indicating the representation is not tightly tied to a single tissue model.
- Latent prediction (JEPA) outperforms pixel-reconstruction masked autoencoding on this task at every label count, particularly when labels are scarce.
Where Pith is reading between the lines
- If the same label-efficiency gain persists on real clinical IQ data, the practical path to quantitative ultrasound could be to pretrain on abundant unannotated acquisitions and fine-tune on a relatively small set of simulated or specially acquired labeled maps.
- Because the encoder is built around the U(1) symmetry of the demodulation phase, it may transfer across transducer geometries, center frequencies, and demodulation references with less fine-tuning than a phase-blind model.
- The Hermitian flash-attention reduction, which exploits the real-valued Hermitian score, could make complex-valued transformers practical for other coherent imaging modalities such as MRI k-space and seismic waveform data.
- A direct test is whether the fourfold label-efficiency gain at 1,000 labels also holds at higher center frequencies (e.g., 5 MHz) and with focused or synthetic-aperture transmit sequences, which the authors explicitly leave as future work.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces IQ-JEPA, a self-supervised joint-embedding predictive architecture for estimating sound-speed maps directly from raw complex IQ channel data. A Hermitian vision transformer—comprising a complex patch embedding, one Hermitian-attention block, and a conjugate-product feed-forward—is pretrained with a masked latent-prediction objective on 63,435 unlabeled Fullwave-simulated plane-wave acquisitions, then fine-tuned on labeled sound-speed maps. The main result is a label-efficiency gain: at 10,000 labels the method reaches 15.60 m/s test MAE versus 23.46 m/s for a supervised Hermitian ViT, with larger relative gains at 1,000 labels and 8.71 m/s at full labels. Additional results include pretraining-data scaling, a masked-autoencoder control, a 2^5 ablation, and transfer of the frozen encoder to attenuation and across phantom families. The authors explicitly scope the work as an in-silico proof of concept.
Significance. If the label-efficiency gains hold under a compute-matched comparison, this is a valuable demonstration that self-supervised latent prediction on raw complex channel data can substantially reduce the label bottleneck in learned quantitative ultrasound. Notable strengths are the fixed held-out test split withheld from pretraining, the three-seed main results, an explicit validation split for pretraining checkpoints, a compute-matched IQ-JEPA-versus-MAE control, and a thorough ablation with a symmetry analysis. The Hermitian flash-attention construction and the U(1)-equivariance/invariance analysis are clear and potentially reusable in other complex-valued transformer applications. The central uncertainty is that the headline label-efficiency gain is measured against a supervised baseline that uses far fewer total optimizer steps, so the causal attribution to the SSL objective is not fully identified. A step-matched supervised baseline and adjusted claims would make the contribution solid.
major comments (2)
- [Section III-A, Table I] The headline label-efficiency comparison is not compute-matched. IQ-JEPA SSL+ft receives 100 epochs of pretraining (~44,600 steps) plus ~7,800 fine-tuning steps at 10k labels, while the supervised Hermitian ViT receives only the ~7,800 fine-tuning steps (~6.7x fewer; much larger ratio at 1k labels). The masked-autoencoder control in Section III-D is compute-matched across objectives but does not provide a supervised baseline at matched total steps, so the portion of the 3-4x label-efficiency gain attributable to the SSL objective rather than to the extra optimizer budget is not identified. Please add a supervised-from-scratch run with the same total step budget (or a step-matched learning curve) and restate the label-efficiency claim relative to it.
- [Section IV-A] Section IV-A concludes that 'self-supervised pretraining moves the sound-speed accuracy more than any other single design choice.' This 'dominant factor' claim rests on the label-efficiency results, which are confounded by total optimizer steps as described above. The compute-matched IQ-JEPA-vs-MAE comparison in Section III-D supports the narrower claim that latent prediction beats pixel reconstruction under the same SSL pretraining budget. Please separate the two claims—objective choice at fixed SSL compute versus SSL over supervised at fixed total compute—and support the latter with a step-matched baseline or soften the conclusion.
minor comments (4)
- [Section II-G] The masking description says the context is the complement of the union of four target blocks (≈28% of patches), then states that for other samples in a batch about 20% of their own target patches remain visible because the context is shared. Clarify whether the 28% context coverage is per-sample or only for the mask-generating sample, and how the shared context interacts with per-sample target masks.
- [Section III-C3, Table II] Table II is single-seed, yet the text says the transfer gap is 'within the seed spread of comparable cells.' No seed spread is reported. Either provide multi-seed results with spread or remove the 'within seed spread' statement.
- [Section III-D] The masked-autoencoder series in Figure 10 and the text is single-seed, unlike the three-seed IQ-JEPA and supervised series. State this in the caption/text and, if feasible, add seeds for the MAE points near the smaller label counts.
- [Section III-B] The 'descriptive exponent' is described as a straight-line fit on log-log axes, but the text also says the curve does not follow a single scaling exponent. Specify which points are used for the fit and add the fit to the figure or its caption.
Circularity Check
No material circularity: the SSL label-efficiency result is an empirical comparison, not a construction; the main caveats are compute matching and shared simulation data, neither of which makes the claim definitional.
full rationale
None of the load-bearing steps reduces to its inputs by construction. The central label-efficiency claim compares a supervised Hermitian ViT with the same encoder pretrained by IQ-JEPA; pretraining is unlabeled, labels enter only at fine-tuning, and the test set is held out from both stages ('Self-supervised pretraining uses only the 63,435 training acquisitions, without labels. The validation and test acquisitions are withheld from it entirely'; 'Sound-speed labels are used only for fine-tuning'). The U(1)-invariance of the conjugate-product feed-forward is a design constraint justified by symmetry and explicitly attributed to an analogy with classical coherence methods, not a fitted parameter renamed as a prediction. The main self-citations ([56,57,60]) support data generation with a physical simulator and numerical phantoms; they are shared by all compared conditions and are externally checkable, so they do not make the SSL claim circular. The one non-circular experimental caveat is that the headline SSL-vs-supervised comparison is matched in labels and fine-tuning epochs but not in total optimizer steps: pretraining adds roughly 44.6k steps at batch 128 before fine-tuning, so part of the reported gain could be additional compute. This is a confound, not a circular equivalence, and the paper's compute-matched masked-autoencoder control (Section III-D) partially addresses the objective's contribution. Overall, the derivation is self-contained with respect to its empirical comparisons.
Axiom & Free-Parameter Ledger
free parameters (6)
- I-JEPA masking schedule (context ≈28%, 4 target blocks of 0.25–0.35 of patches)
- EMA target momentum τ (0.996 → 1.0 over training)
- Pretraining/fine-tuning schedule (100 epochs, LR 1.5e-4, batch 128, AdamW)
- InversionNet learning rate =
1e-3
- Descriptive scaling exponent of pretraining-size curve =
0.16 (Hermitian ViT); 0.21 (real ViT)
- IQ reshape/crop (1,120×630, dropping 10 element columns)
axioms (5)
- domain assumption Fullwave 2.5 multiple-relaxation finite-difference solver is a faithful model of 2D pulse-echo ultrasound (Eq. 1-2).
- standard math Bandlimited RF demodulated at 2.5 MHz without decimation is a lossless complex reparameterization carrying the same information.
- domain assumption Sound speed appears only as phase differences invariant to the global U(1) demodulation phase.
- domain assumption Per-region random assignment of material properties makes sound speed and attenuation nearly uncorrelated (r=0.16).
- domain assumption A fixed per-family 10% test draw from the same simulated phantom families measures the relevant generalization.
Cite this review
Pith. "Pith review of IQ-JEPA: A Joint-Embedding Predictive Architecture with a Hermitian Vision Transformer for Sound Speed and Attenuation Estimation from Ultrasound IQ Data." pith.science (2026). https://pith.science/paper/PXHOHBV6
@misc{pith2026260722351,
author = {Pith},
title = {Pith review of: IQ-JEPA: A Joint-Embedding Predictive Architecture with a Hermitian Vision Transformer for Sound Speed and Attenuation Estimation from Ultrasound IQ Data},
year = {2026},
howpublished = {\url{https://pith.science/paper/PXHOHBV6}},
note = {Machine review of arXiv:2607.22351}
}
read the original abstract
The speed of sound in tissue is a prerequisite for well-focused imaging and has diagnostic value, but recovering it from raw pulse-echo channel data is fundamentally a nonlinear inverse problem. Learned solvers are fast yet label hungry. Simulated sound-speed labels are expensive, while abundant real channel data is unlabeled. We propose IQ-JEPA to exploit both data types. An encoder is pretrained without labels to predict the latent representation of masked in-phase and quadrature (IQ) regions from visible context, then fine-tuned on simulated maps. Sound speed appears in the IQ signal as a phase difference, invariant to the constant phase offset. The encoder is a Hermitian vision transformer that operates on the complex signal directly. Its attention is equivariant to that phase and its conjugate-product feed-forward is invariant to it, so the encoder reads a quantity analogous to the one classical coherence methods use. On 79,293 Fullwave 2.5 simulations at 2.5 MHz, pretraining on the 63,435 unlabeled acquisitions reaches 15.60 m/s at 10,000 labels. This is a roughly threefold gain in label efficiency over supervised training, growing to over fourfold at 1,000 labels. It is about 2.2x below an InversionNet baseline, and 8.71 m/s at full labels. The gain still grows with more unlabeled pretraining data. Our comparisons point to self-supervision as the dominant factor. The same encoder transfers. Its frozen features expose sound speed and attenuation, and cross-distribution pretraining between layered and abdominal phantoms costs little accuracy. We see this as a first step toward a foundation model for quantitative ultrasound.
Figures
Reference graph
Works this paper leans on
-
[1]
Computing speed-of-sound from ultrasound: User- agnostic recovery and a new benchmark,
M. Feigin, D. Freedman, and B. W. Anthony, “Computing speed-of-sound from ultrasound: User- agnostic recovery and a new benchmark,”IEEE Trans. Biomed. Eng., vol. 71, no. 4, pp. 1094–1103, 2024. https://ieeexplore.ieee.org/abstract/document/10294174/
arXiv 2024
-
[2]
Seismic waveform inversion in the frequency domain, part 1: Theory and verification in a physical scale model,
R. G. Pratt, “Seismic waveform inversion in the frequency domain, part 1: Theory and verification in a physical scale model,”Geophysics, vol. 64, no. 3, pp. 888–901, 1999
1999
-
[3]
An overview of full-waveform inversion in exploration geophysics,
J. Virieux and S. Operto, “An overview of full-waveform inversion in exploration geophysics,”Geophysics, vol. 74, no. 6, pp. WCC1–WCC26, 2009. http://dx.doi.org/10. 1190/1.3238367
2009
-
[4]
Forward and inverse scattering in synthetic aperture radar using machine learning,
O. DeGuchy, J. Alvarez, A. D. Kim, R. F. Marcia, and C. Tsogka, “Forward and inverse scattering in synthetic aperture radar using machine learning,” in Applications of Machine Learning 2020, M. E. Zelinski, T. M. Taha, J. Howe, A. A. Awwal, and K. M. Iftekharuddin, Eds., vol. 11511. SPIE, 2020, p. 148. http://dx.doi.org/10.1117/12.2568302
-
[5]
Synthesis of complex-valued InSAR data with a multi-task convolutional neural network,
P. Sibler, F. Sica, and M. Schmitt, “Synthesis of complex-valued InSAR data with a multi-task convolutional neural network,”ISPRS J. Photogramm. Remote Sens., vol. 220, pp. 192–206, 2025. http: //dx.doi.org/10.1016/j.isprsjprs.2024.12.007
-
[6]
Efficient complex-valued vision transformers for MRI classification directly from k-space,
M. Rempe, L. T. Rotkopf, M. Schlimbach, H. Becker, F. Hörst, J. Haubold, P. Dammann, K. Kröninger, and J. Kleesiek, “Efficient complex-valued vision transformers for MRI classification directly from k-space,”arXiv preprint arXiv:2601.18392, 2026
arXiv 2026
-
[7]
Correlations of sound speed with tissue constituents in normal and diffuse liver disease,
T. Lin, J. Ophir, and G. Potter, “Correlations of sound speed with tissue constituents in normal and diffuse liver disease,”Ultrason. Imaging, vol. 9, no. 1, pp. 29–40, Jan
-
[8]
Clinical sound speed measurement in liver and spleen in vivo,
C. F. Chen, D. E. Robinson, L. S. Wilson, K. A. Griffiths, A. Manoharan, and B. D. Doust, “Clinical sound speed measurement in liver and spleen in vivo,” Ultrason. Imaging, vol. 9, no. 4, pp. 221–235, Oct. 1987. http://dx.doi.org/10.1177/016173468700900401
-
[9]
Measurement and use of acoustic nonlinearity and sound speed to estimate composition of excised livers,
C. M. Sehgal, G. M. Brown, R. C. Bahn, and J. F. Greenleaf, “Measurement and use of acoustic nonlinearity and sound speed to estimate composition of excised livers,” 16 Ultrasound Med. Biol., vol. 12, no. 11, pp. 865–874, Nov
-
[10]
Dependence of ultrasonic nonlinear parameter BA on fat,
R. L. Errabolu, C. M. Sehgal, and J. F. Greenleaf, “Dependence of ultrasonic nonlinear parameter BA on fat,”Ultrason. Imaging, vol. 9, no. 3, pp. 180–194, 1 Jul
-
[11]
First-in-human diagnostic study of hepatic steatosis with computed ultrasound tomography in echo mode,
P. Stähli, C. Becchetti, N. Korta Martiartu, A. Berzigotti, M. Frenz, and M. Jaeger, “First-in-human diagnostic study of hepatic steatosis with computed ultrasound tomography in echo mode,”Communications Medicine, vol. 3, no. 1, p. 176, 2023
2023
-
[12]
Ultrasonic Sound Speed Estimation for Liver Fat Quantification: A Review by the AIUM-RSNA QIBA Pulse-Echo Quantitative Ultrasound Initiative,
X. Wang, J. C. Bamber, R. Esquivel-Sirvent, J. Or- machea, P. S. Sidhu, K. E. Thomenius, S. Schoen, Jr, S. Rosenzweig, and T. T. Pierce, “Ultrasonic Sound Speed Estimation for Liver Fat Quantification: A Review by the AIUM-RSNA QIBA Pulse-Echo Quantitative Ultrasound Initiative,”Ultrasound in Medicine & Biology, vol. 49, no. 11, pp. 2327–2335, 2023
2023
-
[13]
https://www.sciencedirect.com/science/article/pii/ 0161734687900046
-
[14]
Speed of sound and shear wave speed for calf soft tissue composition and nonlinearity assessment,
N. Korta Martiartu, D. Nakhostin, L. Ruby, T. Frauenfelder, M. B. Rominger, and S. J. Sanabria, “Speed of sound and shear wave speed for calf soft tissue composition and nonlinearity assessment,”Quant. Imaging Med. Surg., vol. 11, no. 9, pp. 4149–4161, Sep
-
[15]
High Resolution Imaging and Digital Characterization of Skin Pathology By Scanning Acoustic Microscopy,
S. Youssef, “High Resolution Imaging and Digital Characterization of Skin Pathology By Scanning Acoustic Microscopy,” Ph.D. dissertation, University of Windsor, 2018. https://scholar.uwindsor.ca/etd/7439/
2018
-
[16]
In vivo breast sound-speed imaging with ultrasound tomography,
C. Li, N. Duric, P. Littrup, and L. Huang, “In vivo breast sound-speed imaging with ultrasound tomography,” Ultrasound Med. Biol., vol. 35, no. 10, pp. 1615–1628,
-
[17]
Global burden and risk factors of MASLD: trends from 1990 to 2021 and predictions to 2030,
M. Huang, H. Chen, H. Wang, Y . Zhang, L. Li, Y . Lan, and L. Ma, “Global burden and risk factors of MASLD: trends from 1990 to 2021 and predictions to 2030,”Internal and Emergency Medicine, vol. 20, no. 4, pp. 1013–1024, 2025
1990
-
[18]
Sampling variability of liver biopsy in nonalcoholic fatty liver disease,
V . Ratziu, F. Charlotte, A. Heurtier, S. Gombert, P. Giral, E. Bruckert, A. Grimaldi, F. Capron, T. Poynard, and LIDO Study Group, “Sampling variability of liver biopsy in nonalcoholic fatty liver disease,”Gastroenterology, vol. 128, no. 7, pp. 1898–1906, 2005
1906
-
[19]
The diagnostic value of MRI-PDFF in hepatic steatosis of patients with metabolic dysfunction-associated steatotic liver disease: a systematic review and meta-analysis,
Y .-X. Zhang, Y .-P. Feng, C.-L. You, and L.-Y . Zhang, “The diagnostic value of MRI-PDFF in hepatic steatosis of patients with metabolic dysfunction-associated steatotic liver disease: a systematic review and meta-analysis,” BMC Gastroenterology, vol. 25, no. 1, p. 451, 2025
2025
-
[20]
G. Cloutier, F. Destrempes, F. Yu, and A. Tang, “Quantitative ultrasound imaging of soft biological tissues: a primer for radiologists and medical physicists,” Insights Imaging, vol. 12, no. 1, p. 127, 2021. http://dx.doi.org/10.1186/s13244-021-01071-w
-
[21]
The global epidemi- ology of nonalcoholic fatty liver disease (NAFLD) and nonalcoholic steatohepatitis (NASH): a systematic review,
Z. M. Younossi, P. Golabi, J. M. Paik, A. Henry, C. Van Dongen, and L. Henry, “The global epidemi- ology of nonalcoholic fatty liver disease (NAFLD) and nonalcoholic steatohepatitis (NASH): a systematic review,” Hepatology, vol. 77, no. 4, pp. 1335–1347, 2023
2023
-
[22]
The impact of sound speed errors on medical ultrasound imaging,
M. E. Anderson, M. S. McKeag, and G. E. Trahey, “The impact of sound speed errors on medical ultrasound imaging,”J. Acoust. Soc. Am., vol. 107, no. 6, pp. 3540–3548, 2000. http://dx.doi.org/10.1121/1.429422
-
[23]
Aberration correction in diagnostic ultrasound: A review of the prior field and current directions,
R. Ali, T. Brevett, L. Zhuang, H. Bendjador, A. S. Podkowa, S. S. Hsieh, W. Simson, S. J. Sanabria, C. D. Herickhoff, and J. J. Dahl, “Aberration correction in diagnostic ultrasound: A review of the prior field and current directions,”Z. Med. Phys., vol. 33, no. 3, pp. 267–291, 2023. http: //dx.doi.org/10.1016/j.zemedi.2023.01.003
-
[24]
Distributed Aberration Correction Techniques Based on Tomographic Sound Speed Estimates,
R. Ali, T. Brevett, D. Hyun, L. L. Brickson, and J. J. Dahl, “Distributed Aberration Correction Techniques Based on Tomographic Sound Speed Estimates,”IEEE Trans. Ultrason. Ferroelectr. Freq. Control, vol. 69, no. 5, pp. 1714–1726, 2022. http://dx.doi.org/10.1109/TUFFC.2022.3162836
arXiv 2022
-
[25]
M. L. Oelze and J. Mamou, “Review of quantitative ultrasound: Envelope statistics and backscatter coefficient imaging and contributions to diagnostic ultrasound,” IEEE Trans. Ultrason. Ferroelectr. Freq. Control, vol. 63, no. 2, pp. 336–351, 2016. http://dx.doi.org/10.1109/ TUFFC.2015.2513958
arXiv 2016
-
[26]
EASL-EASD-EASO Clinical Practice Guide- lines on the management of metabolic dysfunction- associated steatotic liver disease (MASLD),
European Association for the Study of the Liver (EASL), European Association for the Study of Diabetes (EASD), and European Association for the Study of Obesity (EASO), “EASL-EASD-EASO Clinical Practice Guide- lines on the management of metabolic dysfunction- associated steatotic liver disease (MASLD),”Journal of Hepatology, vol. 81, no. 3, pp. 492–542, 2024
2024
-
[27]
Full-waveform inversion imaging of the human brain,
L. Guasch, O. Calderón Agudo, M.-X. Tang, P. Nachev, and M. Warner, “Full-waveform inversion imaging of the human brain,”NPJ Digit Med, vol. 3, p. 28, 2020. http://dx.doi.org/10.1038/s41746-020-0240-8
-
[28]
Nonlinear waveform inversion for quantitative ultrasound,
A. Shultzman and Y . C. Eldar, “Nonlinear waveform inversion for quantitative ultrasound,”IEEE Trans. Comput. Imaging, vol. 8, pp. 893–904, 2022. https: //ieeexplore.ieee.org/document/9897002
arXiv 2022
-
[29]
M. Jaeger, G. Held, S. Peeters, S. Preisser, M. Grünig, and M. Frenz, “Computed ultrasound tomography in echo mode for imaging speed of sound using pulse-echo sonography: proof of principle,”Ultrasound Med. Biol., vol. 41, no. 1, pp. 235–250, 2015. http://dx.doi.org/10.1016/j.ultrasmedbio.2014.05.019
-
[30]
Improved forward model for quantitative pulse-echo speed-of-sound imaging,
P. Stähli, M. Kuriakose, M. Frenz, and M. Jaeger, “Improved forward model for quantitative pulse-echo speed-of-sound imaging,”Ultrasonics, vol. 108, p. 106168, 2020. https://www.sciencedirect.com/science/ article/pii/S0041624X20301074
2020
-
[31]
Full- waveform inversion, Part 1: Forward modeling,
M. Louboutin, P. Witte, M. Lange, N. Kukreja, F. Luporini, G. Gorman, and F. J. Herrmann, “Full- waveform inversion, Part 1: Forward modeling,”Lead. Edge, vol. 36, no. 12, pp. 1033–1036, 2017. https: //library.seg.org/doi/10.1190/tle36121033.1
-
[32]
Implicit neural representations for speed-of-sound estimation in ultrasound,
M. Byra, P. Jarosik, P. Karwat, Z. Klimonda, and M. Lewandowski, “Implicit neural representations for speed-of-sound estimation in ultrasound,” in2024 IEEE Ultrasonics, Ferroelectrics, and Frequency Control Joint Symposium (UFFC-JS). IEEE, 2024, pp. 1–4. http: //dx.doi.org/10.1109/UFFC-JS60046.2024.10793775
arXiv 2024
-
[33]
InversionNet: An Efficient and Accurate Data-Driven Full Waveform Inversion,
Y . Wu and Y . Lin, “InversionNet: An Efficient and Accurate Data-Driven Full Waveform Inversion,”IEEE Transactions on Computational Imaging, vol. 6, pp. 419– 433, 2020. http://dx.doi.org/10.1109/TCI.2019.2956866
arXiv 2020
-
[34]
OpenFWI: Large-scale Multi- structural Benchmark Datasets for Full Waveform Inversion,
C. Deng, S. Feng, H. Wang, X. Zhang, and others, “OpenFWI: Large-scale Multi- structural Benchmark Datasets for Full Waveform Inversion,”Advances in, 2022. https: //proceedings.neurips.cc/paper_files/paper/2022/hash/ 27d3ef263c7cb8d542c4f9815a49b69b-Abstract-Datasets_ and_Benchmarks.html
2022
-
[35]
Yale/UNC-CH - geophysical waveform inversion,
Y . Lin, L. Lu, W. Reade, A. Howard, M. Cruz, and A. Chow, “Yale/UNC-CH - geophysical waveform inversion,” Kaggle, 2025. https://kaggle.com/competitions/waveform-inversion
2025
-
[36]
Differentiable Beamforming for Ultrasound Autofocusing,
W. Simson, L. Zhuang, S. J. Sanabria, N. Antil, J. J. Dahl, and D. Hyun, “Differentiable Beamforming for Ultrasound Autofocusing,” inMedical Image Computing and Computer Assisted Intervention – MICCAI 2023. 17 Springer Nature Switzerland, 2023, pp. 428–437. http://dx.doi.org/10.1007/978-3-031-43999-5_41
-
[37]
Vishnevskiy, S
V . Vishnevskiy, S. J. Sanabria, and O. Goksel,Image reconstruction via variational network for real-time hand-held sound-speed imaging, ser. Lecture notes in computer science. Springer International Publishing, 2018, pp. 120–128. https://link.springer.com/chapter/10. 1007/978-3-030-00129-2_14
2018
-
[38]
A Deep Learning Framework for Single-Sided Sound Speed Inversion in Medical Ultrasound,
M. Feigin, D. Freedman, and B. W. Anthony, “A Deep Learning Framework for Single-Sided Sound Speed Inversion in Medical Ultrasound,”IEEE Trans. Biomed. Eng., vol. 67, pp. 1142–1151, 2020. http: //dx.doi.org/10.1109/TBME.2019.2931195
arXiv 2020
-
[39]
Training Variational Networks With Multidomain Simulations: Speed-of-Sound Image Reconstruction,
M. Bernhardt, V . Vishnevskiy, R. Rau, and O. Goksel, “Training Variational Networks With Multidomain Simulations: Speed-of-Sound Image Reconstruction,” IEEE Trans. Ultrason. Ferroelectr. Freq. Control, vol. 67, pp. 2584–2594, 2020. http://dx.doi.org/10.1109/TUFFC. 2020.3010186
arXiv 2020
-
[40]
Khun Jush, P
F. Khun Jush, P. M. Dueppenbecker, and A. Maier, Data-driven speed-of-sound reconstruction for medical ultrasound: Impacts of training data format and imperfections on convergence, ser. Lecture notes in computer science. Springer International Publishing, 2021, pp. 140–150. https://link.springer.com/chapter/10. 1007/978-3-030-80432-9_11
2021
-
[41]
1st place solution: Yale/UNC-CH - geophysical waveform inversion,
H. Sheoran, “1st place solution: Yale/UNC-CH - geophysical waveform inversion,” Kaggle, 2025. https: //www.kaggle.com/competitions/waveform-inversion/ writeups/harshit-sheoran-1st-place-solution
2025
-
[42]
Speed-of-sound reconstruction with deep neural networks in pulse-echo mode: Coherency- vs RF-data-based approach,
M. Heller and G. Schmitz, “Speed-of-sound reconstruction with deep neural networks in pulse-echo mode: Coherency- vs RF-data-based approach,” in2023 IEEE International Ultrasonics Symposium (IUS). IEEE, 2023, pp. 1–4. https://ieeexplore.ieee.org/abstract/document/ 10307895/
2023
-
[43]
Abdominal sound speed estimation using neural networks trained on wave propagation physics,
L. Zhuang, W. Simson, O. Ostras, D. Hyun, G. Pinton, and J. Dahl, “Abdominal sound speed estimation using neural networks trained on wave propagation physics,” in 2023 IEEE International Ultrasonics Symposium (IUS). IEEE, Sep. 2023. http://dx.doi.org/10.1109/ius51837. 2023.10308076
arXiv 2023
-
[44]
BERT: Pre-training of deep bidirectional transformers for language understanding,
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” inProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), J. Burstein, C. Doran, and T. Solorio, Eds....
2019
-
[45]
Language models are few-shot learners,
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-V oss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei...
2020
-
[46]
Speed-of-sound mapping for pulse-echo ultrasound raw data using linked-autoencoders,
F. K. Jush, P. Dueppenbecker, and A. K. Maier, “Speed-of-sound mapping for pulse-echo ultrasound raw data using linked-autoencoders,”ML4MHD 2023, Lecture Notes in Computer Science, pp. 103–114, 2023. https://doi.org/10.1007/978-3-031-47679-2_8
-
[47]
Auto-linear phenomenon in subsurface imaging,
Y . Feng, Y . Chen, P. Jin, S. Feng, and Y . Lin, “Auto-linear phenomenon in subsurface imaging,” inProceedings of the 41st International Conference on Machine Learning (ICML), 2024, pp. 13 153–13 174
2024
-
[48]
Unsupervised learning of full-waveform inversion: Connecting CNN and partial differential equation in a loop,
P. Jin, X. Zhang, Y . Chen, S. X. Huang, Z. Liu, and Y . Lin, “Unsupervised learning of full-waveform inversion: Connecting CNN and partial differential equation in a loop,” inInternational Conference on Learning Representations, 2022. https://openreview.net/ forum?id=izvwgBic9q
2022
-
[49]
Seismic foundation model: A next generation deep-learning model in geophysics,
H. Sheng, X. Wu, X. Si, J. Li, S. Zhang, and X. Duan, “Seismic foundation model: A next generation deep-learning model in geophysics,”Geophysics, vol. 90, no. 2, pp. IM59–IM79, 2025. http://dx.doi.org/10.1190/ geo2024-0262.1
2025
-
[50]
SeisLM: A foundation model for seismic waveforms,
T. Liu, J. Münchmeyer, L. Laurenti, C. Marone, M. V . de Hoop, and I. Dokmani ´c, “SeisLM: A foundation model for seismic waveforms,”arXiv [physics.geo-ph],
-
[51]
Masked autoencoders are scalable vision learners,
K. He, X. Chen, S. Xie, Y . Li, P. Dollár, and R. Girshick, “Masked autoencoders are scalable vision learners,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 15 979–15 988
2022
-
[52]
Building blocks for a complex-valued transformer architecture,
F. Eilers and X. Jiang, “Building blocks for a complex-valued transformer architecture,” inICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023, arXiv:2306.09827. http://dx.doi.org/10.1109/icassp49357. 2023.10095349
Pith/arXiv arXiv 2023
-
[53]
Evaluating complex-valued transformers based on various architec- tural parameters for image classification,
M. Sabokpa, N. Mozayani, and M. Garshasbi, “Evaluating complex-valued transformers based on various architec- tural parameters for image classification,”Information Sciences, vol. 751, p. 123562, 2026
2026
-
[54]
From pretraining to privacy: federated ultrasound foundation model with self-supervised learning,
Y . Jiang, C.-M. Feng, J. Ren, J. Wei, Z. Zhang, Y . Hu, Y . Liu, R. Sun, X. Tang, J. Du, X. Wan, Y . Xu, B. Du, X. Gao, G. Wang, S. Zhou, S. Cui, and Z. Li, “From pretraining to privacy: federated ultrasound foundation model with self-supervised learning,”NPJ Digit. Med., vol. 8, no. 1, p. 714, 2025. http://dx.doi.org/10.1038/s41746-025-02085-0
-
[55]
US-JEPA: A joint embedding predictive architecture for medical ultrasound,
A. Radhachandran, V . Ivezi´c, S. Athreya, R. Anilkumar, C. W. Arnold, and W. Speier, “US-JEPA: A joint embedding predictive architecture for medical ultrasound,” arXiv [cs.CV], 2026. http://arxiv.org/abs/2602.19322
arXiv 2026
-
[56]
A fullwave model of the nonlin- ear wave equation with multiple relaxations and relaxing perfectly matched layers for high-order numerical finite difference solutions,
M. Sode and G. Pinton, “A fullwave model of the nonlin- ear wave equation with multiple relaxations and relaxing perfectly matched layers for high-order numerical finite difference solutions,”Physics in Medicine & Biology, 2026, in press
2026
-
[57]
Self-supervised 18 learning from images with a joint-embedding predictive architecture,
M. Assran, Q. Duval, I. Misra, P. Bojanowski, P. Vincent, M. Rabbat, Y . LeCun, and N. Ballas, “Self-supervised 18 learning from images with a joint-embedding predictive architecture,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 15 619–15 629
2023
-
[58]
M. J. Ackerman, “The Visible Human Project,”Proc. IEEE, vol. 86, no. 3, pp. 504–511, Mar. 1998. http://dx.doi.org/10.1109/5.662875
doi:10.1109/5.662875 1998
-
[59]
The visible human male: a technical report,
V . Spitzer, M. J. Ackerman, A. L. Scherzinger, and D. Whitlock, “The visible human male: a technical report,” J. Am. Med. Inform. Assoc., vol. 3, no. 2, pp. 118–130,
-
[60]
L. Zhuang, O. Ostras, M. Sode, W. Simson, D. Hyun, F. Santibanez, J. Dahl, and G. Pinton, “Labeled numerical phantom of abdominal wall for wave- physics based ultrasound imaging: applications to image reconstruction,”IEEE Trans. Ultrasonics, pp. 1–1, 2025. http://dx.doi.org/10.1109/tuson.2025.3638314
arXiv 2025
-
[61]
On the applicability of Kramers- Kronig relations for ultrasonic attenuation obeying a frequency power law,
K. R. Waters, M. S. Hughes, J. Mobley, G. H. Branden- burger, and J. G. Miller, “On the applicability of Kramers- Kronig relations for ultrasonic attenuation obeying a frequency power law,”J. Acoust. Soc. Am., vol. 108, no. 2, pp. 556–563, 2000
2000
-
[62]
The acoustic properties of the epidermis and stratum corneum,
C. Edwards, “The acoustic properties of the epidermis and stratum corneum,” inThe Physical Nature of the Skin, R. M. Marks, S. P. Barton, and C. Edwards, Eds. Dordrecht: Springer Netherlands, 1988, pp. 201–207. https://doi.org/10.1007/978-94-009-1291-5_21
-
[63]
——, “Spatially heterogeneous power-law attenuation with multiple relaxation mechanisms for ultrasound mod- eling,”arXiv preprint arXiv:2606.11103, 2026
Pith/arXiv arXiv 2026
-
[64]
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” inInternational Conference on Learning Representations, 2020. https: //openreview.net/pdf?id=YicbFdNTTy
2020
-
[65]
Deep complex networks,
C. Trabelsi, O. Bilaniuk, Y . Zhang, D. Serdyuk, S. Sub- ramanian, J. F. Santos, S. Mehri, N. Rostamzadeh, Y . Bengio, and C. J. Pal, “Deep complex networks,” in International Conference on Learning Representations (ICLR), 2018
2018
-
[66]
RoFormer: Enhanced transformer with rotary position embedding,
J. Su, Y . Lu, S. Pan, A. Murtadha, B. Wen, and Y . Liu, “RoFormer: Enhanced transformer with rotary position embedding,”arXiv preprint arXiv:2104.09864, 2021, published in Neurocomputing, vol. 568, 127063, 2024
Pith/arXiv arXiv 2021
-
[67]
Complex Transformer: A framework for model- ing complex-valued sequence,
M. Yang, M. Q. Ma, D. Li, Y .-H. H. Tsai, and R. Salakhut- dinov, “Complex Transformer: A framework for model- ing complex-valued sequence,” inIEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020, pp. 4232–4236, arXiv:1910.10202
Pith/arXiv arXiv 2020
-
[68]
A. Kumar, “When do complex-valued neural networks help? A study of representation, geometry, and optimiza- tion,”arXiv preprint arXiv:2605.27673, 2026
Pith/arXiv arXiv 2026
-
[69]
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” inAdvances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, 4-9 December 2017, Long Beach, CA, USA, I. Guyon, U. von Luxburg, S. Bengio, H. M. Wallach, R. Fergus, S....
2017
-
[70]
Complex-valued neural networks: A comprehensive survey,
C. Lee, H. Hasegawa, and S. Gao, “Complex-valued neural networks: A comprehensive survey,”IEEE/CAA Journal of Automatica Sinica, vol. 9, no. 8, pp. 1406–1426, 2022
2022
-
[71]
A. V . Telichko, R. Ali, T. Brevett, H. Wang, J. G. Vilches-Moure, S. U. Kumar, R. Paulmurugan, and J. J. Dahl, “Noninvasive estimation of local speed of sound by pulse-echo ultrasound in a rat model of nonalcoholic fatty liver,”Phys. Med. Biol., vol. 67, 2022. http://dx.doi.org/10.1088/1361-6560/ac4562
-
[72]
Local sound speed estimation for pulse-echo ultrasound in layered media,
R. Ali, A. V . Telichko, H. Wang, U. K. Sukumar, J. G. Vilches-Moure, R. Paulmurugan, and J. J. Dahl, “Local sound speed estimation for pulse-echo ultrasound in layered media,”IEEE Trans. Ultrason. Ferroelectr. Freq. Control, vol. 69, pp. 500–511, 2022. https://ieeexplore.ieee.org/document/9597614
arXiv 2022
-
[73]
Revisiting feature prediction for learning visual representations from video,
A. Bardes, Q. Garrido, J. Ponce, X. Chen, M. Rabbat, Y . LeCun, M. Assran, and N. Ballas, “Revisiting feature prediction for learning visual representations from video,” Transactions on Machine Learning Research, 2024. https://openreview.net/forum?id=QaCCuDfBk2
2024
-
[74]
V-JEPA 2.1: Unlocking dense features in video self- supervised learning,
L. Mur-Labadia, M. Muckley, A. Bar, M. Assran, 19 K. Sinha, M. Rabbat, Y . LeCun, N. Ballas, and A. Bardes, “V-JEPA 2.1: Unlocking dense features in video self- supervised learning,”arXiv preprint arXiv:2603.14482, 2026
Pith/arXiv arXiv 2026
-
[75]
Measurements of ultrasonic pulse arrival time and energy level variations produced by propagation through abdominal wall,
L. M. Hinkelman, D.-L. Liu, L. A. Metlay, and R. C. Waag, “Measurements of ultrasonic pulse arrival time and energy level variations produced by propagation through abdominal wall,”The Journal of the Acoustical Society of America, vol. 95, no. 1, pp. 530–541, 1994
1994
-
[76]
complex-valued-transformer,
P. Wang, “complex-valued-transformer,” https://github. com/lucidrains/complex-valued-transformer, 2023
2023
-
[77]
FlashAttention: Fast and memory-efficient exact attention with IO-awareness,
T. Dao, D. Y . Fu, S. Ermon, A. Rudra, and C. Ré, “FlashAttention: Fast and memory-efficient exact attention with IO-awareness,” inAdvances in Neural Information Processing Systems (NeurIPS), 2022, arXiv:2205.14135
Pith/arXiv arXiv 2022
-
[78]
Complex convolutional neural networks for ultrafast ultrasound imaging reconstruction from in-phase/quadrature signal,
J. Lu, F. Millioz, D. Garcia, S. Salles, D. Ye, and D. Friboulet, “Complex convolutional neural networks for ultrafast ultrasound imaging reconstruction from in-phase/quadrature signal,”IEEE Trans. Ultrason. Ferroelectr. Freq. Control, vol. 69, no. 2, pp. 592–603,
-
[79]
Complex residual attention U-Net for fast ultrasound imaging from a single plane-wave equivalent to diverging wave imaging,
A. Bentaleb, C. Sintes, P.-H. Conze, F. Rousseau, A. Guezou-Philippe, and C. Hamitouche, “Complex residual attention U-Net for fast ultrasound imaging from a single plane-wave equivalent to diverging wave imaging,”Sensors, vol. 24, no. 16, p. 5111, 2024. https://www.mdpi.com/1424-8220/24/16/5111
2024
-
[80]
W. Han, W. Zhou, L. Huang, J. Luo, and B. Peng, “Tissue clutter filtering methods in ultrasound localization microscopy based on complex-valued networks and knowledge distillation,”IEEE Trans. Ultrason. Ferroelectr. Freq. Control, vol. 72, no. 4, pp. 440–453, 2025. http://dx.doi.org/10.1109/tuffc.2025.3544692
arXiv 2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.