REVIEW 5 major objections 7 minor 8 references
Direct PPG-to-BP deep learning hits clinical Grade A and beats every ECG-mediated pipeline on the same data.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-30 23:03 UTC pith:75WVZPYE
load-bearing objection Solid controlled bake-off showing direct PPG→BP beats ECG-mediated stacks on MIMIC, but the Grade A claim and wearable framing rest on an unspecified split that is probably leaking. the 5 major comments →
Blood Pressure Estimation from PPG: A Comparative Study of Direct and ECG-Mediated Deep Learning Pipelines
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Across 1.74 M segments from 3,127 patients, end-to-end PPG→BiLSTM blood-pressure estimation meets BHS Grade A (MAE SBP 4.82 mmHg, DBP 4.31 mmHg) and significantly outperforms every matched PPG→ECG→BiLSTM pipeline, which reach only Grade B; real ECG→BiLSTM remains Grade A, isolating the generation step as the source of the accuracy loss.
What carries the argument
The controlled dual-pipeline design that freezes the identical BiLSTM BP predictor and varies only the input path (raw PPG versus ECG reconstructed by four generators), so any accuracy gap is attributable to the intermediate ECG representation rather than to the BP head.
Load-bearing premise
That correlations and model rankings measured on high-quality ICU arterial-line recordings with strict SNR filters will still hold for everyday wearable PPG under motion and lower sensor quality.
What would settle it
Re-run the identical Direct-versus-Wave-U-Net comparison on ambulatory wearable PPG paired with cuff or arterial reference BP during daily activity; if the ECG-mediated route then matches or beats direct PPG, the central claim fails.
If this is right
- Wearable BP systems can drop the PPG-to-ECG generator, cutting parameters from millions to ~0.15 M and inference latency roughly four-fold.
- Clinical-grade continuous cuffless monitoring becomes feasible with a single lightweight model rather than a two-stage cascade.
- Future generative PPG-to-ECG work should be judged on tasks that actually need ECG morphology, not on BP estimation.
- Demographic subgroup results imply the same direct architecture remains Grade A across age, sex, and hypertensive status on this cohort.
Where Pith is reading between the lines
- If the information-bottleneck explanation is correct, multi-task models that jointly predict BP and other PPG-derived vitals may still outperform any forced ECG intermediate.
- The same controlled freeze-the-head design could be reused to test whether other popular intermediates (PTT features, spectrograms, latent embeddings) help or hurt BP accuracy.
- Personalization with a few cuff calibrations, left untested here, is the most direct route to close the remaining gap to arterial-line precision on wearables.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript reports a controlled comparison on MIMIC-III waveforms (3,127 ICU patients, ~1.74M 2-second segments) between direct PPG→BP estimation and ECG-mediated pipelines (PPG→ECG→BP) using a shared BiLSTM prediction head and four generators (U-Net, Wave-U-Net, Transformer, CycleGAN), plus spectral-domain variants with a diffusion generator. The direct pipeline is reported to achieve "BHS Grade A" (SBP/DBP MAE 4.82/4.31 mmHg), significantly outperforming all mediated pipelines (best: Wave-U-Net path, 5.38/4.89), while real ECG→BiLSTM (4.23/3.87) serves as an upper bound, localizing the degradation to the generation step. A waveform-level correlation analysis (PPG↔ABP |r|=0.247 vs ECG↔ABP r=0.018) is offered as physiological motivation.
Significance. If the numbers hold, this is a useful, cleanly designed negative result for the cuffless-BP community: an apples-to-apples bake-off with a shared BP head, multiple state-of-the-art generators, paired statistical tests, a demographic subgroup table, and — most valuable — a real-ECG upper bound that isolates the generation step as the source of degradation. The scale (1.74M segments) and the controlled design exceed much of the prior literature, and the conclusion (skip the ECG intermediate) is actionable for wearable pipeline design. However, the absolute accuracy claims rest on an unspecified data-splitting unit and a non-standard redefinition of the BHS grading protocol, both of which are load-bearing for the headline "Grade A" claim. No code or preprocessing pipeline is released, which limits reproducibility.
major comments (5)
- [§III.B (Preprocessing Pipeline) and §VI.C/E] The manuscript never states whether the train/validation/test split is performed at the subject level or the segment level. With 256-sample (2 s) windows at 50% overlap, adjacent segments from the same ICU recording share waveform morphology, baseline BP, and hemodynamic state; per-segment z-score normalization (§III.B, stage 4) further removes absolute amplitude, forcing the model to rely on precisely the morphological features that are near-duplicated across overlapping windows. Segment-level splitting is the well-documented failure mode in MIMIC-based cuffless BP work (segment-level splits typically yield 4–6 mmHg MAE; subject-wise splits degrade to 8–15 mmHg absent calibration). The authors must state the splitting unit explicitly; if it is segment-level, the headline Grade A numbers in Table V and the abstract cannot be taken at face value and a subject-wise re-evaluation is require
- [§V (Evaluation Metrics) and Table V] The BHS grading is redefined as MAE/SD thresholds (Grade A: MAE≤5, SD≤8 mmHg). The actual BHS protocol grades devices by cumulative percentages of readings falling within 5/10/15 mmHg of the reference (Grade A requires ≥60% within 5, ≥85% within 10, ≥95% within 15 mmHg). MAE/SD thresholds are closer in spirit to the AAMI criteria, and even those are defined on mean and SD of per-subject error with a subject-count requirement, not pooled segment statistics. As written, the central claim 'BHS Grade A' in the abstract, §I objective (3), and §VII.B is not supported by the cited standard. The authors should either compute and report the actual BHS cumulative-error-band percentages (and state the protocol, including the per-subject measurement requirements they are waiving) or relabel the metric honestly as an MAE/SD-based criterion and drop the BHS designation.
- [§VI.E (Statistical Validation)] The paired t-tests treat all 217,456 test segments as independent observations (t(217455)). Segments from the same patient — especially overlapping windows seconds apart — are strongly correlated, so the effective sample size is far smaller and the reported p-values are uninformative. This is not merely cosmetic: with N this large, p<0.001 is guaranteed by construction, and the reported effect sizes (Cohen's d = 0.168/0.151) already indicate the practical gap is small. The authors should cluster the test by patient (e.g., patient-level bootstrap or mixed-effects analysis) and report cluster-robust significance, and foreground the effect size rather than the p-value when characterizing the direct-vs-mediated gap.
- [§III.C and Table I; framing in §I/§VII.A] The physiological warrant — that PPG↔ABP |r|=0.247 vs ECG↔ABP r=0.018 shows PPG is the better intermediate — is not established by pairwise Pearson correlation of raw waveforms. These correlations conflate sampling-rate/artifact structure with physiological coupling; an r of 0.018 for ECG↔ABP reflects that raw ECG voltage and ABP pressure waveforms are simply different physical quantities, not that ECG carries no BP-relevant information (e.g., heart-rate and timing features from ECG are known BP covariates). The empirical bake-off in Table V stands on its own and is the real contribution; the correlation analysis should be reframed as descriptive motivation rather than evidence that ECG 'discards' BP information, or replaced with a task-relevant analysis (e.g., mutual information with SBP/DBP labels, or feature-level correlations).
- [Abstract and §VII.B; data in §III.A] The abstract concludes that 'accurate continuous BP monitoring can be achieved directly from wearable PPG signals,' but all data are ICU arterial-line ABP with clinical-grade PPG filtered at SNR>15 dB, 2 s windows, no motion artifacts, and no calibration protocol. §VII.C acknowledges this, but the abstract and §VII.B ('supports the feasibility of continuous cuffless BP monitoring using wearable PPG sensors') overstate what the experiment tests. The claims should be scoped to ICU-quality signals, or the abstract should explicitly state that wearable validation is future work rather than a demonstrated capability.
minor comments (7)
- [Fig. 1 caption vs §VI.B] The caption describes the shown ECG as generated by a 'diffusion-based PPG→ECG model,' but the generators in §IV.B/§VI.B are U-Net, Wave-U-Net, Transformer, and CycleGAN; the diffusion model appears only in the spectral experiments (§VI.D). Please correct the caption or clarify which model produced the figure.
- [Table V, spectral and mixed-input rows] The spectral-domain rows report MAE without SD and without BHS grades, and the 'Mixed-Input Pipelines' (PPG+ECG→CNN/BiLSTM) appear without description in §IV — their architectures, training, and provenance are never specified. Either describe them in §IV or move them to an appendix with full detail. The '12–13' and '14–15' ranges should be replaced with point estimates.
- [Index Terms] The index terms are the IEEE template placeholder ('component, formatting, style, styling, insert'). Replace with actual keywords.
- [Table III] It is unclear whether the LSTM [1] row is the authors' re-implementation on their splits or numbers transcribed from the cited paper. If the latter, it is not comparable to Table V (different data/splits) and should be labeled as such.
- [§III.B] Please state how the 94.2% QC pass rate is computed (per segment? per patient?) and report the number of distinct patients in each split, which would also help resolve the splitting-unit question.
- [Reproducibility] No code, hyperparameter tables beyond learning rates, or train/val/test patient lists are released. Given that the paper's value is precisely its controlled comparison, releasing the pipeline (or at least full configs and split manifests) would substantially strengthen it.
- [§VI.C] The 11.6%/13.5% relative-error figures compare the direct pipeline to the best mediated pipeline only; please also report the comparison against the real-ECG upper bound (the direct pipeline is ~14% worse than real ECG on SBP MAE), which is relevant to the information-bottleneck argument in §VII.A.
Circularity Check
No definitional or self-citation circularity; the Grade A vs Grade B claim is an empirical bake-off on held-out ABP labels, not forced by construction.
full rationale
The paper’s load-bearing chain is: (1) pairwise Pearson correlations on MIMIC-III waveforms (Table I) motivate preferring PPG over ECG as an intermediate; (2) generators are trained with reconstruction MSE to map PPG→ECG; (3) an identical BiLSTM BP head is trained under ABP supervision for both the direct PPG→BP path and the mediated PPG→ECG→BP paths; (4) MAE/SD and BHS grades are reported on a held-out test partition. None of these steps equates a reported BP MAE to a fitted input by algebra or definition: generator losses optimize waveform fidelity, not BP error; the BP head is optimized end-to-end on ABP; real-ECG→BiLSTM is reported as an upper bound showing the drop comes from generation error rather than the head. Citations (U-Net, Wave-U-Net, CycleGAN, Transformer, MIMIC-III, prior cuffless BP baselines) are external architectures or datasets, not author-overlapping uniqueness theorems that force the result. Ordinary risks (possible segment-level rather than subject-wise splits, ICU vs wearable domain) affect absolute generalization and statistical independence of the t-tests, but they do not make the direct-vs-mediated ordering or the Grade A numbers true by construction. Therefore the derivation is self-contained empirical comparison with circularity score 0.
Axiom & Free-Parameter Ledger
free parameters (4)
- QC thresholds (SNR>15 dB; SBP 80–200; DBP 40–120) =
SNR>15 dB; SBP 80-200 mmHg; DBP 40-120 mmHg
- Segment length and overlap (256 samples @125 Hz, 50% overlap) =
256 samples (~2 s), 50% overlap
- BiLSTM and generator hyperparameters =
BiLSTM 128-//dropout 0.3; Adam 1e-3 (2e-4 CycleGAN); patience 3; batch 64
- Spectral cutoff 0–20 Hz for rFFT features =
0–20 Hz retained
axioms (5)
- domain assumption Arterial-line ABP maxima/minima in 2 s windows are the correct supervision targets for cuffless SBP/DBP estimation claims aimed at wearables.
- ad hoc to paper BHS Grade A/B can be declared from MAE≤5/10 and SD≤8/15 mmHg alone.
- domain assumption Holding the BP predictor fixed as one BiLSTM makes PPG→BP vs PPG→ECG→BP an apples-to-apples test of the intermediate representation.
- ad hoc to paper Pearson correlation between waveform channels measures which signal is the better intermediate representation for BP estimation.
- standard math Standard encoder-decoder / GAN / Transformer / DDPM training objectives and MIMIC-III synchrony are reliable enough for the comparison.
read the original abstract
Continuous cuffless blood pressure (BP) monitoring is essential for connected health systems and wearable devices, enabling early detection, longitudinal tracking, and personalized management of cardiovascular disease. Many prior approaches attempt to estimate BP indirectly by reconstructing electrocardiography (ECG) from photoplethysmography (PPG), assuming ECG provides a stronger physiological link to BP. However, ECG sensing is less accessible in wearable settings and may introduce unnecessary complexity. In this work, we first perform a large-scale physiological correlation analysis on the MIMIC-III waveform database, revealing that PPG exhibits substantially stronger coupling with arterial blood pressure (ABP) ($|r|=0.247$, $p<0.001$) than ECG does ($r=0.018$, $p=0.187$), challenging the assumption that ECG provides a superior intermediate representation. Motivated by this insight, we conduct a systematic comparison between direct PPG-to-BP prediction and ECG-mediated pipelines using multiple state-of-the-art deep learning models. Across 1.74M segments from 3,127 patients, direct PPG-to-BP prediction achieves British Hypertension Society Grade A performance ($\mathrm{MAE}_{\mathrm{SBP}} = 4.82 mmHg$, $\mathrm{MAE}_{\mathrm{DBP}} = 4.31 mmHg$), outperforming all ECG-mediated approaches, which achieve only Grade B accuracy. Our findings suggest that accurate continuous BP monitoring can be achieved directly from wearable PPG signals, enabling simpler, more efficient pipelines for real-world connected health systems.
Figures
Reference graph
Works this paper leans on
-
[1]
Cuff- less blood pressure estimation algorithms for continuous health-care monitoring,
M. Kachuee, M. M. Kiani, H. Mohammadzade, and M. Shabany, “Cuff- less blood pressure estimation algorithms for continuous health-care monitoring,”IEEE Trans. Biomed. Eng., vol. 64, no. 4, pp. 859–869, Apr. 2017
2017
-
[2]
Deep learning-based cuff-less blood pressure estimation,
O. M. El-Basyouni, M. A. El-Baky, and F. M. El-Feshawy, “Deep learning-based cuff-less blood pressure estimation,”J. Med. Syst., vol. 43, no. 8, Art. no. 245, 2019
2019
-
[3]
MIMIC-III, a freely accessible critical care database,
A. E. W. Johnsonet al., “MIMIC-III, a freely accessible critical care database,”Sci. Data, vol. 3, Art. no. 160035, May 2016
2016
-
[4]
U-Net: Convolutional net- works for biomedical image segmentation,
O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional net- works for biomedical image segmentation,” inMedical Image Comput- ing and Computer-Assisted Intervention – MICCAI 2015(Lecture Notes in Computer Science, vol. 9351). Springer, 2015, pp. 234–241
2015
-
[5]
Wave-U-Net: A multi-scale neural network for end-to-end audio source separation,
D. Stoller, S. Ewert, and S. Dixon, “Wave-U-Net: A multi-scale neural network for end-to-end audio source separation,” inProc. 19th Int. Soc. Music Inf. Retrieval Conf.(ISMIR), Paris, France, Sep. 2018, pp. 334– 340
2018
-
[6]
Unpaired image-to-image translation using cycle-consistent adversarial networks,
J.-Y . Zhu, T. Park, P. Isola, and A. A. Efros, “Unpaired image-to-image translation using cycle-consistent adversarial networks,” inProc. IEEE Int. Conf. Comput. Vis.(ICCV), Oct. 2017, pp. 2223–2232
2017
-
[7]
Attention is all you need,
A. Vaswaniet al., “Attention is all you need,” inAdvances in Neural Information Processing Systems 30(NeurIPS 2017), pp. 5998–6008, 2017
2017
-
[8]
Advancing PPG-based continuous blood pressure monitoring from a generative perspective,
H. Ji and P. Zhou, “Advancing PPG-based continuous blood pressure monitoring from a generative perspective,” inProc. 22nd ACM Conf. Embedded Networked Sensor Systems (SenSys), New York, NY , USA, 2024, pp. 661–674. doi: 10.1145/3666025.3699365
arXiv 2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.