Pith. sign in

REVIEW 4 major objections 5 minor 78 references

VidFormer: A novel end-to-end framework fused by 3DCNN and Transformer for Video-based Remote Physiological Measurement

T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read VidFormer claims to set a new state of the art in video-based remote heart-rate measurement by fusing a 3DCNN with a Transformer, reaching mean absolute errors below 1 bpm on large and motion-rich datasets.

desk verdict Plausible new fusion architecture, but the headline sub-1-bpm results rest on an underspecified evaluation protocol and the manuscript is incomplete. read the letter →

arxiv 2501.01691 v2 pith:ONQKBJ4D submitted 2025-01-03 cs.CV cs.AI

classification cs.CVcs.AI
keywords remotephotoplethysmographyrPPGheartrateestimation3DCNNTransformerdual-branchfusionspatiotemporalattentionvideo-basedvitalsigns
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

VidFormer is an end-to-end framework for remote photoplethysmography (rPPG) that reconstructs blood-volume-pulse signals from ordinary facial videos and then derives heart rate, heart rate variability, and respiration frequency from them. The paper's central claim is that by running a 3DCNN branch for local spatiotemporal features alongside a Transformer branch for global features, with explicit cross-branch information exchange, the model outperforms previously published methods on all five tested public datasets. The headline result is a mean absolute heart-rate error below 1 bpm on the large DEAP dataset and on the motion-heavy ECG-fitness dataset, which the authors attribute to the joint local-global modeling and to an improved skin reflection model that treats the face-to-signal mapping as time-dependent.

What carries the argument

The core mechanism is a dual-branch fusion architecture composed of five modules: a Stem for initial feature extraction, a Local Convolution Branch built on BS-3DCNN blocks augmented by a Global Attention 3DCNN (GA-3DCNN) with separate Spatial Attention and Time Attention, a Global Transformer Branch using a Spatial-Time Multi-headed Self-attention (ST-MHSA) that splits attention into spatial and temporal streams, a CNN-Transformer Interaction Module (CTIM) with Trans-Conv and Conv-Trans blocks that reshape and exchange features between the branches, and an rPPG Generation Module (RGM) that turns each branch's features into BVP estimates. The two outputs are optimized separately with a combined negative Pearson and Smooth L1 loss, and their heart-rate estimates are averaged.

What would settle it

Re-run VidFormer with a strict subject-exclusive split on DEAP and ECG-fitness, using the same window length and step size, and report the resulting MAE. If the mean absolute error rises above 1 bpm or the margin over PhysFormer narrows substantially, the paper's central SOTA claim would be falsified.

Watch

Extended reading notes

Core claim

The paper claims that a dual-branch fusion of 3DCNN and Transformer, named VidFormer, is the first such end-to-end architecture specifically designed for rPPG, and that it achieves superior heart-rate estimation accuracy on UBFC-rPPG, PURE, COHFACE, ECG-fitness, and DEAP. On the author's own terms, the discovery is that separately extracting local features (via a convolutional branch with spatial and temporal attention) and global features (via a Transformer branch with split spatial and temporal self-attention), and then exchanging information between the branches at multiple levels, lets the model reconstruct BVP signals accurately enough to keep mean absolute error below 1 bpm even on datasets with complex lighting, electrode patches, and exercise-induced motion. The improved skin reflection model, which rewrites the dichromatic model so that the mapping from blood-volume changes to skin color depends explicitly on time and pixel location, serves as the design rationale for combining local and global modeling rather than relying on either alone.

Load-bearing premise

The reported superiority assumes that the train/test protocol keeps the same subject's randomly selected video windows out of the test set, but the paper never states whether the split is subject-independent or how the random segment selection is seeded.

Editorial extensions

If this is right

  • If correct, non-contact heart-rate monitoring can hold mean absolute error below 1 bpm on large, varied datasets, not just on small controlled ones.
  • The cross-dataset results suggest the model generalizes across recording conditions, lighting, and subject populations without per-dataset retraining.
  • Ablations indicate that both the local convolutional branch and the global Transformer branch are individually necessary, and that the interaction module substantially improves accuracy over either branch alone.
  • The framework also produces competitive respiratory-frequency and heart-rate-variability estimates, broadening its potential use beyond heart rate.
  • The improved skin reflection model offers a principled reason to pair local and global feature extractors, a design choice other rPPG systems could adopt.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A reader should treat the sub-1 bpm figures as contingent on the evaluation protocol: if the random segment selection allowed windows from the same subject to appear in both training and testing, the model could memorize per-subject skin appearance and inflate accuracy.
  • The split spatial/temporal attention design could transfer to other video-understanding tasks where short-range texture and long-range dynamics both matter.
  • The ethnicity analysis suggests a testable extension: the reported lower accuracy on African subjects could be probed with cameras or preprocessing that linearize sensor response in dark skin tones.
  • A stricter test of the paper's central claim would be to fix the train/test split to be subject-independent and to report the random-segment seed, so other groups can reproduce the exact protocol.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes VidFormer, a dual-branch architecture combining 3DCNN and Transformer branches for video-based remote photoplethysmography (rPPG). The method introduces a 3DCNN branch with spatial and temporal attention (GA-3DCNN), a Transformer branch with separated spatial and temporal multi-head self-attention (ST-MHSA), and a CTIM module for cross-branch information exchange. The authors also present an 'enhanced skin reflection model' in Eq. (2), which reformulates the dichromatic reflection model with reflection terms that depend on the blood volume pulse, pixel location, and time. The paper reports state-of-the-art heart-rate estimation results on five public datasets (UBFC-rPPG, PURE, DEAP, ECG-fitness, COHFACE), with MAE values below 1 bpm on DEAP and ECG-fitness, along with HRV/RF results, ablations, and discussions on ethnicity, makeup, and exercise.

Significance. If the reported results hold under a fair, reproducible evaluation protocol, VidFormer would be a strong empirical contribution to the rPPG literature. The architecture is plausible and the paper covers a broad set of benchmarks and ablations, which is valuable. The paper also clearly attempts to address known limitations of both CNN and Transformer models for this task. However, the load-bearing claim of state-of-the-art performance is currently unverifiable because the train/test protocol is not specified, and the comparison to prior methods is not controlled. The proposed analytical model is not used to derive any testable prediction, so the theoretical contribution is overstated. These issues must be resolved before the empirical claims can be assessed.

major comments (4)
  1. [Section IV-B and IV-C] The intra-dataset evaluation protocol is not specified. Section IV-B states only that "we randomly select one segment and slice it using a window with length of 250 frames and a step size of 50 frames," with no statement about subject-independent train/test splits, the number of windows extracted per video, whether overlapping windows from the same subject can appear in both training and test sets, or the random seed used. Since the central claim is state-of-the-art performance, this omission makes Tables I-V unverifiable and creates a real risk of subject/segment leakage. The authors must specify the exact protocol for every dataset, including subject-independent splits and window sampling details.
  2. [Table I and Section IV-D1] The comparison with prior methods is not controlled. It is unclear whether all numbers in Table I were produced under the same protocol as VidFormer or quoted from original papers that may use different splits and preprocessing. The reported gaps (e.g., MAE 0.42 vs 1.10 bpm on PURE and 0.75 vs 3.03 bpm on DEAP relative to Physformer) are implausibly large without a shared evaluation setup. The authors should either re-run all comparison methods under an identical subject-independent protocol or clearly state the protocol and source for each entry in the table.
  3. [Tables I-X] All tables report single-point metrics without error bars, confidence intervals, or multiple-seed statistics. Given the random segment selection described in Section IV-B, the results may be sensitive to the chosen segments. The paper should report mean and standard deviation over at least three random seeds and, where possible, use paired statistical tests to support the claimed improvements over prior methods.
  4. [Section III-A, Eq. (2)] The claimed 'enhanced skin reflection model' in Eq. (2) is a notational restatement of Eq. (1) with ℓ_s and ℓ_d declared to depend on y_t, ρ, and t. No prediction from this model is derived, no parameter is estimated from it, and it does not constrain the architecture or loss. The statement that VidFormer is 'based on this improved model' is therefore not supported. The authors should either use the model to derive a testable component or remove the claim that the architecture is based on it.
minor comments (5)
  1. [Table I, HRCNN row] The MAE value for HRCNN is printed as '14 , 48' (likely 14.48); the formatting should be corrected.
  2. [Table XII] The column heading 'COHFACE-African-wM' is inconsistent with Section VI-B, which describes the subset as 'COHFACE with makeup'; rename the column to 'COHFACE-wM'.
  3. [Figures 17-19] Figures 17, 18, and 19 contain the literal placeholder string '121312312大苏打撒旦'; these should be replaced with proper figure content or removed.
  4. [Section III-D, Eq. (7)] The description 'P_os is a random number that satisfies a Gaussian distribution with a mean of 0 and a variance of 1' is ambiguous; please specify whether the positional encoding is fixed after initialization, learned, or re-sampled at every forward pass, and state its dimension.
  5. [Section IV-C] The cross-dataset protocol is underspecified; the paper should describe how frame rates, ROIs, and signal sampling rates are aligned across datasets, and how the model is adapted when moving from one dataset to another.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the enhanced skin reflection model is qualitative motivation, and all claimed results are empirical evaluations against external benchmarks.

full rationale

The paper's central claims are empirical: VidFormer is trained end-to-end on five public datasets and compared against external methods in Tables I-V. The 'enhanced skin reflection model' in Eq. (2) is introduced as conceptual motivation for the two-branch architecture; it contains unspecified mapping functions and is never fitted to data, nor is any quantitative prediction derived from it that is then checked against the reported results. The network optimization uses standard Pearson and Smooth L1 losses, and the ablations measure empirical contributions of each module. There are no load-bearing self-citations: the cited works [55] and [57] are used for standard architectural components (video patch embedding) and for generating synthetic PPG signals from ECG, respectively, and neither is a prior result of the present authors invoked to force the main conclusion. The reported state-of-the-art performance may raise reproducibility or evaluation-protocol concerns (e.g., the train/test split is not fully specified), but an unspecified split is a verification gap, not a circular reduction. No step in the paper equates an output to an input by construction, fits a parameter and renames it as a prediction, or imports a uniqueness theorem from self-citation. Accordingly, no circular step is identified.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

The analytical model in Section III-A contributes no derived parameters; it only narratively motivates the architecture. All performance-related numbers come from supervised training, which depends on the listed hyperparameters and domain assumptions. No new physical entities are introduced.

free parameters (6)
  • loss_balance_alpha = 0.5
    Balances negative Pearson loss and Smooth L1 loss in Eq (16); chosen by hand, no sensitivity analysis.
  • cube_patch_size = 25 x 16 x 16
    Patch dimensions for the Transformer branch; set in Section IV-B.
  • window_length_and_step = 250 frames, step 50
    Video slicing parameters; set in Section IV-B; overlap may create train/test leakage risk.
  • frame_resolution = 128 x 128
    All face crops resized to 128x128 before training (Section IV-B).
  • training_hyperparameters = batch size 2, max lr 8e-5, min lr 2e-9, weight decay 5e-4
    Reported for training; not clear if tuned per dataset; no seeds.
  • architecture_depth_N = not reported
    Number of CTIM/transformer blocks is shown as 'N x' in figures but never specified; central to model capacity.
assumptions (5)
  • standard math Dichromatic reflection model C_k(t) = I(t)(vs(t)+vd(t))+vn(t)
    Basis of the skin reflection analysis, cited from [5], [20], [53].
  • domain assumption Under stationary conditions, skin color and blood flow are bijectively related
    Assumed in Section III-A to motivate the mapping between BVP and skin color; violated by real motion and lighting change.
  • ad hoc to paper The reflection mappings ℓs and ℓd depend on BVP y_t(t), pixel location ρ, and time t
    Introduced in Eq (2) as an 'enhanced model' but never validated or used to derive network parameters.
  • domain assumption CNN provides useful inductive bias and Transformer provides global modeling
    Used to justify the dual-branch architecture (Sections I, III-B).
  • domain assumption Ground truth PPG from pulse oximeters is reliable for training and evaluation
    All five datasets use contact PPG/ECG-derived references; no quality filtering discussed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of VidFormer: A novel end-to-end framework fused by 3DCNN and Transformer for Video-based Remote Physiological Measurement." pith.science (2026). https://pith.science/paper/ONQKBJ4D

@misc{pith2026250101691,
  author       = {Pith},
  title        = {Pith review of: VidFormer: A novel end-to-end framework fused by 3DCNN and Transformer for Video-based Remote Physiological Measurement},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ONQKBJ4D}},
  note         = {Machine review of arXiv:2501.01691}
}
read the original abstract

Remote physiological signal measurement based on facial videos, also known as remote photoplethysmography (rPPG), involves predicting changes in facial vascular blood flow from facial videos. While most deep learning-based methods have achieved good results, they often struggle to balance performance across small and large-scale datasets due to the inherent limitations of convolutional neural networks (CNNs) and Transformer. In this paper, we introduce VidFormer, a novel end-to-end framework that integrates 3-Dimension Convolutional Neural Network (3DCNN) and Transformer models for rPPG tasks. Initially, we conduct an analysis of the traditional skin reflection model and subsequently introduce an enhanced model for the reconstruction of rPPG signals. Based on this improved model, VidFormer utilizes 3DCNN and Transformer to extract local and global features from input data, respectively. To enhance the spatiotemporal feature extraction capabilities of VidFormer, we incorporate temporal-spatial attention mechanisms tailored for both 3DCNN and Transformer. Additionally, we design a module to facilitate information exchange and fusion between the 3DCNN and Transformer. Our evaluation on five publicly available datasets demonstrates that VidFormer outperforms current state-of-the-art (SOTA) methods. Finally, we discuss the essential roles of each VidFormer module and examine the effects of ethnicity, makeup, and exercise on its performance.

Figures

Figures reproduced from arXiv: 2501.01691 by the authors.

Figure 1
Figure 1. The potential mapping relationship between each frame in the video [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. The overall framework of VidFormer. VidFormer leverages 3DCNN and Transformers to extract local and global features from input facial videos, [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. The illustration of the GA-3DCNN module. GA-3DCNN incorporates a global attention mechanism to assist BS-3DCNN in focusing on critical [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (12 more)
Figure 4
Figure 4. Figure 4: The schematic diagram of multi-head attention mechanism in Spatial [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: The illustration of BS-3DCNN. BS-3DCNN is designed for feature [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]
Figure 7
Figure 7. Figure 7: The process of cube patch of video data, where [PITH_FULL_IMAGE:figures/full_fig_p006_7.png]
Figure 8
Figure 8. Figure 8: The illustration of Global Transformer Branch. The Global Transformer Branch based on the Transformer architecture is designed to extract global [PITH_FULL_IMAGE:figures/full_fig_p007_8.png]
Figure 9
Figure 9. Figure 9: The overall framework of ST-MHSA. ST-MHSA features a parallel [PITH_FULL_IMAGE:figures/full_fig_p007_9.png]
Figure 11
Figure 11. Figure 11: The overall framework of RGM. Local Convolution Branch and Global Transformer Branch, respectively. Given the invariant shape of the T-dimension across the features within the Local Convolution Branch, we employ a 3-dimensional convolution to reduce the number of feat…
Figure 12
Figure 12. Figure 12: Face images in COHFACE, DEAP and ECG-fitness, (a) denotes the [PITH_FULL_IMAGE:figures/full_fig_p010_12.png]
Figure 13
Figure 13. Figure 13: (a) is the scatter plots between ground-truth HR and estimated HR on UBFC-rPPG, PURE, COHFACE, ECG-fitness and DEAP. The straight red [PITH_FULL_IMAGE:figures/full_fig_p011_13.png]
Figure 14
Figure 14. Figure 14: The visual comparison between estimated rPPG signals (orange curves) and their corresponding ground truth BVP signals (blue curves) on UBFC [PITH_FULL_IMAGE:figures/full_fig_p013_14.png]
Figure 15
Figure 15. Figure 15: The training and testing loss curves of wo-LCB and w-LCB on [PITH_FULL_IMAGE:figures/full_fig_p014_15.png]
Figure 18
Figure 18. Figure 18: The performance of BVP signal reconstruction on subjects with 121312312大苏打撒旦 makeup [PITH_FULL_IMAGE:figures/full_fig_p015_18.png]
Figure 19
Figure 19. Figure 19: The performance of BVP signal reconstruction on subjects with [PITH_FULL_IMAGE:figures/full_fig_p016_19.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

78 extracted references · 74 canonical work pages

  1. [1]

    Deep learning for healthcare applications based on physiological signals: A review,

    O. Faust, Y . Hagiwara, T. J. Hong, O. S. Lih, and U. R. Acharya, “Deep learning for healthcare applications based on physiological signals: A review,” Comput. Methods Programs Biomedicine , vol. 161, pp. 1–13, 2018

  2. [2]

    Heart rate variability: a review,

    U. Rajendra Acharya, K. Paul Joseph, N. Kannathal, C. M. Lim, and J. S. Suri, “Heart rate variability: a review,” Medical Biol. Eng. Comput., vol. 44, pp. 1031–1051, 2006

  3. [3]

    Heart rate variability: measurement and clinical utility,

    R. E. Kleiger, P. K. Stein, and J. T. Bigger Jr, “Heart rate variability: measurement and clinical utility,” Ann. Noninvasive Electrocardiology , vol. 10, no. 1, pp. 88–101, Jan. 2005

  4. [4]

    Blind source separation and independent component analysis: A review,

    S. Choi, A. Cichocki, H.-M. Park, and S.-Y . Lee, “Blind source separation and independent component analysis: A review,” Neural Inf. Process. Lett. Reviews, vol. 6, no. 1, pp. 1–57, Jan. 2005

  5. [5]

    Robust pulse rate from chrominance-based rppg,

    G. De Haan and V . Jeanne, “Robust pulse rate from chrominance-based rppg,” IEEE Trans. Biomed. Eng. , vol. 60, no. 10, pp. 2878–2886, Jun. 2013

  6. [6]

    Improved motion robustness of remote- ppg by using the blood volume pulse signature,

    G. De Haan and A. Van Leest, “Improved motion robustness of remote- ppg by using the blood volume pulse signature,” Physiological Meas. , vol. 35, no. 9, p. 1913, Aug. 2014

  7. [7]

    Contact-free screening of atrial fibrillation by a smartphone using facial pulsatile photoplethysmographic signals,

    B. P. Yan, W. H. Lai, C. K. Chan, S. C.-H. Chan, L.-H. Chan, K.-M. Lam, H.-W. Lau, C.-M. Ng, L.-Y . Tai, K.-W. Yip et al. , “Contact-free screening of atrial fibrillation by a smartphone using facial pulsatile photoplethysmographic signals,” J. Amer. Heart Assoc., vol. 7, no. 8, p. e008585, Apr. 2018

  8. [8]

    Synrhythm: Learning a deep heart rate estimator from general to specific,

    X. Niu, H. Han, S. Shan, and X. Chen, “Synrhythm: Learning a deep heart rate estimator from general to specific,” in 2018 24th Int. Conf. Pattern Recognit. IEEE, 2018, pp. 3580–3585

Show all 78 references
  1. [9]

    Remote heart rate measurement from highly compressed facial videos: an end-to-end deep learning solution with video enhancement,

    Z. Yu, W. Peng, X. Li, X. Hong, and G. Zhao, “Remote heart rate measurement from highly compressed facial videos: an end-to-end deep learning solution with video enhancement,” in Proc. IEEE Int. Conf. Comput. Vis., 2019, pp. 151–160. JOURNAL OF LATEX CLASS FILES, VOL. 18, NO. ...

  2. [10]

    Facial video-based remote physiological measurement via self-supervised learning,

    Z. Yue, M. Shi, and S. Ding, “Facial video-based remote physiological measurement via self-supervised learning,” IEEE Trans. Pattern Anal. Mach. Intell., Jul. 2023

  3. [11]

    Heart rate recovery and treadmill exercise score as predictors of mortality in patients referred for exercise ecg,

    E. O. Nishime, C. R. Cole, E. H. Blackstone, F. J. Pashkow, and M. S. Lauer, “Heart rate recovery and treadmill exercise score as predictors of mortality in patients referred for exercise ecg,” Jama, vol. 284, no. 11, pp. 1392–1398, Sep. 2000

  4. [12]

    Accurate heart rate monitoring during physical exercises using ppg,

    A. Temko, “Accurate heart rate monitoring during physical exercises using ppg,” IEEE Trans. Biomed. Eng. , vol. 64, no. 9, pp. 2016–2024, Mar. 2017

  5. [13]

    Rhythmnet: End-to-end heart rate estimation from face via spatial-temporal representation,

    X. Niu, S. Shan, H. Han, and X. Chen, “Rhythmnet: End-to-end heart rate estimation from face via spatial-temporal representation,” IEEE Trans. Image Process., vol. 29, pp. 2409–2423, Oct. 2019

  6. [14]

    Facial-video-based physiological signal measurement: Recent advances and affective applications,

    Z. Yu, X. Li, and G. Zhao, “Facial-video-based physiological signal measurement: Recent advances and affective applications,” IEEE Signal Process. Mag., vol. 38, no. 6, pp. 50–58, Oct. 2021

  7. [15]

    Meta-rppg: Remote heart rate estima- tion using a transductive meta-learner,

    E. Lee, E. Chen, and C.-Y . Lee, “Meta-rppg: Remote heart rate estima- tion using a transductive meta-learner,” in Proc. Eur. Conf. Comput. Vis. Springer, Aug. 2020, pp. 392–409

  8. [16]

    Dual-gan: Joint bvp and noise modeling for remote physiological measurement,

    H. Lu, H. Han, and S. K. Zhou, “Dual-gan: Joint bvp and noise modeling for remote physiological measurement,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2021, pp. 12 404–12 413

  9. [17]

    Video- based heart rate measurement: Recent advances and future prospects,

    X. Chen, J. Cheng, R. Song, Y . Liu, R. Ward, and Z. J. Wang, “Video- based heart rate measurement: Recent advances and future prospects,” IEEE Trans. Instrum. Meas., vol. 68, no. 10, pp. 3600–3615, Nov. 2018

  10. [18]

    Modeling diffuse reflectance from semi- infinite turbid media: application to the study of skin optical properties,

    G. Zonios and A. Dimou, “Modeling diffuse reflectance from semi- infinite turbid media: application to the study of skin optical properties,” Opt. Exp., vol. 14, no. 19, pp. 8661–8674, 2006

  11. [19]

    Remote plethysmo- graphic imaging using ambient light

    W. Verkruysse, L. O. Svaasand, and J. S. Nelson, “Remote plethysmo- graphic imaging using ambient light.” Opt. Exp. , vol. 16, no. 26, pp. 21 434–21 445, 2008

  12. [20]

    Algorithmic principles of remote ppg,

    W. Wang, A. C. Den Brinker, S. Stuijk, and G. De Haan, “Algorithmic principles of remote ppg,” IEEE Trans. Biomed. Eng. , vol. 64, no. 7, pp. 1479–1491, Sep. 2016

  13. [21]

    Deep learning methods for remote heart rate measurement: A review and future research agenda,

    C.-H. Cheng, K.-L. Wong, J.-W. Chin, T.-T. Chan, and R. H. So, “Deep learning methods for remote heart rate measurement: A review and future research agenda,” Sensors, vol. 21, no. 18, p. 6296, Sep. 2021

  14. [22]

    And-rppg: A novel denoising-rppg network for improving remote heart rate estimation,

    B. Lokendra and G. Puneet, “And-rppg: A novel denoising-rppg network for improving remote heart rate estimation,” Comput. Biol. Medicine , vol. 141, p. 105146, Feb. 2022

  15. [23]

    Non-contact, automated cardiac pulse measurements using video imaging and blind source separation

    M.-Z. Poh, D. J. McDuff, and R. W. Picard, “Non-contact, automated cardiac pulse measurements using video imaging and blind source separation.” Opt. Exp., vol. 18, no. 10, pp. 10 762–10 774, May 2010

  16. [24]

    Measuring pulse rate with a webcam—a non-contact method for evaluating cardiac activity,

    M. Lewandowska, J. Rumi ´nski, T. Kocejko, and J. Nowak, “Measuring pulse rate with a webcam—a non-contact method for evaluating cardiac activity,” in 2011 Federated Conf. Comput. Sci. Inf. Syst. IEEE, Nov. 2011, pp. 405–410

  17. [25]

    Deepphys: Video-based physiological mea- surement using convolutional attention networks,

    W. Chen and D. McDuff, “Deepphys: Video-based physiological mea- surement using convolutional attention networks,” in Proc. Eur. Conf. Comput. Vis., 2018, pp. 349–365

  18. [26]

    Robust heart rate estimation with spatial–temporal attention network from facial videos,

    M. Hu, F. Qian, X. Wang, L. He, D. Guo, and F. Ren, “Robust heart rate estimation with spatial–temporal attention network from facial videos,” IEEE Trans. Cognitive Develop. Syst. , vol. 14, no. 2, pp. 639–647, Feb. 2021

  19. [27]

    The way to my heart is through contrastive learning: Remote photoplethysmography from unlabelled video,

    J. Gideon and S. Stent, “The way to my heart is through contrastive learning: Remote photoplethysmography from unlabelled video,” in Proc. IEEE Int. Conf. Comput. Vis. , 2021, pp. 3995–4004

  20. [28]

    Rtrppg: An ultra light 3dcnn for real-time remote photoplethysmography,

    D. Botina-Monsalve, Y . Benezeth, and J. Miteran, “Rtrppg: An ultra light 3dcnn for real-time remote photoplethysmography,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2022, pp. 2146–2154

  21. [29]

    Transrppg: Remote photoplethys- mography transformer for 3d mask face presentation attack detection,

    Z. Yu, X. Li, P. Wang, and G. Zhao, “Transrppg: Remote photoplethys- mography transformer for 3d mask face presentation attack detection,” IEEE Signal Process. Lett. , vol. 28, pp. 1290–1294, Jun. 2021

  22. [30]

    Continuous gesture segmentation and recognition using 3dcnn and convolutional lstm,

    G. Zhu, L. Zhang, P. Shen, J. Song, S. A. A. Shah, and M. Bennamoun, “Continuous gesture segmentation and recognition using 3dcnn and convolutional lstm,” IEEE Trans. Multimedia, vol. 21, no. 4, pp. 1011– 1021, Sep. 2018

  23. [31]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Adv. Neural Inf. Process. Syst. , vol. 30, 2017

  24. [32]

    Are transformer-based models more robust than cnn-based models?

    Z. Liu, S. Qian, C. Xia, and C. Wang, “Are transformer-based models more robust than cnn-based models?” Neural Netw., vol. 172, p. 106091, Apr. 2024

  25. [33]

    Un- supervised skin tissue segmentation for remote photoplethysmography,

    S. Bobbia, R. Macwan, Y . Benezeth, A. Mansouri, and J. Dubois, “Un- supervised skin tissue segmentation for remote photoplethysmography,” Pattern Recognit. Lett., vol. 124, pp. 82–90, Jun. 2019

  26. [34]

    Non-contact video-based pulse rate measurement on a mobile service robot,

    R. Stricker, S. M ¨uller, and H.-M. Gross, “Non-contact video-based pulse rate measurement on a mobile service robot,” in IEEE Int Symp. Robot Hum. Interactive Commun. IEEE, Oct. 2014, pp. 1056–1062

  27. [35]

    Deap: A database for emotion analysis; using physiological signals,

    S. Koelstra, C. Muhl, M. Soleymani, J.-S. Lee, A. Yazdani, T. Ebrahimi, T. Pun, A. Nijholt, and I. Patras, “Deap: A database for emotion analysis; using physiological signals,” IEEE Trans. Affect. Comput., vol. 3, no. 1, pp. 18–31, Jun. 2011

  28. [36]

    Visual heart rate estimation with convolutional neural network,

    R. ˇSpetl´ık, V . Franc, and J. Matas, “Visual heart rate estimation with convolutional neural network,” in Proc. Brit. Mach. Vis. Conf., Newcastle, 2018, pp. 3–6

  29. [37]

    A reproducible study on remote heart rate measurement,

    G. Heusch, A. Anjos, and S. Marcel, “A reproducible study on remote heart rate measurement,” arXiv preprint arXiv:1709.00962 , 2017

  30. [38]

    Remote photoplethys- mography with constrained ica using periodicity and chrominance constraints,

    R. Macwan, Y . Benezeth, and A. Mansouri, “Remote photoplethys- mography with constrained ica using periodicity and chrominance constraints,” Biomed. Eng. Online , vol. 17, no. 1, pp. 1–22, Feb. 2018

  31. [39]

    Heart rate and heart rate variability from single-channel video and ica integration of multiple signals,

    R. Favilla, V . C. Zuccala, and G. Coppini, “Heart rate and heart rate variability from single-channel video and ica integration of multiple signals,” IEEE J. Biomed. Health Inform., vol. 23, no. 6, pp. 2398–2408, Nov. 2018

  32. [40]

    Non-contact heart rate monitoring by combining convolutional neural network skin detection and remote photoplethysmography via a low-cost camera,

    C. Tang, J. Lu, and J. Liu, “Non-contact heart rate monitoring by combining convolutional neural network skin detection and remote photoplethysmography via a low-cost camera,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops , 2018, pp. 1309–1315

  33. [41]

    A lstm-based re- altime signal quality assessment for photoplethysmogram and remote photoplethysmogram,

    H. Gao, X. Wu, C. Shi, Q. Gao, and J. Geng, “A lstm-based re- altime signal quality assessment for photoplethysmogram and remote photoplethysmogram,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., 2021, pp. 3831–3840

  34. [42]

    Remote heart rate estimation by signal quality attention network,

    H. Gao, X. Wu, J. Geng, and Y . Lv, “Remote heart rate estimation by signal quality attention network,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., 2022, pp. 2122–2129

  35. [43]

    Long short-term memory deep-filter in remote photoplethysmography,

    D. Botina-Monsalve, Y . Benezeth, R. Macwan, P. Pierrart, F. Parra, K. Nakamura, R. Gomez, and J. Miteran, “Long short-term memory deep-filter in remote photoplethysmography,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops , 2020, pp. 306–307

  36. [44]

    rppg-mae: Self- supervised pretraining with masked autoencoders for remote physiolog- ical measurements,

    X. Liu, Y . Zhang, Z. Yu, H. Lu, H. Yue, and J. Yang, “rppg-mae: Self- supervised pretraining with masked autoencoders for remote physiolog- ical measurements,” IEEE Trans. Multimedia , Feb. 2024

  37. [45]

    Demodulation based transformer for rppg generation and heart rate estimation,

    X. Zhang, Z. Xia, L. Liu, and X. Feng, “Demodulation based transformer for rppg generation and heart rate estimation,” IEEE Signal Process. Lett., vol. 30, pp. 1042–1046, Aug. 2023

  38. [46]

    Self-supervised rgb-nir fusion video vision transformer framework for rppg estimation,

    S. Park, B.-K. Kim, and S.-Y . Dong, “Self-supervised rgb-nir fusion video vision transformer framework for rppg estimation,” IEEE Trans. Instrum. Meas., vol. 71, pp. 1–10, Oct. 2022

  39. [47]

    Physformer: Facial video-based physiological measurement with temporal difference transformer,

    Z. Yu, Y . Shen, J. Shi, H. Zhao, P. H. Torr, and G. Zhao, “Physformer: Facial video-based physiological measurement with temporal difference transformer,” in Proc. IEEE/CVF Cof. Comput. Vis Pattern Recognit. , 2022, pp. 4186–4196

  40. [48]

    Shuffle-rppgnet: Efficient network with global context for remote heart rate variability measurement,

    H. Kuang, C. Ao, X. Ma, and X. Liu, “Shuffle-rppgnet: Efficient network with global context for remote heart rate variability measurement,” IEEE Sensors J., vol. 23, May 2023

  41. [49]

    Pulsenet: A multitask learning network for remote heart rate estimation,

    R.-N. Yin, R.-S. Jia, Z. Cui, and H.-M. Sun, “Pulsenet: A multitask learning network for remote heart rate estimation,” Knowl. Based Syst. , vol. 239, p. 108048, Mar. 2022

  42. [50]

    Cross-architecture self-supervised video representation learning,

    S. Guo, Z. Xiong, Y . Zhong, L. Wang, X. Guo, B. Han, and W. Huang, “Cross-architecture self-supervised video representation learning,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2022, pp. 19 270–19 279

  43. [51]

    Vitae: Vision transformer advanced by exploring intrinsic inductive bias,

    Y . Xu, Q. Zhang, J. Zhang, and D. Tao, “Vitae: Vision transformer advanced by exploring intrinsic inductive bias,” Advances Neural Inf. Process. Syst., vol. 34, pp. 28 522–28 535, 2021

  44. [52]

    A random cnn sees objects: One inductive bias of cnn and its applications,

    Y .-H. Cao and J. Wu, “A random cnn sees objects: One inductive bias of cnn and its applications,” in Proc. AAAI Conf. Artif. Intell. , vol. 36, no. 1, 2022, pp. 194–202

  45. [53]

    Dichromatic reflection models for a variety of materials,

    S. Tominaga, “Dichromatic reflection models for a variety of materials,” Color Res. & Application , vol. 19, no. 4, pp. 277–285, 1994

  46. [54]

    Image quality metrics: Psnr vs. ssim,

    A. Hore and D. Ziou, “Image quality metrics: Psnr vs. ssim,” in 2010 20th Int. Conf. Pattern Recognit. IEEE, 2010, pp. 2366–2369

  47. [55]

    Vivit: A video vision transformer,

    A. Arnab, M. Dehghani, G. Heigold, C. Sun, M. Lu ˇci´c, and C. Schmid, “Vivit: A video vision transformer,” in Proc. IEEE/CVF Int. Conf. Comput. Vis., 2021, pp. 6836–6846

  48. [56]

    Pulsegan: Learning to generate realistic pulse waveforms in remote photoplethys- mography,

    R. Song, H. Chen, J. Cheng, C. Li, Y . Liu, and X. Chen, “Pulsegan: Learning to generate realistic pulse waveforms in remote photoplethys- mography,” IEEE J. Biomed. Health Inform. , vol. 25, no. 5, pp. 1373– 1384, Jan. 2021

  49. [57]

    Time-frequency learning frame- work for rppg signal estimation using scalogram based feature map of facial video data,

    M. Das, M. Bhuyan, and L. Sharma, “Time-frequency learning frame- work for rppg signal estimation using scalogram based feature map of facial video data,” IEEE Trans. Instrum. Meas. , Jun. 2023. JOURNAL OF LATEX CLASS FILES, VOL. 18, NO. 9, SEPTEMBER 2020 18

  50. [58]

    Pytorch: An imperative style, high-performance deep learning library,

    A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al. , “Pytorch: An imperative style, high-performance deep learning library,” Advances Neural Inf. Process. Syst. , vol. 32, 2019

  51. [59]

    Decoupled weight decay regularization,

    I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” arXiv preprint arXiv:1711.05101 , 2017

  52. [60]

    Analysing noisy driver physiology real-time using off-the-shelf sensors: Heart rate analysis software from the taking the fast lane project,

    P. Van Gent, H. Farah, N. Van Nes, and B. Van Arem, “Analysing noisy driver physiology real-time using off-the-shelf sensors: Heart rate analysis software from the taking the fast lane project,” J. Open Res. Softw., vol. 7, no. 1, pp. 1–9, 2019

  53. [61]

    Advancements in non- contact, multiparameter physiological measurements using a webcam,

    M.-Z. Poh, D. J. McDuff, and R. W. Picard, “Advancements in non- contact, multiparameter physiological measurements using a webcam,” IEEE Trans. Biomed. Eng. , vol. 58, no. 1, pp. 7–11, Oct. 2010

  54. [62]

    The obf database: A large face video database for remote physiological signal measurement and atrial fibrillation detection,

    X. Li, I. Alikhani, J. Shi, T. Seppanen, J. Junttila, K. Majamaa- V oltti, M. Tulppo, and G. Zhao, “The obf database: A large face video database for remote physiological signal measurement and atrial fibrillation detection,” in 2018 13th IEEE Int. Conf. Autom. Face & Gesture ...

  55. [63]

    Video-based remote physiological measurement via cross-verified feature disentangling,

    X. Niu, Z. Yu, H. Han, X. Li, S. Shan, and G. Zhao, “Video-based remote physiological measurement via cross-verified feature disentangling,” in Comput. Vis. ECCV 2020: Eur. Conf. Springer, 2020, pp. 295–310

  56. [64]

    Contrast-phys: Unsupervised video-based remote physiological measurement via spatiotemporal contrast,

    Z. Sun and X. Li, “Contrast-phys: Unsupervised video-based remote physiological measurement via spatiotemporal contrast,” in Eur. Conf. on Comput. Vis. Springer, Oct. 2022, pp. 492–510

  57. [65]

    A novel algorithm for remote photoplethysmography: Spatial subspace rotation,

    W. Wang, S. Stuijk, and G. De Haan, “A novel algorithm for remote photoplethysmography: Spatial subspace rotation,” IEEE Trans. Biomed. Eng., vol. 63, no. 9, pp. 1974–1984, Dec. 2015

  58. [66]

    Remote heart rate measurement from face videos under realistic situations,

    X. Li, J. Chen, G. Zhao, and M. Pietikainen, “Remote heart rate measurement from face videos under realistic situations,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2014, pp. 4264–4271

  59. [67]

    Cpulse: heart rate estimation from rgb videos under realistic conditions,

    A. D. Mehta and H. Sharma, “Cpulse: heart rate estimation from rgb videos under realistic conditions,” IEEE Trans. Instrum. Meas. , Aug. 2023

  60. [68]

    Learning motion-robust remote photoplethys- mography through arbitrary resolution videos,

    J. Li, Z. Yu, and J. Shi, “Learning motion-robust remote photoplethys- mography through arbitrary resolution videos,” in Proc. AAAI Conf. Artif. Intell., vol. 37, no. 1, 2023, pp. 1334–1342

  61. [69]

    Non-contrastive unsuper- vised learning of physiological signals from video,

    J. Speth, N. Vance, P. Flynn, and A. Czajka, “Non-contrastive unsuper- vised learning of physiological signals from video,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2023, pp. 14 464–14 474

  62. [70]

    Simper: Simple self-supervised learning of periodic targets,

    Y . Yang, X. Liu, J. Wu, S. Borac, D. Katabi, M.-Z. Poh, and D. McDuff, “Simper: Simple self-supervised learning of periodic targets,” arXiv preprint arXiv:2210.03115, 2022

  63. [71]

    A self-supervised learning network for remote heart rate measurement,

    N. Zhang, H.-M. Sun, J.-R. Ma, and R.-S. Jia, “A self-supervised learning network for remote heart rate measurement,” Meas., p. 114379, Mar. 2024

  64. [72]

    Recognizing, fast and slow: Complex emotion recognition with facial expression detection and remote physiological measurement,

    Y .-C. Wu, L.-W. Chiu, C.-C. Lai, B.-F. Wu, and S. S. Lin, “Recognizing, fast and slow: Complex emotion recognition with facial expression detection and remote physiological measurement,” IEEE Trans. Affect. Comput., Mar. 2023

  65. [73]

    Deep facial expression recognition: A survey,

    S. Li and W. Deng, “Deep facial expression recognition: A survey,” IEEE Trans. Affect. Comput., vol. 13, no. 3, pp. 1195–1215, Mar. 2020

  66. [74]

    Atrial fibrillation detection from face videos by fusing subtle variations,

    J. Shi, I. Alikhani, X. Li, Z. Yu, T. Sepp ¨anen, and G. Zhao, “Atrial fibrillation detection from face videos by fusing subtle variations,” IEEE Trans. Circuits Syst. Video Technol., vol. 30, no. 8, pp. 2781–2795, Jul. 2019

  67. [75]

    Vision-based heart rate estimation via a two-stream cnn,

    Z.-K. Wang, Y . Kao, and C.-T. Hsu, “Vision-based heart rate estimation via a two-stream cnn,” in 2019 IEEE Int. Conf. Image Process. IEEE, 2019, pp. 3327–3331

  68. [76]

    A general remote photoplethysmography estimator with spatiotemporal convolutional network,

    S.-Q. Liu and P. C. Yuen, “A general remote photoplethysmography estimator with spatiotemporal convolutional network,” in 2020 15th IEEE Int. Conf. Auto. Face Gesture Recognit. IEEE, 2020, pp. 481–488

  69. [77]

    Camera measurement of physiological vital signs,

    D. McDuff, “Camera measurement of physiological vital signs,” ACM Comput. Surveys, vol. 55, no. 9, pp. 1–40, 2023

  70. [78]

    Estimation of reflectance, transmittance, and absorbance of cosmetic foundation layer on skin using translucency of skin,

    K. Yoshida and N. Okiyama, “Estimation of reflectance, transmittance, and absorbance of cosmetic foundation layer on skin using translucency of skin,” Opt. Exp., vol. 29, no. 24, pp. 40 038–40 050, 2021. Jiachen Li received the B.S. degree in informa- tion engineering from the...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.