{"id":"017c45d8-8f11-49fe-9a45-d91d485522d7","arxiv_id":"2505.01299","paper_version":4,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"Camera-based pulse estimation in a driving simulator reaches about 5 bpm error only after applying a linear correction fitted to the same data, so out-of-sample accuracy is not established.","lead":"Video of a driver's face can be used to estimate pulse rate during a driving simulator session, with a reported error of about 5 beats per minute after adding a video-amplification step. The catch is that the best error numbers come from fitting a correction on the same recordings used to measure the error, so the true real-world accuracy is still open.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline MAE of 5.04 bpm is an in-sample number: the linear correction in §2.7 and the moving-average window in §2.5 were fit against the same Empatica E4 data used to report error, with no held-out split.","rationale":"The reader identified the most load-bearing weakness: the headline error reduction comes from parameters fit to the same data used for evaluation, with no held-out split. My reading agrees. The paper is otherwise transparent about its retrospective design, low cross-correlation between SCLI and BVP, and the Empatica E4's known biases, all of which make external validation more important rather than less. The group-level age difference is a more robust result because it is significant before correction in both B.EVM and A.EVM, but its post-correction p-values in Table 2 should be treated cautiously for the same in-sample reason. Because the issue is addressable by a calibration-test split or by reporting the external-correction numerics, the existing CONDITIONAL verdict stands; no change is needed. A secondary inconsistency worth correcting is the stated number of recordings (79 in the intro and discussion versus 75 implied by the §2.1 counts and exclusions), but it does not alter the main methodological concern.","tokens_in":26597,"tokens_out":4118,"duration_ms":45124,"concrete_test":"Run leave-one-participant-out (or k-fold) cross-validation over all recordings: fit the linear correction (§2.7) and select the moving-average window width (§2.5) using only training folds, then compute MAE/RMSE on held-out folds for B.EVM and A.EVM. Also report numeric MAE/RMSE when applying only the external Medarević et al. [66] correction (a=0.32, b=-30.42). If held-out or external MAE remains near 5 bpm, the concern does not land; if it reverts toward the uncorrected 6.5–10.5 bpm range, the headline accuracy is an in-sample artifact.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim depends on parameters that are fit to the evaluation data. In §2.7, the linear correction y = a*x + b (B.EVM: a=0.94, b=-69.41; A.EVM: a=0.96, b=-74.01) is fit to the differences between video-derived PR and Empatica E4 PR on the same recordings whose errors are then reported after correction in Figure 5. In §2.5, the moving-average window width is selected per sequence to minimize MAE against the same reference and then averaged into a single width. No cross-validation, held-out participant split, or nested tuning is reported, so the reported MAE/RMSE are best-case in-sample optima rather than predictive accuracies. The external correction from Medarević et al. [66] is a legitimate out-of-sample anchor, but the paper never reports numeric MAE/RMSE for the bottom panel of Figure 5 using that correction, only that errors decrease. The age-group finding has some pre-correction support (B.EVM p=0.04, A.EVM p=0.01 in Table 2), but the strengthened p-values and effect sizes in the p.f. columns are computed after applying the same in-sample correction, so they are partly entangled with the fitted parameters. The feasibility conclusion that a webcam pipeline can substitute for a wrist sensor at moderate error is therefore under-supported until the correction generalizes to unseen recordings.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents a contactless pulse rate (PR) estimation pipeline for driving-simulator scenarios. The pipeline uses YOLO-based face detection, eye exclusion to form a facial skin ROI, optional Eulerian Video Magnification (EVM), PCA-based extraction of a light-intensity signal, and a modified Pan–Tompkins peak detector with a moving-average smoothing step. The authors compare video-derived PR against simultaneously recorded Empatica E4 reference values, report MAE/RMSE before and after a linear correction, test whether EVM improves accuracy, evaluate the method on two public datasets, and compare younger versus older drivers. The headline result is an MAE of 5.04 bpm and RMSE of 6.38 bpm after EVM and linear correction, with the claim that this demonstrates feasibility of rPPG-based PR monitoring in driving simulators. The paper also includes a time-complexity analysis of EVM and an independent assessment of Empatica E4 bias using Faros 360 data.","tokens_in":26898,"tokens_out":5407,"duration_ms":50958,"significance":"If the reported accuracy were out-of-sample, the study would be a useful feasibility demonstration for contactless driver monitoring in a realistic, retrospective driving-simulator corpus. The manuscript has several strengths: it uses a relatively large participant sample (79 recordings from 65 participants) with a wide PR range, makes the dataset publicly available on Zenodo, explicitly compares EVM and non-EVM processing, provides execution-time measurements, and uses an independent Empatica E4 versus Faros 360 dataset to assess reference-device bias. The age-group difference in PR is supported even before the linear correction, which is a meaningful result. However, the central quantitative accuracy claim currently rests on parameters selected on the same data used to report the errors, so the numbers should be interpreted as in-sample optima until cross-validation or an external correction is reported numerically.","major_comments":[{"comment":"The headline accuracy numbers (MAE 5.04 ± 0.37 bpm, RMSE 6.38 ± 0.51 bpm for A.EVM after correction) are computed on the same data used to choose two sets of parameters. Section 2.5 states that the moving-average window width is selected per overlapping sequence to minimize MAE against the Empatica E4 reference and the per-sequence optima are then averaged; Section 2.7 fits the linear correction parameters a and b to the same B.EVM/A.EVM differences that are subsequently corrected and reported in Figure 5. No cross-validation, leave-one-subject-out split, or nested tuning is described. The 6.48-to-5.04 bpm improvement therefore reflects a fitted optimum on the evaluation data, not a predictive accuracy. The paper should provide out-of-sample estimates, for example participant-level cross-validation for both the window width and the linear correction, or re-label the headline numbers as in-sample and report the external-correction results as the predictive estimate.","section":"Section 2.5, Section 2.7, Figure 5"},{"comment":"The age-group finding is not an artifact of the correction: the pre-correction p-values (B.EVM p=0.04, A.EVM p=0.01) are already significant and the reference data show p<0.001. However, the p.f. and p.f.f. columns and the Cliff's delta values (0.29 to 0.38) are computed after applying the in-sample linear correction, so these strengthened results are entangled with fitted parameters. The paper should report effect sizes for the uncorrected data with confidence intervals and treat the post-correction values as supporting, not primary, evidence.","section":"Section 3, Table 2"},{"comment":"The external correction derived from Medarević et al. [66] is the only genuinely out-of-sample correction in the paper, yet the manuscript only states that 'errors decrease' and gives no numeric MAE, RMSE, AAE, SAE, or ARE for this correction. Because the external fit parameters (a=0.32, b=-30.42) differ substantially from the in-sample fits, the resulting error values are important for judging generalization; please report them explicitly.","section":"Section 2.7, Figure 5 (bottom panel)"}],"minor_comments":[{"comment":"The definition of ARE contains a double summation with N in both the inner and outer sums; this is malformed and should be corrected.","section":"Equation (4)"},{"comment":"The text refers to upper, middle, and bottom panels, but the caption mentions only upper and lower panels; align the caption with the three-panel layout.","section":"Figure 5 caption"},{"comment":"The Discussion refers to 'Renne et al.'; the cited author is Renner et al. Please correct the name.","section":"Discussion, reference [1]"},{"comment":"The abstract states that EVM adds 'about 20 s for 30 s sequence', while Table 3 reports 23.26 ± 0.86 s for single-core execution and 16.61 ± 1.62 s for four-core execution on a 30-s video; please clarify which configuration the abstract figure refers to.","section":"Abstract and Table 3"},{"comment":"The row 'Our approach applied to the second dataset presented in [61]' leaves the MAE and RMSE cells empty; either fill them in or state explicitly why they are unavailable.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"This manuscript states that the final publication is available at MDPI Applied Sciences (https://www.mdpi.com/2076-3417/15/17/9512). If the current submission is to a different venue, the editor should verify that this does not constitute duplicate publication of already-published work. The main technical concern, namely the in-sample fitting of the linear correction and the moving-average window width, is correctable through cross-validation or by reporting only the external-correction results as predictive accuracy."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHonest rPPG application paper, but the headline accuracy number is circular. The MAE of 5.04 bpm comes from a linear correction fit to the same data used to report the error, and the moving-average window is also tuned against the same reference. No cross-validation. The age-group finding is the more robust result and survives without the correction.\n\nWhat's actually new: the retrospective application to 79 driving-simulator recordings from 65 participants, the EVM time-complexity measurements, and the younger-vs-older driver comparison. The authors clearly state that the pipeline is not novel, which I appreciate. The dataset is shared on Zenodo, and they run their pipeline on two public datasets as an external check—both are good practices.\n\nThe problems are concentrated in the quantitative claim. Section 2.7 fits y = a*x + b (a=0.94, b=-69.41 for B.EVM; a=0.96, b=-74.01 for A.EVM) to the differences between video-derived PR and Empatica E4 on the same recordings whose errors are then reported after correction. Section 2.5 selects the moving-average window width per sequence to minimize MAE against the same reference and then averages over sequences. That makes the reported MAE/RMSE in-sample optima. The external correction from Medarević et al. is a sensible anchor, but no numeric MAE/RMSE is given for that bottom panel of Figure 5, only that errors decrease. So it doesn't back the 5.04 figure.\n\nThe age-group comparison is more solid. In Table 2, the uncorrected B.EVM and A.EVM p-values are 0.04 and 0.01, and the reference is <0.001. The strengthened p-values in the p.f. columns are partly entangled with the in-sample correction, but the group difference is present before correction. That is the paper's most defensible contribution.\n\nOne softer issue: the correlation between extracted SCLI and reference BVP is remarkably low (Pearson/Spearman ~0.08–0.09). The authors acknowledge it but don't fully explain why peak-based PR can still land within a few bpm when waveform correlation is near zero. The PR error rests entirely on peak-picking, and that relationship deserves more scrutiny.\n\nWho is this for? People working on contactless driver monitoring or rPPG in non-ideal conditions. The dataset and the age-group result are useful; the MAE headline should not be quoted as predictive accuracy until a held-out evaluation is done. I would send this to peer review and ask for a calibration-test split or cross-validation, plus numeric results for the external correction.","headline":"Honest rPPG application paper, but the headline accuracy number is circular; the age-group finding is the more robust result.","tokens_in":27488,"tokens_out":3576,"would_cite":true,"duration_ms":33962,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A webcam-based video pipeline can estimate a driver's pulse rate in a driving simulator to within about 5 beats per minute of a wrist-worn reference sensor, the paper claims.","keywords":["remote photoplethysmography","Eulerian video magnification","pulse rate estimation","driving simulator","contactless driver monitoring","motion artifacts","facial skin color variation","Empatica E4"],"falsifier":"Run the published pipeline unchanged—same face detector, same 433 ms moving average for the after-EVM case, same linear correction $y = 0.96x - 74.01$—on a fresh set of driving-simulator videos with simultaneous Empatica E4 recordings, and compute MAE on those held-out sequences. If the error is close to the uncorrected 10.55 bpm rather than 5.04 bpm, the correction was an in-sample artifact; if MAE stays near 5 bpm, the central accuracy claim generalizes.","tokens_in":26382,"feed_emoji":"📷","tokens_out":8296,"duration_ms":75952,"temperature":0.7,"pith_summary":"This paper claims that remote photoplethysmography (rPPG)—reading pulse from subtle facial color changes in ordinary webcam video—is feasible inside a motion-based driving simulator. On recordings from 65 drivers, the authors' video pipeline, with Eulerian video magnification (EVM) and a linear correction applied afterward, reached a mean absolute error of 5.04 bpm and a root mean squared error of 6.38 bpm against the Empatica E4 wrist sensor. The same pipeline also detected a statistically significant pulse-rate difference between younger and older drivers, matching the reference sensor. The claim matters because contactless monitoring could replace wrist-worn devices in driving-simulator studies and eventually support in-car fatigue or stress monitoring, all from a low-cost camera.","feed_headline":"Webcam pulse reading lands within 5 bpm of wrist sensor","feed_subtitle":"A video-based rPPG pipeline tracks drivers' heart rate in a simulator and detects age-group differences.","key_machinery":"The load-bearing machinery is the Signal of Change in Light Intensity (SCLI), extracted as the first principal component of the mean red, green, and blue pixel values over a face-without-eyes region from every video frame, followed by a modified Pan–Tompkins algorithm for peak detection. Eulerian video magnification (EVM)—an algorithm that amplifies tiny temporal color changes in video—is applied as an optional preprocessing step. A linear correction $y = a x + b$ (with $a = 0.96$, $b = -74.01$ for the EVM case) is then applied to video-derived pulse-rate values to remove a systematic deviation that grows with pulse rate; this correction is grounded in the observation that the Empatica E4 itself shows a similar linear bias against an ECG-based reference in independent data [66].","core_discovery":"The paper's central claim is that a complete video-only pipeline—face detection, facial-skin region-of-interest extraction, optional EVM, PCA-based signal extraction, and a modified Pan–Tompkins peak detector—can turn ordinary webcam footage of a person driving a simulator into a usable pulse-rate estimate. With EVM and a linear correction fitted to their data, the pipeline's mean absolute error against the Empatica E4 wristband was 5.04 ± 0.37 bpm (RMSE 6.38 ± 0.51 bpm); without EVM the MAE was 6.48 ± 0.41 bpm. The paper also claims that the video-derived pulse rate preserves a meaningful physiological signal: it separates younger drivers (mean 80.99 bpm) from older drivers (mean 73.12 bpm) with statistical significance, as the reference sensor does. Cross-correlation between the video signal and the wrist blood-volume-pulse waveform is very low (0.08–0.09), so the claim is about heart-rate estimation at the segment level, not waveform fidelity.","pith_inferences":["If the correction generalizes, the same calibration approach could be used to cross-calibrate webcam rPPG against any wrist-worn PPG sensor, letting laboratories pool data collected with different wearables.","The low waveform correlation coexisting with usable beat-rate estimates suggests that segment-averaged pulse rate, not pulse waveform, is what this method reliably delivers; a testable extension is whether heart-rate variability features derived from the same signal also survive the motion-heavy setting.","A natural next experiment is to compare the video pipeline directly against ECG (not wrist PPG) in the same simulator; the paper's use of the Faros 360 data points toward this, and it would separate correction of sensor bias from correction of video error.","For driver monitoring, the method's practical ceiling may be set by head motion and lighting; the authors' own improvement list—standardized lighting, explicit skin segmentation, and higher-resolution cameras—gives a concrete checklist for pushing MAE below the roughly 2 bpm best cases they observed."],"forward_implications":["A single low-cost webcam can replace a wrist-worn PPG sensor for group-level pulse-rate monitoring in driving-simulator studies, such as comparing younger and older drivers.","EVM's accuracy gain is small relative to its cost—about 20 extra seconds per 30-second sequence—so practical quasi-real-time deployments may skip EVM and still retain most of the benefit.","The linear-bias correction transfers conceptually to other reference sensors: because the wrist sensor itself overestimates pulse rate at higher rates, some of the raw error in video-versus-reference comparisons is sensor bias rather than video error.","Video-derived pulse rate can support age-related driver workload and stress studies, provided per-individual errors are tolerated."],"supporting_citations":[{"why":"Supplies the Eulerian video magnification algorithm and code used to amplify subtle facial color changes before pulse extraction.","marker":"[13]"},{"why":"Documents a prior driving-simulator attempt that found webcam EVM-based pulse measurement unreliable, providing the feasibility challenge the paper addresses.","marker":"[1]"},{"why":"Supplies the driving-simulator dataset and technical recording details on which the video analysis was performed.","marker":"[18]"},{"why":"Supplies the deep-learning face detector used to locate faces in each video frame.","marker":"[28]"},{"why":"Supplies the modified Pan–Tompkins algorithm used for peak detection in the SCLI signal.","marker":"[50]"},{"why":"Establishes the Empatica E4 wristband as the reference device for blood volume pulse and pulse-rate measurements.","marker":"[25]"},{"why":"Motivates the linear correction by showing that Empatica E4 deviations grow at higher pulse rates.","marker":"[64]"},{"why":"Provides the independent Empatica E4 versus Faros 360 data used to assess reference-sensor bias and alternative correction parameters.","marker":"[66]"},{"why":"Supplies publicly available controlled-condition rPPG datasets against which the authors benchmark their pipeline.","marker":"[61]"}],"fun_headline_variants":["Webcam pulse in driving sim within 5 bpm of wrist sensor","EVM cost vs benefit: 20s extra for 1.4 bpm gain in rPPG","Video pulse tracking separates young and old drivers in sim","No-touch pulse reading for drivers within 5 bpm of wristband","Webcam pulse rate in driving sim: EVM gains 1.4 bpm but costs 20s"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported 5.04 bpm error assumes that the linear correction parameters and the moving-average window width, both chosen to minimize error on the same recordings, will also reduce error on new recordings rather than merely fitting noise.","fun_headline_variants_meta":{"raw":{"variants":["Webcam pulse in driving sim within 5 bpm of wrist sensor","EVM cost vs benefit: 20s extra for 1.4 bpm gain in rPPG","Video pulse tracking separates young and old drivers in sim","No-touch pulse reading for drivers within 5 bpm of wristband","Webcam pulse rate in driving sim: EVM gains 1.4 bpm but costs 20s"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000872,"raw_usage":{"total_tokens":3809,"prompt_tokens":1013,"completion_tokens":2796,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":629,"completion_tokens_details":{"reasoning_tokens":2687}},"tokens_in":629,"tokens_out":2796,"duration_ms":20288,"temperature":1.0,"reasoning_tokens":2687,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:21:08.576169+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the published pipeline unchanged—same face detector, same 433 ms moving average for the after-EVM case, same linear correction $y = 0.96x - 74.01$—on a fresh set of driving-simulator videos with simultaneous Empatica E4 recordings, and compute MAE on those held-out sequences. If the error is close to the uncorrected 10.55 bpm rather than 5.04 bpm, the correction was an in-sample artifact; if MAE stays near 5 bpm, the central accuracy claim generalizes.","supporting_citations":[{"cited_title":"Analysis and detection R -peak detection using Modified Pan-Tompkins algorithm","cited_arxiv_id":null,"evidence_quote":"Supplies the modified Pan–Tompkins algorithm used for peak detection in the SCLI signal."},{"cited_title":"Validation of the Empatica E4 wristband","cited_arxiv_id":null,"evidence_quote":"Establishes the Empatica E4 wristband as the reference device for blood volume pulse and pulse-rate measurements."},{"cited_title":"Investigating sources of inaccuracy in wearable optical heart rate sensors","cited_arxiv_id":null,"evidence_quote":"Motivates the linear correction by showing that Empatica E4 deviations grow at higher pulse rates."},{"cited_title":"Distress Detection in VR environment using Empatica E4 wristband and Bittium Faros 360","cited_arxiv_id":null,"evidence_quote":"Provides the independent Empatica E4 versus Faros 360 data used to assess reference-sensor bias and alternative correction parameters."}],"review_version":1}