{"id":"c84741c3-be5d-45d5-aca7-dc82743986f8","arxiv_id":"2512.03988","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"HEART-Watch is a 40-participant dataset of synchronized 4-minute Google Pixel Watch ECG, PPG, and accelerometer signals with chest ECG and cuff blood pressure across sitting, standing, and walking.","lead":"Researchers collected four minutes of heart signals (ECG, PPG, motion) from a Google Pixel Watch plus a chest ECG and blood pressure readings from 40 adults while sitting, standing, and walking, and are releasing it as a benchmark dataset. The dataset is meant to help developers test smartwatch heart-monitoring algorithms on a more diverse group of people than earlier studies.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Constant time-shift synchronization cannot correct linear clock drift; residual ECG misalignment may grow over 4-minute recordings.","rationale":"The reader's weakest_assumption already points to the constant-shift synchronization as a fragile step, and the reader's CONDITIONAL verdict captures that concern. My stress-test refines the issue: even without non-linear drift, the measured mean sampling-rate difference (Table 3) implies linear drift that a constant time-shift cannot fix. This is a correctness risk to the dataset's central reuse value for PTT/HRV. However, the concern is addressable by reporting residual alignment statistics or by providing a corrected synchronization model; it does not negate the dataset's novelty or demographic contribution. Therefore the existing CONDITIONAL verdict remains appropriate. I do not find a stronger internal inconsistency: the novelty claim about 4-minute smartwatch ECGs appears supported by the literature table, and the signal quality validation is reasonable for a dataset paper. The access-controlled availability is a real limitation but secondary to the timing issue because a DUA-based dataset can still be used by researchers. The concrete test I propose would settle whether the synchronization concern actually lands. Agreement is 'partial' because the reader emphasized non-linear drift and motion-artifact bias, whereas I identify the more basic mathematical mismatch between a constant shift and observed linear drift.","tokens_in":17418,"tokens_out":4062,"duration_ms":40486,"concrete_test":"Using the provided synchronize.ipynb and one sitting and one walking session per participant, compute per-beat R-peak time differences between chest and watch ECG over the full 4-minute recording after applying the published constant-shift correction. Fit a linear regression of these differences versus time. If the slope is significantly non-zero (e.g., >1 ms/min) or the residual standard deviation exceeds ~5 ms, the constant-shift model is insufficient. Also compare the mean absolute R-peak misalignment in the first 30 s versus the last 30 s; a substantial increase indicates unmodeled drift.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The dataset's core value is tight cross-modal time alignment, claimed in Section 3.6. The pipeline aligns via Unix timestamps, then 'estimating a constant time domain shift based on mean R-peak time differences' and 'shifting the watch signals accordingly to account for potential clock drift.' A constant shift can remove a fixed offset, but it cannot remove clock drift, which is time-varying. The sampling-rate validation (Table 3) shows exactly this risk: watch ECG runs at ~250.93–250.96 Hz while chest ECG runs at ~250.11–250.14 Hz, a relative difference of ~0.3%. Over a 4-minute session this difference alone produces a cumulative timing discrepancy on the order of 700–800 ms if unmodeled. The mean-shift correction centers average error but leaves a linearly growing residual. Section 5 acknowledges 'clock drifts and jitters' but only as a qualitative concern, and Section 4.2 checks mean rates, not residual alignment over time. Figure 5 displays only a 15-second window from one participant, which cannot expose long-range drift. For downstream PTT or HRV analyses over full 4-minute recordings, uncorrected drift of this magnitude would corrupt inter-beat intervals and pulse-transit-time estimates. The paper also does not report any quantitative synchronization error after correction (e.g., residual R-peak time differences across all participants and states). Thus the central claim of a synchronized multimodal dataset is not yet fully supported; the published validation does not rule out substantial temporal misalignment in the very data users would analyze.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces HEART-Watch, a multimodal physiological dataset collected from a Google Pixel Watch 2 (ECG, PPG, accelerometer) alongside reference chest ECG and intermittent cuff blood pressure from 40 healthy adults across sitting, standing, and walking states. The central claim is that HEART-Watch is the first research dataset to provide consumer-grade smartwatch ECG recordings of 4 minutes across these physical states, with synchronized raw signals and rich demographic diversity. The paper describes the acquisition hardware, protocol, data organization, synchronization procedure, and preliminary validation including sampling-rate stability, signal-quality statistics, ECG morphology comparison with Bland-Altman analysis, heart-rate estimation examples, and blood-pressure summaries. The dataset is made available under a data-use agreement through the authors' institution.","tokens_in":17716,"tokens_out":4155,"duration_ms":41640,"significance":"If the synchronization and data-sharing concerns are adequately addressed, HEART-Watch would be a valuable community resource. Its strengths include a diverse participant cohort (age 19-75, eight racial categories, varied body types), four-minute smartwatch ECG recordings across postural and walking states, simultaneous reference chest ECG and cuff BP, and a preprocessing pipeline whose code is promised alongside the data. The paper also provides reproducible validation scripts for sampling-rate and signal-strength checks. The clear documentation of sensor rates, weak-signal rates, and ECG morphology comparisons is a positive feature. The dataset directly targets a recognized gap in consumer-wearable cardiovascular research: the scarcity of open, raw, multimodal, prolonged smartwatch ECG data with demographic coverage. The validity of the central use case, however, depends on the temporal alignment between watch and chest signals, and that alignment is not quantitatively demonstrated in the current manuscript.","major_comments":[{"comment":"The synchronization procedure is not adequately validated for the dataset's core claim. Section 3.6 states that after Unix-timestamp alignment, a 'constant time domain shift based on mean R-peak time differences' is applied 'to account for potential clock drift.' A constant shift can remove a fixed offset, but it cannot compensate for clock drift, which is time-varying. Table 3 shows that the smartwatch ECG sampling rate is ~250.93-250.96 Hz while the chest ECG is ~250.11-250.14 Hz across the physical states. Over a 4-minute recording, this ~0.3% relative rate difference implies a cumulative timing error on the order of 700-800 ms if the two clocks drift linearly and are not otherwise corrected. The published validation reports only mean sampling rates (Section 4.2) and shows a 15-second example in Figure 5, neither of which can expose drift over full sessions. Section 5 mentions 'risk o","section":"Section 3.6 / Table 3 / Section 4.2 / Figure 5"}],"minor_comments":[{"comment":"The dataset is described as being available through a data-use agreement, but no persistent identifier, version number, or repository DOI is given. Adding a DOI or institutional repository link would improve reproducibility and citation.","section":"Section 7"},{"comment":"The PR-interval comparison shows a consistent overestimation bias (10-18 ms) with wide limits of agreement. This is presented as a finding, but the abstract/conclusion could briefly note this as a limitation of smartwatch ECG morphology analysis, since the dataset is intended for cardiovascular biomarker development.","section":"Section 4.4"},{"comment":"The statement that ppg_watch_two 'should be used for pulse waveform analysis' is based on 'internal testing' without quantitative supporting results. A brief summary of the testing (e.g., signal-to-noise ratio comparison) would make the recommendation more transparent.","section":"Section 3.6.3"},{"comment":"Minor typographical issues: Section 3.5 has 'each participant s’' and Section 2/3 contain 'arrythmias' and 'heterogenicity' (e.g., Section 3.6.2, Section 2). A careful proofreading pass is recommended.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The dataset is potentially a strong contribution, and the central novelty claim is plausible. The main blocker is the synchronization validation: the authors must provide a quantitative residual temporal-alignment analysis or explicitly model clock drift. If the residual error is shown to be small (e.g., within a few milliseconds), I would support acceptance. The paper would also benefit from a repository DOI to make the resource independently verifiable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is a legitimately useful dataset: prolonged (4-minute) smartwatch ECG across sitting/standing/walking, synchronized with PPG, ACC, chest ECG, and cuff BP in 40 adults with real demographic spread. Second, the weakest link is the synchronization validation. They estimate a constant time shift from mean R-peak differences and call that an account for clock drift, but a constant shift does not correct linear drift. The mean sampling rates in Table 3 show about a 0.3% relative difference between watch and chest ECG; over four minutes that is potentially several hundred milliseconds of accumulated misalignment. The paper doesn't report residual R-peak alignment errors across participants or states, and Figure 5 is a 15-second window from one participant. That matters for anyone doing PTT or HRV on the full recordings.\n\nWhat's genuinely new: their Table 1 makes the case that existing smartwatch ECG datasets are short (30-50 s) or non-existent, and only two provide smartwatch ECG at all. HEART-Watch's 4-minute recordings across physical states, with raw signals and eight racial categories, is a real contribution. The internal checks are decent for a dataset paper: weak-signal rates are quantified, the Bland-Altman QRS/PR comparisons are honest, and the HR estimation example shows how the modalities can be combined. The PR overestimation is reported rather than hidden, which is good practice.\n\nThe soft spots beyond sync: the data is behind a data-use agreement and only on request; that's normal for health data, but the code availability 'on request' is weaker. The constant-shift calibration itself is fine as a correction step, but not as evidence that the alignment is good. I'd want residual error distributions and maybe a drift model. The stress-test note is right: this is a load-bearing issue for time-critical analyses, not a nitpick.\n\nWho's this for: anyone building or benchmarking wrist-ECG/PPG algorithms. The paper deserves a serious referee. I'd send it out, but with a requirement to expand the synchronization validation and ideally release the code and a small sample of data openly.\n\nRecommendation: accept if the authors quantify residual misalignment over full sessions and, while they're at it, correct the language about drift.","headline":"HEART-Watch fills a real gap with 4-minute consumer smartwatch ECG across postures, but the time-alignment is only roughly validated and the dataset is currently hard to access.","tokens_in":18227,"tokens_out":4326,"would_cite":false,"duration_ms":39735,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper presents HEART-Watch, the first research dataset with four-minute consumer smartwatch ECG recordings in sitting, standing, and walking states, synchronized with PPG, accelerometer, chest ECG, and cuff blood pressure from 40 divers","keywords":["smartwatch ECG","photoplethysmography","accelerometer","cardiovascular dataset","pulse transit time","heart rate variability","demographic diversity","signal synchronization"],"falsifier":"Measure beat-by-beat R-peak (the sharp spike marking each heartbeat) time differences between the smartwatch and chest ECG across each four-minute file; if the difference drifts by more than a few milliseconds over the recording, the constant-shift model fails. A device-independent check would be to compare a simultaneous physical event (for example a tap or cuff inflation) detected in the accelerometer and chest ECG, avoiding any reliance on R-peak detection.","tokens_in":17314,"feed_emoji":"⌚","tokens_out":7756,"duration_ms":62953,"temperature":0.7,"pith_summary":"HEART-Watch is a research dataset built around a Google Pixel Watch 2, recording four minutes of smartwatch ECG, photoplethysmography (PPG), and accelerometer data simultaneously with a reference chest ECG while each of 40 healthy adults sits, stands, and walks. The authors claim this is the first dataset to offer consumer smartwatch ECG recordings of that length across those three physical states, and one of the first to pair them with raw synchronized signals and detailed demographics. The resource matters because smartwatch algorithms are typically trained on short, homogeneous, or aggregate-only recordings; HEART-Watch provides prolonged raw signals with older adults, multiple self-identified racial groups, and varied body types to support realistic and fair validation. Initial checks show stable sampling rates, mostly strong signal quality, close QRS timing between watch and chest ECG, and a consistent PR-interval overestimation attributed to attenuated P-waves.","feed_headline":"First 4-minute smartwatch ECG across sitting, standing, walking","feed_subtitle":"Synchronized wrist ECG, PPG, and accelerometer data help validate heart-rate and blood-pressure algorithms.","key_machinery":"The load-bearing object is the synchronized multimodal recording pipeline. Raw streams are stored as separate CSVs; the synchronized version zero-centers ECG and PPG, linearly interpolates all signals to 250 Hz, aligns them by Unix timestamps, and then applies a constant time-domain shift estimated from the mean R-peak time difference between the smartwatch and chest ECG to correct clock offset and drift. This constant shift is what makes inter-modality timing analyses—pulse transit time, beat-to-beat heart rate, HRV—possible. The four-minute smartwatch ECG, captured by having participants hold a finger on the watch crown, is the novel signal the pipeline was designed to preserve and is the","core_discovery":"HEART-Watch is a multimodal physiological dataset collected from a Google Pixel Watch 2 worn by 40 healthy adults (23 female, 17 male, ages 19–75). Each participant completed four-minute recordings in sitting, standing, and walking states, with the smartwatch capturing ECG, PPG, and accelerometer signals while a reference chest ECG was recorded with gel electrodes; five cuff blood-pressure readings with concurrent biosignals were taken during transitions between states. The dataset provides raw sensor values rather than proprietary aggregates, and the synchronized version aligns all modalities to 250 Hz. The paper's central claim is that this is the first research dataset to provide consumer","pith_inferences":["Because the dataset's four-minute windows exceed the conventional 30-second smartwatch ECG clip, it likely supports frequency-domain HRV metrics that earlier consumer-watch datasets cannot; the paper notes the four-minute choice but does not itself compute those metrics.","The reported PR-interval overestimation may extend to the BP-measurement recordings; checking whether P-wave bias varies by age or body type could refine the delineation algorithm.","Since the first PPG channel is described as a probable reference channel, future work could test whether combining it with the second channel reduces motion artifacts during walking; the paper identifies the channels but does not perform this comparison.","A practical consequence for users: before computing pulse transit time, verify that R-peak delay residuals do not drift over the four minutes; if they do, per-segment re-alignment would be needed."],"forward_implications":["Researchers can use the synchronized files to benchmark heart-rate and HRV algorithms under realistic sitting, standing, and walking motion, including degraded smartwatch ECG segments that earlier datasets excluded.","Because each state lasts four minutes, frequency-domain HRV metrics like low- and very-low-frequency power can be computed, which 30-second smartwatch ECG recordings cannot support.","The combination of wrist ECG, wrist PPG, and accelerometer with chest ECG and cuff BP enables pulse-transit-time and cuffless blood-pressure studies from consumer hardware.","Demographic breadth (age 19–75, eight self-identified racial groups, varied height/weight) permits fairness and age-stratified comparisons of wearable algorithms.","The inclusion of weak or invalid smartwatch ECG segments supports training algorithms to be resilient to real-world signal loss."],"fun_headline_variants":["Multimodal smartwatch health dataset from 40 diverse adults","Pixel Watch data: ECG, PPG, and motion across three states","First consumer smartwatch dataset with synchronized heart signals","Wrist ECG and PPG from sitting, standing, and walking","HEART-Watch: diverse smartwatch cardiovascular dataset for benchmarking"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The synchronization assumes that a single constant time shift, estimated from average R-peak delays between the watch and chest ECG, fully corrects clock offset and drift for the entire four-minute recording.","fun_headline_variants_meta":{"raw":{"variants":["Multimodal smartwatch health dataset from 40 diverse adults","Pixel Watch data: ECG, PPG, and motion across three states","First consumer smartwatch dataset with synchronized heart signals","Wrist ECG and PPG from sitting, standing, and walking","HEART-Watch: diverse smartwatch cardiovascular dataset for benchmarking"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000157,"raw_usage":{"total_tokens":1045,"prompt_tokens":715,"completion_tokens":330,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":459,"completion_tokens_details":{"reasoning_tokens":246}},"tokens_in":459,"tokens_out":330,"duration_ms":3783,"temperature":1.0,"reasoning_tokens":246,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T18:39:00.337645+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure beat-by-beat R-peak (the sharp spike marking each heartbeat) time differences between the smartwatch and chest ECG across each four-minute file; if the difference drifts by more than a few milliseconds over the recording, the constant-shift model fails. A device-independent check would be to compare a simultaneous physical event (for example a tap or cuff inflation) detected in the accelerometer and chest ECG, avoiding any reliance on R-peak detection.","supporting_citations":[],"review_version":1}