{"id":"7d957f45-6f6e-4a9b-b8ee-0ce278bdbfa8","arxiv_id":"2508.03698","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"ETRI Lifelog Dataset 2024 contains continuous passive sensing data from phones, smartwatches, and sleep sensors, paired with self-reported fatigue, stress, and sleep quality surveys before and after sleep.","lead":"Researchers collected continuous smartphone, smartwatch, and sleep-sensor data plus daily stress, fatigue, and sleep quality surveys to create the ETRI Lifelog Dataset 2024. A portion of the anonymized dataset is public, and the authors suggest it can support machine learning models for sleep and stress prediction.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Dataset's 'continuous 24/7' claim is unverified; without wear-time coverage statistics, the core promise of comprehensive daily-life sensing may not hold.","rationale":"The reader's weakest assumption pointed to participant compliance and data loss; my concern is the same load-bearing issue, made concrete. The abstract promises a comprehensive 24-hour continuous record, but provides no evidence that such coverage was actually achieved. Device non-wear due to charging, forgetting, or discomfort is endemic in lifelog studies, and without wear-time statistics or missing-data handling the central value proposition of the dataset is unverified. Since the full paper is not available, I cannot rule out that the paper itself includes coverage analyses; however, the abstract's strong claim requires this validation. A conditional verdict is appropriate: accept if the public data and/or the full paper provide wear-time coverage metrics consistent with the 'continuous 24 hours' statement, and reject or downgrade the claim if coverage is substantially lower. This is not an allegation of misconduct; it is a standard data-quality requirement for datasets that advertise themselves as continuous and minimally intrusive. The proposed concrete test directly checks the claim on the released data, which is the most decisive available check. I agree with the reader that this is the weakest point; my verdict adjustment reflects that the paper should be conditioned on demonstrating coverage.","tokens_in":668,"tokens_out":2348,"duration_ms":29612,"concrete_test":"Download the publicly released subset of the ETRI Lifelog Dataset 2024 (or request through the data-use agreement) and compute, for each participant-day: (1) the fraction of 24 hours with at least one sample from the smartwatch, smartphone, or sleep sensor; (2) the distribution of gaps longer than 30 minutes; and (3) the time difference between each before/after-sleep survey response and the corresponding device-derived sleep window. If median per-participant coverage falls below 80% of hours, or if surveys are frequently more than 1 hour outside the device-defined sleep period, the continuous and minimally intrusive claim should be revised.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that passive smartphone, smartwatch, and sleep sensors collected data continuously for 24 hours a day with minimal interference, enabling quantitative analysis of daily behavior and sleep. This requires participants to wear or carry the devices during waking hours and overnight, and for the hardware to have sufficient battery and storage to avoid large gaps. Typical wearable deployments lose substantial segments due to charging, device removal, and sensor errors. The abstract (the only text available) reports no wear-time statistics, no per-participant coverage, no handling of missingness, and no validation of survey timing (before/after sleep) against actual device-derived sleep onset. If large fractions of participant-days lack sensor data, the 'continuous' record is not real and any behavioral conclusions drawn from the dataset are biased. The load-bearing assumption, therefore, is that device wear is near-ubiquitous and that the public release includes enough metadata to identify and correct non-wear periods. This is not demonstrated by the abstract alone.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The abstract describes the ETRI Lifelog Dataset 2024, collected using smartphones, smartwatches, and sleep sensors to passively capture continuous 24-hour daily behavior and sleep across multiple days, alongside subjective surveys of fatigue, stress, and sleep quality administered immediately before and after sleep. The paper is said to introduce the dataset's structure and potential machine-learning applications such as sleep-quality and stress prediction, with a portion of the anonymized data made publicly available. Because only the abstract was provided for review, my assessment is necessarily limited to the claims and framing in that abstract.","tokens_in":855,"tokens_out":3200,"duration_ms":40971,"significance":"If the full dataset delivers what the abstract promises, it would be a valuable community resource: multimodal passive sensing aligned with temporally anchored subjective sleep/stress reports, partially released with anonymization, would support reproducibility in daily-life sensing and sleep research. The explicit pairing of device-based and survey-based measures around sleep is a strength, as is the stated intent to make data public. However, the significance hinges on data quality and completeness, particularly the \"continuous 24-hour\" collection claim and the \"immediately before and after sleep\" survey timing; neither is evidenced in the abstract, and both are known failure points in wearable deployments.","major_comments":[{"comment":"The abstract's central claim that data were \"collected passively and continuously for 24 hours a day\" is a quantitative assertion, but no wear-time coverage statistics, per-participant completeness rates, battery/charging interruption handling, or non-wear detection strategy are reported anywhere in the abstract. Without these, the continuity claim is unverifiable and the risk of substantial missingness from device removal, charging, or sensor errors is not addressed. The full paper must report the distribution of wear time across participants and days, define valid-day inclusion criteria, and provide missing-data metadata in the public release.","section":"Abstract"},{"comment":"The claim that surveys were \"conducted immediately before and after sleep\" requires validation of self-report timing against device-derived sleep periods, especially because participants may complete surveys late or early. The paper should report compliance rates, the median and IQR of the lag between survey completion and device-detected sleep onset/offset, and how the before/after windows were defined. Without such validation, the subjective-anchoring feature that distinguishes this dataset is not trustworthy.","section":"Abstract"},{"comment":"The word \"comprehensive\" in describing the lifelog dataset is an evaluative claim that is unsupported by the abstract, which provides no information on sensor modalities' completeness, device dropout rates, or the fraction of participant-days containing all data streams. The full paper should define the intended observation period, report coverage per modality, and state the inclusion/exclusion criteria for participants and days used in any analyses; otherwise \"comprehensive\" cannot be evaluated by readers.","section":"Abstract"}],"minor_comments":[{"comment":"The phrase \"passively and continuously\" is ambiguous: it does not specify which smartphone/smartwatch sensors are involved (e.g., accelerometer, PPG, GPS, screen state) or whether any user interaction is required during the day. Please enumerate the sensor modalities and their sampling rates in the abstract or dataset description.","section":"Abstract"},{"comment":"The claim of \"minimal interference to participants' usual behavior\" is presented as a fact, but it would be strengthened by including participant-reported burden measures or by comparing early versus later days of the study to detect reactivity effects.","section":"Abstract"},{"comment":"The statement that \"a portion of the data has been anonymized and made publicly available\" would benefit from specifying which modalities are in the public portion and what anonymization procedures were applied, since this directly affects reproducibility.","section":"Abstract"}],"recommendation":"uncertain","confidential_remarks":"This review is based on the abstract only; the full text was not available. The major concerns are therefore about unverified claims in the abstract rather than confirmed errors in the paper. If the full manuscript contains the requested wear-time coverage statistics, survey-timing validation, and missing-data handling, the dataset contribution may well be publishable. I recommend that the editor obtain the full text before making a final decision. The choice of 'uncertain' reflects that the evidence needed for a definitive verdict is absent from this review package."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a dataset paper, and the dataset itself is a legitimate new artifact. The design is not novel — multimodal passive sensing plus self-reports is standard in lifelogging — but the combination of smartphone, smartwatch, sleep sensor, and pre/post sleep surveys, with a portion publicly released, is a useful resource for sleep and stress prediction research. That alone justifies reading it.\n\nWhat it does well: clear description of the sensing stack and the survey timing. The abstract is honest about the data being a 'foundational resource' rather than a tested set of predictions. There is no derivation or fitting, so circularity burden is near zero. The public release of part of the data is a real plus.\n\nThe soft spot is the load-bearing claim of 'continuously for 24 hours a day' with 'minimal interference.' That is not demonstrated in the abstract, and typical wearable studies lose significant segments to charging, removal, and hardware errors. The stress-test note is right: without per-participant wear-time coverage, handling of non-wear periods, and validation of survey timing against device-derived sleep onset, the core promise of the dataset is unverified. That said, this is an abstract-only review. Dataset papers normally defer these numbers to the full text, so the absence here is a flag for reviewers, not proof of failure.\n\nThe 'comprehensive' descriptor is a bit strong for what we see, but it's a minor overstatement, not a fatal issue.\n\nThe citation pattern is unknown, but the authors do not seem to be claiming a new method. The paper is for researchers who need a continuous lifelog dataset with both objective and subjective sleep markers. I'd bring it to a reading group if anyone works on passive sensing, and I'd probably cite it if I needed such a dataset.\n\nMy recommendation: send it to peer review. The dataset could be useful, and the key questions — wear-time coverage, missingness handling, survey validation, and what exactly is public — are exactly what a good referee can check. If the full paper reports those numbers honestly, the contribution stands.","headline":"A useful new partially public lifelog dataset with pre/post sleep surveys; the 'continuous 24/7' claim is the main thing to verify against wear-time statistics.","tokens_in":1359,"tokens_out":2306,"would_cite":true,"duration_ms":25372,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new lifelog dataset captures 24/7 daily behavior through passive smartphone, smartwatch, and sleep-sensor sensing, paired with pre- and post-sleep surveys on fatigue, stress, and sleep quality.","keywords":["lifelog dataset","continuous sensing","smartphone sensing","smartwatch","sleep sensors","sleep quality","stress prediction","self-report surveys"],"falsifier":"An independent audit of the raw sensor streams that finds substantial gaps—for example, more than a few hours of missing data per day for a large share of participants—would contradict the claimed continuous 24-hour capture, as would a near-zero correlation between self-reported sleep quality and objective sleep metrics from the sensors.","tokens_in":520,"feed_emoji":"😴","tokens_out":5756,"duration_ms":59721,"temperature":0.7,"pith_summary":"This paper introduces a new lifelog dataset built from continuous, passive sensing of daily life: smartphones, smartwatches, and sleep sensors running around the clock, complemented by short self-report surveys administered right before and after sleep. The authors' aim is to provide a quantitative, multi-day record of daily behaviors and sleep activities that imposes minimal burden on participants. They argue that such a dataset can serve as a basis for understanding daily life and lifestyle patterns, and they highlight potential applications, including machine-learning models that predict sleep quality and stress. A portion of the data has been anonymized and released publicly, which would let other researchers build on the resource.","feed_headline":"Lifelog dataset pairs 24/7 sensing with sleep surveys","feed_subtitle":"Smartphone, smartwatch, and sleep-sensor streams plus before- and after-sleep surveys are now partly public.","key_machinery":"The load-bearing mechanism is the synchronized pairing of two measurement streams: continuous passive sensor data from devices participants already carry or wear, and survey responses timed to the moments just before and just after sleep. The continuous stream supplies objective traces of activity, location, and sleep; the surveys capture subjective states that sensors cannot measure directly. Aligning these streams in time is what lets the dataset link daily behavior to self-reported fatigue, stress, and sleep quality.","core_discovery":"The central claim is that a comprehensive, low-burden record of human daily experience can be captured by combining passive 24-hour sensing from smartphone, smartwatch, and sleep-sensor hardware with subjective self-reports of fatigue, stress, and sleep quality collected immediately before and after sleep. The paper presents the structure of this dataset, collected over multiple days, and positions it as a resource for studying daily behavior and sleep. It further argues that the paired objective and subjective measurements are suitable inputs for machine-learning models aimed at predicting sleep quality and stress.","pith_inferences":["An extension of this line of work would be to test whether adding continuously collected self-reports during the day, not just around sleep, improves prediction of stress or fatigue beyond the current design.","If sensor data and surveys disagree—for example, when a participant reports good sleep but actigraphy shows fragmented sleep—the dataset could support studies of the gap between perceived and measured experience.","The anonymized public subset could be used to estimate the minimum device-wear compliance needed for reliable daily-life predictions, guiding future study designs.","This dataset could serve as a calibration resource for researchers developing new wearable-based inference models, provided the public subset preserves the diversity of sensors and participants."],"forward_implications":["The public portion of the dataset gives researchers a ready-made training set for predicting sleep quality and stress from passive sensor streams.","The 24-hour, multi-day design supports studies of day-to-day routines and sleep regularity rather than single-day snapshots.","The combination of objective sensor data and self-reported surveys lets models be validated against both measured and perceived states.","The dataset offers a common benchmark for comparing lifelog collection protocols and prediction methods."],"supporting_citations":[],"fun_headline_variants":["24/7 sensing plus sleep surveys: ETRI lifelog goes public","Dataset merges passive sensing with before/after sleep surveys","ETRI lifelog: phone, watch, sleep sensors meet daily self-reports","Multi-day lifelog data pairs continuous sensing with stress surveys","Public lifelog dataset captures daily behavior and sleep quality"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The dataset's value depends on participants actually wearing the devices throughout the day and night, completing the before- and after-sleep surveys honestly and on time, and the sensors not dropping long stretches of data.","fun_headline_variants_meta":{"raw":{"variants":["24/7 sensing plus sleep surveys: ETRI lifelog goes public","Dataset merges passive sensing with before/after sleep surveys","ETRI lifelog: phone, watch, sleep sensors meet daily self-reports","Multi-day lifelog data pairs continuous sensing with stress surveys","Public lifelog dataset captures daily behavior and sleep quality"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000536,"raw_usage":{"total_tokens":2512,"prompt_tokens":816,"completion_tokens":1696,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":432,"completion_tokens_details":{"reasoning_tokens":1604}},"tokens_in":432,"tokens_out":1696,"duration_ms":13702,"temperature":1.0,"reasoning_tokens":1604,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:20:21.078903+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An independent audit of the raw sensor streams that finds substantial gaps—for example, more than a few hours of missing data per day for a large share of participants—would contradict the claimed continuous 24-hour capture, as would a near-zero correlation between self-reported sleep quality and objective sleep metrics from the sensors.","supporting_citations":[],"review_version":1}