{"id":"79e17bf4-3b1e-410f-b21a-6aae9619867b","arxiv_id":"2508.05845","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"By treating untracked periods as latent skips, SkipTrack estimates cycle length and covariate effects with less bias than methods that pre-specify skips.","lead":"SkipTrack is a Bayesian model that treats missed period logs as hidden events instead of known facts, reducing bias in app-based menstrual cycle studies. It is aimed at making large mobile-health cohorts give trustworthy estimates of cycle length and regularity.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"SkipTrack's superiority claim rests on the posterior being able to distinguish a skipped period from a genuinely long inter-bleed interval; no ground-truth skip labels or identifiability analysis are visible, so the separation is the load-bearing unvalidated assumption.","rationale":"Read in good faith: SkipTrack addresses a genuine and important problem—users do skip logging periods, and observed cycle lengths are inflated as a result. The proposed Bayesian machinery is a plausible way to propagate uncertainty about skips. The abstract's claim is not internally inconsistent, and the simulation comparison is relevant. The load-bearing question is whether the data can actually identify the skip/long-cycle split. The reader's weakest-assumption analysis identified exactly this point: no ground-truth skip labels validate the decomposition, and the prior on true cycle length is doing essential work. My reading agrees. I add that even a successful simulation may not resolve the concern if the simulation generates skips under the model's own prior, so an adversarial long-cycle/no-skip simulation and prior sensitivity analysis are the concrete tests. Because the full text is too corrupted to verify whether the authors already performed such analyses, and because the central claim remains unverified, I do not change the reader's UNVERDICTED status. No objection is raised about the authors' conduct; the issue is the availability and robustness of evidence for the key identification step.","tokens_in":13948,"tokens_out":6568,"duration_ms":80157,"concrete_test":"Run an adversarial simulation in the dangerous regime: generate true cycle lengths from a heavy-tailed distribution (e.g., gamma with substantial mass above 50 days), set the true skip probability to zero, fit SkipTrack, and measure the posterior expected number of skips and the false-skip rate. If false skips are common or posterior intervals exclude zero skips, the model cannot separate long cycles from skipped cycles. Also vary the prior mean/variance of true cycle length across plausible values and record how much the posterior skip probabilities and the AWHS age/BMI/race coefficients change; large shifts would indicate the headline associations are prior-driven. If any AWHS participants have independent ground-truth skip labels (e.g., user-confirmed missed periods or a second tracking source), compare posterior skip probabilities against those labels and report calibration and discr","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim—that SkipTrack beats methods that fix skip status—requires that the latent skip indicators be identifiable from observed inter-bleed intervals. But an observed interval of, say, 60 days can be one 60-day cycle or two 30-day cycles with one unlogged start. The model can separate these only through the prior on true cycle length and the skip probability. The legible portions of the manuscript do not report a recovery analysis comparing posterior skip probabilities to known skip labels in the difficult regime where long cycles occur without skips, nor a prior sensitivity analysis. If long-cycle and skip explanations yield similar observed-data likelihoods, the posterior over skips—and therefore every covariate effect on cycle length and regularity—is driven by the prior rather than the data. The simulations compare against methods that specify skips a priori, but if the data-generating process uses the same prior family the model assumes, the comparison does not establish external validity. In the Apple Women's Health Study application, the reported associations for age, BMI, and race/ethnicity inherit this untested decomposition. A secondary concern is the implicit ignorability assumption: if women with irregular cycles are more likely to miss logging a period, skip status is informative about the very outcome being modeled, and no auxiliary data or sensitivity analysis addresses this.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes SkipTrack, a Bayesian hierarchical model for menstrual cycle length and regularity in large mobile health cohorts. The model treats potentially skipped cycle-tracking events as latent indicators and jointly estimates true cycle length, skip probability, and covariate effects (age, BMI, race/ethnicity). The abstract claims that, in simulations, SkipTrack outperforms methods that specify skip status a priori, which are said to suffer from estimation bias and overconfidence. The model is then applied to the Apple Women's Health Study to estimate associations between demographic covariates and menstrual cycle outcomes.","tokens_in":14298,"tokens_out":3565,"duration_ms":40734,"significance":"If the latent-skip decomposition is identifiable, the framework would be a valuable contribution to the analysis of self-tracked menstrual cycle data, where unlogged period starts can inflate observed cycle lengths and lead to overconfident estimates. The paper addresses a real and timely problem in digital cohort research. However, the visible evidence does not currently support the abstract's central claim: the separation between skipped cycles and genuinely long cycles is not validated against known skip labels, the simulation design appears to generate data from the same model family as SkipTrack, and the real-data application has no external ground truth for skip status. The contribution is promising but the validation is incomplete.","major_comments":[{"comment":"The central claim that SkipTrack 'accounts for the uncertainty of possible skips' requires that the latent skip indicators are identifiable from observed inter-bleed intervals. An observed interval of, say, 60 days could be one 60-day cycle or two 30-day cycles with one unlogged period start. The manuscript does not report a recovery analysis comparing posterior skip probabilities to known skip labels in settings where long cycles and skipped cycles coexist, nor a prior-sensitivity analysis for the cycle-length and skip-probability priors. Without that, every covariate effect on cycle length and regularity inherits the prior's decomposition. Please add simulations that vary the true long-cycle rate and skip rate and report posterior classification accuracy, coverage, and calibration of the skip indicators.","section":"Abstract and model specification"},{"comment":"The simulation comparison appears to generate data from the same model family as SkipTrack, so the comparison with a-priori skip rules may simply reflect a correctly specified model beating misspecified competitors. This is a form of circular validation. Add robustness simulations generated under a different process—for example, skip probability depending on current cycle length, previous skip status, or user-level random effects—and show whether the claimed bias and coverage advantages persist. This is needed for the abstract's 'superiority' claim to be credible.","section":"Simulation study"},{"comment":"The real-data associations for age, BMI, and race/ethnicity are presented as being closer to the underlying biology than prior estimates, but there is no external validation of skip status (e.g., hormone-based cycle phase, follow-up surveys, or comparison with self-reported regularity). Because the model's skip decomposition is untested, these associations should be framed as model-dependent. At minimum, run a prior-sensitivity analysis over the skip model and report how the covariate associations change.","section":"Application to Apple Women's Health Study"},{"comment":"The model assumes that, conditional on covariates, skipping a tracked period is not informative about the true cycle length or regularity under study. If users with irregular cycles are more likely to miss logging a period, the posterior over skip indicators—and therefore the covariate effects—can be biased. The manuscript neither states this ignorability assumption nor reports sensitivity analyses. Please state the assumption explicitly and assess robustness, e.g., by letting skip probability depend on true cycle length or on a user-level random effect.","section":"Model assumptions"}],"minor_comments":[{"comment":"The study name is given as 'Apple Women's Healthy Study' in the abstract but 'Apple Women's Health Study' in the main text. Please use the correct name consistently.","section":"Abstract"},{"comment":"Several symbols in the model equations are not defined near first use in the legible portions of the manuscript. A single, self-contained notation table would improve readability.","section":"Notation"},{"comment":"The simulation tables report point estimates only. Please include Monte Carlo standard errors, number of replicates, and the width/coverage of the estimated intervals, especially for the competing methods.","section":"Simulation tables"},{"comment":"The figures are difficult to interpret without clearer labeling of credible intervals and, where applicable, posterior probabilities of skip indicators. Consider adding a panel that shows posterior skip probabilities versus true skip indicators in the simulation.","section":"Figures"}],"recommendation":"major_revision","confidential_remarks":"The manuscript addresses a practically important problem and the model is well motivated. The difficulty is that the central advantage—separating skipped cycles from long cycles—is not yet supported by the evidence. The requested additions (identifiability/recovery simulations, misspecification robustness, and stated ignorability with sensitivity analysis) are within the scope of a major revision and should not require a fundamentally new model. I would support publication after those additions are made and the claims are appropriately tempered."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know: this paper attacks a real problem. In menstrual cycle tracking apps, a user who forgets to log one period turns two cycles into one apparently long cycle, which biases covariate effect estimates. Modeling skip status as a latent variable inside a Bayesian hierarchical framework is a reasonable and fairly new move, at least compared with methods that fix skip status ahead of time. That part deserves credit.\n\nThe problem is well motivated, and the application to the Apple Women's Health Study is a legitimate real-data test. The abstract's promise of time-varying effects on both length and regularity is a natural extension of existing hierarchical cycle-length models, and if the simulations are honestly built, the comparison against a priori skip rules would be meaningful evidence.\n\nBut the central claim has a load-bearing soft spot. The model separates a 60-day observed inter-bleed interval into either one long cycle or two 30-day cycles with a skipped period only through priors on cycle length and skip probability. Nothing in the abstract shows the posterior can actually distinguish these, and no ground-truth skip labels or recovery analysis is mentioned. The simulation design likely generates data from the same family of models, so the comparison may be testing self-consistency rather than external validity. Also, the model implicitly assumes skip behavior is unrelated to the cycle irregularity being studied. If women with irregular cycles are more likely to skip logging, every covariate effect inherits that bias. These are not minor concerns; they go to the core of what the paper claims.\n\nI should note that the copy of the manuscript we received is badly corrupted, so I could not verify whether the full text addresses these points. If a clean version exists, a referee should check identifiability and prior sensitivity first.\n\nNet: this is a paper for researchers in mobile health statistics and reproductive epidemiology. It deserves a serious referee, but the referee needs to push for an identifiability analysis, prior sensitivity checks, and preferably some external validation—even a small labeled set of cycles or diary data would help. I would not cite it in its current form, but I would follow the revised version with interest.","headline":"Useful idea in search of validation: SkipTrack treats skip status as latent, but its superiority claim rests on simulations that can't yet be checked and an identifiability assumption the paper needs to defend.","tokens_in":14770,"tokens_out":2103,"would_cite":false,"duration_ms":25910,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"SkipTrack claims that accounting for unlogged periods—instead of assuming each gap between logged periods is one cycle—reduces bias and overconfidence in app-based estimates of how age, BMI, and race/ethnicity relate to menstrual cycle leng","keywords":["menstrual cycle length","cycle regularity","Bayesian hierarchical model","self-tracked mobile health","skipped period logging","missing data","covariate effects","menstrual health"],"falsifier":"Find app users with independent confirmation of every period (daily hormone or temperature monitoring). For a user with a confirmed true 60-day cycle and no missed logging, the model should put most posterior mass on one 60-day cycle; if it instead assigns high probability to two 30-day cycles with a skip, the skip-correction mechanism is being driven by the prior, not the data.","tokens_in":13891,"feed_emoji":"🩸","tokens_out":5962,"duration_ms":57332,"temperature":0.7,"pith_summary":"Self-tracked menstrual cycle data have a hidden-data problem: when a user fails to log a period, the gap between logged periods is not one cycle but two or more. This paper presents SkipTrack, a Bayesian hierarchical model that treats the number of hidden cycles in each gap as unknown and averages over the possibilities while estimating how age, body mass index, and race/ethnicity shape cycle length and regularity. The paper argues, with simulations, that methods which assume skip status is known are more prone to biased estimates and overconfident intervals, and it applies SkipTrack to a large US mobile-health cohort. If the claim holds, previously reported covariate associations in app-based cycle data may need revisiting, because part of what looked like longer or more irregular cycles could be logging gaps.","feed_headline":"Bayesian model turns skipped period logs into honest cycle estimates","feed_subtitle":"Simulations show methods that assume every period was logged get biased, overconfident results; SkipTrack averages over possible skips.","key_machinery":"The latent skip-expansion mechanism: each observed interval between logged periods is decomposed as the sum of $k\\ge 1$ unobserved true cycle lengths, where $k$ is a latent variable with a prior favoring a single cycle but allowing skips. Posterior inference averages over all possible skip configurations instead of conditioning on a fixed one, and it is this averaging that underwrites the paper's claim of reduced bias and calibrated uncertainty.","core_discovery":"The paper's central claim is that the gap between two logged period start dates cannot be taken at face value as one menstrual cycle. SkipTrack instead writes each observed gap as a sum of one or more unobserved true cycle lengths, with the number of unlogged cycles in each gap treated as a latent variable to be inferred. A Bayesian hierarchical regression then relates the underlying cycle length and regularity to covariates, propagating uncertainty about skips into the estimates. In simulations, the paper reports that competing approaches that fix whether a skip occurred show estimation bias and overconfidence, while SkipTrack recovers the target effects; the same model, applied to a large","pith_inferences":["The prior on true cycle length is doing the heavy lifting in separating 'one 60-day cycle' from 'two 30-day cycles with a missed log'; the reported associations could shift under different priors, so sensitivity analysis is a natural next check.","If users are more likely to skip logging when their cycles are already irregular, the model's implicit assumption that logging and cycle physiology are independent could itself create or mask associations; linking logging behavior to cycle outcomes would test this.","A direct validation would compare posterior skip probabilities against independently confirmed period dates (hormonal or temperature markers) in a subset of participants; miscalibration would challenge the method.","The same 'gap may hide multiple events' structure applies to other self-tracked symptom diaries, such as headaches or asthma attacks, wherever a missing entry makes one observed interval ambiguous."],"forward_implications":["App-based studies that take each logged gap as one cycle will systematically inflate cycle-length estimates and narrow uncertainty; SkipTrack is designed to avoid both.","Reported associations of age, BMI, and race/ethnicity with cycle length and regularity from SkipTrack come with intervals that reflect uncertainty about skipped logs, not just sampling noise.","The hierarchical regression supports time-varying effects, so the same framework can trace how cycle regularity changes across the reproductive lifespan while skip uncertainty is propagated."],"supporting_citations":[],"fun_headline_variants":["Bayesian model recovers cycle length from skipped period logs","SkipTrack: Bayesian fix for missing period logs in health apps","Skipped period logs? New Bayesian model handles the uncertainty","How to get honest cycle estimates from messy app data: Bayesian model","Bayesian model accounts for unlogged cycles in period tracker data"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The model can only tell a skipped cycle from a genuinely long one through its prior distribution on cycle length, and it assumes skipping is unrelated to the irregularity being studied; if either gives way, the corrected estimates collapse.","fun_headline_variants_meta":{"raw":{"variants":["Bayesian model recovers cycle length from skipped period logs","SkipTrack: Bayesian fix for missing period logs in health apps","Skipped period logs? New Bayesian model handles the uncertainty","How to get honest cycle estimates from messy app data: Bayesian model","Bayesian model accounts for unlogged cycles in period tracker data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00066,"raw_usage":{"total_tokens":2836,"prompt_tokens":709,"completion_tokens":2127,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":453,"completion_tokens_details":{"reasoning_tokens":2041}},"tokens_in":453,"tokens_out":2127,"duration_ms":15977,"temperature":1.0,"reasoning_tokens":2041,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T23:06:43.881207+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Find app users with independent confirmation of every period (daily hormone or temperature monitoring). For a user with a confirmed true 60-day cycle and no missed logging, the model should put most posterior mass on one 60-day cycle; if it instead assigns high probability to two 30-day cycles with a skip, the skip-correction mechanism is being driven by the prior, not the data.","supporting_citations":[],"review_version":1}