Pith. sign in

REVIEW 4 major objections 4 minor 29 references

SkipTrack: A Bayesian Hierarchical Model for Self-tracked Menstrual Cycle Length and Regularity in Large Mobile Health Cohorts

T0 review · 4 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read SkipTrack claims that accounting for unlogged periods—instead of assuming each gap between logged periods is one cycle—reduces bias and overconfidence in app-based estimates of how age, BMI, and race/ethnicity relate to menstrual cycle leng

desk verdict Useful idea in search of validation: SkipTrack treats skip status as latent, but its superiority claim rests on simulations that can't yet be checked and an identifiability assumption the paper needs to defend. read the letter →

arxiv 2508.05845 v1 pith:VYDBBQYS submitted 2025-08-07 stat.AP

classification stat.AP MSC 62F1562P10
keywords menstrualcyclelengthregularityBayesianhierarchicalmodelself-trackedmobilehealthskippedperiodloggingmissingdatacovariateeffects
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Self-tracked menstrual cycle data have a hidden-data problem: when a user fails to log a period, the gap between logged periods is not one cycle but two or more. This paper presents SkipTrack, a Bayesian hierarchical model that treats the number of hidden cycles in each gap as unknown and averages over the possibilities while estimating how age, body mass index, and race/ethnicity shape cycle length and regularity. The paper argues, with simulations, that methods which assume skip status is known are more prone to biased estimates and overconfident intervals, and it applies SkipTrack to a large US mobile-health cohort. If the claim holds, previously reported covariate associations in app-based cycle data may need revisiting, because part of what looked like longer or more irregular cycles could be logging gaps.

What carries the argument

The latent skip-expansion mechanism: each observed interval between logged periods is decomposed as the sum of $k\ge 1$ unobserved true cycle lengths, where $k$ is a latent variable with a prior favoring a single cycle but allowing skips. Posterior inference averages over all possible skip configurations instead of conditioning on a fixed one, and it is this averaging that underwrites the paper's claim of reduced bias and calibrated uncertainty.

What would settle it

Find app users with independent confirmation of every period (daily hormone or temperature monitoring). For a user with a confirmed true 60-day cycle and no missed logging, the model should put most posterior mass on one 60-day cycle; if it instead assigns high probability to two 30-day cycles with a skip, the skip-correction mechanism is being driven by the prior, not the data.

Watch

Extended reading notes

Core claim

The paper's central claim is that the gap between two logged period start dates cannot be taken at face value as one menstrual cycle. SkipTrack instead writes each observed gap as a sum of one or more unobserved true cycle lengths, with the number of unlogged cycles in each gap treated as a latent variable to be inferred. A Bayesian hierarchical regression then relates the underlying cycle length and regularity to covariates, propagating uncertainty about skips into the estimates. In simulations, the paper reports that competing approaches that fix whether a skip occurred show estimation bias and overconfidence, while SkipTrack recovers the target effects; the same model, applied to a large

Load-bearing premise

The model can only tell a skipped cycle from a genuinely long one through its prior distribution on cycle length, and it assumes skipping is unrelated to the irregularity being studied; if either gives way, the corrected estimates collapse.

Editorial extensions

If this is right

  • App-based studies that take each logged gap as one cycle will systematically inflate cycle-length estimates and narrow uncertainty; SkipTrack is designed to avoid both.
  • Reported associations of age, BMI, and race/ethnicity with cycle length and regularity from SkipTrack come with intervals that reflect uncertainty about skipped logs, not just sampling noise.
  • The hierarchical regression supports time-varying effects, so the same framework can trace how cycle regularity changes across the reproductive lifespan while skip uncertainty is propagated.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The prior on true cycle length is doing the heavy lifting in separating 'one 60-day cycle' from 'two 30-day cycles with a missed log'; the reported associations could shift under different priors, so sensitivity analysis is a natural next check.
  • If users are more likely to skip logging when their cycles are already irregular, the model's implicit assumption that logging and cycle physiology are independent could itself create or mask associations; linking logging behavior to cycle outcomes would test this.
  • A direct validation would compare posterior skip probabilities against independently confirmed period dates (hormonal or temperature markers) in a subset of participants; miscalibration would challenge the method.
  • The same 'gap may hide multiple events' structure applies to other self-tracked symptom diaries, such as headaches or asthma attacks, wherever a missing entry makes one observed interval ambiguous.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes SkipTrack, a Bayesian hierarchical model for menstrual cycle length and regularity in large mobile health cohorts. The model treats potentially skipped cycle-tracking events as latent indicators and jointly estimates true cycle length, skip probability, and covariate effects (age, BMI, race/ethnicity). The abstract claims that, in simulations, SkipTrack outperforms methods that specify skip status a priori, which are said to suffer from estimation bias and overconfidence. The model is then applied to the Apple Women's Health Study to estimate associations between demographic covariates and menstrual cycle outcomes.

Significance. If the latent-skip decomposition is identifiable, the framework would be a valuable contribution to the analysis of self-tracked menstrual cycle data, where unlogged period starts can inflate observed cycle lengths and lead to overconfident estimates. The paper addresses a real and timely problem in digital cohort research. However, the visible evidence does not currently support the abstract's central claim: the separation between skipped cycles and genuinely long cycles is not validated against known skip labels, the simulation design appears to generate data from the same model family as SkipTrack, and the real-data application has no external ground truth for skip status. The contribution is promising but the validation is incomplete.

major comments (4)
  1. [Abstract and model specification] The central claim that SkipTrack 'accounts for the uncertainty of possible skips' requires that the latent skip indicators are identifiable from observed inter-bleed intervals. An observed interval of, say, 60 days could be one 60-day cycle or two 30-day cycles with one unlogged period start. The manuscript does not report a recovery analysis comparing posterior skip probabilities to known skip labels in settings where long cycles and skipped cycles coexist, nor a prior-sensitivity analysis for the cycle-length and skip-probability priors. Without that, every covariate effect on cycle length and regularity inherits the prior's decomposition. Please add simulations that vary the true long-cycle rate and skip rate and report posterior classification accuracy, coverage, and calibration of the skip indicators.
  2. [Simulation study] The simulation comparison appears to generate data from the same model family as SkipTrack, so the comparison with a-priori skip rules may simply reflect a correctly specified model beating misspecified competitors. This is a form of circular validation. Add robustness simulations generated under a different process—for example, skip probability depending on current cycle length, previous skip status, or user-level random effects—and show whether the claimed bias and coverage advantages persist. This is needed for the abstract's 'superiority' claim to be credible.
  3. [Application to Apple Women's Health Study] The real-data associations for age, BMI, and race/ethnicity are presented as being closer to the underlying biology than prior estimates, but there is no external validation of skip status (e.g., hormone-based cycle phase, follow-up surveys, or comparison with self-reported regularity). Because the model's skip decomposition is untested, these associations should be framed as model-dependent. At minimum, run a prior-sensitivity analysis over the skip model and report how the covariate associations change.
  4. [Model assumptions] The model assumes that, conditional on covariates, skipping a tracked period is not informative about the true cycle length or regularity under study. If users with irregular cycles are more likely to miss logging a period, the posterior over skip indicators—and therefore the covariate effects—can be biased. The manuscript neither states this ignorability assumption nor reports sensitivity analyses. Please state the assumption explicitly and assess robustness, e.g., by letting skip probability depend on true cycle length or on a user-level random effect.
minor comments (4)
  1. [Abstract] The study name is given as 'Apple Women's Healthy Study' in the abstract but 'Apple Women's Health Study' in the main text. Please use the correct name consistently.
  2. [Notation] Several symbols in the model equations are not defined near first use in the legible portions of the manuscript. A single, self-contained notation table would improve readability.
  3. [Simulation tables] The simulation tables report point estimates only. Please include Monte Carlo standard errors, number of replicates, and the width/coverage of the estimated intervals, especially for the competing methods.
  4. [Figures] The figures are difficult to interpret without clearer labeling of credible intervals and, where applicable, posterior probabilities of skip indicators. Consider adding a panel that shows posterior skip probabilities versus true skip indicators in the simulation.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the model's simulation validation is an internal consistency check, and the identifiability concerns are modeling limitations rather than derivation-circular steps.

full rationale

The paper's main chain is: (i) propose a hierarchical model with latent skip indicators; (ii) simulate data and compare parameter recovery against methods that fix skip status a priori; (iii) apply the model to the Apple Women's Health Study. None of these steps defines the target quantity in terms of the input, fits a parameter and then renames it as a prediction, or imports a load-bearing conclusion from a self-citation. In a simulation study, choosing a data-generating process that matches the model family is a deliberate and transparent condition; the resulting superiority over misspecified competitors is an expected property of the assumed DGP, not a hidden circular derivation of real-world validity. The potential non-identifiability between skipped cycles and genuinely long cycles is a substantive assumption about the prior, and the lack of ground-truth skip labels is a limitation of external validation; however, stating that the model 'accounts for the uncertainty' of skips is not equivalent to assuming that the posterior split is identified. The abstract explicitly limits the superiority claim to simulations. No load-bearing self-citations, imported uniqueness theorems, ansatz-smuggling citations, or renaming of known results are visible in the legible portions. Thus the derivation chain is self-contained and not circular.

Assumptions & free parameters 3 free parameters · 4 assumptions · 1 invented entities

The model's ability to detect skipped cycles rests entirely on its prior for true cycle length and on the assumption that unlogged cycles behave like whole cycles. With only the abstract readable, the quantitative prior choices and identifiability guarantees could not be audited.

free parameters (3)
  • True cycle length distribution (mean, variance, possibly time-varying)
    The model assigns a parametric distribution to unobserved true cycle lengths; these parameters are estimated from observed gaps and drive the skip-versus-long-cycle split. Values are not stated in the abstract.
  • Per-cycle skip probability or skip propensity
    The latent skip mechanism is identified only through the observed distribution of intervals between logged bleed starts. Not quantified in the abstract.
  • Covariate regression coefficients (age, BMI, race/ethnicity)
    Reported association estimates from the Apple Women's Health Study application are fitted values from data, not predictions with independent validation.
assumptions (4)
  • domain assumption Observed gaps between logged bleed starts are integer multiples of a single true cycle length, i.e., only whole cycles are skipped
    Required for the skip variable to be identifiable. Partial or fragmentary logging would break the integer-multiple structure and is not addressed in the abstract.
  • domain assumption The parametric family chosen for true cycle length is correctly specified
    If the true distribution has heavier tails than assumed, long cycles will be misclassified as skipped cycles. The abstract reports no sensitivity analysis.
  • domain assumption Skip behavior is conditionally independent of cycle characteristics and recorded covariates (ignorable missingness)
    If users skip logging precisely because their cycles are irregular, estimates of regularity and covariate effects will be biased. The abstract does not state or test this assumption.
  • standard math Bayesian posterior computation (MCMC or variational) converges and priors are proper enough for identifiability
    Standard computational requirement for hierarchical Bayesian models; not stated in the abstract but relied on by the method.
invented entities (1)
  • Latent skip indicator (unobserved skipped cycles)
    purpose: Partitions a long observed gap into one or more true cycles plus an unlogged cycle
    Skips are never directly observed in app data; the model infers them from the assumed prior. The abstract describes no external handle such as diary validation that would ground the latent variable.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SkipTrack: A Bayesian Hierarchical Model for Self-tracked Menstrual Cycle Length and Regularity in Large Mobile Health Cohorts." pith.science (2026). https://pith.science/paper/VYDBBQYS

@misc{pith2026250805845,
  author       = {Pith},
  title        = {Pith review of: SkipTrack: A Bayesian Hierarchical Model for Self-tracked Menstrual Cycle Length and Regularity in Large Mobile Health Cohorts},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VYDBBQYS}},
  note         = {Machine review of arXiv:2508.05845}
}
read the original abstract

Menstrual cycle length and regularity are important vital signs with implications for a variety of acute and chronic health conditions. Large datasets derived from cycle-tracking mobile health apps are being used to investigate the effects of various covariates on menstrual cycle length and regularity. One limitation on these analyses is that recorded cycle lengths can be incorrectly inflated if users skip tracking any cycle related bleeding days in the app. Here we present SkipTrack, a novel Bayesian hierarchical framework for examining baseline and time-varying effects on menstrual cycle length and regularity while accounting for the uncertainty of possible skips in cycle tracking. In simulations we demonstrate the superiority of the SkipTrack model by showing that competing methods which specify cycle skips a priori are more susceptible to issues of estimation bias and overconfidence than the SkipTrack model. Finally, we apply the SkipTrack framework to data from the Apple Women's Healthy Study, a US-based digital cohort (consent provided at study enrollment) to examine patterns of association between age, BMI and race/ethnicity, and menstrual cycle length and regularity.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

29 extracted references · 28 canonical work pages

  1. [1]

    write newline

    " write newline "" before.all 'output.state := FUNCTION format.url url empty "" url if FUNCTION article output.bibitem format.authors "author" output.check author format.key output output.year.check new.block format.title "title" output.check new.block crossref missing format.jour.vol output format.article.crossref output.nonnull format.pages output if ne...

  2. [2]

    barticle [author] Agarwal , Sanjay K S. K. , Chapron , Charles C. , Giudice , Linda C L. C. , Laufer , Marc R M. R. , Leyland , Nicholas N. , Missmer , Stacey A S. A. , Singh , Sukhbir S S. S. Taylor , Hugh S H. S. ( 2019 ). Clinical diagnosis of endometriosis: a call to action . American journal of obstetrics and gynecology 220 354--e1 . barticle

  3. [3]

    barticle [author] Bricker , Jonathan B J. B. , Sridharan , Vasundhara V. , Zhu , Yifan Y. , Mull , Kristin E K. E. , Heffner , Jaimee L J. L. , Watson , Noreen L N. L. , McClure , Jennifer B J. B. Di , Chongzhi C. ( 2018 ). Trajectories of 12-month usage patterns for two smoking cessation websites: Exploring how users engage over time . Journal of Medical...

  4. [4]

    barticle [author] Bull , Jonathan R J. R. , Rowland , Simon P S. P. , Scherwitzl , Elina Berglund E. B. , Scherwitzl , Raoul R. , Danielsson , Kristina Gemzell K. G. Harper , Joyce J. ( 2019 ). Real-world menstrual cycle characteristics of more than 600,000 menstrual cycles . NPJ digital medicine 2 83 . barticle

  5. [5]

    George , Edward I E

    barticle [author] Casella , George G. George , Edward I E. I. ( 1992 ). Explaining the Gibbs sampler . The American Statistician 46 167--174 . barticle

  6. [6]

    Greenberg , Edward E

    barticle [author] Chib , Siddhartha S. Greenberg , Edward E. ( 1995 ). Understanding the Metropolis-Hastings algorithm . The American Statistician 49 327--335 . barticle

  7. [7]

    , Damokosh , Andrew I A

    barticle [author] Cho , Sung-Il S.-I. , Damokosh , Andrew I A. I. , Ryan , Louise M L. M. , Chen , Dafang D. , Hu , Ye A Y. A. , Smith , Thomas J T. J. , Christiani , David C D. C. Xu , Xiping X. ( 2001 ). Effects of exposure to organic solvents on menstrual cycle length . Journal of occupational and environmental medicine 43 567--575 . barticle

  8. [8]

    , Laufer , Marc R M

    barticle [author] Diaz , AMRL A. , Laufer , Marc R M. R. , Breech , Lesley L L. L. et al. ( 2006 ). Menstruation in girls and adolescents: using the menstrual cycle as a vital sign. Pediatrics 118 2245--2250 . barticle

Show all 29 references
  1. [9]

    ( 2024 )

    bmanual [author] Duttweiler , Luke L. ( 2024 ). skipTrack: A Bayesian Hierarchical Model that Controls for Non-Adherence in Mobile Menstrual Cycle Tracking R package version 0.1.0 . bmanual

  2. [10]

    SkipTrack: A Bayesian Hierarchical Model for Self-tracked Menstrual Cycle Length and Regularity in Large Mobile Health Cohorts

    barticle [author] Duttweiler , Luke L. ( 2025 ). Supplement to "SkipTrack: A Bayesian Hierarchical Model for Self-tracked Menstrual Cycle Length and Regularity in Large Mobile Health Cohorts" . barticle

  3. [11]

    , Mahalingaiah , Shruthi S

    barticle [author] Duttweiler , Luke L. , Mahalingaiah , Shruthi S. Coull , Brent B. ( 2024 ). skipTrack: An R package for Identifying Skips in Self-Tracked Mobile Menstrual Cycle Data . Journal of Open Source Software 9 6928 . barticle

  4. [12]

    , Chen , Won Sun W

    barticle [author] Flitcroft , Leah L. , Chen , Won Sun W. S. Meyer , Denny D. ( 2020 ). The demographic representativeness and health outcomes of digital health station users: longitudinal study . Journal of Medical Internet Research 22 e14977 . barticle

  5. [13]

    barticle [author] Gibson , Elizabeth A E. A. , Li , Huichu H. , Fruh , Victoria V. , Gabra , Malaika M. , Asokan , Gowtham G. , Jukic , Anne Marie Z A. M. Z. , Baird , Donna D D. D. , Curry , Christine L C. L. , Fischer-Colbrie , Tyler T. , Onnela , Jukka-Pekka J.-P. et al. ( ...

  6. [14]

    , Thalabard , Jean-Christophe J.-C

    barticle [author] Giorgis-Allemand , Lise L. , Thalabard , Jean-Christophe J.-C. , Rosetta , Lyliane L. , Siroux , Val \'e rie V. , Bouyer , Jean J. Slama , R \'e my R. ( 2020 ). Can atmospheric pollutants influence menstrual cycle function? Environmental pollution 257 113605 ...

  7. [15]

    barticle [author] Grieger , Jessica A J. A. Norman , Robert J R. J. ( 2020 ). Menstrual cycle length and patterns in a global cohort of women using a mobile phone app: retrospective cohort study . Journal of Medical Internet Research 22 e17109 . barticle

  8. [16]

    barticle [author] Hammer , Karissa C K. C. , Veiga , Alexis A. Mahalingaiah , Shruthi S. ( 2020 ). Environmental toxicant exposure and menstrual cycle length . Current Opinion in Endocrinology, Diabetes and Obesity 27 373--379 . barticle

  9. [17]

    barticle [author] Jukic , Anne Marie Z A. M. Z. , Steiner , Anne Z A. Z. Baird , Donna D D. D. ( 2015 ). Lower plasma 25-hydroxyvitamin D is associated with irregular menstrual cycles in a cross-sectional study . Reproductive Biology and Endocrinology 13 1--6 . barticle

  10. [18]

    , Urteaga , I \ n igo I

    barticle [author] Li , Kathy K. , Urteaga , I \ n igo I. , Wiggins , Chris H C. H. , Druet , Anna A. , Shea , Amanda A. , Vitzthum , Virginia J V. J. Elhadad , No \'e mie N. ( 2020 ). Characterizing physiological and symptomatic variation in menstrual cycles using self-tracked...

  11. [19]

    , Urteaga , I \ n igo I

    barticle [author] Li , Kathy K. , Urteaga , I \ n igo I. , Shea , Amanda A. , Vitzthum , Virginia J V. J. , Wiggins , Chris H C. H. Elhadad , No \'e mie N. ( 2022 ). A predictive model for next cycle start date that accounts for adherence in menstrual self-tracking . Journal o...

  12. [20]

    , Gibson , Elizabeth A E

    barticle [author] Li , Huichu H. , Gibson , Elizabeth A E. A. , Jukic , Anne Marie Z A. M. Z. , Baird , Donna D D. D. , Wilcox , Allen J A. J. , Curry , Christine L C. L. , Fischer-Colbrie , Tyler T. , Onnela , Jukka-Pekka J.-P. , Williams , Michelle A M. A. , Hauser , Russ R....

  13. [21]

    , Gold , Ellen B E

    barticle [author] Liu , Yan Y. , Gold , Ellen B E. B. , Lasley , Bill L B. L. Johnson , Wesley O W. O. ( 2004 ). Factors affecting menstrual cycle characteristics . American journal of epidemiology 160 131--140 . barticle

  14. [22]

    , Fruh , Victoria V

    barticle [author] Mahalingaiah , Shruthi S. , Fruh , Victoria V. , Rodriguez , Erika E. , Konanki , Sai Charan S. C. , Onnela , Jukka-Pekka J.-P. , de Figueiredo Veiga , Alexis A. , Lyons , Genevieve G. , Ahmed , Rowana R. , Li , Huichu H. , Gallagher , Nicola N. et al. ( 2022...

  15. [23]

    , Jasienska , Grazyna G

    barticle [author] Merklinger-Gruchala , Anna A. , Jasienska , Grazyna G. Kapiszewska , Maria M. ( 2017 ). Effect of air pollution on menstrual cycle length—a prognostic factor of women’s reproductive health . International Journal of Environmental Research and Public Health 14...

  16. [24]

    , Harlow , Siob \'a n D S

    barticle [author] Paramsothy , Pangaja P. , Harlow , Siob \'a n D S. D. , Elliott , Michael R M. R. , Yosef , Matheos M. , Lisabeth , Lynda D L. D. , Greendale , Gail A G. A. , Gold , Ellen B E. B. , Crawford , Sybil L S. L. Randolph Jr , John F J. F. ( 2015 ). Influence of ra...

  17. [25]

    , Wand , Matt P M

    bbook [author] Ruppert , David D. , Wand , Matt P M. P. Carroll , Raymond J R. J. ( 2003 ). Semiparametric regression 12 . Cambridge university press . bbook

  18. [26]

    , Li , Cheng C

    barticle [author] Srivastava , Sanvesh S. , Li , Cheng C. Dunson , David B D. B. ( 2018 ). Scalable Bayes via barycenter in Wasserstein space . Journal of Machine Learning Research 19 1--35 . barticle

  19. [27]

    bbook [author] Strauss , Jerome F J. F. , Barbieri , Robert L R. L. , Dokras , Anuja A. , Williams , Carmen J C. J. Williams , S Zev S. Z. ( 2023 ). Yen & Jaffe's Reproductive Endocrinology-E-Book: Physiology, Pathophysiology, and Clinical Management . Elsevier Health Sciences . bbook

  20. [28]

    barticle [author] Vollmar , Ana K Rosen A. K. R. , Mahalingaiah , Shruthi S. Jukic , Anne Marie A. M. ( 2024 ). The Menstrual Cycle as a Vital Sign: a comprehensive review . F&S Reviews 100081 . barticle

  21. [29]

    , Asokan , Gowtham G

    barticle [author] Wang , Zifan Z. , Asokan , Gowtham G. , Onnela , Jukka-Pekka J.-P. , Baird , Donna D D. D. , Jukic , Anne Marie Z A. M. Z. , Wilcox , Allen J A. J. , Curry , Christine L C. L. , Fischer-Colbrie , Tyler T. , Williams , Michelle A M. A. , Hauser , Russ R. et al...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.