{"id":"a32dca5e-0b16-4414-930e-1f7fd6ee3bad","arxiv_id":"1908.09830","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Using last-crossing-time measures on 18 months of GPS data, the paper finds that human mobility patterns stabilize only after roughly 15 to 37 weeks of monitoring, far longer than the 14 days previously recommended.","lead":"GPS studies usually run for a week or two, but this paper's data on 185 people over 18 months suggests that mobility patterns need many weeks, often over 30, to stabilize. The authors built statistical measures of when an individual's movement pattern stops changing and used them to estimate how long GPS monitoring must last.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'at least 15 weeks' conclusion is anchored to the final 18-month estimate; if that endpoint is not the stable long-run pattern, the recommended durations are window artifacts.","rationale":"The reader's weakest assumption identifies exactly the load-bearing vulnerability: every LCT definition in the paper measures error relative to the final value of the process over the entire observed window. The empirical claim that GPS monitoring must last 'at least 15 weeks' only has design implications if that final value is a reliable proxy for the individual's stable long-run mobility pattern. The paper does not test this, and the retrospective nature of LCT makes the reported durations sensitive to the arbitrary 18-month horizon. The proposed split-sample test would directly settle whether the endpoint is stable: if LCTs computed against a disjoint later period (or against a truncated 12-month window) differ materially from the Table 1 values, then the central recommendation is window-dependent rather than a property of the individuals' mobility. The theoretical results on estimator consistency are separate and do not address this empirical anchoring problem, so the paper remains conditionally acceptable pending verification of the stability assumption.","tokens_in":15133,"tokens_out":3512,"duration_ms":36965,"concrete_test":"Using the MDC data, for each participant compute the weekly activity-distribution estimate \\hat{\\bar{\\pi}}(D) from the first D weeks and the estimate \\hat{\\bar{\\pi}}_{\\text{later}} from the disjoint later period (weeks D+1 to Dmax). Define a validation LCT as the last D such that ||\\hat{\\bar{\\pi}}(D) - \\hat{\\bar{\\pi}}_{\\text{later}}||_1 > \\gamma, and compare its mean and median with the reported LCT-distribution values in Table 1. Also repeat the full analysis using only the first 12 months of data. If the validation LCTs are substantially larger than the reported means, or if the 12-month analysis shifts the Table 1 entries by more than a few weeks, the reported 'minimum required length' is an artifact of the retrospective endpoint.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim depends on last-crossing-time measures defined against the final value of the entire observation window: Eq. (6) uses Z(tmax−tmin) and Eq. (14) uses \\hat{\\bar{\\pi}}(Dmax). These are retrospective summaries of the specific 18-month MDC window, not estimates of a stable long-run mobility pattern. If the activity distribution or velocity process drifts over the study period, or if the final 18-month estimate is not close to the individual's long-run equilibrium, then the LCT can be small simply because the endpoint itself has moved; conversely, a genuinely unstable individual can have a large LCT that is censored at the horizon. The paper provides no check that the endpoint is stable: no split-half analysis, no comparison against a holdout period, and no report of how many participants' LCTs hit Dmax. Table 1's means (30, 37, 18 weeks) and the Discussion's 'at least 15 weeks' are also not reconciled, but the deeper problem is that the LCT is not a forward-looking design parameter unless the endpoint is validated as the stable pattern.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a statistical framework for determining the minimum required length of GPS monitoring for human mobility studies. It defines last-crossing-time (LCT) measures based on the average velocity process and on activity distributions over a spatial grid, introduces ordinary and conservative proportional time estimators for activity distributions, and proves their asymptotic equivalence and consistency under assumptions (S1)-(S3). The method is applied to the Nokia Mobile Data Challenge, with GPS data from 185 individuals over about 18 months. The empirical results give mean LCT values of roughly 30 weeks for velocity, 37 weeks for the full activity distribution, and 18 weeks for the 0.2-level set of important places. The authors conclude that GPS monitoring should last at least 15 weeks, about seven times longer than the 14-day recommendation of Zenk et al., and that study duration should depend on demographic group.","tokens_in":15329,"tokens_out":4157,"duration_ms":42603,"significance":"If the empirical claims held, the paper would provide a valuable theoretical basis for a design question that has so far been addressed mainly empirically: how long GPS monitoring must last. The theoretical results are a genuine contribution: the consistency and asymptotic equivalence proofs for the proportional time estimators are clean under assumptions (S1)-(S3), and the LCT formalism gives a concrete, interpretable target for stabilization. The paper also makes constructive use of a publicly available longitudinal GPS dataset and makes falsifiable quantitative predictions (e.g., 30, 37, and 18 weeks in Table 1). However, the central empirical conclusion depends on an assumption about the endpoint of the observation window that is not tested, and the headline 'at least 15 weeks' is not directly supported by the table of aggregate results. The framework is promising, but the central design recommendation needs substantially more validation before it can be accepted.","major_comments":[{"comment":"All three LCT measures are defined relative to the final value of the process over the entire observed window: Eq. (6) uses Z(tmax-tmin), Eq. (14) uses \\hat{\\bar{\\pi}}(Dmax), and Eq. (17) uses Lα(Dmax). These are retrospective summaries of the particular 18-month MDC window, not estimates of a stable long-run mobility pattern. The paper's design recommendation assumes that the final estimate is close to each individual's stable mobility pattern, but it provides no check of this assumption: there is no split-half validation, no holdout-period comparison, no assessment of drift or nonstationarity over the study period, and no report of how many participants' LCTs are censored at the horizon. Without such checks, a short LCT can occur simply because the endpoint itself has moved, and a long LCT for a genuinely unstable individual may be truncated at Dmax. The 'minimum required length' conclusion is therefore contaminated by the arbitrary length of the observation window unless endpoint stability is established.","section":"Sections 2.1 and 2.3, Eqs. (6), (14), and (17)"},{"comment":"The aggregate results in Table 1 report mean LCTs of 30.04 weeks (velocity), 37.18 weeks (distribution), and 17.69 weeks (0.2-level set), yet the Discussion concludes that GPS monitoring needs to be done for 'at least 15 weeks.' This number does not follow from Table 1: it is closer to the subgroup-specific value for middle-age participants shown in Figure 4 for the level-set measure, and it is far below the mean for the other two measures. A minimum study duration should be derived from an explicit design criterion, typically an upper quantile of the distribution across participants, not from subgroup means. The paper should reconcile Table 1 with the Discussion, report the quantiles of the LCT distributions, and state how much censoring occurs at Dmax.","section":"Section 3, Table 1, and Section 4, Discussion"},{"comment":"The claim that study duration should differ by demographic group is based on visual comparison of mean LCT curves with 90% confidence intervals. No formal test, effect size, or adjustment for repeated measures or multiple comparisons is provided, and the confidence intervals are not defined (standard error? bootstrap? adjusted for within-person correlation?). Since demographic differentiation is presented as a second substantive contribution, it needs supporting inference rather than descriptive curves alone.","section":"Section 3, Figure 4"}],"minor_comments":[{"comment":"The text says the window is partitioned into '40002 square grid cells' while later text uses N = 4000^2; 4000^2 is 16,000,000, not 4,000. Please clarify the grid resolution and the total number of cells.","section":"Section 3, first paragraph of Application"},{"comment":"The sentence 'We denote by π(d) the activity distribution from Eq. (12) associated with time period D' conflates the true activity distribution from Eq. (9) with the estimator from Eq. (12). Please separate the estimand from the estimator.","section":"Section 2.3, text before Eq. (13)"},{"comment":"The denominator ||Lα(Dmax)|| can be zero for some participants or some values of α; please state a convention (e.g., define the LCT as 0 in that case) so that the ratio is well defined.","section":"Section 2.3, Eq. (17)"},{"comment":"The estimator \\hat V_k(τ) sums distances between consecutive observation times with t_{k,i+1} ≤ τ, which omits the partial segment between the last recorded time before τ and τ itself; please state this explicitly or use interpolation.","section":"Section 2.1, Eq. (5)"},{"comment":"There is a typo: 'histrogram' should be 'histogram'.","section":"Figure 2 caption"}],"recommendation":"major_revision","confidential_remarks":"The paper is essentially a preprint from 2019 with a clean theoretical core but an empirical central claim that is not yet supported. The main fix is feasible within the manuscript's scope: the authors need to validate the endpoint assumption (e.g., with split-half or rolling-window analyses), report censoring, and align the Discussion's headline number with the actual results. If the endpoint validation fails, the empirical conclusion would need to be substantially weakened. The theoretical consistency results are sound and should not be held hostage to the empirical issue, but the manuscript as a whole needs another round of revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know before you read this one. First, the statistical framework is genuinely new: last-crossing-time processes for mobility stability, plus ordinary and conservative proportional-time estimators of activity distributions with clean consistency proofs (Theorems 2.1 and 2.2). Second, the paper's headline empirical claim—'GPS monitoring needs to be done for at least 15 weeks'—is not actually supported by the paper's own Table 1, and the last-crossing-time approach is retrospective in a way that makes the recommended durations fragile.\n\nWhat the paper does well: it identifies a real gap in GPS study design, uses a rich 18-month public dataset (MDC), stratifies by demographic group, and is unusually honest about its assumptions (straight-line distances, dropped long trips, grid-cell discretization). The consistency theorems are modest but correctly proved, and the distinction between ordinary and conservative estimators is a nice practical touch.\n\nThe soft spots are real. Table 1 reports mean LCTs of 30 weeks (velocity), 37 weeks (distribution), and 18 weeks (level set at α=0.2). The Discussion says 'at least 15 weeks' without reconciling that number with these means; it appears to come from a subgroup analysis (middle-aged participants in Figure 4), but the text never connects them. More importantly, LCT is defined relative to the final estimate over the entire observed window—Eq. (6) uses Z(tmax−tmin) and Eq. (14) uses \\hat{\\bar{\\pi}}(Dmax). That makes it a retrospective summary of this particular 18-month dataset, not a forward-looking design parameter. If the endpoint has drifted or is still changing, the LCT can be small because the endpoint moved, or large because it never settled. The authors give no split-half check, no holdout comparison, and do not report how many participants' LCTs hit the observation horizon (the average observation length is 55 weeks, so censoring is plausibly widespread). Without such a check, the 'minimum required length' is at risk of being an artifact of the observation window. The spatial window that drops long trips is acknowledged and could bias stability estimates downward; key tuning parameters (γ, α, grid size) get no sensitivity analysis.\n\nWho is this for? Researchers designing GPS-based health or mobility studies, and anyone working on activity-space measures. The estimators and the conceptual framework deserve a serious referee. But the paper needs a major revision: reconcile the headline number with the tables, and validate that the endpoint used in the LCT is actually stable—for example, by recomputing LCTs on shorter windows and comparing, or by reporting censoring rates. I'd send it out, with the expectation of heavy revision.","headline":"New stability measures for GPS mobility are worth knowing, but the 'at least 15 weeks' headline is not backed by the paper's own tables and the LCT approach needs a stability check before its durations are used.","tokens_in":15863,"tokens_out":3277,"would_cite":true,"duration_ms":32143,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that GPS monitoring must last at least 15 weeks—roughly seven times the current 14-day recommendation—to capture temporally stable human mobility patterns, with average stabilization times of 30 weeks for velocity and 37…","keywords":["density estimation","global positioning systems (GPS)","human mobility","spatiotemporal trajectories","temporal dynamics","last crossing time","activity distribution","GPS monitoring duration"],"falsifier":"Recompute the three last-crossing-time measures on a GPS panel that runs three years or more; if median LCT-velocity rises above 30 weeks or a large share of participants still have last crossing times at the end of the observation window, the paper's stabilization times are censored by the 18-month horizon.","tokens_in":14913,"feed_emoji":"📍","tokens_out":6563,"duration_ms":59266,"temperature":0.7,"pith_summary":"This paper asks how long a GPS-based study must run before the mobility pattern it records stops changing in any systematic way. The authors build a statistical framework around last crossing times: the last moment at which an estimate made from the first part of a person's record still differs from the estimate made from the entire record by more than a tolerance. Applied to 18 months of phone GPS data from 185 people in Switzerland, the framework says monitoring should last at least 15 weeks, roughly seven times the widely cited 14-day minimum, with average stabilization times of 30 weeks for average velocity, 37 weeks for weekly activity distributions, and 18 weeks for the set of places where a person spends most time. The paper also argues that the required duration differs by demographic group, with younger participants needing longer observation than older ones.","feed_headline":"Fifteen weeks of GPS data, not two, capture stable mobility","feed_subtitle":"New last-crossing-time measures show common 7-to-14-day GPS studies capture activity patterns that are still unstable.","key_machinery":"The engine is the last crossing time (LCT) process. For a process $Z(\\tau)$ summarizing the trajectory up to time $\\tau$—average velocity or the estimated weekly activity distribution—the LCT is the largest $\\tau$ at which the absolute percentage error $\\varphi(Z;\\tau) = |Z(\\tau) - Z(T)|/Z(T)$ still exceeds a threshold $\\gamma$; it says how long one must observe before the prefix estimate stops disagreeing with the full-record estimate. The supporting machinery includes two estimators of the distribution of time spent in spatial grid cells, the ordinary and conservative proportional time estimators, which the paper proves are asymptotically equivalent and consistent, and a ranking-based $\\alpha$-level set that focuses the stability check on the grid cells where a person actually spends most of their time.","core_discovery":"The central claim is that human mobility patterns, as recorded by GPS, take far longer to become temporally stable than previous empirical work suggested. Using the last crossing time of the absolute percentage error with threshold $\\gamma = 0.2$, the authors find mean stabilization times of 30.04 weeks for average velocity, 37.18 weeks for the weekly activity distribution, and 17.69 weeks for the 0.2-level set of important places, and conclude that GPS monitoring needs to last at least 15 weeks—about seven times the 14 days recommended by the earlier study. A second claim is that stability is demographic: older adults' mobility stabilizes sooner (about 10 weeks for the level-set measure), middle-aged adults need about 15 weeks, and younger adults about 20 weeks, so study designs should set durations by demographic group rather than a single universal window.","pith_inferences":["Because every LCT is measured against the estimate from the entire 18-month window, the quoted stabilization times are best read as lower bounds tied to that window; a longer observation study could push them upward if truly stable patterns emerge only after more than 18 months.","The same last-crossing-time template could be applied to other longitudinal behavioral records (e.g., daily retail visits, app usage, or mobility from call detail records) by replacing the velocity or activity distribution process with the relevant summary statistic.","The demographic pattern suggests an adaptive design in which monitoring continues until a participant-specific LCT drops below a threshold, which could reduce participant burden while preserving stability guarantees.","Because the consistency theorems assume increasingly dense sampling, a testable implication is that much denser GPS recording could shorten the required calendar duration to reach the same stability level."],"forward_implications":["Seven- to fourteen-day GPS studies, common in health and social research, are likely to record mobility patterns that are still changing; conclusions drawn from them may not reflect stable long-run behavior.","A minimum monitoring length of about 15 weeks is needed for stable estimates; studies targeting full weekly activity distributions should plan for roughly 37 weeks, while studies that only need the main places of activity can use about 18 weeks.","Differential study durations by demographic group are warranted: older adults stabilize faster, so shorter monitoring may suffice, while younger adults need longer windows.","The ordinary and conservative proportional time estimators give consistent recovery of activity distributions even when GPS sampling is irregular, so researchers can use them to correct for non-uniform observation times.","The level-set measure provides a way to separate stability of core places from stability of rarely visited places, so researchers can tailor the observation window to the spatial resolution their research question needs."],"supporting_citations":[{"why":"Supplies the 14-day minimum GPS monitoring recommendation that the paper's 15-week finding directly contradicts and extends.","marker":"[43]"},{"why":"Earlier estimate (17 weeks from social media data) that the paper builds on and gives a theoretical framework to.","marker":"[25]"},{"why":"Describes the Lausanne data collection campaign that produced the MDC GPS data.","marker":"[18]"},{"why":"Introduces the Mobile Data Challenge and its dataset used for the empirical analysis.","marker":"[23]"},{"why":"Documents the MDC smartphone dataset, the source of the 185-participant, 18-month GPS records.","marker":"[24]"},{"why":"Provides the ranking distribution concept used to define alpha-level sets of important places.","marker":"[7]"},{"why":"Supplies the curve-length definition used in the average velocity process.","marker":"[9]"},{"why":"Used for Great Circle distance calculations between consecutive GPS locations.","marker":"[4]"}],"fun_headline_variants":["GPS studies need 15 weeks, not 2, for stable mobility","Mobility stability takes 15 weeks, not a fortnight","Demographics change GPS monitoring duration needs","Stable mobility needs 15 weeks of GPS data","Seven times longer: GPS monitoring for stability"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the estimate computed from the entire 18-month window equals each person's stable long-run mobility pattern; if the window is too short, every reported minimum duration is a lower bound rather than a true stabilization time.","fun_headline_variants_meta":{"raw":{"variants":["GPS studies need 15 weeks, not 2, for stable mobility","Mobility stability takes 15 weeks, not a fortnight","Demographics change GPS monitoring duration needs","Stable mobility needs 15 weeks of GPS data","Seven times longer: GPS monitoring for stability"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000133,"raw_usage":{"total_tokens":1100,"prompt_tokens":875,"completion_tokens":225,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":491,"completion_tokens_details":{"reasoning_tokens":149}},"tokens_in":491,"tokens_out":225,"duration_ms":2507,"temperature":1.0,"reasoning_tokens":149,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:18:44.154660+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the three last-crossing-time measures on a GPS panel that runs three years or more; if median LCT-velocity rises above 30 weeks or a large share of participants still have last crossing times at the end of the observation window, the paper's stabilization times are censored by the 18-month horizon.","supporting_citations":[{"cited_title":"Zenk, S.A","cited_arxiv_id":null,"evidence_quote":"Supplies the 14-day minimum GPS monitoring recommendation that the paper's 15-week finding directly contradicts and extends."},{"cited_title":"Lee, A.W","cited_arxiv_id":null,"evidence_quote":"Earlier estimate (17 weeks from social media data) that the paper builds on and gives a theoretical framework to."},{"cited_title":"Kiukkonen, J","cited_arxiv_id":null,"evidence_quote":"Describes the Lausanne data collection campaign that produced the MDC GPS data."},{"cited_title":"Laurila, D","cited_arxiv_id":null,"evidence_quote":"Introduces the Mobile Data Challenge and its dataset used for the empirical analysis."},{"cited_title":"Laurila, D","cited_arxiv_id":null,"evidence_quote":"Documents the MDC smartphone dataset, the source of the 185-participant, 18-month GPS records."},{"cited_title":"Chen, Generalized cluster trees and singular measures , Annals of Statistics 47 (2019), pp","cited_arxiv_id":null,"evidence_quote":"Provides the ranking distribution concept used to define alpha-level sets of important places."},{"cited_title":"Courant and F","cited_arxiv_id":null,"evidence_quote":"Supplies the curve-length definition used in the average velocity process."},{"cited_title":"Bivand, E","cited_arxiv_id":null,"evidence_quote":"Used for Great Circle distance calculations between consecutive GPS locations."}],"review_version":1}