{"id":"799eb2dc-c2cf-43b5-a4c5-0d918bdd94e6","arxiv_id":"2411.17393","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A time-dependent Poisson-gamma recruitment model is equipped with homogeneity tests and moving-window re-estimation, evaluated on simulated clinical trial data.","lead":"This paper extends a standard statistical model for forecasting how many patients a clinical trial will recruit, allowing recruitment rates to change over time and adding tests for detecting such changes. It also proposes analytic approximations and a moving-window strategy for re-forecasting during a trial.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Lemma 3.1's Poisson-gamma sum approximation is unproven for time-dependent rates; its tail accuracy directly controls the analytic predictive bounds and the PG homogeneity test.","rationale":"The reader's verdict (CONDITIONAL, high confidence) identifies Lemma 3.1 as the weakest assumption, and my reading agrees: the entire analytic forecasting methodology and the Poisson-gamma homogeneity test depend on an approximation that is asserted rather than established for time-dependent rates. I examined whether a more concrete internal flaw existed, such as the Section 5 moving-window example missing the true completion date by about 100 days; that is a real practical weakness, but the theoretical bottleneck is the unproven PG-sum approximation. The paper does provide some independent support through Monte Carlo calibration of the PG test in Tables 4 and 5, and the type I error bounds are reported as near 0.07 rather than 0.10, which suggests conservative behavior in the tested scenarios. However, those simulations use relatively large numbers of centres and do not stress the approximation in small samples or strongly time-varying regimes. A focused quantitative check of the approximation's tail accuracy is therefore the single most decisive test: if the approximation fails there, the analytic predictive bounds and the PG test's P-values are not trustworthy, and the paper's central claim would need substantial qualification. If the approximation passes the check, the conditional verdict should stand as an accept-with-revision. Since the reader already conditioned on this issue, no verdict change is warranted.","tokens_in":20609,"tokens_out":5205,"duration_ms":54525,"concrete_test":"Under the Section 5 simulation setup with N = 2, 3, 5 centres, exponential-decay r(t), and staggered activation dates, compute the exact distribution of n(Is,t) at an interim time (e.g., day 200) by conditioning on the gamma rates and summing Poisson variables, or by a 10^7-replicate Monte Carlo. Compare the exact 0.1, 0.5, and 0.9 quantiles and the achieved coverage of the nominal 80% predictive interval with the PG(A,B) approximation in Section 3. If the achieved coverage deviates from 0.80 by more than 0.05, or the nominal 0.1 lower/upper tail probabilities are off by more than 0.02, Lemma 3.1 is not reliable for time-dependent rates and the analytic bounds and PG test P-values need explicit error control.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Lemma 3.1 asserts that n(Is,t), a sum of independent time-dependent PG processes, is well approximated by PG(A(Is,t), B(Is,t)) with moment-matched parameters. The only support is a citation to numerical work in [9] for the homogeneous case; no proof, error bound, or numerical check is given for general r(t). When rates are time-dependent, the cumulative windows R_i(0,t,u_i) differ across centres because activation times differ, so the random cumulative rate Λ(Is,t) is a sum of independent gamma variables with unequal scales. Matching the first two moments of the count distribution to a single PG variable does not control higher moments or tail quantiles. Yet Section 3 uses quantiles of this approximate PG for 80% predictive bounds, and Section 4.5 uses the same approximation to compute PG test P-values. For small numbers of centres or strongly time-varying r(t), the tail mismatch could make predictive intervals miscalibrated and PG test P-values unreliable. The paper contains no evidence that the approximation holds in the time-dependent setting, so the central analytic machinery rests on an unverified assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a time-dependent extension of the Poisson-gamma (PG) recruitment model in which each centre's recruitment rate is a gamma-distributed baseline multiplied by a common non-negative rate-shape function r(t). It proposes approximating the aggregate count over a set of centres by a single moment-matched PG variable, derives analytic predictive bounds from that approximation, and develops three tests for rate homogeneity: an exact conditional binomial test, a Poisson parametric test, and a PG parametric test. The paper derives asymptotic sample-size relations for the Poisson tests, calibrates critical values and evaluates power by Monte Carlo, and recommends a moving-window parameter re-estimation strategy at interim analyses. A single simulated artificial trial is used to illustrate the forecasting approach.","tokens_in":20775,"tokens_out":9224,"duration_ms":86634,"significance":"If the moment-matched PG approximation is reliable for time-dependent rates, the analytic predictive bounds and the PG homogeneity test would be practically useful because they avoid Monte Carlo for country/global forecasts. The paper has clear strengths: Lemma 4.1's conditional binomial test is exact under the stated Poisson model; the variance calculations for the non-parametric and parametric Poisson test statistics appear correct; and the power analysis is unusually honest, showing that controlling the expected P-value yields only around 70% power and then searching for sample sizes that give 80%. The simulation-based calibration of the critical value a(delta) is a constructive way to handle discreteness. The main weakness is that the central PG-sum approximation is asserted rather than established for the time-dependent case, and the practical forecasting comparison rests on a single simulated trajectory.","major_comments":[{"comment":"The approximation of n(Is,t) by PG(A(Is,t), B(Is,t)) is the load-bearing device of the paper: it is used in Section 3 for the analytic predictive bounds and in Section 4.5 for the PG test P-values. As written, Lemma 3.1 is not proved; the text cites numerical calculations in [9] for the homogeneous case and asserts an extension to time-dependent rates with different gamma parameters. When r(t) is time-dependent and activation times differ, the cumulative rates Lambda_i are gamma variables with unequal scales, and matching two moments of the count does not control higher moments or tail quantiles. I ask for either a proof/error bound for the time-dependent case or a systematic numerical verification covering small numbers of centres (N = 2, 3, 5, 10) and strongly time-varying r(t) (e.g., the 2.5x-to-0.2x exponential decline used in Section 5). The verification should report tail discrepancies, such as maximum absolute differences between the true and approximate CDFs, and the achieved coverage of the 80% predictive intervals, because these quantities drive the paper's practical claims.","section":"Section 3, Lemma 3.1"},{"comment":"The calibrated critical values for the PG test in Table 4 are approximately 0.066-0.074, not 0.1, whereas for the Poisson test in Table 1 they are approximately 0.092-0.100. The paper does not comment on this discrepancy. This matters because the practical recommendation in Section 6 is to use P-values with a threshold delta; if a practitioner uses the uncalibrated threshold 0.1, the Type I error will differ from the nominal delta, and the powers reported in Tables 4-5 are computed with the calibrated a(delta), not with 0.1. Please state the uncalibrated Type I error, explain the source of the deviation (discreteness, parameter estimation, or the PG approximation), and either recommend calibration or show that the deviation is negligible in the intended operating range.","section":"Section 4.6, Tables 4 and 5"},{"comment":"The claim that a moving-window re-estimation strategy improves forecasting is supported by only one simulated trajectory. In that example the moving-window forecast (day 287) is only 18 days closer to the actual completion (day 384) than the all-data forecast (day 269), and both forecasts are far from the truth because the continuing decline is not extrapolated. The text acknowledges this, but it still concludes that the moving-window approach is an improvement. Please replace this with a repeated-simulation study reporting a forecast error metric (e.g., median/mean absolute error of completion time or coverage of predictive intervals) over multiple replications and over several scenarios with different rate-decline shapes and window lengths. This is needed to substantiate the interim re-projection recommendation that is part of the paper's stated contribution.","section":"Section 5"}],"minor_comments":[{"comment":"The sentence 'Otherwise, the mean rate in [a,b] is smaller' is not logically implied by PUpper > delta; the lower P-value PLow must be used, and the criterion should be stated as 'if PLow <= delta, conclude the rate in [a,b] is smaller'.","section":"Section 4.2, after Eq. (11)"},{"comment":"The paper relies on [14] (to appear) for the core time-dependent PG methodology and on [9] for the PG-sum approximation. For a self-contained manuscript, please summarize the relevant results from [14] or provide a published/available version, and clarify what part of Lemma 3.1 is proven versus numerically verified in [9].","section":"Section 1 and reference [14]"},{"comment":"Terminology is inconsistent between 'centres' and 'sites' in the text and in figures; please harmonize the wording, and clarify in Section 3.2.1 that the indexing in 'veclaik[k] = 0 for k < vecu[i]' refers to positions on the daily simulation grid.","section":"Throughout"},{"comment":"Reference [31] gives page numbers '85-506', which appear to be a typo; please verify the correct pagination or article numbers.","section":"Reference list"}],"recommendation":"major_revision","confidential_remarks":"The central concern is the unverified Lemma 3.1 for time-dependent rates; this is fixable by adding a proper numerical verification or an error bound. If the authors can point to an existing proof in [9] or [14] for the general gamma-sum case, the revision can be lighter, but the paper as submitted does not make that case. The forecasting example in Section 5 should also be expanded into a repeated-simulation study, since the current single trajectory does not support the moving-window recommendation. The exact conditional binomial test and the honest power analysis are genuine positives."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the quick read. The genuinely new thing is the Poisson-gamma homogeneity test in Section 4.5 plus its power analysis, and that analysis is more honest than most: it plainly says that targeting expected P-value ≤ 0.1 only gets ~70% power and then computes what you need for 80%. That part is worth a referee's time.\n\nThe rest is a competent repackaging of the first author's earlier PG machinery. The binomial test in Section 4.1 is exact under H0; the variance calculations in Sections 4.2 and 4.4 check out. The paper is also fair about the competing time-dependent models (Urbas et al., Turchetta et al.).\n\nThe main soft spot is Lemma 3.1. The sum of independent PG processes with time-dependent rates and staggered centre activations is approximated by a single PG variable via moment matching, and the support is only a citation to numerical work for the homogeneous case. That approximation feeds directly into the analytic predictive bounds in Section 3 and the P-values of the PG test in Section 4.5. For small numbers of centres or strongly varying r(t), the tail mis-calibration could bite. This needs either a proof sketch or a numerical study across a grid of rate-shape functions and centre counts.\n\nThe simulation example is also oversold. At the day-200 interim, the all-data forecast says completion on day 269 and the moving-window says 287; the true value is 384. The text calls this 'an improvement' — it is an improvement between two forecasts that are both badly wrong. The method that assumes r(t) known gives day 391, but r(t) is not known in practice. The paper acknowledges these limits, but Section 5 frames the moving-window result too generously.\n\nMissing for a methods paper aimed at practitioners: no code, no real-data validation.\n\nAll in, this is a worthwhile subfield contribution. The new test is real and the power analysis is usable. Send it to peer review, but with the expectation that the authors either validate Lemma 3.1 under time-dependence or soften the claims, and that they rewrite the moving-window example. I would bring it to a reading group, mostly to discuss the approximation.","headline":"New PG homogeneity test with honest power analysis; the time-dependent PG sum approximation is unproven and the moving-window example is oversold.","tokens_in":21353,"tokens_out":5429,"would_cite":true,"duration_ms":50952,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62P10","60G55","62F03"],"pacs":[],"model":"deepseek-v4-flash","headline":"A Poisson-gamma approximation lets trial planners forecast recruitment analytically even when rates change over time.","keywords":["patient recruitment forecasting","Poisson-gamma model","time-dependent rates","homogeneity testing","clinical trials","interim analysis","moving window","mixed Poisson process"],"falsifier":"A simulation study where the true rate function r(t) (e.g., a steep exponential decay) is combined with a small number of centres (say, 2-3) and strongly time-varying rates; if the observed quantiles of the total recruitment deviate from the negative-binomial quantiles predicted by Lemma 3.1 by more than the claimed $10^{{-4}}$ error, the approximation fails.","tokens_in":20338,"feed_emoji":"📈","tokens_out":1516,"duration_ms":15659,"temperature":0.7,"pith_summary":"This paper extends the standard Poisson-gamma (PG) recruitment model for clinical trials to allow time-dependent recruitment rates, and shows that the total recruitment process across centres can still be approximated by a single PG variable with moment-matched parameters. It introduces a set of homogeneity tests—non-parametric Poisson, parametric Poisson, and PG-based—to detect whether rates have changed between two intervals, and recommends a moving-window re-estimation strategy for interim prediction when rates are declining. If correct, trial operations can obtain analytic means and predictive bounds for patient recruitment without Monte Carlo simulation, provided the shape of the rate function r(t) is known.","feed_headline":"One approximation forecasts trial recruitment as rates change","feed_subtitle":"Poisson-gamma sums give analytic bounds even with time-dependent rates, plus tests to detect when rates shift.","key_machinery":"The key object is the Poisson-gamma (PG) approximation of the sum of independent PG processes with time-dependent rates. The approximation matches the mean and variance of the cumulative rate of the country-level process to those of a single gamma distribution, yielding closed-form predictive mean and negative-binomial quantile bounds (via qnbinom). This approximation is what allows analytic forecasting without Monte Carlo simulation and what underlies the PG homogeneity test in Section 4.5.","core_discovery":"The central claim is Lemma 3.1: the distribution of the total recruitment process n(Is, t) over a country or region, where individual centres follow a PG process with time-dependent rates, can be well approximated by a PG random variable PG(A(Is, t), B(Is, t)) with moment-matched parameters A = $E^{2}$/$S^{2}$ and B = E/$S^{2}$. Building on this, the paper shows that the Poisson and Poisson-gamma homogeneity tests can detect time-dependent recruitment rates at an interim analysis, and that a moving-window re-estimation strategy improves forecasting when rates change. The authors further provide analytic formulas for the number of centres needed to detect a given proportional rate difference with a given confidence level, and demonstrate through simulation that when rates are declining steadily, the moving-window approach outperforms the standard all-data approach, while the best predictions come from knowing the rate function r(t).","pith_inferences":["The moment-matched PG approximation could also be applied to other doubly stochastic Poisson processes where the mixing distribution is not gamma, by matching two moments to a gamma surrogate; the accuracy would then depend on the tail behaviour of the true mixing distribution.","The authors' recommendation to use a moving window of 2–4 months could be tested against adaptive window-length selection based on the observed rate change magnitude, e.g., choosing the window that minimizes prediction error in a rolling validation.","The paper's homogeneity tests are conditional on a fixed schedule of centre initiations; an extension could treat centre activation times as random, which would change the probability p in the binomial test and likely require a different test statistic.","The PG test's requirement of many more centres to achieve 80% power (up to ~2000 for dramatic rate declines) suggests that for typical late-phase trials, the non-parametric Poisson test may be the only practical interim check, and the PG test should be reserved for trials with very large numbers of centres."],"forward_implications":["If correct, trial planners can obtain analytic predictive means and confidence bounds for patient recruitment at country and global levels without running Monte Carlo simulations, as long as the rate function r(t) is known.","The homogeneity tests provide an interim stage check for time-dependent recruitment rates, enabling operational decisions about whether to switch from a standard PG model to a moving-window or time-dependent model.","The formulas for the required number of centres to detect a given proportional rate difference (e.g., N = 2 z_delta^2 / (m1 L) * (1+q)/(1-q)^2 for the non-parametric Poisson test) give trial planners a direct tool for designing monitoring schemes.","The moving-window re-estimation strategy, using only the most recent data window for parameter estimation, is shown in simulations to improve prediction accuracy when rates are declining steadily, compared to using all historical data.","If the true rate function r(t) is known and correctly estimated, the time-dependent maximum likelihood approach yields substantially better predictions than both the all-data and moving-window approaches, as shown in the simulation example."],"supporting_citations":[{"why":"Provides the numerical evidence that the PG approximation of sums of PG processes is accurate even for a small number of centres (N=2,3), which Lemma 3.1 extends to the time-dependent case.","marker":"[9]"},{"why":"This is the companion paper by the first author that develops the basic methodology for a PG model with time-dependent rates; the present paper extends its results.","marker":"[14]"},{"why":"Introduces the original Poisson-gamma recruitment model and the maximum likelihood estimation technique used throughout this paper for fitting parameters.","marker":"[1]"},{"why":"Provides the standard PG model estimation and prediction techniques (including the analytic predictive bounds) that this paper extends to time-dependent rates and moving-window re-estimation.","marker":"[5]"},{"why":"Proposes a more general time-dependent PG model using maximum likelihood, which the authors' model is similar to; used as a comparison baseline.","marker":"[31]"}],"fun_headline_variants":["Time-dependent Poisson-gamma model improves trial recruitment forecasts","Detect shifting recruitment rates with new Poisson-gamma homogeneity tests","Moving-window re-estimation boosts clinical trial recruitment predictions","Analytic approximation handles time-varying patient recruitment rates","Interim analysis detects recruitment rate shifts via Poisson-gamma tests"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The analytic forecasting and the PG homogeneity test depend on Lemma 3.1, which asserts that the sum of independent PG processes with time-dependent rates is well approximated by a single PG variable with moment-matched parameters; the accuracy of this approximation for small numbers of centres or strongly time-varying rates is not proven here, only asserted based on numerical evidence for the homogeneous case.","fun_headline_variants_meta":{"raw":{"variants":["Time-dependent Poisson-gamma model improves trial recruitment forecasts","Detect shifting recruitment rates with new Poisson-gamma homogeneity tests","Moving-window re-estimation boosts clinical trial recruitment predictions","Analytic approximation handles time-varying patient recruitment rates","Interim analysis detects recruitment rate shifts via Poisson-gamma tests"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000157,"raw_usage":{"total_tokens":1254,"prompt_tokens":1008,"completion_tokens":246,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":624,"completion_tokens_details":{"reasoning_tokens":166}},"tokens_in":624,"tokens_out":246,"duration_ms":3856,"temperature":1.0,"reasoning_tokens":166,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:09:18.352216+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A simulation study where the true rate function r(t) (e.g., a steep exponential decay) is combined with a small number of centres (say, 2-3) and strongly time-varying rates; if the observed quantiles of the total recruitment deviate from the negative-binomial quantiles predicted by Lemma 3.1 by more than the claimed $10^{{-4}}$ error, the approximation fails.","supporting_citations":[{"cited_title":"Anisimov and M","cited_arxiv_id":null,"evidence_quote":"Provides the numerical evidence that the PG approximation of sums of PG processes is accurate even for a small number of centres (N=2,3), which Lemma 3.1 extends to the time-dependent case."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This is the companion paper by the first author that develops the basic methodology for a PG model with time-dependent rates; the present paper extends its results."},{"cited_title":"Anisimov and V","cited_arxiv_id":null,"evidence_quote":"Introduces the original Poisson-gamma recruitment model and the maximum likelihood estimation technique used throughout this paper for fitting parameters."},{"cited_title":"Anisimov, Statistical modeling of clinical trials (recruitment and randomization),Communications in Statistics - Theory and Methods, 40, 19-20, 2011, pp","cited_arxiv_id":null,"evidence_quote":"Provides the standard PG model estimation and prediction techniques (including the analytic predictive bounds) that this paper extends to time-dependent rates and moving-window re-estimation."},{"cited_title":"Urbas, C","cited_arxiv_id":null,"evidence_quote":"Proposes a more general time-dependent PG model using maximum likelihood, which the authors' model is similar to; used as a comparison baseline."}],"review_version":1}