{"id":"517f6483-8136-4364-8e79-b47035ac4023","arxiv_id":"1908.04053","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Estimating the mortality rate ratio from prevalence and incidence data is numerically unstable below age 55, so the paper recommends estimating the mortality rate difference instead.","lead":"This paper estimates two measures of excess death in diabetes from public prevalence and incidence data. It finds the mortality rate ratio estimate is wildly unstable at younger ages and recommends reporting the mortality rate difference instead.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The recommendation to prefer Δm is not yet supported: its apparent stability may be an artifact of the same two-survey logit-linear smoothing that makes R unstable, and no external or numerical validation of Δm is provided.","rationale":"The reader's weakest assumption identifies exactly the load-bearing premise: the direct method uses logit-linear prevalence fits at two time points, derives ∂p from the difference of those fits, and treats 2012 incidence as constant. Equations (2a) and (2b) both depend on ∂p, so if the smoothing is poor, the instability of R and the apparent stability of Δm could both be artifacts. The paper deserves credit for the bootstrap-based demonstration that R is unstable in Table 1; that negative finding is solid. The positive recommendation to prefer Δm, however, is not supported by an equivalent quantitative demonstration. The funnel in Figure 1 even suggests that Δm uncertainty also increases at younger ages, but without numeric intervals or an external benchmark it is impossible to judge whether the estimates are 'sensible' beyond visual inspection. The proposed additional-time-point check would distinguish a robust property of the underlying data from a smoothing artifact. If the check confirms the current result, the existing CONDITIONAL verdict is appropriate; if not, the recommendation would need to be tempered.","tokens_in":3861,"tokens_out":4747,"duration_ms":52105,"concrete_test":"Recompute Δm and R using at least three prevalence surveys between 2009 and 2015 (or annual data if available), with a nonparametric smoother and finite-difference ∂p, and with incidence allowed to vary by year. If Δm at ages 30–50 remains stable with overlapping confidence intervals while R remains unstable, the paper's recommendation is supported. If Δm shifts substantially or its intervals widen to include implausible values, the apparent stability in Figure 1 is an artifact of the two-survey logit-linear assumption. As a second check, report the numeric 2.5, 50, and 97.5 percentiles of Δm from the 2000 replicates at ages 15–50, and compare them with external estimates of excess mortality for type 2 diabetes.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim has two parts. The first, that R from Eq. (2b) is unstable, is convincingly supported by Table 1 and by the bootstrap replicates. The second, that Δm from Eq. (2a) yields sensible results, rests entirely on Figure 1 and on a smoothing assumption that is not validated. Eq. (2a) and Eq. (2b) share the term A = i − ∂p/(1−p), and both divide by prevalence p. Eq. (2b) has the additional denominator m − A, which can pass through zero and makes R especially fragile; but Eq. (2a) is not automatically safe. At young ages p is small, so relative errors in the fitted logit-linear prevalence or in the two-survey time-derivative ∂p are amplified by 1/p. The manuscript gives no numeric table for Δm, no coverage check for the bootstrap intervals, and no comparison with external estimates of diabetes excess mortality. The visual stability of Δm in Figure 1 could therefore be a property of the 2009–2015 logit-linear fit combined with the 2012 incidence, rather than an intrinsic advantage of Δm. Because the stated aim is to extend reliable estimation below age 50, the missing validation of Δm is load-bearing.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies estimation of excess mortality in chronic diseases from aggregated age-specific prevalence and incidence data, using German statutory health insurance claims data on diabetes (roughly 70 million people, 2009–2015) as an example. Building on prior work relating the temporal change of age-specific prevalence to incidence and mortality, the author applies two approaches: a direct method using Equations (2a) and (2b), and an indirect method that solves the prevalence PDE for candidate mortality rate ratios R and compares the computed 2015 prevalence with observed values. The main empirical finding is that direct estimates of the mortality rate ratio R from Equation (2b) are numerically unstable and implausible below age 55, including negative values in Table 1, whereas estimates of the mortality rate difference Δm from Equation (2a) appear sensible in Figure 1. The paper concludes that future work should estimate Δm rather than R.","tokens_in":4220,"tokens_out":2697,"duration_ms":31419,"significance":"If the claims are substantiated, the paper would provide a practical recommendation for epidemiological surveillance based on claims data and extend excess-mortality estimation to ages below 50, which is currently considered unreliable. The paper has clear strengths: it uses a large real-world dataset, implements a bootstrap procedure with 2000 replicates, and gives a transparent table showing the instability of direct R estimates. The negative and extreme R values in Table 1 convincingly demonstrate that Equation (2b) is numerically fragile in this application. However, the central recommendation to replace R with Δm rests on the visual stability of Δm in a single example, with no numerical table, no uncertainty diagnostics, no sensitivity analysis for the smoothing assumptions, and no external validation. The paper is therefore more convincing as a cautionary demonstration of instability than as a validated recommendation for a new estimand.","major_comments":[{"comment":"The central claim that Δm yields sensible results is not supported by quantitative evidence. Figure 1 shows medians and percentile bands, but no numerical table of Δm is provided, no bootstrap coverage or calibration check is reported, and no comparison is made with external estimates of diabetes excess mortality from the literature. Because Equation (2a) shares the term A = i − ∂p/(1−p) with Equation (2b) and also divides by p, the apparent stability of Δm may be a consequence of the same logit-linear two-survey smoothing and the constant-2012-incidence assumption rather than an intrinsic advantage of Δm. The reader cannot assess whether the authors' recommendation is robust without these additional analyses.","section":"Direct methods, Eq. (2a)"},{"comment":"The temporal derivative ∂p is estimated from only two prevalence surveys (2009 and 2015) after fitting logit-linear models, and the incidence i is taken from 2012 and treated as fixed over the whole period. This approximation is load-bearing for both Equations (2a) and (2b), yet no sensitivity analysis is presented. For example, the author does not test alternative smoothing functions, alternative knots, or a time-varying incidence rate, nor does the manuscript report the fitted coefficients or residuals of the logit-linear prevalence models. Without such checks, the observed contrast between unstable R and stable Δm could be an artifact of the smoothing, which would undermine the paper's main conclusion.","section":"Direct methods, paragraph 2"},{"comment":"The L-curve in Figure 2 is presented as evidence that the inverse problem is ill-posed, and this is used to support the recommendation to prefer Δm. However, the indirect procedure itself has several unexamined assumptions: the choice of knots at ages 30, 60, and 90, the piecewise linear interpolation of log R, and the extrapolation below 30 and above 90. No uncertainty quantification is provided for the indirect estimates, and the L-curve's shape is a heuristic diagnostic, not a validation of the alternative estimand. The connection between the ill-posedness of the indirect inversion and the instability of direct estimates of R would need a more formal argument or at least a simulation study to be load-bearing.","section":"Indirect method, Figure 2"},{"comment":"Table 1 shows extreme median R values (e.g., 850 at ages 15–19, −3409 at ages 25–29) and very wide percentile intervals, which clearly indicate instability. However, the bootstrap procedure assumes a binomial error model for prevalence and a Poisson error model for incidence, and the manuscript does not state how the resampling handles the fact that the input data are already smoothed regressions rather than raw counts. If the resampling is applied to the fitted values rather than to the raw data, the uncertainty may be underestimated; if applied to raw data, the regression fits should be repeated within each replicate. The description should be clarified, and the sensitivity of the conclusions to this choice should be discussed.","section":"Results, Table 1"}],"minor_comments":[{"comment":"The caption contains an incomplete citation: the text ends with 'L-curve []' and the reference placeholder is not filled.","section":"Figure 2 caption"},{"comment":"There are typographical artifacts in the text, such as 'preval ence' and 'incid ence', which should be corrected.","section":"Title and Introduction"},{"comment":"Reference [Pol01] is cited in the text as '[Pol]' rather than '[Pol01]'; please make the citation consistent.","section":"References"},{"comment":"The phrase 'sampling uncertainties' is used for variability in claims data, but claims data are not necessarily a random sample; the binomial and Poisson error models are assumptions that should be explicitly justified.","section":"Direct methods"},{"comment":"The manuscript does not report the number of candidate vectors x used to produce Figure 2, nor how the candidate values were generated; this information is needed for reproducibility.","section":"Indirect method"},{"comment":"The conclusion that 'estimates for the rate difference Δm yield sensible results' is stated as a finding, but it is based on one dataset and one figure; adding a numeric summary or a discussion of limitations would make the conclusion more measured.","section":"Conclusion"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is very short and overlaps substantially with the author's previous work, including several self-citations. The editor may wish to consider whether the incremental contribution -- mainly the empirical demonstration that R is unstable in this German example -- is sufficient for a standalone paper, or whether the Δm recommendation needs a much more careful validation study before publication. The lack of any external comparison or simulation check for Δm is the main reason I do not recommend acceptance at this stage."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a short methods note that does one thing well: Table 1 makes a convincing case that the mortality rate ratio R, estimated from Equations (2a) and (2b) applied to German diabetes claims data, is numerically unstable below age 55, sometimes turning negative. The bootstrap step is a reasonable way to propagate sampling uncertainty, and the paper is honest in showing the implausible R values rather than sweeping them aside. The L-curve from the indirect method is a nice illustration of why the inverse problem is ill-posed, though it is more suggestive than definitive.\n\nWhat is actually new is modest. The instability of R at young ages was already reported in the cited prior simulation study [Bri19], and Equations (1)–(2) come from [Bri14, Bri16]. The new elements are the real-world application to about 70 million insured Germans and the L-curve picture. That is an incremental extension, not a new technique.\n\nThe soft spot is exactly the one flagged in the stress-test note: the recommendation to prefer the rate difference Δm rests on Figure 1 alone. Δm and R share the term A = i − ∂p/(1−p), and both divide by prevalence p, so relative errors in the fitted logit-linear prevalence or in the two-survey time derivative ∂p are amplified at young ages where p is small. R has the extra denominator m − A, which can cross zero and makes R especially fragile, so it is plausible that Δm is better behaved. But plausible is not the same as shown. The paper gives no numeric table for Δm, no coverage checks for the bootstrap intervals, and no comparison with external estimates of diabetes excess mortality. The visual stability of Δm could be an artifact of the smooth 2009–2015 logit-linear fit combined with the constant 2012 incidence. The logit-linear assumption and the time-constant incidence are both load-bearing and unvalidated.\n\nNone of this is fatal. The instability result stands. But the central recommendation—that surveillance should prefer Δm—needs more support. A serious referee should ask for sensitivity analyses (e.g., alternative smoothing of prevalence), numeric Δm results with uncertainty, and external validation.\n\nWho gets value from this paper: epidemiologists working with aggregated prevalence/incidence data for chronic disease surveillance. It deserves peer review despite its gaps, because the instability of R is a real finding and the paper honestly raises an important practical question about which estimand to report.\n\nMy recommendation: send it to review, but expect or require substantial revision before acceptance. The missing validation of Δm is addressable, not a fundamental flaw.","headline":"Useful short methods note: convincingly shows R is unstable in claims data, but the case for preferring Δm needs validation.","tokens_in":4662,"tokens_out":1637,"would_cite":false,"duration_ms":18608,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"When prevalence is low, the mortality rate ratio estimated from prevalence and incidence data becomes numerically unstable, so excess mortality should be reported as a rate difference instead.","keywords":["excess mortality","mortality rate ratio","rate difference","prevalence","incidence","illness-death model","inverse problem","claims data"],"falsifier":"Simulate the illness-death model with known age-specific mortality rates and a known mortality rate ratio $R$ that is constant across age, generate prevalence and incidence from the model, then feed them into Equations (2a) and (2b); if the ratio estimate remains stable below age 55 in that controlled setting, the German-data instability comes from the smoothing or sampling, not from the inverse problem itself.","tokens_in":3655,"feed_emoji":"🩺","tokens_out":11458,"duration_ms":108791,"temperature":0.7,"pith_summary":"The paper investigates why estimates of excess mortality in chronic diseases, computed from age-specific prevalence and incidence data, become unreliable at younger ages. Working with German claims data covering about 70 million people and diabetes as the example, it shows that the mortality rate ratio $R=m_1/m_0$ can take impossible negative values below age 55, while the mortality rate difference $\\Delta m=m_1-m_0$ remains plausible. The paper points to the ill-posedness of the underlying inverse problem as the likely reason and recommends that excess mortality be reported as the rate difference. The practical payoff would be extending usable excess-mortality estimates to ages below 50, where claims data are increasingly available.","feed_headline":"Mortality ratios fail below age 55; estimate rate differences instead","feed_subtitle":"German diabetes claims data show the ratio is unstable, while the rate difference stays sensible.","key_machinery":"The central object is the illness-death equation for age-specific prevalence $p$, written as $\\partial p=(1-p)(i-p\\,\\Delta m)$ in terms of the rate difference and equivalently in terms of the ratio $R$. Inverting that equation gives the direct estimators $\\Delta m=\\{i-\\partial p/(1-p)\\}/p$ and $R=1+(1/p)\\{i-\\partial p/(1-p)\\}/\\{m-i+\\partial p/(1-p)\\}$. The argument turns on the divisor in the $R$ formula: when the combination $m-i+\\partial p/(1-p)$ is near zero, small sampling variation in the inputs is amplified, and the outer factor $1/p$ amplifies it further when prevalence is low. The L-curve appearing in the indirect estimation is used as evidence that the difficulty is an ill-posed inverse problem rather than a simple data flaw.","core_discovery":"The paper's central claim is that two mathematically equivalent ways of quantifying excess mortality from aggregated data behave very differently in practice. The direct formula for the mortality rate ratio $R$ produces implausible, even negative, estimates for ages below 55 when applied to German diabetes claims data, whereas the direct formula for the rate difference $\\Delta m$ gives stable, sensible values. Because a ratio of two positive mortality rates cannot be negative, the negative values signal numerical instability in the inversion, and the paper points to an L-curve typical of ill-posed inverse problems as evidence for that diagnosis. The recommendation is to estimate and report $\\Delta m$ rather than $R$ when working from prevalence and incidence data.","pith_inferences":["Not in the paper: the practical trigger for instability is probably low prevalence, because the factor $1/p$ in Equation (2b) amplifies small input errors when $p$ is small; a testable prediction is that chronic conditions with substantial prevalence at young ages will yield stable ratio estimates.","Not in the paper: a regularized or constrained version of the direct ratio estimator, for example one that forces $R\\ge 1$ or smooths the denominator before division, might recover a usable ratio; the paper itself stops at recommending $\\Delta m$.","Not in the paper: the boundary where $R$ becomes unreliable could be mapped by simulation with known mortality rates, giving applied researchers an explicit age-and-prevalence threshold for when a ratio may safely be quoted."],"forward_implications":["Epidemiological surveillance based on claims data can report excess mortality down to younger ages if it uses the rate difference instead of the rate ratio.","Studies that currently report only the mortality rate ratio should add the rate difference, because the ratio can be unstable in low-prevalence age groups.","The indirect parameter-fitting approach, judged by its L-curve, offers a way to check whether a direct estimate is trustworthy.","The age range of usable excess-mortality estimates can be extended below the previous practical limit of 50 years, at least for rate differences."],"supporting_citations":[{"why":"Derives the partial differential equation for prevalence change that the direct and indirect estimators invert.","marker":"[Bri14, Bri16]"},{"why":"Proposes the excess-mortality estimation method and reports the growing bias at younger ages that this paper investigates.","marker":"[Bri19]"},{"why":"Establishes the illness-death inverse problem and its ill-posed character, the explanation offered for the instability.","marker":"[Bri18]"},{"why":"Supplies the age-specific diabetes prevalence and incidence data used in the example.","marker":"[Gof17]"},{"why":"Supplies the general mortality rates used for the ratio estimates and in the indirect partial differential equation solution.","marker":"[Fed19]"},{"why":"Introduces the L-curve method that the paper uses as evidence of ill-posedness.","marker":"[Han99]"},{"why":"Motivates the Gompertz-Makeham form of the mortality rates used to parameterize the indirect estimator.","marker":"[Mis13]"}],"fun_headline_variants":["Ratio unstable under 55; estimate rate difference instead","Mortality ratio fails young; use rate difference for stable excess mortality","For chronic disease excess mortality ratio fails under 55; use difference","German diabetes data: ratio unstable below 55; difference stable"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The direct estimates assume that logit-transformed prevalence surveys from 2009 and 2015 accurately describe the true temporal change in prevalence over the whole period, with the 2012 incidence curve treated as fixed.","fun_headline_variants_meta":{"raw":{"variants":["Ratio unstable under 55; estimate rate difference instead","Mortality ratio fails young; use rate difference for stable excess mortality","For chronic disease excess mortality ratio fails under 55; use difference","German diabetes data: ratio unstable below 55; difference stable"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000776,"raw_usage":{"total_tokens":3333,"prompt_tokens":746,"completion_tokens":2587,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":362,"completion_tokens_details":{"reasoning_tokens":2516}},"tokens_in":362,"tokens_out":2587,"duration_ms":18718,"temperature":1.0,"reasoning_tokens":2516,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:52:52.811698+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate the illness-death model with known age-specific mortality rates and a known mortality rate ratio $R$ that is constant across age, generate prevalence and incidence from the model, then feed them into Equations (2a) and (2b); if the ratio estimate remains stable below age 55 in that controlled setting, the German-data instability comes from the smoothing or sampling, not from the inverse problem itself.","supporting_citations":[],"review_version":1}