REVIEW 4 major objections 4 minor 22 references
Cancer incidence estimation from mortality data: a validation study within a population-based cancer registry
T0 review · 4 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper claims that the IMR method—estimating cancer incidence from mortality data and an incidence-to-mortality ratio—is a valid tool, validated against a ten-year observed series from a population-based registry, and that a…
desk verdict Genuine multi-year validation of the IMR method, but the reported accuracy is inflated by selecting the best scenario on the same data used for evaluation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the incidence-to-mortality ratio (IMR)—the number of new cancer cases per cancer death in a given year, age group, and site—combined with a two-stage estimation pipeline. First, an age-period-cohort model (the NORDPRED method) predicts the number of cancer deaths in the target year from a mortality time series; second, a generalized linear mixed model with a Poisson distribution, age splines, and a second-degree polynomial for year estimates the IMR, which is then projected under five scenarios: constant at the last value, constant at the mean of the last three years, constant at the mean of the last five years, linear trend, and quadratic trend. An iterative back-casting procedure reproduces the real-world delay in data availability and generates expected incidence for each year 2004–2013, and the mean absolute percentage error (MAPE) between expected and observed cases serves as the goodness-of-fit indicator that selects the best scenario per site. This design lets the method's assumptions be tested directly against observed registry data.
What would settle it
Run the same iterative back-casting procedure on another high-quality registry with a long historical series, for a site whose IMR is known to have changed in a step-like manner (e.g., prostate cancer around the introduction of screening). If the best-scenario estimate deviates from observed cases by more than the paper's reported site-level range of roughly −13% to +14%, or if the MAPE indicator selects a scenario while year-to-year errors exceed 25% for non-screening sites, the paper's validity claim would be contradicted. A simpler observation: a single site whose true IMR trajectory is, say, an inverted U over the projection window—captured by none of the five scenarios—would produce systematic bias that the GOF indicator cannot detect.
Extended reading notes
Core claim
The central claim is that the incidence-to-mortality ratio (IMR) method—which estimates new cancer cases from death counts and a modelled ratio of incidence to mortality—is a valid tool for estimating cancer incidence, and that this validity can be demonstrated by comparing estimated cases with observed cases over a continuous ten-year historical series. The paper reports that under the best scenario for each site, the relative difference between estimated and observed total cancer cases was −0.3% for men (23,126 estimated vs 23,197 observed) and −0.5% for women (16,574 vs 16,651), with site-level relative deviations ranging from −9.1% (rectal cancer in men) to +11.6% (prostate cancer in men), and from −12.6% (lung cancer in women) to +14.3% (ovarian cancer). The constant-IMR assumption gave the best fit for colon, rectal, lung, bladder, and stomach cancers in men and colon, rectal, breast, and corpus uteri in women; the linear assumption was best for prostate cancer in men and lung, ovary, and other sites in women. The method's weakest performance is for prostate and female breast cancer, where screening-induced sudden changes in the IMR produce errors up to 136% in a single year, and the paper explicitly recommends new strategies for such sites.
Load-bearing premise
The load-bearing premise is that the five assumed shapes for the incidence-to-mortality ratio—constant, mean of the last three or five years, linear, or quadratic—cover the true way that ratio changes over time for each cancer site, so that if the real ratio follows a different pattern, the method's estimates will be biased no matter how well the mortality model performs.
Editorial extensions
If this is right
- For regions without a cancer registry but with reliable mortality and population data, the IMR method can produce total cancer incidence estimates within about half a percent over a decade, with most site-specific estimates within 10%.
- The MAPE goodness-of-fit indicator gives future users an objective, data-driven way to choose among IMR trend assumptions instead of relying on subjective judgement.
- Using the mean of the last three to five IMR values instead of the single most recent value generally improves the fit, while quadratic extrapolation performs worst and should be avoided.
- Cancers affected by screening programmes—prostate, breast, ovary, and female lung—are the method's known weak points, and estimates for these sites should be interpreted with caution or modelled separately.
- The method's validity is conditional on high-quality input data (incidence, mortality, and population), so the same accuracy cannot be assumed in settings with poorer data.
Reading between the lines
- The validation is retrospective: the best scenario is chosen after seeing the observed data, so the reported accuracy is likely an upper bound on what a prospective user would achieve when the true IMR trajectory is unknown.
- The near-perfect ten-year totals arise partly because yearly over- and under-estimates cancel out; decision-makers relying on single-year estimates for screening-affected sites should expect much larger errors than the decade totals suggest.
- A natural extension would be to apply the same back-casting validation to registries in other regions with different survival and screening patterns, which would test whether the method's accuracy transfers or is specific to a particular health system.
- The method's failure on screening-affected sites suggests a pragmatic hybrid: use the IMR method for sites with stable IMR, and switch to incidence-based projection or screening-adjusted models for sites with known early-detection programmes.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper validates the incidence-to-mortality ratio (IMR) method for estimating cancer incidence from mortality data, using a 10-year historical series (2004–2013) from the Granada population-based cancer registry. Mortality data from 1982–2010 are used to derive mortality projections via NORDPRED-style APC models, and IMR projections are obtained from Bayesian GLMMs for each of five assumed IMR scenarios (constant last value, mean of last 3/5 years, linear, quadratic). A goodness-of-fit indicator based on the mean absolute percentage error (MAPE) is used to select the best scenario per site and sex. The paper reports that, for the best scenario per site, the overall relative deviation is −0.3% in men and −0.5% in women, with MAPE values of 6.34% and 3.85%, respectively, and concludes that the IMR method is a valid tool for cancer incidence estimation when data quality is high.
Significance. This is the first validation of the IMR method over a multi-year time series against a high-quality population-based registry, rather than a single-year comparison. The rolling iterative procedure that mimics real-world prediction delays is a genuine strength, as is the use of a Bayesian framework and the explicit comparison of multiple IMR assumptions. However, the quantitative evidence for the 'valid tool' claim is weakened by the fact that the best scenario for each site is selected on exactly the same validation period used to compute the reported MAPE and deviations, making those error metrics post-selection estimates. If the authors add a temporally separated or cross-validated assessment of the scenario-selection procedure, the result would be considerably more convincing; as it stands, the central claim is plausible but not yet supported at the level the paper asserts.
major comments (4)
- [Goodness-of-fit assessment; Table 1] The reported MAPE values and relative deviations are optimistically biased because the best scenario for each site is chosen by minimizing MAPE over the same 2004–2013 period used to report those errors. With five scenarios and ten annual observations per site, the minimum of five MAPE estimates is not an unbiased estimator of the expected error in a new application. The paper should provide an unbiased evaluation, for example by holding out the last two to three years for assessment, by using a nested or rolling-origin cross-validation scheme, or by reporting the MAPE averaged over all scenarios in addition to the best-scenario MAPE.
- [Assumptions regarding the mortality-to-incidence ratios] The five assumed IMR scenarios (C1, C3, C5, L, Q) are asserted to span plausible trajectories, but the paper's own results show that for prostate cancer in men and for breast, ovary, and female lung cancer, none of the scenarios provides a satisfactory fit, with annual deviations up to 136% for prostate cancer in 2013 (Section 'Results', Additional file 1). This means the conclusion 'valid tool' is conditional on the assumption that the true IMR trajectory belongs to the family of five scenarios, and that assumption is already violated for several sites studied here. The paper should state this limitation more prominently and avoid a blanket validity claim for all cancer sites.
- [Validity assessment; Table 2] The small overall relative deviations (−0.3% men, −0.5% women) are largely driven by cancellation of offsetting annual errors, not by consistent accuracy. Table 2 shows annual relative deviations for all sites combined ranging from −7.7% to +24.7% in men, with the largest error in the final year (2013). The conclusion that the method provides 'good reliability' should be based on a measure that does not hide yearly inaccuracy, such as the distribution of annual absolute errors or prediction intervals for each year, rather than the cumulative relative deviation.
- [Goodness-of-fit assessment; Eq. (3)] The MAPE-based GOF indicator is promoted as an objective way to select the best IMR scenario, but the paper never assesses the selection performance of this indicator. Since the indicator is fitted and evaluated on the same data, its ability to identify the correct scenario in a prospective setting is unknown. A validation study that applies the GOF rule to the first part of the series and then checks accuracy on the last part is needed to support the claim that the GOF indicator 'can help select the best assumption' (Discussion).
minor comments (4)
- [Methods, Eq. (2)] Equation (2) is typeset as '100: Expected − Observed / Observed', which is ambiguous; it should be written as 100 × (Expected − Observed)/Observed, and the direction of the sign (positive for overestimation) is stated only in the text, not with the equation.
- [Methods, Eq. (1)] Equation (1) contains garbled notation ('D /C1 p + Pp + Cc /C0/C1 5') that appears to be an OCR artifact; the model formula should be presented in standard APC notation so that the drift, period, and cohort terms are unambiguous.
- [Results, Table 1] The row for 'All sites' does not report a 'Best scenario' column value, which is confusing because the all-sites estimates must be based on some combination of site-specific scenarios; please clarify whether the all-sites MAPE is computed from the sum of best-scenario sites or from a separately chosen overall scenario.
- [Discussion] The Discussion mentions that the method is less precise when case numbers are small or IMR changes suddenly, and this is a welcome limitation statement. However, the same paragraph also asserts that the method is 'valid for most cancer sites' without quantifying 'most' or listing which sites are excluded; a more precise enumeration would help the reader assess the scope of validity.
Circularity Check
No significant circularity: the incidence estimates are derived from mortality and IMR data and benchmarked against external registry-observed cases.
full rationale
The paper's derivation chain is not circular. Estimated incidence cases are produced from (i) NORDPRED mortality predictions (Eq. 1) and (ii) GLMM-based IMR projections under five explicit scenarios, then compared with observed registry cases via Eqs. (2) and (3). The observed 2004-2013 incidence is used only as the benchmark, not as an input to the estimation model: for each target year, the IMR model is fitted to data ending five years before the target year and mortality data ending two to three years before, reproducing the real-world data delay. The GOF indicator (Eq. 3) selects the best scenario on the validation window, which makes the reported best-scenario MAPE a within-sample selection-minimum and therefore somewhat optimistic; however, this is a methodological caveat about post-selection, not a definitional circularity, because the estimated case counts are not derived from the observed case counts by construction and the central conclusion rests on an external comparison against a population-based registry. Self-citations (e.g., reference 7) are used to describe prior applications of the method, not to justify the validity claim, and the external benchmark is independent. No step in the derivation reduces to its own inputs.
Assumptions & free parameters
free parameters (4)
- Age spline knots in IMR model =
10th percentile, first tertile, second tertile, 90th percentile of the mortality pool
- Degree of year polynomial in IMR model =
2 (quadratic)
- Time windows for projections =
20-year mortality series, 15-year IMR series, 3-year and 5-year projections
- Best IMR scenario per site =
C1, C3, C5, or L depending on site (see Table 1)
assumptions (5)
- domain assumption Granada registry and official mortality data are complete and correctly coded.
- ad hoc to paper The five IMR trend scenarios span the plausible true behavior of the IMR for each site.
- domain assumption NORDPRED age-period-cohort model with a power link accurately forecasts mortality three years ahead.
- domain assumption Incident cases follow a Poisson distribution with deaths as the offset in the GLMM.
- ad hoc to paper Bayesian priors used in OpenBUGS are adequate, though not reported.
Cite this review
Pith. "Pith review of Cancer incidence estimation from mortality data: a validation study within a population-based cancer registry." pith.science (2026). https://pith.science/paper/3H22ZMKD
@misc{pith2026241112784,
author = {Pith},
title = {Pith review of: Cancer incidence estimation from mortality data: a validation study within a population-based cancer registry},
year = {2026},
howpublished = {\url{https://pith.science/paper/3H22ZMKD}},
note = {Machine review of arXiv:2411.12784}
}
read the original abstract
We assessed the validity of one of the most frequently used methods to estimate cancer incidence, on the basis of cancer mortality data and the incidence-to-mortality ratio IMR, the IMR method. Using the previous 15 year cancer mortality time series, we derived the expected yearly number of cancer cases in the period 2004 to 2013 for six cancer sites for each sex. Generalized linear mixed models, including a polynomial function for the year of death and smoothing splines for age, were adjusted. Models were fitted under a Bayesian framework based on Markov chain Monte Carlo methods. The IMR method was applied to five scenarios reflecting different assumptions regarding the behavior of the IMR. We compared incident cases estimated with the IMR method to observed cases diagnosed in 2004 to 2013 in Granada. A goodness-of-fit GOF indicator was formulated to determine the best estimation scenario. The relative differences between the observed and predicted numbers of cancer cases were less than 10 percent for most cancer sites. The constant assumption for the IMR trend provided the best GOF for colon, rectal, lung, bladder, and stomach cancers in men and colon, rectum, breast, and corpus uteri in women.
Reference graph
Works this paper leans on
-
[1]
dos Santos Silva I. Cancer epidemiology: principles and methods, Chapter 4: Measures of occurrence of disease and other health related events. Lyon: International Agency for Research on Cancer; 1999. p. 60 – 3. Available from: http://publications.iarc.fr/Non-Series-Publications/Other-Non-Series-Publica tions/Cancer-Epidemiology-Principles-And-Methods-1999...
work page 1999
-
[2]
Evaluation of data quality in the cancer registry: principles and methods
Bray F, Parkin DM. Evaluation of data quality in the cancer registry: principles and methods. Part I: comparability, validity and timeliness. Eur J Cancer. 2009;45:747– 55 Available from: http://linkinghub.elsevier.com/ retrieve/pii/S0959804908009209. [cited 2018 May 28]
work page 2009
-
[3]
Evaluation of data quality in the cancer registry: principles and methods Part II
Parkin DM, Bray F. Evaluation of data quality in the cancer registry: principles and methods Part II. Completeness. Eur J Cancer 2009;45:756 – 764. Available from: http://www.ncbi.nlm.nih.gov/pubmed/19128954. [cited 2018 May 28]
-
[4]
Cancer incidence in five continents, Vol
Forman D, Bray F, Brewster DH, Gombe Mbalawa C, Kohler B, Piñeros M, Znaor A, Zanetti R and Ferlay J E. Cancer incidence in five continents, Vol. XI. Lyon Int. Agency Res. Cancer. 2017. Available from: http://ci5.iarc.fr. [cited 2018 Jun 13]
work page 2017
-
[5]
ECIS - European Cancer Information System
Joint Research Center-European Commission. ECIS - European Cancer Information System. 2018. Available from: https://ecis.jrc.ec.europa.eu/. [cited 2018 Jun 13]
work page 2018
-
[6]
Cancer today (powered by GLOBOCAN 2018)
Ferlay J, Ervik M, Lam F, Colombet M, Mery L, Piñeros M, Znaor A, Soerjomataram I BF. Cancer today (powered by GLOBOCAN 2018). IARC CancerBase No. 15. 2018. Available from: http://publications.iarc.fr/577. [cited 2018 Dec 23]
work page 2018
-
[7]
Cancer incidence in Spain, 2015
Galceran J, Ameijide A, Carulla M, Mateos A, Quirós JR, Rojas D, et al. Cancer incidence in Spain, 2015. Clin Transl Oncol. 2017;19:799 – 825 Available from: http://link.springer.com/10.1007/s12094-016-1607-9. [cited 2017 Jan 19]
-
[8]
Cancer incidence and mortality patterns in Europe: estimates for 40 countries in 2012
Ferlay J, Steliarova-Foucher E, Lortet-Tieulent J, Rosso S, Coebergh JWW, Comber H, et al. Cancer incidence and mortality patterns in Europe: estimates for 40 countries in 2012. Eur J Cancer. 2013;49:1374 – 403 Available from: http://www.ncbi.nlm.nih.gov/pubmed/23485231. [cited 2016 Nov 17]
Show all 22 references
-
[9]
An assessment of GLOBOCAN methods for deriving national estimates of cancer incidence
Antoni S, Soerjomataram I, Møller B, Bray F, Ferlay J. An assessment of GLOBOCAN methods for deriving national estimates of cancer incidence. Bull World Health Organ. 2016;94:174 – 84. Available from: https://www. ncbi.nlm.nih.gov/pmc/articles/PMC4773935/ . https://doi.org/10....
2016 doi
-
[10]
Available from: https://www.registrocancergrana da.es/
Granada Cancer Registry. Available from: https://www.registrocancergrana da.es/. [cited 2019 May 11]
2019
-
[11]
Situación del cáncer en España: incidencia
López-Abente G, Pollán M, Aragonés N, Pérez Gómez B, Hernández Barrera V, Lope V, et al. Situación del cáncer en España: incidencia. An Sist Sanit Navar. 2004;27:165 – 73 Available from: http://scielo.isciii.es/scielo.php?script= sci_arttext&pid=S1137-66272004000300001&lng=en&...
2004
-
[12]
Estimación de la incidencia de cáncer en España: período 1993 – 1996
Moreno V, González JR, Soler M, Bosch FX, Kogevinas M, Borràs JM. Estimación de la incidencia de cáncer en España: período 1993 – 1996. Gac Sanit. 2001;15:380 – 8 Available from: http://linkinghub.elsevier.com/retrieve/ pii/S0213911101715919
1993
-
[13]
Continuous Register Statistics
Spanish National Institute of Statistics. Continuous Register Statistics. Available from: https://ine.es/dyngs/INEbase/es/operacion.htm?c=Esta distica_C&cid=1254736176951&menu=resultados&idp=1254735572981. [cited 2018 Jun 13]
2018
-
[14]
Government of Spain
Ministry of Health. Government of Spain. Mortality by cause of death. Available from: https://pestadistico.inteligenciadegestion.mscbs.es/ publicoSNS/S/mortalidad-por-causa-de-muerte. [cited 2018 Jun 13]
2018
-
[15]
ICD-10: International Statistical Classification of diseases and related health problems: 10th revision
World Health Organization (WHO). ICD-10: International Statistical Classification of diseases and related health problems: 10th revision. 1990
1990
-
[16]
Estimates of cancer incidence and mortality in Europe in 2008
Ferlay J, Parkin DM, Steliarova-Foucher E. Estimates of cancer incidence and mortality in Europe in 2008. Eur J Cancer. 2010;46(4):765 – 81. Available from: https://linkinghub.elsevier.com/retrieve/pii/S0959804909009265. https://doi. org/10.1016/j.ejca.2009.12.014
2008 doi
-
[17]
Prediction of cancer incidence in the Nordic countries: empirical comparison of different approaches
Møller B, Fekjaer H, Hakulinen T, Sigvaldason H, Storm HH, Talbäck M, et al. Prediction of cancer incidence in the Nordic countries: empirical comparison of different approaches. Stat Med. 2003;22:2751 – 66 Available from: http://www.ncbi.nlm.nih.gov/pubmed/12939784
2003
-
[18]
Making BUGS open
Thomas A, O ’Hara B, Ligges U, Sturtz S. Making BUGS open. R News. 2006;6: 12– 7 Available from: https://cran.r-project.org/doc/Rnews/. Redondo-Sánchez et al. Population Health Metrics (2021) 19:18 Page 9 of 10
2021
-
[19]
The BUGS project: evolution, critique and future directions
Lunn D, Spiegelhalter D, Thomas A, Best N. The BUGS project: evolution, critique and future directions. Stat Med. 2009;28(25):3049 – 67. Available from: http://doi.wiley.com/10.1002/sim.3680
2009 doi
-
[20]
Available from: http://openbugs.net/w/FrontPage
OpenBUGS. Available from: http://openbugs.net/w/FrontPage. [cited 2019 May 11]
2019
-
[21]
Simple versus complex models: evaluation, accuracy, and combining
Ahlburg DA. Simple versus complex models: evaluation, accuracy, and combining. Math Popul Stud. 1995;5:281 – 90 Available from: http://www.ta ndfonline.com/doi/abs/10.1080/08898489509525406. [cited 2019 May 11]
1995 doi
-
[22]
Cancer screening in Spain
Ascunce N, Salas D, Zubizarreta R, Almazan R, Ibanez J, Ederra M, et al. Cancer screening in Spain. Ann Oncol. 2010;21:iii43 – 51 Available from: http://www.ncbi.nlm.nih.gov/pubmed/20427360. [cited 2018 Dec 7]. Publisher’sN o t e Springer Nature remains neutral with regard to ...
2021
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.