Pith. sign in

REVIEW 4 major objections 4 minor 22 references

Cancer incidence estimation from mortality data: a validation study within a population-based cancer registry

T0 review · 4 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read The paper claims that the IMR method—estimating cancer incidence from mortality data and an incidence-to-mortality ratio—is a valid tool, validated against a ten-year observed series from a population-based registry, and that a…

desk verdict Genuine multi-year validation of the IMR method, but the reported accuracy is inflated by selecting the best scenario on the same data used for evaluation. read the letter →

arxiv 2411.12784 v1 pith:3H22ZMKD submitted 2024-11-19 q-bio.QM

classification q-bio.QM
keywords CancerincidenceIncidence-to-mortalityratioEstimationvalidationGoodness-of-fitMeanabsolutepercentageerrorPopulation-basedregistryBayesiangeneralizedlinearmixedmodelNORDPRED
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that a widely used shortcut for estimating cancer incidence—the incidence-to-mortality ratio (IMR) method, which derives new cancer cases from death records and a modelled ratio of incidence to mortality—produces trustworthy estimates when checked against a decade of real registry data. Using a 15-year mortality series and five different assumptions about how the IMR behaves over time, the authors back-cast yearly incidence for 2004–2013 in Granada, Spain, and compare each estimate to the cases the local population-based registry actually recorded. For most of the 14 site–sex combinations, the best-fitting scenario stays within 10% of observed cases per year, and the overall ten-year totals differ by only −0.3% in men and −0.5% in women. The paper also proposes a goodness-of-fit score based on the mean absolute percentage error (MAPE) to choose among IMR scenarios on statistical grounds rather than subjective judgement. A sympathetic reader would take the conclusion as: for high-quality mortality and population data, the IMR method is a valid tool for total and site-specific incidence estimation, with the caveat that cancers affected by screening programs (prostate, breast, ovary, female lung) need separate handling.

What carries the argument

The machinery is the incidence-to-mortality ratio (IMR)—the number of new cancer cases per cancer death in a given year, age group, and site—combined with a two-stage estimation pipeline. First, an age-period-cohort model (the NORDPRED method) predicts the number of cancer deaths in the target year from a mortality time series; second, a generalized linear mixed model with a Poisson distribution, age splines, and a second-degree polynomial for year estimates the IMR, which is then projected under five scenarios: constant at the last value, constant at the mean of the last three years, constant at the mean of the last five years, linear trend, and quadratic trend. An iterative back-casting procedure reproduces the real-world delay in data availability and generates expected incidence for each year 2004–2013, and the mean absolute percentage error (MAPE) between expected and observed cases serves as the goodness-of-fit indicator that selects the best scenario per site. This design lets the method's assumptions be tested directly against observed registry data.

What would settle it

Run the same iterative back-casting procedure on another high-quality registry with a long historical series, for a site whose IMR is known to have changed in a step-like manner (e.g., prostate cancer around the introduction of screening). If the best-scenario estimate deviates from observed cases by more than the paper's reported site-level range of roughly −13% to +14%, or if the MAPE indicator selects a scenario while year-to-year errors exceed 25% for non-screening sites, the paper's validity claim would be contradicted. A simpler observation: a single site whose true IMR trajectory is, say, an inverted U over the projection window—captured by none of the five scenarios—would produce systematic bias that the GOF indicator cannot detect.

Watch

Extended reading notes

Core claim

The central claim is that the incidence-to-mortality ratio (IMR) method—which estimates new cancer cases from death counts and a modelled ratio of incidence to mortality—is a valid tool for estimating cancer incidence, and that this validity can be demonstrated by comparing estimated cases with observed cases over a continuous ten-year historical series. The paper reports that under the best scenario for each site, the relative difference between estimated and observed total cancer cases was −0.3% for men (23,126 estimated vs 23,197 observed) and −0.5% for women (16,574 vs 16,651), with site-level relative deviations ranging from −9.1% (rectal cancer in men) to +11.6% (prostate cancer in men), and from −12.6% (lung cancer in women) to +14.3% (ovarian cancer). The constant-IMR assumption gave the best fit for colon, rectal, lung, bladder, and stomach cancers in men and colon, rectal, breast, and corpus uteri in women; the linear assumption was best for prostate cancer in men and lung, ovary, and other sites in women. The method's weakest performance is for prostate and female breast cancer, where screening-induced sudden changes in the IMR produce errors up to 136% in a single year, and the paper explicitly recommends new strategies for such sites.

Load-bearing premise

The load-bearing premise is that the five assumed shapes for the incidence-to-mortality ratio—constant, mean of the last three or five years, linear, or quadratic—cover the true way that ratio changes over time for each cancer site, so that if the real ratio follows a different pattern, the method's estimates will be biased no matter how well the mortality model performs.

Editorial extensions

If this is right

  • For regions without a cancer registry but with reliable mortality and population data, the IMR method can produce total cancer incidence estimates within about half a percent over a decade, with most site-specific estimates within 10%.
  • The MAPE goodness-of-fit indicator gives future users an objective, data-driven way to choose among IMR trend assumptions instead of relying on subjective judgement.
  • Using the mean of the last three to five IMR values instead of the single most recent value generally improves the fit, while quadratic extrapolation performs worst and should be avoided.
  • Cancers affected by screening programmes—prostate, breast, ovary, and female lung—are the method's known weak points, and estimates for these sites should be interpreted with caution or modelled separately.
  • The method's validity is conditional on high-quality input data (incidence, mortality, and population), so the same accuracy cannot be assumed in settings with poorer data.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The validation is retrospective: the best scenario is chosen after seeing the observed data, so the reported accuracy is likely an upper bound on what a prospective user would achieve when the true IMR trajectory is unknown.
  • The near-perfect ten-year totals arise partly because yearly over- and under-estimates cancel out; decision-makers relying on single-year estimates for screening-affected sites should expect much larger errors than the decade totals suggest.
  • A natural extension would be to apply the same back-casting validation to registries in other regions with different survival and screening patterns, which would test whether the method's accuracy transfers or is specific to a particular health system.
  • The method's failure on screening-affected sites suggests a pragmatic hybrid: use the IMR method for sites with stable IMR, and switch to incidence-based projection or screening-adjusted models for sites with known early-detection programmes.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper validates the incidence-to-mortality ratio (IMR) method for estimating cancer incidence from mortality data, using a 10-year historical series (2004–2013) from the Granada population-based cancer registry. Mortality data from 1982–2010 are used to derive mortality projections via NORDPRED-style APC models, and IMR projections are obtained from Bayesian GLMMs for each of five assumed IMR scenarios (constant last value, mean of last 3/5 years, linear, quadratic). A goodness-of-fit indicator based on the mean absolute percentage error (MAPE) is used to select the best scenario per site and sex. The paper reports that, for the best scenario per site, the overall relative deviation is −0.3% in men and −0.5% in women, with MAPE values of 6.34% and 3.85%, respectively, and concludes that the IMR method is a valid tool for cancer incidence estimation when data quality is high.

Significance. This is the first validation of the IMR method over a multi-year time series against a high-quality population-based registry, rather than a single-year comparison. The rolling iterative procedure that mimics real-world prediction delays is a genuine strength, as is the use of a Bayesian framework and the explicit comparison of multiple IMR assumptions. However, the quantitative evidence for the 'valid tool' claim is weakened by the fact that the best scenario for each site is selected on exactly the same validation period used to compute the reported MAPE and deviations, making those error metrics post-selection estimates. If the authors add a temporally separated or cross-validated assessment of the scenario-selection procedure, the result would be considerably more convincing; as it stands, the central claim is plausible but not yet supported at the level the paper asserts.

major comments (4)
  1. [Goodness-of-fit assessment; Table 1] The reported MAPE values and relative deviations are optimistically biased because the best scenario for each site is chosen by minimizing MAPE over the same 2004–2013 period used to report those errors. With five scenarios and ten annual observations per site, the minimum of five MAPE estimates is not an unbiased estimator of the expected error in a new application. The paper should provide an unbiased evaluation, for example by holding out the last two to three years for assessment, by using a nested or rolling-origin cross-validation scheme, or by reporting the MAPE averaged over all scenarios in addition to the best-scenario MAPE.
  2. [Assumptions regarding the mortality-to-incidence ratios] The five assumed IMR scenarios (C1, C3, C5, L, Q) are asserted to span plausible trajectories, but the paper's own results show that for prostate cancer in men and for breast, ovary, and female lung cancer, none of the scenarios provides a satisfactory fit, with annual deviations up to 136% for prostate cancer in 2013 (Section 'Results', Additional file 1). This means the conclusion 'valid tool' is conditional on the assumption that the true IMR trajectory belongs to the family of five scenarios, and that assumption is already violated for several sites studied here. The paper should state this limitation more prominently and avoid a blanket validity claim for all cancer sites.
  3. [Validity assessment; Table 2] The small overall relative deviations (−0.3% men, −0.5% women) are largely driven by cancellation of offsetting annual errors, not by consistent accuracy. Table 2 shows annual relative deviations for all sites combined ranging from −7.7% to +24.7% in men, with the largest error in the final year (2013). The conclusion that the method provides 'good reliability' should be based on a measure that does not hide yearly inaccuracy, such as the distribution of annual absolute errors or prediction intervals for each year, rather than the cumulative relative deviation.
  4. [Goodness-of-fit assessment; Eq. (3)] The MAPE-based GOF indicator is promoted as an objective way to select the best IMR scenario, but the paper never assesses the selection performance of this indicator. Since the indicator is fitted and evaluated on the same data, its ability to identify the correct scenario in a prospective setting is unknown. A validation study that applies the GOF rule to the first part of the series and then checks accuracy on the last part is needed to support the claim that the GOF indicator 'can help select the best assumption' (Discussion).
minor comments (4)
  1. [Methods, Eq. (2)] Equation (2) is typeset as '100: Expected − Observed / Observed', which is ambiguous; it should be written as 100 × (Expected − Observed)/Observed, and the direction of the sign (positive for overestimation) is stated only in the text, not with the equation.
  2. [Methods, Eq. (1)] Equation (1) contains garbled notation ('D /C1 p + Pp + Cc /C0/C1 5') that appears to be an OCR artifact; the model formula should be presented in standard APC notation so that the drift, period, and cohort terms are unambiguous.
  3. [Results, Table 1] The row for 'All sites' does not report a 'Best scenario' column value, which is confusing because the all-sites estimates must be based on some combination of site-specific scenarios; please clarify whether the all-sites MAPE is computed from the sum of best-scenario sites or from a separately chosen overall scenario.
  4. [Discussion] The Discussion mentions that the method is less precise when case numbers are small or IMR changes suddenly, and this is a welcome limitation statement. However, the same paragraph also asserts that the method is 'valid for most cancer sites' without quantifying 'most' or listing which sites are excluded; a more precise enumeration would help the reader assess the scope of validity.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the incidence estimates are derived from mortality and IMR data and benchmarked against external registry-observed cases.

full rationale

The paper's derivation chain is not circular. Estimated incidence cases are produced from (i) NORDPRED mortality predictions (Eq. 1) and (ii) GLMM-based IMR projections under five explicit scenarios, then compared with observed registry cases via Eqs. (2) and (3). The observed 2004-2013 incidence is used only as the benchmark, not as an input to the estimation model: for each target year, the IMR model is fitted to data ending five years before the target year and mortality data ending two to three years before, reproducing the real-world data delay. The GOF indicator (Eq. 3) selects the best scenario on the validation window, which makes the reported best-scenario MAPE a within-sample selection-minimum and therefore somewhat optimistic; however, this is a methodological caveat about post-selection, not a definitional circularity, because the estimated case counts are not derived from the observed case counts by construction and the central conclusion rests on an external comparison against a population-based registry. Self-citations (e.g., reference 7) are used to describe prior applications of the method, not to justify the validity claim, and the external benchmark is independent. No step in the derivation reduces to its own inputs.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The central claim rests on standard epidemiological assumptions about data quality and on the ad hoc set of five IMR scenarios. No new entities are introduced. The main free parameters are modeling choices: spline knots, polynomial degree, projection windows, and the per-site scenario selection.

free parameters (4)
  • Age spline knots in IMR model = 10th percentile, first tertile, second tertile, 90th percentile of the mortality pool
    Chosen by hand to flexibly model age effects; not fitted to the target outcome but affects all estimates.
  • Degree of year polynomial in IMR model = 2 (quadratic)
    Assumed functional form for the time trend of the incidence-to-mortality ratio.
  • Time windows for projections = 20-year mortality series, 15-year IMR series, 3-year and 5-year projections
    Chosen to mimic administrative reporting delays; results may depend on these window lengths.
  • Best IMR scenario per site = C1, C3, C5, or L depending on site (see Table 1)
    Selected as the scenario minimizing MAPE on the same 2004-2013 validation data, so the reported best-scenario error is partly in-sample.
assumptions (5)
  • domain assumption Granada registry and official mortality data are complete and correctly coded.
    The validity comparison treats observed incident cases and deaths as ground truth, as described in Methods, Study population.
  • ad hoc to paper The five IMR trend scenarios span the plausible true behavior of the IMR for each site.
    Scenarios C1, C3, C5, L, and Q are introduced in Methods, Assumptions regarding the mortality-to-incidence ratios; if the true IMR follows none of these, predictions fail, as seen for prostate and breast cancer.
  • domain assumption NORDPRED age-period-cohort model with a power link accurately forecasts mortality three years ahead.
    Used in the first stage to predict deaths, as stated in Methods, Estimation method and equation (1).
  • domain assumption Incident cases follow a Poisson distribution with deaths as the offset in the GLMM.
    Stated in Methods, Estimation method; a standard assumption that is not tested against alternative error distributions.
  • ad hoc to paper Bayesian priors used in OpenBUGS are adequate, though not reported.
    The paper states Bayesian MCMC but gives no prior distributions or convergence diagnostics, so this remains an unverified modeling choice.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Cancer incidence estimation from mortality data: a validation study within a population-based cancer registry." pith.science (2026). https://pith.science/paper/3H22ZMKD

@misc{pith2026241112784,
  author       = {Pith},
  title        = {Pith review of: Cancer incidence estimation from mortality data: a validation study within a population-based cancer registry},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3H22ZMKD}},
  note         = {Machine review of arXiv:2411.12784}
}
read the original abstract

We assessed the validity of one of the most frequently used methods to estimate cancer incidence, on the basis of cancer mortality data and the incidence-to-mortality ratio IMR, the IMR method. Using the previous 15 year cancer mortality time series, we derived the expected yearly number of cancer cases in the period 2004 to 2013 for six cancer sites for each sex. Generalized linear mixed models, including a polynomial function for the year of death and smoothing splines for age, were adjusted. Models were fitted under a Bayesian framework based on Markov chain Monte Carlo methods. The IMR method was applied to five scenarios reflecting different assumptions regarding the behavior of the IMR. We compared incident cases estimated with the IMR method to observed cases diagnosed in 2004 to 2013 in Granada. A goodness-of-fit GOF indicator was formulated to determine the best estimation scenario. The relative differences between the observed and predicted numbers of cancer cases were less than 10 percent for most cancer sites. The constant assumption for the IMR trend provided the best GOF for colon, rectal, lung, bladder, and stomach cancers in men and colon, rectum, breast, and corpus uteri in women.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

22 extracted references · 20 canonical work pages

  1. [1]

    Cancer epidemiology: principles and methods, Chapter 4: Measures of occurrence of disease and other health related events

    dos Santos Silva I. Cancer epidemiology: principles and methods, Chapter 4: Measures of occurrence of disease and other health related events. Lyon: International Agency for Research on Cancer; 1999. p. 60 – 3. Available from: http://publications.iarc.fr/Non-Series-Publications/Other-Non-Series-Publica tions/Cancer-Epidemiology-Principles-And-Methods-1999...

  2. [2]

    Evaluation of data quality in the cancer registry: principles and methods

    Bray F, Parkin DM. Evaluation of data quality in the cancer registry: principles and methods. Part I: comparability, validity and timeliness. Eur J Cancer. 2009;45:747– 55 Available from: http://linkinghub.elsevier.com/ retrieve/pii/S0959804908009209. [cited 2018 May 28]

  3. [3]

    Evaluation of data quality in the cancer registry: principles and methods Part II

    Parkin DM, Bray F. Evaluation of data quality in the cancer registry: principles and methods Part II. Completeness. Eur J Cancer 2009;45:756 – 764. Available from: http://www.ncbi.nlm.nih.gov/pubmed/19128954. [cited 2018 May 28]

  4. [4]

    Cancer incidence in five continents, Vol

    Forman D, Bray F, Brewster DH, Gombe Mbalawa C, Kohler B, Piñeros M, Znaor A, Zanetti R and Ferlay J E. Cancer incidence in five continents, Vol. XI. Lyon Int. Agency Res. Cancer. 2017. Available from: http://ci5.iarc.fr. [cited 2018 Jun 13]

  5. [5]

    ECIS - European Cancer Information System

    Joint Research Center-European Commission. ECIS - European Cancer Information System. 2018. Available from: https://ecis.jrc.ec.europa.eu/. [cited 2018 Jun 13]

  6. [6]

    Cancer today (powered by GLOBOCAN 2018)

    Ferlay J, Ervik M, Lam F, Colombet M, Mery L, Piñeros M, Znaor A, Soerjomataram I BF. Cancer today (powered by GLOBOCAN 2018). IARC CancerBase No. 15. 2018. Available from: http://publications.iarc.fr/577. [cited 2018 Dec 23]

  7. [7]

    Cancer incidence in Spain, 2015

    Galceran J, Ameijide A, Carulla M, Mateos A, Quirós JR, Rojas D, et al. Cancer incidence in Spain, 2015. Clin Transl Oncol. 2017;19:799 – 825 Available from: http://link.springer.com/10.1007/s12094-016-1607-9. [cited 2017 Jan 19]

  8. [8]

    Cancer incidence and mortality patterns in Europe: estimates for 40 countries in 2012

    Ferlay J, Steliarova-Foucher E, Lortet-Tieulent J, Rosso S, Coebergh JWW, Comber H, et al. Cancer incidence and mortality patterns in Europe: estimates for 40 countries in 2012. Eur J Cancer. 2013;49:1374 – 403 Available from: http://www.ncbi.nlm.nih.gov/pubmed/23485231. [cited 2016 Nov 17]

Show all 22 references
  1. [9]

    An assessment of GLOBOCAN methods for deriving national estimates of cancer incidence

    Antoni S, Soerjomataram I, Møller B, Bray F, Ferlay J. An assessment of GLOBOCAN methods for deriving national estimates of cancer incidence. Bull World Health Organ. 2016;94:174 – 84. Available from: https://www. ncbi.nlm.nih.gov/pmc/articles/PMC4773935/ . https://doi.org/10....

  2. [10]

    Available from: https://www.registrocancergrana da.es/

    Granada Cancer Registry. Available from: https://www.registrocancergrana da.es/. [cited 2019 May 11]

  3. [11]

    Situación del cáncer en España: incidencia

    López-Abente G, Pollán M, Aragonés N, Pérez Gómez B, Hernández Barrera V, Lope V, et al. Situación del cáncer en España: incidencia. An Sist Sanit Navar. 2004;27:165 – 73 Available from: http://scielo.isciii.es/scielo.php?script= sci_arttext&pid=S1137-66272004000300001&lng=en&...

  4. [12]

    Estimación de la incidencia de cáncer en España: período 1993 – 1996

    Moreno V, González JR, Soler M, Bosch FX, Kogevinas M, Borràs JM. Estimación de la incidencia de cáncer en España: período 1993 – 1996. Gac Sanit. 2001;15:380 – 8 Available from: http://linkinghub.elsevier.com/retrieve/ pii/S0213911101715919

  5. [13]

    Continuous Register Statistics

    Spanish National Institute of Statistics. Continuous Register Statistics. Available from: https://ine.es/dyngs/INEbase/es/operacion.htm?c=Esta distica_C&cid=1254736176951&menu=resultados&idp=1254735572981. [cited 2018 Jun 13]

  6. [14]

    Government of Spain

    Ministry of Health. Government of Spain. Mortality by cause of death. Available from: https://pestadistico.inteligenciadegestion.mscbs.es/ publicoSNS/S/mortalidad-por-causa-de-muerte. [cited 2018 Jun 13]

  7. [15]

    ICD-10: International Statistical Classification of diseases and related health problems: 10th revision

    World Health Organization (WHO). ICD-10: International Statistical Classification of diseases and related health problems: 10th revision. 1990

  8. [16]

    Estimates of cancer incidence and mortality in Europe in 2008

    Ferlay J, Parkin DM, Steliarova-Foucher E. Estimates of cancer incidence and mortality in Europe in 2008. Eur J Cancer. 2010;46(4):765 – 81. Available from: https://linkinghub.elsevier.com/retrieve/pii/S0959804909009265. https://doi. org/10.1016/j.ejca.2009.12.014

  9. [17]

    Prediction of cancer incidence in the Nordic countries: empirical comparison of different approaches

    Møller B, Fekjaer H, Hakulinen T, Sigvaldason H, Storm HH, Talbäck M, et al. Prediction of cancer incidence in the Nordic countries: empirical comparison of different approaches. Stat Med. 2003;22:2751 – 66 Available from: http://www.ncbi.nlm.nih.gov/pubmed/12939784

  10. [18]

    Making BUGS open

    Thomas A, O ’Hara B, Ligges U, Sturtz S. Making BUGS open. R News. 2006;6: 12– 7 Available from: https://cran.r-project.org/doc/Rnews/. Redondo-Sánchez et al. Population Health Metrics (2021) 19:18 Page 9 of 10

  11. [19]

    The BUGS project: evolution, critique and future directions

    Lunn D, Spiegelhalter D, Thomas A, Best N. The BUGS project: evolution, critique and future directions. Stat Med. 2009;28(25):3049 – 67. Available from: http://doi.wiley.com/10.1002/sim.3680

  12. [20]

    Available from: http://openbugs.net/w/FrontPage

    OpenBUGS. Available from: http://openbugs.net/w/FrontPage. [cited 2019 May 11]

  13. [21]

    Simple versus complex models: evaluation, accuracy, and combining

    Ahlburg DA. Simple versus complex models: evaluation, accuracy, and combining. Math Popul Stud. 1995;5:281 – 90 Available from: http://www.ta ndfonline.com/doi/abs/10.1080/08898489509525406. [cited 2019 May 11]

  14. [22]

    Cancer screening in Spain

    Ascunce N, Salas D, Zubizarreta R, Almazan R, Ibanez J, Ederra M, et al. Cancer screening in Spain. Ann Oncol. 2010;21:iii43 – 51 Available from: http://www.ncbi.nlm.nih.gov/pubmed/20427360. [cited 2018 Dec 7]. Publisher’sN o t e Springer Nature remains neutral with regard to ...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.