{"id":"63e80f01-cab4-4cd3-ac45-cf9e5919e5bb","arxiv_id":"2411.12784","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"The IMR method, which derives cancer incidence from mortality data, produced total incidence estimates within 6% (men) and 4% (women) of observed cases in Granada over 2004 to 2013.","lead":"This study checked whether a widely used method that estimates cancer incidence from death records works, by comparing its predictions for 2004 to 2013 against actual registered cases in Granada, Spain. Overall yearly estimates were within about 4 to 6 percent of the real counts for most cancer types, so the method looks usable for regions without cancer registries.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Best-scenario selection on the validation data biases the reported MAPE; a temporal hold-out is needed to support the IMR method's validity claim.","rationale":"The reader's verdict is CONDITIONAL and identifies the IMR scenario coverage as the weakest assumption. My concern is closely related but more specific: even if the five scenarios were a plausible family, the method's reported accuracy is inflated by selecting the best scenario on the same data used to evaluate it. This is an internal methodological issue rather than an external assumption about future IMR trajectories. The reader did flag 'selection of the best scenario on the same validation data used to report accuracy' in the rationale, so there is partial agreement. Because the paper is a validation study, the central claim depends on the credibility of the validation protocol; post-selection inference undermines that credibility but does not necessarily invalidate the method entirely. A temporal hold-out test would settle whether the reported accuracy is genuine. This does not change the CONDITIONAL verdict, so verdict_should_be is UNCHANGED. I am not proposing a stronger verdict because the core comparison (estimated versus observed incidence over a decade) is a genuine out-of-sample exercise for each individual year, and the concern is testable rather than fatal.","tokens_in":10033,"tokens_out":3266,"duration_ms":32204,"concrete_test":"Perform a temporal two-fold cross-validation: use 2004-2008 to select the best scenario for each site via the GOF/MAPE indicator, then compute the MAPE and relative deviation on the hold-out years 2009-2013 for those pre-selected scenarios. Compare these hold-out MAPEs with the Table 1 values; if the hold-out MAPE for all sites exceeds the reported 6.34%/3.85% by a substantial margin (e.g., greater than 10% for men or women), the headline accuracy is partly an artifact of within-validation scenario selection.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central evidence for validity is the accuracy of the 'best scenario' per cancer site: Table 1 reports MAPE values such as 6.34% for men and 3.85% for women for all sites, and relative deviations of -0.3% and -0.5%. These figures are obtained after selecting, for each site, the scenario (C1, C3, C5, L, or Q) that minimizes MAPE on exactly the same 2004-2013 validation period used to compute those errors. With ten annual observations per site and five candidate scenarios, the minimum MAPE is an optimistically biased estimate of the expected error when the GOF indicator is applied in a new setting. The paper's own results show the problem: for prostate cancer, even the best scenario has a 27.37% MAPE and a 136% relative deviation in 2013, so the aggregate -0.3% is not a reliable measure of general predictive performance. The proposed GOF indicator is promoted as an objective way to select scenarios, but its selection accuracy is never assessed on data separate from the data used to choose and evaluate it. Thus the quantitative support for 'valid tool' is a post-selection estimate, not an unbiased validation.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper validates the incidence-to-mortality ratio (IMR) method for estimating cancer incidence from mortality data, using a 10-year historical series (2004–2013) from the Granada population-based cancer registry. Mortality data from 1982–2010 are used to derive mortality projections via NORDPRED-style APC models, and IMR projections are obtained from Bayesian GLMMs for each of five assumed IMR scenarios (constant last value, mean of last 3/5 years, linear, quadratic). A goodness-of-fit indicator based on the mean absolute percentage error (MAPE) is used to select the best scenario per site and sex. The paper reports that, for the best scenario per site, the overall relative deviation is −0.3% in men and −0.5% in women, with MAPE values of 6.34% and 3.85%, respectively, and concludes that the IMR method is a valid tool for cancer incidence estimation when data quality is high.","tokens_in":10273,"tokens_out":2733,"duration_ms":26915,"significance":"This is the first validation of the IMR method over a multi-year time series against a high-quality population-based registry, rather than a single-year comparison. The rolling iterative procedure that mimics real-world prediction delays is a genuine strength, as is the use of a Bayesian framework and the explicit comparison of multiple IMR assumptions. However, the quantitative evidence for the 'valid tool' claim is weakened by the fact that the best scenario for each site is selected on exactly the same validation period used to compute the reported MAPE and deviations, making those error metrics post-selection estimates. If the authors add a temporally separated or cross-validated assessment of the scenario-selection procedure, the result would be considerably more convincing; as it stands, the central claim is plausible but not yet supported at the level the paper asserts.","major_comments":[{"comment":"The reported MAPE values and relative deviations are optimistically biased because the best scenario for each site is chosen by minimizing MAPE over the same 2004–2013 period used to report those errors. With five scenarios and ten annual observations per site, the minimum of five MAPE estimates is not an unbiased estimator of the expected error in a new application. The paper should provide an unbiased evaluation, for example by holding out the last two to three years for assessment, by using a nested or rolling-origin cross-validation scheme, or by reporting the MAPE averaged over all scenarios in addition to the best-scenario MAPE.","section":"Goodness-of-fit assessment; Table 1"},{"comment":"The five assumed IMR scenarios (C1, C3, C5, L, Q) are asserted to span plausible trajectories, but the paper's own results show that for prostate cancer in men and for breast, ovary, and female lung cancer, none of the scenarios provides a satisfactory fit, with annual deviations up to 136% for prostate cancer in 2013 (Section 'Results', Additional file 1). This means the conclusion 'valid tool' is conditional on the assumption that the true IMR trajectory belongs to the family of five scenarios, and that assumption is already violated for several sites studied here. The paper should state this limitation more prominently and avoid a blanket validity claim for all cancer sites.","section":"Assumptions regarding the mortality-to-incidence ratios"},{"comment":"The small overall relative deviations (−0.3% men, −0.5% women) are largely driven by cancellation of offsetting annual errors, not by consistent accuracy. Table 2 shows annual relative deviations for all sites combined ranging from −7.7% to +24.7% in men, with the largest error in the final year (2013). The conclusion that the method provides 'good reliability' should be based on a measure that does not hide yearly inaccuracy, such as the distribution of annual absolute errors or prediction intervals for each year, rather than the cumulative relative deviation.","section":"Validity assessment; Table 2"},{"comment":"The MAPE-based GOF indicator is promoted as an objective way to select the best IMR scenario, but the paper never assesses the selection performance of this indicator. Since the indicator is fitted and evaluated on the same data, its ability to identify the correct scenario in a prospective setting is unknown. A validation study that applies the GOF rule to the first part of the series and then checks accuracy on the last part is needed to support the claim that the GOF indicator 'can help select the best assumption' (Discussion).","section":"Goodness-of-fit assessment; Eq. (3)"}],"minor_comments":[{"comment":"Equation (2) is typeset as '100: Expected − Observed / Observed', which is ambiguous; it should be written as 100 × (Expected − Observed)/Observed, and the direction of the sign (positive for overestimation) is stated only in the text, not with the equation.","section":"Methods, Eq. (2)"},{"comment":"Equation (1) contains garbled notation ('D /C1 p + Pp + Cc /C0/C1 5') that appears to be an OCR artifact; the model formula should be presented in standard APC notation so that the drift, period, and cohort terms are unambiguous.","section":"Methods, Eq. (1)"},{"comment":"The row for 'All sites' does not report a 'Best scenario' column value, which is confusing because the all-sites estimates must be based on some combination of site-specific scenarios; please clarify whether the all-sites MAPE is computed from the sum of best-scenario sites or from a separately chosen overall scenario.","section":"Results, Table 1"},{"comment":"The Discussion mentions that the method is less precise when case numbers are small or IMR changes suddenly, and this is a welcome limitation statement. However, the same paragraph also asserts that the method is 'valid for most cancer sites' without quantifying 'most' or listing which sites are excluded; a more precise enumeration would help the reader assess the scope of validity.","section":"Discussion"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the scope of Population Health Metrics and addresses a practical estimation problem. The main weakness is methodological: the post-selection evaluation of the best IMR scenario inflates the reported accuracy. This is fixable with a temporal hold-out or a report of all-scenario results, so I do not recommend rejection. The paper would also benefit from releasing the analysis code or the aggregated mortality/IMR data to support reproducibility, since the data availability statement says only 'upon reasonable request'."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Georges,\n\nThis paper is a genuine validation study, not a methods toy. It takes the IMR method (mortality plus incidence-to-mortality ratio, used in GLOBOCAN/EUCAN) and checks it against ten years of observed incident cases from the Granada registry, using an iterative design that mimics the real data-delay situation. That is the new thing: Antoni et al. did a single-year check; this one runs 2004-2013 and compares five assumptions about the IMR trend. The paper is also honest about where the method breaks down: prostate and breast, with their screening-driven IMR fluctuations, get errors of 27% MAPE and 12-15% respectively, and the 2013 male total spikes 25% because of prostate. The authors deserve credit for reporting those results rather than hiding them.\n\nThe main soft spot is the selection of the 'best scenario.' For each site, they pick the scenario (C1, C3, C5, L, Q) with the lowest MAPE over 2004-2013, and then report that same MAPE as the error of the method. That makes the headline numbers - MAPE 6.3% men, 3.9% women, total relative deviation -0.3% - optimistic estimates of predictive performance. The selection is on the same data used for evaluation. The bias is not enormous, but it's real, and it matters because the authors also propose their GOF indicator as a tool for choosing scenarios in new settings. They never test whether the GOF-based selection works out-of-sample. A temporal hold-out - fit and select on, say, 2004-2008, evaluate on 2009-2013 - would settle it, and it wouldn't cost much given the data are already there. There are also smaller issues: no confidence or credible intervals for MAPE or relative deviations, model priors under-specified, and code/data only 'available from the corresponding author upon request,' which is weak for a methods-validation paper.\n\nNone of this breaks the central conclusion. For total cancer incidence, the IMR method tracks observed counts within single-digit percentage points in most years, and the paper is upfront that per-site estimates for screening-affected cancers need caution. But 'valid tool' is a qualified claim, not a clean bill of health. The paper is worth a serious referee: the empirical design is meaningful and the results are useful for anyone working in cancer surveillance in regions without registries. My recommendation: send it to review, but ask the authors to validate the GOF-based scenario choice on held-out data and to add uncertainty measures. With those changes, the paper becomes a solid reference.","headline":"Genuine multi-year validation of the IMR method, but the reported accuracy is inflated by selecting the best scenario on the same data used for evaluation.","tokens_in":10841,"tokens_out":2628,"would_cite":true,"duration_ms":24442,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that the IMR method—estimating cancer incidence from mortality data and an incidence-to-mortality ratio—is a valid tool, validated against a ten-year observed series from a population-based registry, and that a…","keywords":["Cancer incidence","Incidence-to-mortality ratio","Estimation validation","Goodness-of-fit","Mean absolute percentage error","Population-based cancer registry","Bayesian generalized linear mixed model","NORDPRED"],"falsifier":"Run the same iterative back-casting procedure on another high-quality registry with a long historical series, for a site whose IMR is known to have changed in a step-like manner (e.g., prostate cancer around the introduction of screening). If the best-scenario estimate deviates from observed cases by more than the paper's reported site-level range of roughly −13% to +14%, or if the MAPE indicator selects a scenario while year-to-year errors exceed 25% for non-screening sites, the paper's validity claim would be contradicted. A simpler observation: a single site whose true IMR trajectory is, say, an inverted U over the projection window—captured by none of the five scenarios—would produce systematic bias that the GOF indicator cannot detect.","tokens_in":9831,"feed_emoji":"📊","tokens_out":8676,"duration_ms":74933,"temperature":0.7,"pith_summary":"This paper tries to establish that a widely used shortcut for estimating cancer incidence—the incidence-to-mortality ratio (IMR) method, which derives new cancer cases from death records and a modelled ratio of incidence to mortality—produces trustworthy estimates when checked against a decade of real registry data. Using a 15-year mortality series and five different assumptions about how the IMR behaves over time, the authors back-cast yearly incidence for 2004–2013 in Granada, Spain, and compare each estimate to the cases the local population-based registry actually recorded. For most of the 14 site–sex combinations, the best-fitting scenario stays within 10% of observed cases per year, and the overall ten-year totals differ by only −0.3% in men and −0.5% in women. The paper also proposes a goodness-of-fit score based on the mean absolute percentage error (MAPE) to choose among IMR scenarios on statistical grounds rather than subjective judgement. A sympathetic reader would take the conclusion as: for high-quality mortality and population data, the IMR method is a valid tool for total and site-specific incidence estimation, with the caveat that cancers affected by screening programs (prostate, breast, ovary, female lung) need separate handling.","feed_headline":"Cancer incidence estimated from death records lands within 0.5%","feed_subtitle":"Ten-year registry check confirms the IMR estimation method; screening-affected cancers remain the weak spot.","key_machinery":"The machinery is the incidence-to-mortality ratio (IMR)—the number of new cancer cases per cancer death in a given year, age group, and site—combined with a two-stage estimation pipeline. First, an age-period-cohort model (the NORDPRED method) predicts the number of cancer deaths in the target year from a mortality time series; second, a generalized linear mixed model with a Poisson distribution, age splines, and a second-degree polynomial for year estimates the IMR, which is then projected under five scenarios: constant at the last value, constant at the mean of the last three years, constant at the mean of the last five years, linear trend, and quadratic trend. An iterative back-casting procedure reproduces the real-world delay in data availability and generates expected incidence for each year 2004–2013, and the mean absolute percentage error (MAPE) between expected and observed cases serves as the goodness-of-fit indicator that selects the best scenario per site. This design lets the method's assumptions be tested directly against observed registry data.","core_discovery":"The central claim is that the incidence-to-mortality ratio (IMR) method—which estimates new cancer cases from death counts and a modelled ratio of incidence to mortality—is a valid tool for estimating cancer incidence, and that this validity can be demonstrated by comparing estimated cases with observed cases over a continuous ten-year historical series. The paper reports that under the best scenario for each site, the relative difference between estimated and observed total cancer cases was −0.3% for men (23,126 estimated vs 23,197 observed) and −0.5% for women (16,574 vs 16,651), with site-level relative deviations ranging from −9.1% (rectal cancer in men) to +11.6% (prostate cancer in men), and from −12.6% (lung cancer in women) to +14.3% (ovarian cancer). The constant-IMR assumption gave the best fit for colon, rectal, lung, bladder, and stomach cancers in men and colon, rectal, breast, and corpus uteri in women; the linear assumption was best for prostate cancer in men and lung, ovary, and other sites in women. The method's weakest performance is for prostate and female breast cancer, where screening-induced sudden changes in the IMR produce errors up to 136% in a single year, and the paper explicitly recommends new strategies for such sites.","pith_inferences":["The validation is retrospective: the best scenario is chosen after seeing the observed data, so the reported accuracy is likely an upper bound on what a prospective user would achieve when the true IMR trajectory is unknown.","The near-perfect ten-year totals arise partly because yearly over- and under-estimates cancel out; decision-makers relying on single-year estimates for screening-affected sites should expect much larger errors than the decade totals suggest.","A natural extension would be to apply the same back-casting validation to registries in other regions with different survival and screening patterns, which would test whether the method's accuracy transfers or is specific to a particular health system.","The method's failure on screening-affected sites suggests a pragmatic hybrid: use the IMR method for sites with stable IMR, and switch to incidence-based projection or screening-adjusted models for sites with known early-detection programmes."],"forward_implications":["For regions without a cancer registry but with reliable mortality and population data, the IMR method can produce total cancer incidence estimates within about half a percent over a decade, with most site-specific estimates within 10%.","The MAPE goodness-of-fit indicator gives future users an objective, data-driven way to choose among IMR trend assumptions instead of relying on subjective judgement.","Using the mean of the last three to five IMR values instead of the single most recent value generally improves the fit, while quadratic extrapolation performs worst and should be avoided.","Cancers affected by screening programmes—prostate, breast, ovary, and female lung—are the method's known weak points, and estimates for these sites should be interpreted with caution or modelled separately.","The method's validity is conditional on high-quality input data (incidence, mortality, and population), so the same accuracy cannot be assumed in settings with poorer data."],"supporting_citations":[{"why":"Defines the adapted mortality-to-incidence ratio estimation method that the study validates.","marker":"[16]"},{"why":"Supplies the NORDPRED age-period-cohort model used to predict cancer deaths in the first stage.","marker":"[17]"},{"why":"Provides the previous one-year comparison of estimation methods that this study extends to a ten-year series.","marker":"[9]"},{"why":"Shows a prior application of the IMR method for national incidence estimates, establishing the method's practical use.","marker":"[7]"},{"why":"Provides the observed incidence cases from the population-based registry used as the validation reference.","marker":"[10]"},{"why":"Supplies the population denominators needed to compute rates and expected counts.","marker":"[13]"},{"why":"Supplies the official cancer mortality data that feeds the estimation.","marker":"[14]"},{"why":"Justifies the use of mean absolute percentage error as the goodness-of-fit indicator.","marker":"[21]"}],"fun_headline_variants":["Death records yield cancer incidence within 0.5% overall","IMR method validated: cancer estimates match registry closely","Mortality-based cancer counts pass 10-year registry test","Most cancers estimated well from death data, but not all","Death data estimates cancer well except screening cancers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the five assumed shapes for the incidence-to-mortality ratio—constant, mean of the last three or five years, linear, or quadratic—cover the true way that ratio changes over time for each cancer site, so that if the real ratio follows a different pattern, the method's estimates will be biased no matter how well the mortality model performs.","fun_headline_variants_meta":{"raw":{"variants":["Death records yield cancer incidence within 0.5% overall","IMR method validated: cancer estimates match registry closely","Mortality-based cancer counts pass 10-year registry test","Most cancers estimated well from death data, but not all","Death data estimates cancer well except screening cancers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000964,"raw_usage":{"total_tokens":4151,"prompt_tokens":1038,"completion_tokens":3113,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":654,"completion_tokens_details":{"reasoning_tokens":3035}},"tokens_in":654,"tokens_out":3113,"duration_ms":24287,"temperature":1.0,"reasoning_tokens":3035,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T17:35:17.575213+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same iterative back-casting procedure on another high-quality registry with a long historical series, for a site whose IMR is known to have changed in a step-like manner (e.g., prostate cancer around the introduction of screening). If the best-scenario estimate deviates from observed cases by more than the paper's reported site-level range of roughly −13% to +14%, or if the MAPE indicator selects a scenario while year-to-year errors exceed 25% for non-screening sites, the paper's validity claim would be contradicted. A simpler observation: a single site whose true IMR trajectory is, say, an inverted U over the projection window—captured by none of the five scenarios—would produce systematic bias that the GOF indicator cannot detect.","supporting_citations":[{"cited_title":"Prediction of cancer incidence in the Nordic countries: empirical comparison of different approaches","cited_arxiv_id":null,"evidence_quote":"Supplies the NORDPRED age-period-cohort model used to predict cancer deaths in the first stage."},{"cited_title":"An assessment of GLOBOCAN methods for deriving national estimates of cancer incidence","cited_arxiv_id":null,"evidence_quote":"Provides the previous one-year comparison of estimation methods that this study extends to a ten-year series."},{"cited_title":"Cancer incidence in Spain, 2015","cited_arxiv_id":null,"evidence_quote":"Shows a prior application of the IMR method for national incidence estimates, establishing the method's practical use."},{"cited_title":"Available from: https://www.registrocancergrana da.es/","cited_arxiv_id":null,"evidence_quote":"Provides the observed incidence cases from the population-based registry used as the validation reference."},{"cited_title":"Continuous Register Statistics","cited_arxiv_id":null,"evidence_quote":"Supplies the population denominators needed to compute rates and expected counts."},{"cited_title":"Government of Spain","cited_arxiv_id":null,"evidence_quote":"Supplies the official cancer mortality data that feeds the estimation."},{"cited_title":"Simple versus complex models: evaluation, accuracy, and combining","cited_arxiv_id":null,"evidence_quote":"Justifies the use of mean absolute percentage error as the goodness-of-fit indicator."}],"review_version":1}