{"id":"b22e756b-7388-48f5-bb7f-9611c466f1c4","arxiv_id":"1908.01446","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Forecasting life-table death counts with compositional data analysis and exponential smoothing gives lower point and interval errors than Lee-Carter, Hyndman-Ullah, and random-walk benchmarks on Australian data, and it supports temporary annuity pricing.","lead":"A statistical method that treats life-table deaths as proportions that must sum to a fixed total produced more accurate forecasts of Australia's age distribution of deaths than three standard mortality forecasting benchmarks. The authors then used the forecast death counts to estimate prices for temporary annuities at different ages and maturities.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"CoDa's recommended L=6 is selected on the same holdout used for evaluation, so its out-of-sample superiority may be an artifact of test-set tuning.","rationale":"The reader's weakest assumption focused on the LC/HU conversion approximation, which is indeed a real limitation but is acknowledged in the paper and could be corrected. My read identifies a more load-bearing issue: the recommended configuration (L=6 with ETS) appears to be selected using the same holdout that is then used to demonstrate superiority. This violates the independence of the evaluation and directly affects the central claim that CoDa outperforms the benchmarks. The paper otherwise provides a competent adaptation of CoDa, a clear expanding-window evaluation, and reproducible R code, which I credit. However, because the L selection is not validated on separate data, the comparison is conditional on fixing that methodological gap. Since the reader already issued a CONDITIONAL verdict, my concern does not change the verdict, but it sharpens the condition: the authors should demonstrate that the CoDa advantage persists when L is chosen without peeking at the holdout.","tokens_in":19964,"tokens_out":2358,"duration_ms":26515,"concrete_test":"Re-run the expanding-window evaluation of Section 5.1 with L selected only from the training portion of each window (e.g., via rolling-origin cross-validation on the training data, or by fixing L=6 before any holdout data is examined). Then compare the resulting CoDa-ETS forecasts with the LC, HU, and naive benchmarks on the untouched holdout. If the CoDa advantage shrinks or disappears, the headline claim of superiority is not robust to honest model selection.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that CoDa with ETS and L=6 outperforms HU, LC, and naive random walks for forecasting the age distribution of death counts. The load-bearing condition is that the comparison is a genuine out-of-sample evaluation. Section 5.1 describes an expanding-window holdout used to compute MAPE and interval scores in Table 1. Section 5.2 and the conclusion recommend L=6 because it 'generally provides more accurate point and interval forecast accuracies' than the CPV-selected L=1 or L=2. But the CPV criterion (Section 3.1) selects L=1 for females and L=2 for males; L=6 is introduced only as a sensitivity analysis and is then chosen as the final configuration because it performs best on the same holdout data that is later used to claim superiority. This is test-set model selection: the holdout has already influenced the choice of L, so the reported error metrics for the recommended CoDa-ETS(L=6) configuration are not independent. The margin over HU and LC could be inflated by selection bias. The paper does not describe any separate validation set or nested cross-validation to prevent this. A secondary concern is the acknowledged q=1-exp(-m) conversion for LC/HU (Section 3.3.1, 5.2), which may disadvantage those benchmarks, but the L selection issue is more direct and is not explicitly acknowledged as a limitation.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a compositional data analysis (CoDa) approach for forecasting the age distribution of life-table death counts. Using Australian sex-specific period life-table death counts from 1921 to 2014, the method applies a centred log-ratio transformation, principal component analysis, and exponential smoothing (ETS) forecasts of principal component scores, with bootstrap prediction intervals. The authors compare point and interval forecast accuracy against the Lee-Carter method, the Hyndman-Ullah method, and random-walk benchmarks in an expanding-window holdout evaluation, and then use the forecasts to price single-premium temporary immediate annuities. The central claim is that CoDa with ETS and L=6 principal components outperforms the benchmarks for one- to 20-step-ahead forecasts.","tokens_in":20258,"tokens_out":3526,"duration_ms":37849,"significance":"If the comparative claim holds, the paper offers a useful and practical extension of compositional-data mortality forecasting: it adapts the Hyndman-Ullah functional-data approach to the constrained space of life-table death counts, adds a bootstrap interval procedure, and connects improved death-count forecasts to annuity pricing. The strengths are the use of public HMD data, standard evaluation criteria (MAPE and Gneiting-Raftery interval scores), a clear actuarial application, and the provision of R code and data files. The main weakness is that the recommended configuration (L=6, ETS) is selected by inspecting the same holdout sample that is later used to establish superiority, which undermines the independence of the headline out-of-sample comparison. The benchmark conversion from mortality rates to death counts is also a plausible source of unfair disadvantage to Lee-Carter and Hyndman-Ullah.","major_comments":[{"comment":"The recommended configuration L=6 is selected on the same expanding-window holdout used to report forecast superiority. Section 3.1 states that the cumulative-percentage-of-variance criterion selects L=1 for females and L=2 for males; Section 5.2 then recommends L=6 because it 'generally provides more accurate point and interval forecast accuracies' than the CPV-selected choices, based on Table 1. Since Table 1 is computed from the same holdout used to support the abstract's 'outperforms' claim, the reported error metrics for CoDa-ETS(L=6) are not independent test-set results. The authors should either use a separate validation period for model selection, implement nested cross-validation, or at minimum report the comparison both for the CPV-selected L and for the sensitivity L=6 while clearly flagging which numbers are selection-influenced. Without this, the size of the reported improvement over Hyndman-Ullah and Lee-Carter may be inflated by test-set tuning.","section":"Sections 5.1 and 5.2; Table 1"},{"comment":"The comparison may unfairly disadvantage the Lee-Carter and Hyndman-Ullah benchmarks through the conversion of forecast central mortality rates to death counts using q = 1 - exp(-m). As the paper itself acknowledges in Section 5.2, this approximation is likely to be poor at high ages, and the authors suggest it may explain the benchmarks' worse performance. Because the central claim is improved forecast accuracy over these benchmarks, the conversion assumption is load-bearing. The authors should use the exact life-table relation available from HMD (or an approximation with an explicit separation factor) and quantify how sensitive the rankings in Table 1 are to this choice. If the conversion is retained, a sensitivity analysis comparing alternative conversions is needed to establish that the CoDa advantage is not an artifact of the benchmark implementation.","section":"Sections 3.3.1 and 5.2"},{"comment":"Interval forecast performance is assessed only through the interval score, with no empirical coverage check for the proposed bootstrap prediction intervals. The interval score rewards narrow intervals but does not, by itself, reveal systematic under- or over-coverage; Table 3 shows very narrow 95% intervals for annuity prices, and it is unclear whether nominal 80% and 95% coverage are achieved in the holdout. The authors should report empirical coverage rates alongside interval scores. In addition, the MAPE and interval-score differences between methods are reported as point values without any measure of uncertainty (e.g., standard errors or Diebold-Mariano-type tests), so the statistical significance of the claimed superiority is not established.","section":"Sections 3.2 and 5.1; Tables 1 and 3"}],"minor_comments":[{"comment":"The column headings 'K = 6' should read 'L = 6' for consistency with the main text's notation for the number of principal components.","section":"Supplementary Tables 4 and 5"},{"comment":"The term 'multiple log-ratio' appears to be a typo for 'multiplicative log-ratio'; please correct it.","section":"Section 3.1, item on log-ratio transformations"},{"comment":"The text in Section 6 refers to 'Table 4' and 'Table 5' for female annuity prices, but the tables in the main text are numbered 2 and 3; the table numbering should be made consistent.","section":"Section 6 and Tables 2-5"},{"comment":"The caption of Figure 2 says 'although we use L = 6 for fitting', while Section 4's R-squared discussion uses the CPV-selected L=1 and L=2; please clarify which value of L is used in the figure and in the goodness-of-fit calculations.","section":"Section 4 and Figure 2"},{"comment":"The text states that 'using the CPV criterion, we select L = 1 for both female and male data', but Section 3.1 reports L=1 for females and L=2 for males; this inconsistency should be resolved.","section":"Section 4"}],"recommendation":"major_revision","confidential_remarks":"The paper's core idea is sound and the applied context is relevant, but the current version's headline claim rests on a holdout-based model selection choice (L=6) and on a benchmark mortality-rate-to-death-count conversion that the authors themselves flag as potentially unfavorable. Both issues are fixable within the scope of a revision, so I do not recommend rejection, but the authors should be asked to supply an independent evaluation protocol, benchmark sensitivity checks, and interval coverage diagnostics before the comparative claims can be accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Hi [Colleague],\n\nQuick take: this is a competent applied extension of compositional data analysis to mortality forecasting, with a real chance of being useful, but the headline result is weaker than it looks because the recommended L=6 is selected on the same holdout that is later used to claim superiority. That is not a fatal flaw, but it is exactly the kind of test-set tuning that inflates apparent out-of-sample performance, and the paper does not report any separate validation or nested model selection.\n\nWhat is actually new: Oeppen (2008) and Bergeron-Boucher et al. (2017) introduced CoDa for life-table death counts with a single component (or somewhat limited component choice). Here the authors allow multiple principal components, use the automatic ETS search of Hyndman and Khandakar (2008) to forecast scores, propose a bootstrap interval scheme that samples both forecast errors and model residuals, and demonstrate the method on temporary annuity pricing. The empirical comparison on Australian data (1921-2014) is not in the prior work, and the annuity application is a natural and useful extension. The paper also shares R code and data, which is appreciated.\n\nThe main soft spot is the L selection. In Section 3.1, L is chosen by CPV (L=1 for females, L=2 for males). Section 5.2 then says L=6 'generally provides more accurate' errors, and the conclusion recommends it. But the evidence for L=6 comes from the same expanding-window holdout that produced Table 1. So the reported MAPE and interval scores for the recommended configuration are not independent; they are in-sample relative to model selection. The margin over HU and LC could shrink under a fairer protocol. This is fixable: use the first part of the data to choose L, then evaluate on the rest, or use nested CV. I'd ask the authors to do that before accepting the headline claim.\n\nSecondary issues: the paper does not report empirical coverage of the bootstrap intervals, only interval scores; coverage would reassure that the proposed interval method is not just sharp but calibrated. The single-country design limits generalizability, though that is common in this literature. The q=1-exp(-m) conversion for LC/HU is acknowledged by the authors as a possible reason for their poorer performance at high ages; it is a standard approximation, so it is not damning, but it is one more reason to be cautious about the magnitude of the claimed superiority.\n\nOverall, the paper is worth a serious referee. The method is promising, the writing is clear, and the application is relevant. But the L-selection issue needs to be addressed before the main claim can be taken at face value. I would send it to review with a request for nested validation or a separate tuning period, and for a coverage analysis of the intervals.\n\nBest,\n[You]","headline":"A useful CoDa extension whose headline comparison is undermined by test-set selection of L=6.","tokens_in":20763,"tokens_out":3809,"would_cite":true,"duration_ms":35366,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Modelling age-specific life-table death counts as compositional data yields more accurate one- to 20-year-ahead forecasts of the death distribution and improves temporary annuity prices.","keywords":["compositional data analysis","life-table death counts","log-ratio transformation","principal component analysis","mortality forecasting","temporary annuity pricing","bootstrap prediction intervals","Lee-Carter method"],"falsifier":"Re-run the expanding-window forecast comparison with Lee-Carter and Hyndman-Ullah converted to death counts using the exact life-table identity $q_x = m_x / (1 + (1-a_x)m_x)$ (or a direct estimate of $q_x$) instead of the exponential approximation, and compare high-age errors separately; if CoDa's advantage over these benchmarks shrinks or reverses at ages above about 80, the reported superiority is an artifact of the conversion.","tokens_in":19769,"feed_emoji":"📊","tokens_out":6067,"duration_ms":58550,"temperature":0.7,"pith_summary":"This paper argues that the age distribution of life-table death counts should be forecast as compositional data: vectors of age-specific death counts that must sum to a fixed radix, transformed into unconstrained coordinates before forecasting. On Australian female and male life-table death counts from 1921 to 2014, the paper claims this approach, using principal component analysis with an automatic exponential-smoothing forecast of the component scores and retaining six components, gives more accurate one- to 20-step-ahead point and interval forecasts than the Lee-Carter, Hyndman-Ullah, and two random-walk benchmarks. The improved death-count forecasts feed directly into survival probabilities, and the paper shows they produce single-premium temporary annuity prices with bootstrap prediction intervals for ages 60 to 105 and maturities up to 30 years. A demographer or actuary who trusts the comparison would adopt the method for mortality forecasting and annuity pricing.","feed_headline":"Compositional-data forecasts beat Lee-Carter on deaths","feed_subtitle":"Australian life-table data show one- to 20-year death-count forecasts improve, sharpening annuity prices.","key_machinery":"The central object is the vector of age-specific life-table death counts, treated as compositional data in the simplex: $K$ positive components summing to the life-table radix, here 100,000. The machinery is the centred log-ratio transformation $z_{t,x}=\\ln(f_{t,x}/g_t)$, where $f_{t,x}$ is the de-centred proportion and $g_t$ the geometric mean over ages, followed by principal component analysis of the transformed matrix, automatic ETS forecasting of the resulting principal component scores, and inverse transformation back to death-count space. The paper also builds interval forecasts by bootstrapping $h$-step-ahead score forecast errors and model residuals. Choosing $L=6$ principal components rather than the one or two selected by cumulative variance is what the forecast comparison shows to be most accurate.","core_discovery":"The central claim is that forecasting the distribution of life-table deaths in its natural constrained sample space beats forecasting mortality rates and converting. The method models the relative proportions $d_{t,x}/\\sum_x d_{t,x}$, applies a centred log-ratio transformation to break the sum constraint, estimates principal components and scores, forecasts each score with an ETS model selected automatically, then back-transforms and rescales to the 100,000 radix. With $L=6$ components and ETS, the paper reports the smallest MAPE and mean interval score among the CoDa, Hyndman-Ullah, Lee-Carter, and random-walk-with/without-drift methods, with the gap typically widening at longer horizons. It also proposes a nonparametric bootstrap that resamples multi-step score forecast errors and model residuals to build prediction intervals for death counts, survival probabilities, and annuity prices. The conclusion is that CoDa with ETS and $L=6$ is the recommended method for forecasting the age distribution of death counts and for pricing temporary immediate annuities.","pith_inferences":["A testable extension is to repeat the comparison on other countries in the Human Mortality Database; if the CoDa-ETS-$L=6$ advantage is not universal, the recommendation should be qualified by mortality regime.","The comparison's dependence on the $q_x=1-e^{-m_x}$ conversion for Lee-Carter and Hyndman-Ullah suggests that a head-to-head using exact life-table conversion could narrow the gap; until then, the claimed superiority at high ages should be read cautiously.","Because compositional forecasts automatically enforce positivity and summation, they may be especially useful for small populations or subpopulations where benchmark methods produce implausible negative or non-integer death counts.","The same log-ratio machinery could be applied to other constrained demographic distributions, such as cause-of-death composition or migration age profiles, with analogous forecasting and interval-score comparisons."],"forward_implications":["If the central claim holds, demographers can derive survival probabilities and life expectancy directly from death-count forecasts that respect the radix constraint, avoiding inconsistencies from separately forecast mortality rates.","Actuaries can price temporary immediate annuities with point estimates and 95% prediction intervals; the paper's tables show prices for ages 60 to 105 and maturities 5 to 30 at a 3% interest rate.","The CoDa method's advantage grows with forecast horizon, so it is most valuable for long-term contracts, which the paper notes become almost lifetime annuities at high entry ages.","Because the method allows the rate of mortality improvement to vary over time, it captures shifts in the age-at-death distribution that fixed-coefficient Lee-Carter extrapolations miss.","The proposed bootstrap interval construction outperforms the existing CoDa bootstrap of Bergeron-Boucher et al. (2017) in the reported interval-score comparisons."],"supporting_citations":[{"why":"Supplies the original CoDa forecasting framework for life-table death counts that this paper adapts and extends.","marker":"Oeppen (2008)"},{"why":"Establishes the CoDa mortality-forecasting approach and provides the existing bootstrap method used as a comparison.","marker":"Bergeron-Boucher et al. (2017)"},{"why":"Justifies log-ratio transformations for compositional data, the technical foundation of the method.","marker":"Aitchison (1986)"},{"why":"Defines the benchmark Lee-Carter method whose forecast accuracy is compared against CoDa.","marker":"Lee and Carter (1992)"},{"why":"Defines the benchmark method being adapted and supplies the principal-component-plus-time-series-forecast structure.","marker":"Hyndman and Ullah (2007)"},{"why":"Provides the automatic ETS model-selection algorithm used to forecast principal component scores.","marker":"Hyndman and Khandakar (2008)"},{"why":"Supplies the Australian age- and sex-specific life-table death counts used in the empirical comparison.","marker":"Human Mortality Database (2019)"},{"why":"Provides the interval score used to compare interval forecast accuracy across methods.","marker":"Gneiting and Raftery (2007)"},{"why":"Supplies the temporary annuity pricing calculation that the mortality forecasts are applied to.","marker":"Shang and Haberman (2017)"}],"fun_headline_variants":["Compositional data beats Lee-Carter on death forecasts","CoDa outperforms Lee-Carter on death-count distribution","Death distribution forecasts improved by compositional method","Forecasting death counts: CoDa trumps Lee-Carter approach","Better annuity prices via compositional death forecasts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that converting Lee-Carter and Hyndman-Ullah mortality-rate forecasts to death counts via $q_x = 1 - e^{-m_x}$ is a fair implementation of those benchmarks; the paper itself notes this approximation may explain their poorer performance at high ages, so if it biases the benchmarks downward, the CoDa method's superiority would be partly an artifact.","fun_headline_variants_meta":{"raw":{"variants":["Compositional data beats Lee-Carter on death forecasts","CoDa outperforms Lee-Carter on death-count distribution","Death distribution forecasts improved by compositional method","Forecasting death counts: CoDa trumps Lee-Carter approach","Better annuity prices via compositional death forecasts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000179,"raw_usage":{"total_tokens":1265,"prompt_tokens":878,"completion_tokens":387,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":494,"completion_tokens_details":{"reasoning_tokens":312}},"tokens_in":494,"tokens_out":387,"duration_ms":4244,"temperature":1.0,"reasoning_tokens":312,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:13:12.999128+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the expanding-window forecast comparison with Lee-Carter and Hyndman-Ullah converted to death counts using the exact life-table identity $q_x = m_x / (1 + (1-a_x)m_x)$ (or a direct estimate of $q_x$) instead of the exponential approximation, and compare high-age errors separately; if CoDa's advantage over these benchmarks shrinks or reverses at ages above about 80, the reported superiority is an artifact of the conversion.","supporting_citations":[],"review_version":1}