{"id":"f9bcd102-4651-4b41-a330-92316a2b7530","arxiv_id":"2501.11514","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"PoPStat is a correlation between disease mortality and KL divergence from a per-disease optimized reference population pyramid, but the optimization makes the reported associations an artifact.","lead":"The paper proposes a new metric, PoPStat, that links a country's population pyramid shape to disease-specific mortality rates. A generalist might read it as a potential policy tool, but the metric is built by choosing the reference pyramid that maximizes the correlation, which undermines the claim.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Brute-force reference selection (Eq. 2) makes PoPStat the maximum of 180 correlations; without a selection correction, the claimed outperformance over fixed indicators is not an independent estimate.","rationale":"The reader identified the same load-bearing weakness: the reference country is tuned to maximize correlation with the outcome, making PoPStat a maximum over 180 candidates. This is indeed the most serious threat to the paper's central claim. The paper's own Eq. 2 describes a brute-force optimization with no penalty or correction. The comparison to fixed indicators in §2.5 and the reported 'outperformance' are therefore not valid as evidence. A permutation or split-sample test would determine whether the observed correlations exceed what selection alone can produce. Because the paper provides no such correction, the conclusion that PoPStat outperforms traditional indicators is not established. The epidemiological narrative in the Discussion is plausible and consistent with existing literature, and the KL-divergence-based metric itself is a reasonable descriptive tool, but the statistical inference under the optimization is the load-bearing flaw. The verdict of REJECT is appropriate on this basis, and no additional concern is needed to change it.","tokens_in":14158,"tokens_out":2213,"duration_ms":27970,"concrete_test":"Run a permutation test: shuffle mortality rates across countries (holding population pyramids fixed), and for each shuffled dataset recompute the maximum over the 180 references in Eq. 2. Build the null distribution of max correlations. If the observed PoPStat for a disease (e.g., -0.84 for NCDs) falls within the upper tail of this null, the reported association is explainable by selection alone. Additionally, perform a training/test split: choose the reference on a random half of countries, compute PoPStat on the held-out half, and compare it to the fixed indicators (median age, GDP, HDI) on that same held-out half. If the out-of-sample PoPStat no longer beats the fixed indicators, the central outperformance claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that PoPStat outperforms median age, GDP, and HDI rests on Eq. 2, where the reference country is chosen by maximizing the correlation with the outcome, ln S. For each disease, PoPStat is therefore the maximum of 180 Pearson correlations, one per candidate reference. This is a form of selection on the dependent variable: the reference pyramid is selected because it makes the association with that specific disease mortality as large as possible. The comparison in §2.5 is then unfair, because median age, GDP, and HDI are fixed indicators with no analogous optimization. Under the null of no true association, the expected maximum of 180 correlations is non-negligible (about 0.24 for n=180 if correlations were independent; they are not, so the effective null is more complex). Thus the reported values (e.g., -0.84 for NCDs, 0.50 for CMNN, 0.29 for injuries) include an inflation term from selection. The nominal p-values in Tables 1–2 are not corrected for this multiplicity and therefore overstate significance. The paper itself acknowledges the brute-force search but never adjusts for its consequences. Without a selection-aware correction or out-of-sample validation, the outperformance claim collapses because a maximum over 180 candidates is not directly comparable to a single fixed indicator.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces two metrics, PoPDivergence (a KL-divergence between a country's population pyramid and a reference pyramid) and PoPStat (the Pearson correlation between PoPDivergence and the natural log of disease-specific mortality). The reference pyramid is selected per disease by brute-force optimization (Eq. 2) to maximize the correlation with mortality. Using GBD 2021 mortality data and UN WPP 2024 population data, the authors report PoPStat values for 371 diseases across countries and claim that PoPStat outperforms median age, GDP per capita, and HDI in explaining mortality for most diseases. They also interpret the signs of PoPStat in terms of expansive versus constrictive pyramids and link their findings to the epidemiological transition model.","tokens_in":14403,"tokens_out":4211,"duration_ms":46931,"significance":"If the central claim were valid, a scalar population-structure metric that outperforms standard demographic and development indicators would be a useful tool for mortality forecasting and health-policy targeting. The paper draws on comprehensive public datasets and provides code and scripts on GitHub, which is a strength for reproducibility. However, the statistical validity of the headline claim is undermined by the reference-selection procedure, which fits the metric to the outcome. The paper would be valuable if the authors could demonstrate, through out-of-sample validation or selection-corrected inference, that PoPStat retains predictive superiority over fixed indicators.","major_comments":[{"comment":"The reference population is chosen by maximizing the correlation between PoPDivergence and log mortality for each disease. Thus the reported PoPStat is the maximum over roughly 180 candidate correlations. Under a null hypothesis of no true association, the maximum of many noisy correlations is expected to be far from zero, so the reported magnitudes (e.g., –0.846 for NCDs in Table 1) are inflated by selection. The p-values and confidence intervals in Tables 1–2 are nominal and do not account for the arg max over references; they are therefore not valid for the optimized PoPStat. The authors must provide a selection-corrected null distribution (e.g., a permutation test that repeats the reference optimization) or out-of-sample correlation estimates before any claim about the strength of association can be accepted.","section":"§2.3, Eq. (2)"},{"comment":"The comparison with median age, GDP per capita, and HDI is not balanced because PoPStat is allowed to select the best reference population per disease, whereas these comparator indicators are fixed with no analogous optimization. The observed 'outperformance' is at least partly a byproduct of overfitting the reference choice. To support the claim that PoPStat outperforms traditional indicators, the authors should evaluate all indicators out-of-sample: split the sample, choose the reference on a training set, and compute correlations on a validation set, while also permitting the comparators a comparable level of flexibility (e.g., transformations). Without such a procedure, the headline claim in the abstract is not established.","section":"§2.5 and Abstract"},{"comment":"There is an apparent inconsistency between the stated objective, arg max_{ω} Cor[ln S, PoPDivergence(ω)], and the reported negative PoPStat values (e.g., –0.846 for NCDs with Japan as reference). If the objective is to maximize the signed correlation, the selected reference should yield the largest positive correlation; if instead the objective is to maximize the absolute correlation, then the sign of PoPStat is an artifact of the chosen reference and cannot be directly interpreted as evidence that mortality is concentrated in constrictive or expansive pyramids. The paper should clarify which objective is used and, accordingly, temper the interpretive statements in §4.3 about 'constrictive' versus 'expansive' burden.","section":"§2.3, Eq. (2) and Table 1"},{"comment":"The reported 'explained variance' values, such as 81.4% for NCDs, are presented as if they were the R-squared of a pre-specified model. Because the reference is selected to maximize the correlation, these values are in-sample fitting results rather than unbiased estimates of explanatory power. The authors should either use cross-validated R-squared or explicitly label these as in-sample descriptive quantities without inferential claims about the proportion of mortality variation explained by population structure.","section":"§3 and Table 1"}],"minor_comments":[{"comment":"The abstract states that the metrics were applied across 204 countries, while the Methods section reports mortality data for 180 countries. This inconsistency should be reconciled.","section":"Abstract and §2.1"},{"comment":"PoPDivergence is defined as a KL divergence, which is asymmetric in its arguments. The paper does not discuss how the choice of reference affects the distribution of PoPDivergence values across countries, although this is relevant to the interpretability of the metric's direction.","section":"Eq. (1)"},{"comment":"The p-value for respiratory infections and tuberculosis is reported as 0.055 in the text, which is not significant at conventional levels. The paper should state explicitly that this association is not statistically significant (after multiple-comparison considerations) or refrain from discussing it as a 'weak association' without this caveat.","section":"Table 1"},{"comment":"The phrase 'PoPStat of (0.291, Singapore, < 0.001)' omits the 'p' before the p-value; this format should be made consistent throughout.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The core statistical issue—reference selection by arg max—is genuine and load-bearing. I recommend major revision rather than outright rejection because the problem is addressable with out-of-sample validation or a selection-corrected permutation test. If the authors can show that PoPStat retains its predictive advantage under such a procedure, the paper could make a meaningful contribution. If they cannot, the outperformance claim should be withdrawn. The availability of code and the use of comprehensive public data are strengths that would make a corrected version worth publishing."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the headline: the key claim that PoPStat outperforms median age, GDP, and HDI does not survive contact with the method. Eq. 2 chooses the reference country by maximizing the correlation with that disease's mortality. So for each disease, PoPStat is the maximum of 180 Pearson correlations. That's selection on the dependent variable. The p-values and CIs in Tables 1 and 2 are computed as if no selection happened. Under the null, the expected maximum of 180 correlations is around 0.24 if they were independent, and they're not. So values like 0.29 for injuries are within the null range. The comparison with fixed indicators is also unfair: median age, GDP, and HDI don't get an optimization step. The 'outperformance' is at least partly a mathematical artifact.\n\nWhat's genuinely new: condensing the age-sex distribution into a KL divergence from a reference is a reasonable modeling idea. The authors apply it cleanly to 371 diseases, use public GBD and UN data, and provide code. The epidemiological discussion is well-read, and the qualitative patterns they describe—NCDs concentrated in constrictive pyramids, maternal/neonatal in expansive ones—are consistent with the literature. That part reads fine.\n\nThe soft spots are the central ones. Besides the selection problem, the paper doesn't adjust for the fact that the reference country is fitted per disease. The reported correlations, especially the strong negative ones for NCDs, are upper bounds, not unbiased estimates. There's also a minor point: r^2 values are described as 'explaining' variance, which is standard shorthand but still overstates when the correlation is fitted. No correction, no out-of-sample validation.\n\nWho gets value? Demographers and epidemiologists might want to build on the KL-divergence idea. But as presented, the evidence doesn't support the conclusions. I'd send it to peer review only with a strong request to fix the selection problem—pre-registered references, cross-validation, or selection-adjusted inference. If the authors can show the metric works out-of-sample, it could be a useful tool. As it stands, I wouldn't cite it.","headline":"PoPStat's outperformance claim is an artifact of brute-force reference selection; the underlying KL-divergence idea is reasonable but the evidence is fitted.","tokens_in":14953,"tokens_out":2887,"would_cite":false,"duration_ms":34734,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a single scalar measure of population-pyramid shape, PoPStat, explains disease-specific mortality better than median age, GDP per capita, or the Human Development Index for most of the 371 diseases studied.","keywords":["population pyramids","PoPStat","PoPDivergence","demographic transition","epidemiological transition","disease-specific mortality","Kullback-Leibler divergence","reference country tuning"],"falsifier":"Compute PoPStat with the reference country fixed before seeing mortality data, or use a holdout set of countries excluded from the reference search, and check whether the correlation remains as high as reported; alternatively, permute mortality labels across countries to build a null distribution for the maximum-of-180 correlation.","tokens_in":13986,"feed_emoji":"📊","tokens_out":10799,"duration_ms":96729,"temperature":0.7,"pith_summary":"This paper introduces a single number, PoPStat, meant to capture how tightly a country's population pyramid shape is bound to death rates from a given disease. The authors compress each country's age-sex distribution into a divergence distance from a reference pyramid, then correlate those distances with logged mortality across countries. They report that this score beats standard indicators such as median age, GDP per capita, and the Human Development Index for most of the 371 diseases examined. If correct, the metric would give health planners a demographic early-warning signal and would provide quantitative support for the idea that disease burdens shift in a predictable way as populations age.","feed_headline":"A pyramid-shape score beats GDP, median age, and HDI on mortality","feed_subtitle":"One pyramid-shape distance from a reference country condenses age-sex structure into a mortality predictor.","key_machinery":"The central object is PoPDivergence, defined as the Kullback-Leibler divergence between a country's age-sex distribution $P$ and a reference distribution $Q$: $D_{\\mathrm{KL}}(P \\| Q) = \\sum_i P(i) \\log(P(i)/Q(i))$. PoPStat is then the Pearson correlation between these divergence values and the natural log of cause-specific death rates. The reference $Q$ is not fixed a priori: a brute-force search over all 180 candidate countries selects the reference that maximizes that correlation for each disease. This reference tuning lets a single scalar order countries along the demographic transition, but it also means the reported correlation is the maximum of 180 candidate correlations rather than an independently specified one.","core_discovery":"On the paper's own terms, the central discovery is that the full age-sex shape of a population, rather than its median age or wealth, carries much of the demographic signal in disease-specific mortality. For non-communicable diseases the association is strong and negative, with a PoPStat of $-0.84$ using Japan as the optimized reference: mortality concentrates in constrictive, aging pyramids. Communicable, maternal, neonatal, and nutritional diseases show a moderate positive association ($0.50$, Singapore reference), with the burden in expansive, young pyramids, while injuries are only weakly tied to pyramid shape ($0.29$). The metric also identifies causes that are demographically 'free,' such as diabetes, respiratory infections, cirrhosis, self-harm, and interpersonal violence, all of which show weak PoPStat values. The authors interpret the overall pattern as empirical confirmation of the epidemiological transition model, with the optimized reference country marking the demographic archetype that best orders all countries along the transition.","pith_inferences":["Because the reference country is chosen by maximizing the correlation on the same data, the reported PoPStat values are probably optimistic; a permutation test or cross-validated reference choice would give unbiased estimates of how much population shape actually explains.","A direct test would fix the reference country from one disease (for instance, Japan for NCDs) and apply it to another cause or a later year, checking whether the pre-specified correlation holds out of sample.","The same divergence machinery could be applied to other population-level exposures, such as income distributions, education structures, or urban-rural splits, to see whether a single shape scalar is generally more informative than summary statistics.","For countries with missing death data, PoPStat suggests an imputation strategy based on pyramid shape alone, but any prediction interval must account for the selection bias in the reference choice."],"forward_implications":["Non-communicable disease mortality is largely a demographic phenomenon: aging, constrictive pyramids carry the burden, so countries mid-transition can expect rising NCD loads as their pyramids constrict.","Communicable, maternal, and neonatal mortality tracks expansive pyramids, so population structure alone flags where these burdens concentrate.","Because PoPDivergence orders countries along the demographic transition in a single cross-section, it offers a shortcut for studying demographic and epidemiological change without decades of longitudinal data.","The same divergence score can be correlated with any social, behavioral, or economic outcome, extending the metric beyond mortality.","Diseases with weak PoPStat values, such as diabetes, respiratory infections, and injuries, are exactly where demographic structure explains little and policy must look to other determinants."],"supporting_citations":[{"why":"Frames the epidemiological transition model that the paper's results are claimed to support.","marker":"[15]"},{"why":"A systematic review of epidemiological transition theory whose reported discordance with evidence this paper argues its findings address.","marker":"[19]"},{"why":"Provides the multi-country cause-specific mortality data used as the outcome variable.","marker":"[31]"},{"why":"Defines the expansive, stationary, and constrictive pyramid types used to label population structures.","marker":"[13]"},{"why":"Supplies the classic observation of a curvilinear income-mortality relationship, motivating the comparison against GDP.","marker":"[17]"}],"fun_headline_variants":["Pyramid-shape metric beats GDP, age, HDI on mortality","PoPStat: age-sex shape outranks classic health indicators","New metric links population pyramid shape to disease mortality","Age structure score outperforms GDP and HDI for most diseases"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reference population is chosen by brute-force optimization to maximize the correlation with mortality, so PoPStat is a maximum of many candidate correlations and its reported strength is not an independent estimate.","fun_headline_variants_meta":{"raw":{"variants":["Pyramid-shape metric beats GDP, age, HDI on mortality","PoPStat: age-sex shape outranks classic health indicators","New metric links population pyramid shape to disease mortality","Age structure score outperforms GDP and HDI for most diseases"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000253,"raw_usage":{"total_tokens":1618,"prompt_tokens":1055,"completion_tokens":563,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":671,"completion_tokens_details":{"reasoning_tokens":492}},"tokens_in":671,"tokens_out":563,"duration_ms":6782,"temperature":1.0,"reasoning_tokens":492,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T18:08:58.610593+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute PoPStat with the reference country fixed before seeing mortality data, or use a holdout set of countries excluded from the reference search, and check whether the correlation remains as high as reported; alternatively, permute mortality labels across countries to build a null distribution for the maximum-of-180 correlation.","supporting_citations":[{"cited_title":"The Epidemiologic Transition: A Theory of the Epidemiology of Population Change","cited_arxiv_id":null,"evidence_quote":"Frames the epidemiological transition model that the paper's results are claimed to support."},{"cited_title":"The Global Burden of Unintentional Injuries and an Agenda for Progress","cited_arxiv_id":null,"evidence_quote":"A systematic review of epidemiological transition theory whose reported discordance with evidence this paper argues its findings address."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the multi-country cause-specific mortality data used as the outcome variable."},{"cited_title":"Types and Significance of Population Pyramids","cited_arxiv_id":null,"evidence_quote":"Defines the expansive, stationary, and constrictive pyramid types used to label population structures."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the classic observation of a curvilinear income-mortality relationship, motivating the comparison against GDP."}],"review_version":1}