{"id":"580b992e-adda-43e1-bb3f-c7e855b0211a","arxiv_id":"2506.23158","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"An eight-variable, weight-free frailty index built from hospital and prescription records predicts death, disability, dementia, and other adverse outcomes in Italian adults over 65.","lead":"Researchers built a frailty score for older adults using only routine health records, combining eight pieces of information without giving any of them a statistical weight. It is meant to let health agencies quickly identify the most vulnerable elderly people for targeted care and planning.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The across-place validation is asserted but not reported: Piedmont results appear only as spider charts, while the reported 2018/2019 AUCs come from selection-overfit, largely overlapping cohorts; Table 8 also contains identical entries for death and hospitalisation, so the cross-population claim…","rationale":"The reader's weakest_assumption focuses on criterion validity: whether the six adverse outcomes adequately represent the latent construct of frailty. That is a real conceptual concern, but it is inherited from the standard criterion-validity framework the paper explicitly adopts (Rockwood [18]) and is not easily testable with the available administrative data. I find a more concrete and more immediately load-bearing problem in the evidence for the 'across time and place' component of the central claim. The only external place validation is mentioned in one sentence and a figure that reports no numerical results. The temporal validation uses a largely overlapping cohort with variables fixed from the primary analysis, and the main Table 8 contains an apparent transcription error for death and hospitalisation. Because the paper explicitly claims regeneration across populations as a design goal and lists external validity as a strength, the absence of any reproducible external result means the central claim is currently under-supported rather than demonstrated. This does not mean the method is wrong; it means the evidence base is incomplete, which is exactly the kind of issue that warrants a conditional rather than an accept verdict. I therefore agree with the reader's conditional verdict and recommend no change, though my emphasis is on external validation and reporting integrity rather than on the construct-validity assumption. A single concrete test—obtaining and independently checking the Piedmont AUCs—would settle whether the cross-population claim holds.","tokens_in":21330,"tokens_out":7991,"duration_ms":98439,"concrete_test":"Contact the authors and the Piedmont Regional Epidemiology Service to obtain the exact AUC values, sample sizes, and variable-algorithm definitions behind Figure 4, then independently recompute the FI on Piedmont data using the published eight algorithms and compare per-outcome AUCs (with DeLong CIs) against Table 8. If the Piedmont AUCs are unavailable, are computed with different variable definitions, or fall materially outside the Padua ranges for death and high-priority ER access, the central claim of regeneration across populations is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's stated goal (aim 4) is that the Frailty Index 'regenerates when applied to different populations,' and the abstract claims performance 'across time and place.' The fully reported evidence does not establish this. The 2018 AUCs are in-sample: the eight variables were selected to maximize the mean AUC on the same 2018 outcome year (Methods, 'Identification of the variables that compose the Frailty Index'). The 2019 cohort is a temporal replication, but it shares 205,004 of 213,689/216,757 subjects with the 2018 cohort, and the variable set was fixed from the 2018 analysis, so it is not an independent population in the sense required for cross-population regeneration. The only independent external validation is the Piedmont analysis mentioned in Results ('Reproducibility'), but the main text reports only that 'the results obtained were consistent' and refers to Figure 4 as 'spider charts'; no AUC values, cohort definitions, or variable algorithms from Piedmont are given. This is a black-box assertion, not a reproducible validation. In addition, Table 8 reports identical AUCs and confidence intervals for death and hospitalisation in both cohorts (0.854 and 0.664 with identical 95% CIs), which is implausible for cohorts of 213,689 and 216,757 and indicates a reporting error in the main validation table. Finally, two FI components (disability and total number of hospitalisations) are also counted as validation outcomes in the prevalence-based Tables 4, 5, and 7, making part of the descriptive validation tautological. The load-bearing condition for the central claim—that the index predicts frailty outcomes across independent populations—is therefore not currently supported by the reported evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a Frailty Index (FI) for adults aged 65 and older built from Italian administrative health data. Candidate variables are reduced from 75 through prevalence filters, protective-effect exclusions, stepwise logistic regression on balanced subsamples, and a forward POSET-based aggregation step that maximizes the mean AUC over six outcomes (death, femur fracture, hospitalisation, disability onset, dementia onset, and high-priority emergency-room access). The resulting FI comprises eight variables (age, disability, total number of hospitalisations, mental disorders, neurological diseases, heart failure, kidney failure, cancer) and is aggregated without weights using POSET average ranks normalized to [0,1]. The paper reports AUCs for the six outcomes in 2018 and 2019 cohorts, robustness of variable selection, associations with chronic diseases, comorbidity, and socioeconomic deprivation, and an external application in Piedmont. The central claims are that the index is parsimonious, weight-free, based only on routine data, and regenerates across time and place.","tokens_in":21630,"tokens_out":7093,"duration_ms":72162,"significance":"If the claims are fully supported, the paper makes a useful contribution to the electronic-frailty-index literature by showing that a parsimonious, weight-free index based on eight routinely collected administrative variables can stratify older populations and predict multiple adverse outcomes. Strengths include the explicit avoidance of regression weights, the use of multiple outcomes in variable selection, an out-of-time 2019 cohort, extensive robustness analyses of variable selection, and the provision of algorithms for the eight FI components in the Supplementary Material. However, as detailed below, the reported evidence does not yet establish cross-population regeneration, the main validation table contains an apparent reporting error, and part of the descriptive validation is circular because two FI components are also counted as outcomes. With these issues corrected, the approach could be a valuable addition to the frailty measurement toolkit.","major_comments":[{"comment":"The 2018 and 2019 columns report exactly the same AUC and 95% CI for Death (0.854, 0.850–0.858) and Hospitalisation (0.664, 0.661–0.667). With cohorts of 213,689 and 216,757 subjects sharing 205,004 individuals, identical intervals to three decimal places is effectively impossible. This indicates a reporting error; the temporal comparison in this table must be rerun and corrected, and the statement that 'the only significantly different AUCs are related to the outcome of disability onset' needs to be re-assessed after recomputation.","section":"Results, 'Reproducibility', Table 8"},{"comment":"The 2018 AUCs are in-sample because the eight variables were selected on the same 2016–2017 predictors and 2018 outcomes using a mean-AUC criterion. The 2019 cohort is not an independent test of cross-population regeneration: 205,004 of 213,689/216,757 subjects are common to both cohorts, the variable set was fixed by the 2018 analysis, and the paper's only external validation (Piedmont, Figure 4) reports no AUCs, cohort definitions, or variable algorithms. The abstract's 'across time and place' claim and Aim 4 ('regenerates when applied to different populations') therefore exceed the reported evidence. The authors should either provide full quantitative results for the Piedmont analysis (cohort definition, outcome definitions, AUCs with confidence intervals, and variable algorithms) or weaken the claims to temporal replication with a largely overlapping cohort.","section":"Methods 'Identification of the variables' and Results 'Reproducibility'"},{"comment":"Disability and total number of hospitalisations are components of the FI (Methods, 'Index construction') and are also counted as outcomes in their prevalence form in these tables. The strong gradients, such as 70.11% disability prevalence in the highest quartile (Table 4) and 97.36% in the top 1% (Table 7), are partly by construction. The text explicitly switches from incidence to prevalence 'to adequately represent those who are already disabled', but this makes the descriptive validation of these two outcomes non-independent. Please report the incidence-only versions of these tables or exclude input variables from the outcome definitions.","section":"Results, Tables 4, 5, and 7"},{"comment":"The six outcomes define frailty by assumption; they were selected via a literature review and factor/graphical analyses that are only briefly described, and the variable selection step optimizes prediction of these same outcomes. The FI is therefore, by construction, a risk score for these six events. The paper correctly notes that all administrative-data frailty measures follow criterion validity, but the reader is given no external anchor (for example, a subsample with a Fried phenotype or a deficit-index FI) to assess whether the score captures frailty rather than a bundle of healthcare-use risks. This issue should be discussed explicitly, or a small validation sub-study should be added.","section":"Methods, 'Choice of adverse outcomes'"}],"minor_comments":[{"comment":"The forward algorithm is incompletely specified: Step 1 says 'the two variables are chosen' but does not state how all pairs are screened, and the stopping rule 'none of the remaining variables leads to further improvement' has no numerical threshold. Please specify the exact criterion and the set of candidate pairs.","section":"Methods, 'Index construction with POSET'"},{"comment":"The reduction from 75 candidate variables to 47 after prevalence and protective-effect exclusions, and then to 15 after stepwise regression, is not reported in detail. Please provide the excluded variables and the stepwise model details, including entry/exit criteria and the number of balanced samples used.","section":"Methods, 'Identification of the variables'"},{"comment":"Algorithms are provided for the eight FI variables only; the other 67 candidate variable algorithms are 'available upon request'. For reproducibility and independent validation, the full set of variable algorithms should be published or made publicly accessible.","section":"Supplementary Materials, Table S.1"},{"comment":"There is inconsistent terminology: 'nervous system diseases' appears in the abstract and Supplementary Table S.1, while 'neurological diseases' is used elsewhere in the text. Please align the terminology throughout.","section":"Abstract and Table S.1"},{"comment":"The claim that 'the only significantly different AUCs are related to the outcome of disability onset' is not accompanied by p-values or an adjustment for multiple comparisons; adding these would strengthen the cross-cohort comparison.","section":"Results, Table 8"},{"comment":"The statement that 'the percentage of those who change the value of the FI ... is 0.03%' refers to a sensitivity analysis that is not described in the Methods. Please provide the method and, ideally, the result in the main text.","section":"Discussion, 'Strengths and limitations'"},{"comment":"Figure 4 (spider charts) cannot be evaluated without numeric values. Please also provide the underlying table of AUCs and confidence intervals for the Piedmont analysis, along with cohort and outcome definitions.","section":"Results, 'Reproducibility'"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is honest about the in-sample nature of the 2018 AUCs in one sentence, but the abstract and the stated Aim 4 claim cross-population regeneration that the reported evidence does not support. The Table 8 duplicate entries are almost certainly a copy-paste or computation error, and the reproducibility section relies on an unreported Piedmont analysis. These issues, together with the circularity in Tables 4, 5, and 7, require substantive revision before the paper can be considered for publication. If the Piedmont results cannot be supplied in quantitative detail, the claims should be restricted to temporal replication with overlapping cohorts."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a legitimate incremental paper. The POSET framework is their own prior work, but the forward selection to eight variables, the 2019 out-of-time replication, and the robustness resampling are real additions. If you need a weight-free, eight-variable frailty index from administrative data, this is a sensible option with credible AUCs.\n\nThe paper is clearly written and the methods are reproducible in principle: variable algorithms for the eight FI components are in the supplementary, and the stepwise selection on balanced subsamples is described well enough to re-run. The 2019 cohort is a genuine temporal validation, even though it overlaps heavily with 2018 (205,004 of roughly 215k subjects). The robustness analyses (leave-one-fold-out, resampling at each iteration) show the selected variable set is stable. The descriptive results—quartile gradients, chronic disease associations, deprivation gradient—are consistent and supportive.\n\nThe soft spots are real but mostly fixable. The 2018 AUCs are in-sample: the same outcomes were used to select the variables and to evaluate them. That is acknowledged in one sentence but not treated as a limitation. The 2019 AUCs are cleaner, but they are not an independent population. The paper's fourth aim—'regenerates when applied to different populations'—rests on the Piedmont analysis, and the text only says results were consistent and shows spider charts. No AUCs, no cohort definition, no variable algorithms for Piedmont. That is a black-box assertion, not a validation. The authors need to report those numbers or soften the claim.\n\nTable 8 has identical AUCs and confidence intervals for death and hospitalisation in both cohorts (0.854 and 0.664 with identical 95% CIs). For cohorts of 200k+ with slightly different compositions, identical confidence intervals to three decimals are implausible. This is likely a transcription or reporting error, but it sits in the main validation table, so it has to be fixed. Also, disability and hospitalisation are both FI components and counted as outcomes in the prevalence tables, so part of the descriptive validation is circular by construction.\n\nThe paper does not compare against existing electronic frailty indices, so its 'good performance' is not anchored to a benchmark. That is a missed opportunity, not a fatal flaw. Code and data are not public, but the administrative data are not shareable, so that is understandable.\n\nWho is this for? Health service researchers and local health authorities who want a simple, weight-free index from routine data. It deserves a serious referee, but only with the Piedmont results reported, Table 8 corrected, and the in-sample selection issue discussed. I'd take it in for revision rather than desk-reject.","headline":"A useful incremental frailty index with a solid out-of-time check, but the cross-population claim rests on unreported Piedmont results and a likely reporting error in Table 8.","tokens_in":22234,"tokens_out":2100,"would_cite":true,"duration_ms":21699,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62P10","06A06","62H30"],"pacs":[],"model":"deepseek-v4-flash","headline":"Eight routine health-record variables, aggregated by ranking instead of weights, predict death, dementia, disability, femur fracture, and high-priority emergency admission in older adults, with hospitalisation the weak spot.","keywords":["Frailty","Administrative data","Population health","Partially Ordered Sets","Risk stratification","Frailty Index","Average Rank","Adverse outcomes"],"falsifier":"Take a new cohort with the same administrative variables and a face-to-face clinical frailty assessment, then ask whether the top decile of the eight-variable POSET ranking overlaps substantially with the clinically frail group; if the overlap is no better than chance, the index is measuring administrative outcomes, not frailty.","tokens_in":21078,"feed_emoji":"🩺","tokens_out":13201,"duration_ms":128062,"temperature":0.7,"pith_summary":"Frailty is a hidden vulnerability that shows up as a higher chance of bad health events, so it can be measured indirectly through the outcomes it predicts. The paper tries to build a frailty score that a local health authority could compute from administrative records it already holds, with no surveys and no clinical exams. It claims that eight variables—age, disability, number of hospitalisations, mental disorders, neurological diseases, heart failure, kidney failure, and cancer—are enough, and that ranking people by their combination of these variables, instead of weighting them, produces an index that predicts death, dementia onset, disability onset, femur fracture, and high-priority emergency admission in older populations. If the claim holds, health planners get a cheap, portable way to identify who is most frail and to act before the worst outcomes occur.","feed_headline":"Eight health-data variables flag frailest older adults","feed_subtitle":"A ranking-based score from routine health data predicts frailty's worst outcomes with no weights or extra exams.","key_machinery":"The engine of the index is the Average Rank from partially ordered set (POSET) theory. A person's profile is the vector of values on the eight variables; one profile dominates another if it is no better on any variable and strictly worse on at least one. Each profile's Average Rank is its normalised position in the partial order of all profiles observed in the population, so a higher rank means the profile is dominated by fewer people and dominates more people. This mechanism does the aggregation without weights or regression coefficients, and because the order is recomputed from the profiles present in each new population, the index regenerates instead of carrying fixed coefficients. The forward selection procedure—choosing the first two variables that maximise mean AUC, then adding variables while they improve mean AUC—uses this same ranking mechanism to decide which variables belong in the final set.","core_discovery":"The central claim is that a valid frailty index can be made from eight variables in routine administrative data by ordering people rather than scoring them. The paper starts with 75 candidate frailty markers, reduces them to 15 by repeated stepwise logistic regressions across six adverse outcomes, and then applies a forward partially ordered set (POSET) selection that adds a variable only when it raises the mean area under the ROC curve over all six outcomes. The final index assigns each person the Average Rank of their profile in the partial order of all observed profiles; profiles dominate when they are no better on any variable and strictly worse on at least one. In two cohorts of more than 200,000 adults aged 65 and older, the eight-variable index predicts death with AUC 0.854, high-priority emergency access around 0.805–0.812, dementia onset around 0.805–0.806, disability onset 0.749–0.792, and femur fracture 0.758–0.765, while hospitalisation trails at 0.664, which the authors attribute to hospitalisation being a less specific event. The same eight variables are selected in the 2019 cohort and under resampling, and the authors report similar performance in another Italian regional population.","pith_inferences":["Beyond the paper's six outcomes, one would expect the same POSET ranking to order other stress-related events, such as falls, institutionalisation, or post-operative complications, because the index is not tuned to a single endpoint.","The portability of the eight variables depends on coding conventions; the paper itself notes that different algorithms for identifying a disease from administrative flows can change who is counted as affected, so regions with different coding practices may need to re-derive the variable definitions before comparing FI values.","A natural extension is to attach confidence intervals to the Average Ranks; the paper reports that this is not yet available, and without it one cannot formally test whether a change in an individual's frailty over time is real or a by-product of a changing population profile.","Because variable selection was guided by the six outcomes, the index is best understood as a measure of health-frailty risk as those outcomes define it; if a health authority's priority outcome differs, the eight-variable set might have to be re-selected."],"forward_implications":["A local health authority could compute the index from hospital discharge records, drug claims, ticket exemptions, home care registries, psychiatric services, and emergency-room data, without collecting any new information from patients.","The index is stable enough for monitoring: the same eight variables were selected in the 2019 cohort, and Frailty Index values for people present in both cohorts correlate at 0.88, with near-perfect stability for people whose profile did not change.","Targeted action is possible: among the most frail 1% of the 2018 cohort, 97.4% had disability, 51.7% were hospitalised, and 34.8% died in the following year, so the top of the ranking is a small group with very high event rates.","The index's weak spot is hospitalisation, which the authors say is too nonspecific to serve as a frailty signal; users who care about admissions should not expect this index to separate them well.","The index can be applied in a new population without re-estimating regression weights; the only carried-over assumption is that the same eight variables define the profiles."],"supporting_citations":[{"why":"Introduces the POSET average-rank method that the paper extends to construct a weight-free frailty index.","marker":"[31, 32]"},{"why":"Provides the criterion-validity rationale that lets the authors define frailty through prediction of adverse outcomes.","marker":"[18]"},{"why":"The literature review of more than one hundred studies from which the six representative outcomes were selected.","marker":"[33]"},{"why":"The nonparametric method used to calculate confidence intervals for the reported AUC values.","marker":"[34]"}],"fun_headline_variants":["Eight routine health variables predict frailty outcomes","Ranking, not scoring: eight admin data points flag frailty","Frailty index from routine data: just rank people, no weights","Eight administrative variables predict death, dementia, disability","No weights needed: routine data ranking spots frail seniors"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The index's validity rests on the assumption that the six chosen adverse events—death, high-priority emergency access, hospitalisation, disability onset, dementia onset, and femur fracture—together capture what frailty is; if they miss the core of frailty, the ranking measures risk of those events rather than frailty itself.","fun_headline_variants_meta":{"raw":{"variants":["Eight routine health variables predict frailty outcomes","Ranking, not scoring: eight admin data points flag frailty","Frailty index from routine data: just rank people, no weights","Eight administrative variables predict death, dementia, disability","No weights needed: routine data ranking spots frail seniors"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000832,"raw_usage":{"total_tokens":3717,"prompt_tokens":1112,"completion_tokens":2605,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":728,"completion_tokens_details":{"reasoning_tokens":2526}},"tokens_in":728,"tokens_out":2605,"duration_ms":18291,"temperature":1.0,"reasoning_tokens":2526,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:47:45.441079+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a new cohort with the same administrative variables and a face-to-face clinical frailty assessment, then ask whether the top decile of the eight-variable POSET ranking overlaps substantially with the clinically frail group; if the overlap is no better than chance, the index is measuring administrative outcomes, not frailty.","supporting_citations":[{"cited_title":"What would make a definition of frailty successful? Age Ageing","cited_arxiv_id":null,"evidence_quote":"Provides the criterion-validity rationale that lets the authors define frailty through prediction of adverse outcomes."},{"cited_title":"Cos’è la fragilità dell’anziano e come può essere identificata (What frailty in older adults is and how it can be identified)","cited_arxiv_id":null,"evidence_quote":"The literature review of more than one hundred studies from which the six representative outcomes were selected."}],"review_version":1}