{"id":"d3491ea5-f959-41fb-aa29-847234c54b06","arxiv_id":"2412.04275","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A scoping review categorizes 18 papers on methods to improve generalisability and transportability of clinical prediction models into data-driven and knowledge-driven approaches.","lead":"This paper reviews 18 studies on methods that help clinical prediction models keep working in new populations. It sorts them into data-driven and knowledge-driven approaches and says the next step is to compare them fairly.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central taxonomy depends on a title-based search with 80% recall on the authors' own validation set; missed synonym-rich literature could change the claimed two-family structure.","rationale":"The reader's weakest assumption is exactly where I land: the search strategy is the single load-bearing element because every subsequent claim—the two-family taxonomy, the tailored/not-tailored contrast, and the statement that no comparative evaluation exists—is relative to the included set. I agree with the reader rather than introducing a separate objection; the same evidence (the missed Bellamy et al. paper, the admitted exclusion of Debray et al. and transfer-learning work) supports a conditional verdict. The review is otherwise careful: it reports its search string, PRISMA flow, and limitations honestly, and the taxonomy is descriptive of the papers that were included. I would not reject the paper on this basis, because a scoping review may legitimately summarize a search-defined literature, but the conclusion overstates the field-level dichotomy unless search recall can be demonstrated on a broader gold standard. The '7 of 19' arithmetic slip should also be corrected but is not load-bearing. The proposed concrete test directly checks recall and taxonomy fit; if it passes, the concern is resolved and the conditional can be lifted.","tokens_in":20264,"tokens_out":4206,"duration_ms":44731,"concrete_test":"Construct a gold-standard reference set of 40–60 methodological papers on developing CPMs that transport or generalize, drawn from recent reviews of transfer learning, domain adaptation, invariant risk minimization, distributional robustness, and causal transportability (including Debray et al. [4] and Van Calster et al. [43]). Run the review's exact title-based search plus Geersing filters against this set, compute recall, and classify every missed paper into Table 1's categories. If recall falls well below the 4/5 observed on the authors' own seed papers, or if any missed paper does not fit the data-driven/knowledge-driven dichotomy, the central taxonomy is incomplete and the conclusion should be weakened to a provisional map of one search-defined subset.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The review's central claim—that methodologies aiding generalisability/transportability of CPMs fall into two distinct families, data-driven and knowledge-driven—inherits all of its coverage from a search restricted to titles containing generalisab*, generalizab*, transportab* or 'dataset shift' (Search Strategy section). This filter omits the standard vocabularies of large adjacent literatures: domain shift, domain generalization, distributional robustness, invariant risk minimization, transfer learning, and external validity. The paper itself concedes the point: Debray et al.'s framework for developing generalizable models was excluded because it did not match the keywords, and transfer-learning approaches were not captured by the search terms. The strongest internal evidence of incompleteness is the validation set: only four of the five known seed papers were retrieved, and Bellamy et al. was recovered only by manual addition after failing both database search and snowballing. A hand-picked set with 80% recall does not support the inference that the taxonomy is complete; the two-family structure may be an artifact of which terminologies the search happened to admit. Because the taxonomy, the data-requirement contrast, and the claimed absence of comparative evaluation are all stated relative to the included set, this search gap is load-bearing rather than cosmetic.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a scoping review of methodology for improving the generalisability and transportability of clinical prediction models (CPMs). The authors searched MEDLINE, Embase, medRxiv, and arXiv up to September 2023 using title-based search terms (generalisab*, generalizab*, transportab*, 'dataset shift') plus Geersing filters, screened 1,761 records, and included 18 papers. They categorize the identified methods into data-driven versus knowledge-driven families and, within each, whether the method is tailored to a specific target population. The paper summarises assumptions, advantages, disadvantages, and applications of each included method, and offers future research directions such as comparative evaluations on simulated and real data and integration of the two families.","tokens_in":20531,"tokens_out":3389,"duration_ms":37555,"significance":"If the taxonomy is accepted, this review provides a useful map of an emerging and fragmented literature, and it makes a plausible conceptual distinction between methods that exploit data heterogeneity or weighting to improve robustness and methods that rely on causal structure or domain knowledge. The review also clearly identifies the absence of comparative evaluations and the need for methods that require less target-population data. The paper's strengths include a documented search strategy, a PRISMA flow diagram, a structured extraction of assumptions and limitations in supplementary material, and an explicit discussion of search limitations. The main value is as a synthesis and agenda-setting review for clinical prediction methodology; its central claim, however, depends on the completeness and representativeness of the included literature.","major_comments":[{"comment":"The central claim that methodologies for aiding generalisability/transportability fall into two families is load-bearing, yet the evidence for completeness is weak. The search was restricted to titles containing generalisab*, generalizab*, transportab* or 'dataset shift', and the paper concedes that one of the five validation papers (Bellamy et al.) was not found by the search or snowballing, and that Debray et al.'s framework was excluded because it did not match keywords. This means the search has demonstrated imperfect recall even on a hand-picked set; it may have missed large adjacent literatures using terms such as domain shift, domain generalization, distributional robustness, invariant risk minimization, or external validity. The paper should either rerun the search with additional synonym-based terms and a broader validation set, or explicitly reframe the taxonomy as provisional and limited to the included papers rather than as a complete classification of the field.","section":"Search Strategy; Results; Discussion"},{"comment":"Screening and data extraction were conducted by a single author (KP) with no independent second reviewer or adjudication process. For a review whose core output is a classification and summary of included methods, single-reviewer screening is a substantial risk to reliability, as title/abstract decisions determine the inclusion set on which the taxonomy is built. The manuscript should report whether any reliability checks (e.g., dual screening on a subset, or verification of extraction by a second author) were performed, or justify why single-reviewer screening is appropriate for this scoping review.","section":"Information sources; Data extraction"},{"comment":"The exclusion of transfer learning and adjustment methods is not consistently applied. The scope states that methods that adjust an already-developed model are excluded, but density ratio/importance weighting methods (Gao et al., Steingrimsson et al., Dockès et al.) are included in Table 1 even though they are commonly classified as transfer-learning or domain-adaptation techniques. If these are included because they are used during development, the boundary should be stated more precisely; if they are excluded in other guises, the review risks including members of a family while claiming that family is out of scope, which weakens the exhaustiveness of the two-family taxonomy.","section":"Scope of review; Results; Table 1"},{"comment":"The manuscript states 'Fewer studies (7 of 19) focused on knowledge-driven methodologies', but the review includes 18 papers, not 19. More substantively, Piccininni et al. and Liu et al. are placed in both the data-driven and knowledge-driven categories, so the two families are not mutually exclusive. The paper should provide explicit criteria for assigning a method to one or both families, and should discuss whether the overlap indicates that the data-driven/knowledge-driven distinction is better viewed as a continuum or as complementary axes rather than a clean dichotomy.","section":"Knowledge-driven methodology that are not tailored for target population; Table 1"}],"minor_comments":[{"comment":"The database name 'medRxiv' is misspelled as 'merRxiv' in the Information sources section.","section":"Information sources"},{"comment":"Table 1 contains stray timestamps '26/11/2024 12:10:00' in two cells; these artifacts should be removed.","section":"Table 1"},{"comment":"The text says '7 of 19'; the correct total is 18 included papers. Also, 'Parent Child - PB' should likely be 'Parent-Child (PC)' for consistency.","section":"Knowledge-driven methodology that are not tailored for target population"},{"comment":"There is a typo, 'Causaility-aware', which should read 'Causality-aware'.","section":"Knowledge-driven methodology that is tailored for a specific target population"},{"comment":"The phrase 'To compliment the database search' should be 'To complement the database search'.","section":"Snowballing (Citation search)"}],"recommendation":"major_revision","confidential_remarks":"This is a useful and clearly written scoping review, but its central taxonomy is only as credible as the completeness of its search. The authors themselves document a validation failure and several known omissions. I would support publication after substantial revision that either broadens the search and validates recall against a larger set of known methods, or tempers the claims to describe the taxonomy of the included papers rather than the field. The single-reviewer screening process is also a concern that should be addressed or transparently mitigated. The topic is within the scope of stat.ME and would be of interest to methodologies working on clinical prediction models."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper is a scoping review of methods that aim to make clinical prediction models generalisable or transportable at the development stage, not by updating after deployment. The main contribution is a two-axis taxonomy: data-driven vs knowledge-driven, and tailored to a specific target population vs not. That is a useful map, and the paper does a decent job summarizing 18 papers, including assumptions and examples. The discussion of future comparative benchmarks is sensible. The authors are honest about limitations, which counts for a lot.\n\nWhere it gets shaky: the search strategy is narrow. Search terms were restricted to titles containing generalisab*/generalizab*/transportab*/dataset shift, plus Geersing filters. That omits large adjacent literatures—domain shift, domain generalization, distributional robustness, invariant risk minimization, transfer learning—and the authors know it. One of their five validation papers (Bellamy) was missed both by the database search and snowballing and only made it in because they added it manually. So the claim that the taxonomy reflects the structure of the field is weaker than the paper presents. The review effectively maps a subset of the literature, not the whole terrain. The stress-test note calling this load-bearing is right, though I would temper it: the two-family split (more data vs more causal knowledge) is plausible and probably robust, but the tailored/not-tailored axis and the absence of comparative evaluation are claims about the included set, not the full field.\n\nOther issues: only one author screened and extracted, which increases error risk. The paper cites \"7 of 19\" knowledge-driven papers, but only 18 papers were included; that is an internal inconsistency that needs fixing. It also excludes transfer learning as an \"adjustment\" while including density-ratio weighting, which is a transfer-learning technique—the line is fuzzy and should be explained.\n\nShould a serious editor send it to review? Yes. It is a genuinely useful synthesis for applied researchers and methodologists, the PRISMA flow is reported, and the limitations section is more candid than most. A revision should broaden the search or at least justify the exclusion of synonym-rich fields, correct the count, and consider a second screener. I would accept it with major revisions.","headline":"A useful but search-restricted taxonomy of development-stage methods for CPM generalisability and transportability; worth reviewing, but the completeness claims need major revision.","tokens_in":21002,"tokens_out":1915,"would_cite":true,"duration_ms":20148,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This scoping review claims that development-phase methods for generalisable or transportable clinical prediction models split into data-driven and knowledge-driven families, with no head-to-head evaluation yet.","keywords":["clinical prediction models","generalisability","transportability","dataset shift","causal inference","scoping review","density ratio weighting","selection diagram"],"falsifier":"A replication of the search with additional title and abstract terms such as 'domain adaptation', 'domain generalization', 'invariant learning', 'distribution shift', 'external validity', and 'transfer learning' in the same databases up to September 2023, followed by full-text screening against the review's inclusion and exclusion criteria, would test the taxonomy's completeness. If any resultant paper proposes a development-phase method for generalisable or transportable clinical prediction models that fits neither the data-driven nor the knowledge-driven category, or neither the tailored nor the not-tailored axis, the review's classification is incomplete.","tokens_in":20081,"feed_emoji":"🩺","tokens_out":8775,"duration_ms":75281,"temperature":0.7,"pith_summary":"This scoping review tries to establish that methodology for making clinical prediction models generalisable or transportable at the development stage falls into two distinct families: data-driven methods that enhance robustness through heterogeneous data, ensemble learning, or density-ratio weighting, and knowledge-driven methods that encode causal structure through selection diagrams, graph surgery, invariant predictors, or counterfactual adjustment. The review also classifies each method by whether it requires data from the specific target population. If the taxonomy is right, developers can choose an approach based on whether they have target covariates and causal knowledge, and researchers can see that the two families have not yet been compared on common benchmarks. The paper argues this matters because most models are developed without prioritising transportability, and a better map of methods would reduce the need for post-hoc updating.","feed_headline":"Review: methods for generalisable prediction models form two families","feed_subtitle":"Scoping review of 18 methods finds two families and a missing head-to-head comparison.","key_machinery":"The central analytic object is the two-way classification (data-driven vs knowledge-driven × not tailored vs tailored) constructed from the 18 included papers. Within that taxonomy, the load-bearing technical apparatus is the density ratio $r(x)=P_{target}(x)/P_{source}(x)$ for data-driven tailoring, and, for knowledge-driven methods, the selection diagram and its associated transport formula, e.g., $P(Y=1 \\mid X=x, S=0) = \\sum_z P(Y=1 \\mid X=x, S=1, Z=z)P(Z=z)$ under S-admissibility, together with the graph surgery estimator that turns interventional distributions into observable ones. These devices let a model be transported or made invariant to shifts without re-estimation on the target population.","core_discovery":"The paper's central claim is that the emerging literature on development-phase methods for generalisability and transportability of clinical prediction models can be organised by two independent axes: data-driven versus knowledge-driven, and whether the method is tailored to a specific target population. Data-driven methods include training on heterogeneous multi-centre data (internal-external cross-validation), ensemble models, density-ratio or importance weighting, adversarial validation, and measurement-error alignment; knowledge-driven methods include selection diagrams and transport formulas, graph surgery estimators, shortcut-predictor removal, distributionally robust and worst-case risk minimisation with mutable/immutable sets, Markov-blanket or parent-child predictor selection, and counterfactual adjustment for anticausal (diagnostic) tasks. The review states that no included study has comparatively evaluated these methods on shared simulated or real datasets, and it identifies integration of the two families as an open direction.","pith_inferences":["A direct test of the taxonomy would be to run the 18 methods on a common battery of simulated shifts (covariate, prior, concept, measurement error) and real EHR datasets; the paper's own call for such benchmarks implies that none exists.","Because one known seed paper (Bellamy et al.) escaped the search and snowballing, the review's map is probably incomplete; adding domain-adaptation and invariant-learning search terms could surface methods that fit neither family.","The tailored-versus-universal axis suggests a potential bridge to transfer learning: density-ratio weighting is already used in domain adaptation, so formal connections could unify the two literatures.","A head-to-head comparison might find that knowledge-driven methods degrade gracefully under correct causal assumptions but fail catastrophically when those assumptions are wrong, while data-driven methods trade off performance on any single population for robustness across many."],"forward_implications":["If the taxonomy is correct, a model developer can choose a method family based on whether covariate data from the target population is available and whether reliable causal knowledge exists.","Because the review finds that density-ratio weighting corrects covariate shift but not concept shift, these methods cannot rescue models when the predictor-outcome relationship itself changes.","The absence of comparative evaluation means that, as of the review's search date, no evidence indicates which method family performs best for a given type of dataset shift.","Knowledge-driven methods all rest on correctly specified causal structures, so misspecification of the selection diagram or the mutable/immutable split is a shared failure mode.","Integrating data-driven heterogeneity with causal invariant selection is identified by the paper as a promising direction for future methodology."],"supporting_citations":[{"why":"Bellamy et al. is one of the five seed papers used to define the search terms; its absence from search results motivated manual inclusion, exposing a gap in the search strategy.","marker":"[22]"},{"why":"Piccininni et al. provides the seed paper for DAG-based predictor selection and causal thinking in clinical risk prediction.","marker":"[23]"},{"why":"Subbaswamy et al. supplies the seed paper for the graph surgery estimator and shift-stable predictive models.","marker":"[19]"},{"why":"Fehr et al. is the seed paper for a causal framework for assessing transportability of clinical prediction models.","marker":"[20]"},{"why":"de Jong et al. is the seed paper for internal-external cross-validation and developing models from heterogeneous pooled data.","marker":"[24]"},{"why":"Pearl and Bareinboim provide the formal transportability formula and S-admissibility conditions that underpin the knowledge-driven methods.","marker":"[37]"},{"why":"Geersing et al. supplies the search filter for identifying prediction modelling studies, which the review combines with generalisability title terms.","marker":"[25]"}],"fun_headline_variants":["Prediction model methods split into two families, review finds","Scoping review: two families of generalisability methods","No head-to-head tests of prediction model transportability methods","18 methods for generalisable models: two axes, no comparisons","Review: data-driven vs knowledge-driven for model transportability"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The search strategy, based on titles containing generalisab*, generalizab*, transportab* or 'dataset shift' plus the Geersing filters, captures the full universe of relevant methodology papers, even though the review itself found that one known seed paper (Bellamy et al.) was missed by the search and only included manually.","fun_headline_variants_meta":{"raw":{"variants":["Prediction model methods split into two families, review finds","Scoping review: two families of generalisability methods","No head-to-head tests of prediction model transportability methods","18 methods for generalisable models: two axes, no comparisons","Review: data-driven vs knowledge-driven for model transportability"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000657,"raw_usage":{"total_tokens":3013,"prompt_tokens":956,"completion_tokens":2057,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":572,"completion_tokens_details":{"reasoning_tokens":1975}},"tokens_in":572,"tokens_out":2057,"duration_ms":15986,"temperature":1.0,"reasoning_tokens":1975,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T21:32:34.787241+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication of the search with additional title and abstract terms such as 'domain adaptation', 'domain generalization', 'invariant learning', 'distribution shift', 'external validity', and 'transfer learning' in the same databases up to September 2023, followed by full-text screening against the review's inclusion and exclusion criteria, would test the taxonomy's completeness. If any resultant paper proposes a development-phase method for generalisable or transportable clinical prediction models that fits neither the data-driven nor the knowledge-driven category, or neither the tailored nor the not-tailored axis, the review's classification is incomplete.","supporting_citations":[{"cited_title":"From development to deployment: dataset shift, causality, and shift - stable models in health AI","cited_arxiv_id":null,"evidence_quote":"Bellamy et al. is one of the five seed papers used to define the search terms; its absence from search results motivated manual inclusion, exposing a gap in the search strategy."},{"cited_title":"The Clinician and Dataset Shift in Artificial Intelligence","cited_arxiv_id":null,"evidence_quote":"Piccininni et al. provides the seed paper for DAG-based predictor selection and causal thinking in clinical risk prediction."},{"cited_title":"Evaluation of clinical prediction models (part 1): from development to external validation","cited_arxiv_id":null,"evidence_quote":"Subbaswamy et al. supplies the seed paper for the graph surgery estimator and shift-stable predictive models."},{"cited_title":"Evaluation of clinical prediction models (part 2): how to undertake an external validation study","cited_arxiv_id":null,"evidence_quote":"Fehr et al. is the seed paper for a causal framework for assessing transportability of clinical prediction models."},{"cited_title":"Prognostic models will be victims of their own success, unless…","cited_arxiv_id":null,"evidence_quote":"de Jong et al. is the seed paper for internal-external cross-validation and developing models from heterogeneous pooled data."},{"cited_title":"Directed acyclic graphs and causal thinking in clinical risk prediction modeling","cited_arxiv_id":null,"evidence_quote":"Pearl and Bareinboim provide the formal transportability formula and S-admissibility conditions that underpin the knowledge-driven methods."},{"cited_title":"A review of statistical updating methods for clinical prediction models","cited_arxiv_id":null,"evidence_quote":"Geersing et al. supplies the search filter for identifying prediction modelling studies, which the review combines with generalisability title terms."}],"review_version":1}