{"id":"08832d88-f8ed-4f46-9d22-fe721ac9da44","arxiv_id":"2602.17041","paper_version":3,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":1,"one_line_summary":"Marginal odds ratios and hazard ratios from MAIC/STC are population-specific even under the shared effect modifier assumption, so applying them to a decision-relevant population is an implicit transport step needing extra assumptions.","lead":"This paper shows that population-adjusted indirect treatment comparisons usually estimate effects for the comparator trial's patients, not for the patients a health system cares about, and that the shared-effect-modifier assumption does not fix that. It gives formulas and simulations showing when effects can and cannot be carried across populations, with direct implications for health technology assessment.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"SPFA is silently required for the paper's conditional-transportability results, so Section 6.4/Table 1 overstate what SEMA alone achieves.","rationale":"The reader's weakest_assumption identified the same structural reliance on SPFA in Eq. (A1). My analysis confirms this is the weakest point, but I assess it as more than a caveat: the paper's positive claim that conditional effects are directly transportable under SEMA and scale alignment is false without SPFA, and this omission appears in the main-text conditions (Section 6.4) and Table 1, not just the appendix. The central negative claim about marginal non-collapsible measures remains valid, so the paper should not be rejected; however, a conditional acceptance with revisions to explicitly include SPFA in all statements of conditional transportability is appropriate. The concrete simulation proposed would directly test whether the concern lands.","tokens_in":31075,"tokens_out":12135,"duration_ms":103151,"concrete_test":"Amend Eq. (A1) to allow treatment-specific prognostic functions: g(E[Y_t|X]) = m_t(x)+δ_t+φ_t(x)1{t≠A}, with SEMA (φ_B=φ_C=φ). Re-run the log-OR simulation of Section 7.2 with m_B(x)=β0+β1 x and m_C(x)=β0+β1 x + c x (c≠0), same φ(x), and target populations with varying μ_X. If the conditional log OR for B vs C varies with μ_X, Proposition A1(i)'s conclusion fails without SPFA, confirming that SPFA is a required condition for the conditional-transportability claims in Section 6.4 and Table 1.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The formal results Propositions A1–A2 and B1–B3 are built on Eq. (A1): g(E[Y_t|X]) = m(x)+δ_t+φ(x)1{t≠A}, with m(x) common across treatments (SPFA). The proof of Proposition A1(i) uses m(x) cancellation to conclude ΔCond_BC(x)=δ_B−δ_C. Without SPFA, if m_B(x)≠m_C(x), then ΔCond_BC(x)=(m_B(x)−m_C(x))+(δ_B−δ_C), which depends on x; the population-average conditional effect then varies across populations with different covariate distributions, even under SEMA and scale alignment. SEMA (same effect modifiers for B and C relative to A) does not imply m_B=m_C. Section 9.3 itself states SPFA is 'more restrictive and more difficult to justify empirically.' Yet Section 6.4 states 'population-average conditional effects are directly transportable if SEMA is met and effect modifiers and the effect measure both operate on the linear predictor scale—regardless of whether the measure is collapsible,' without listing SPFA. Table 1's column for ΔCond_BC transportability under SEMA likewise omits SPFA. Appendix A's note acknowledges SPFA but then claims 'the only structural assumptions required are additivity on the g-scale and sharing of φ(x) across active treatments under SEMA,' which is inconsistent with the model A1. This is load-bearing because the paper's decision framework encourages analysts to prefer conditional estimands as more transportable; if SPFA is also required, the guidance is incomplete and could lead to invalid transport when prognostic effects differ across active treatments. The central negative claim about marginal non-collapsible measures is unaffected; the concern is about the positive conditional-transportability results.","agreement_with_reader":"agree"},"referee_report":null,"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe paper is worth your time, but read the conditional-transportability claims with a grain of salt. The core message — that in anchored MAIC/STC the active-to-active effect is identified in the comparator population, and moving it to a decision population is a separate, assumption-laden transport step — is correct and important. The two-step decomposition is genuinely useful, and the simulations cleanly show that marginal ORs and HRs are not directly transportable even under SEMA. That negative result lands against current NICE and EU HTA guidance, and it is the main contribution.\n\nWhat is actually new: the explicit two-step formalization and the anchored active–active propositions (A1–A2, B1–B3) characterizing when conditional vs. marginal contrasts are directly transportable. The algebra checks out, and the illustrative examples are transparent population-level calculations with code in the supplement. The distinction between collapsibility and transportability in Appendix D is also well put.\n\nThe soft spot is real. The conditional-transportability proofs rest on Equation (A1), which assumes a single prognostic function m(x) common to all treatments — that is SPFA. Without SPFA, the conditional B-vs-C contrast at x is (m_B(x)-m_C(x)) + (δ_B−δ_C), which need not be constant in x. SEMA does not imply m_B=m_C. So Section 6.4 and Table 1 overstate what SEMA alone achieves for conditional estimands. Appendix A acknowledges the common m(x) as \"standard,\" but then claims the only structural assumptions are additivity and shared φ(x); that claim is inconsistent. This is not fatal for the central negative result about non-collapsible marginal measures — that holds regardless — but it is load-bearing for the paper's positive guidance that conditional estimands on the linear predictor scale are directly transportable under SEMA. The guidance needs a caveat listing SPFA as a required assumption, and the \"regardless of collapsibility\" phrasing should be qualified.\n\nNovelty is moderate: Remiro-Azócar and Phillippo et al. already established the population dependence of non-collapsible marginal measures. The paper's contribution is the explicit transport decomposition and the HTA decision framing, which is enough to be a useful methodological paper.\n\nWho it's for: statisticians and HTA methodologists who use or review MAIC/STC, and anyone writing guidance about indirect comparisons. It deserves a serious referee; the SPFA issue is fixable in revision.\n\nOverall: accept with revisions.","headline":"Useful two-step transportability frame and a correct warning about marginal OR/HR, but the conditional-transportability guidance quietly assumes SPFA and overstates what SEMA alone buys you.","tokens_in":31974,"tokens_out":3255,"would_cite":true,"duration_ms":28594,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Population-adjusted indirect comparisons identify comparator-population effects, and the shared effect modifier assumption alone does not make them transportable to other populations.","keywords":["transportability","population-adjusted indirect comparisons","matching-adjusted indirect comparison","simulated treatment comparison","collapsibility","shared effect modifier assumption","estimands","health technology assessment"],"falsifier":"Simulate an anchored comparison with SEMA and scale alignment holding by construction for a marginal odds ratio or hazard ratio, letting the index and comparator populations differ in the mean of an effect modifier; if the true marginal active-active effect is identical in every population, the paper's central claim is false, and if it varies, the claim is confirmed.","tokens_in":30937,"feed_emoji":"📊","tokens_out":5063,"duration_ms":45898,"temperature":0.7,"pith_summary":"Health technology assessment often relies on indirect comparisons when no head-to-head trials exist. The paper argues that the common adjustment methods—matching-adjusted and simulated treatment comparisons—only pin down the treatment effect in the comparator trial's population, not in any other population. It shows that the shared effect modifier assumption, widely cited as justifying transport of these effects, is not enough: for non-collapsible measures such as odds ratios and hazard ratios, the marginal effect changes with the covariate distribution even when the assumption holds. If accepted, this reframing changes how PAIC results should be reported and used in cost-effectiveness modeling: population-specific estimates need an explicit additional transport step before being applied to decision-relevant populations.","feed_headline":"Odds and hazard ratios stay tied to their trial population","feed_subtitle":"Even shared effect modifiers don't make marginal effects portable; only collapsible, scale-aligned measures travel.","key_machinery":"The central object is the estimand-based decomposition of a pairwise anchored indirect comparison into two transport steps: first, conditional transport of the index-trial effect to the comparator population; second, implicit direct transport of the resulting active-to-active effect to the decision population. The technical workhorse is the structural model g(E[Y_t|X]) = m(x) + δ_t + φ(x)1{t≠A}, with a shared prognostic function m and shared effect modification φ under SEMA. Propositions A1–A2 establish when contrasts are invariant under covariate-distribution shifts; Propositions B1–B3 establish contrast-induced direct collapsibility, for example the log risk ratio case where the active-to-","core_discovery":"The paper's central claim is that pairwise MAIC and STC identify active-to-active contrasts defined in the comparator population, and that these contrasts are not generally portable. SEMA ensures only that effect-modification terms cancel on a chosen scale; it does not make the marginal estimand invariant. Direct transportability of a marginal effect requires three joint conditions: SEMA; effect modification and the effect measure on the same linear predictor scale; and a collapsible measure. For conditional effects, SEMA plus scale alignment suffice. The paper proves this in formal propositions and demonstrates with simulations that mean differences and log risk ratios transport under SEMA,","pith_inferences":["Beyond the paper: the same two-step logic applies to any indirect comparison or network meta-analysis applied in an economic model; a marginal hazard ratio from an NMA is also population-specific, so the transport-bias concern is broader than MAIC/STC.","Beyond the paper: divergent MAIC results in different sponsors' submissions may be explained as different estimands anchored to different comparator populations, suggesting that consistency checks should compare target-population standardized effects rather than raw estimates.","Beyond the paper: a practical test for HTA reviews would be to request the covariate distributions of the comparator trial and run a model-based re-standardization to the decision population; if the effect changes materially, the submission should report both estimates.","Beyond the paper: the scale-alignment condition suggests an actionable design choice—specify effect modification on the same scale as the decision-relevant effect measure, or choose a collapsible measure—to make transportability claims more defensible."],"forward_implications":["Pairwise MAIC/STC results on odds-ratio or hazard-ratio scales should be labelled as comparator-population estimands; applying them to the index or real-world population requires explicit re-standardization or new assumptions.","HTA guidance that treats SEMA as sufficient for transporting marginal relative effects is too permissive; the transport step fails exactly for the non-collapsible measures most commonly submitted.","Cost-effectiveness models that feed MAIC/STC hazard or odds ratios into a decision model defined for another population risk transport bias that is not captured in the statistical confidence intervals.","For collapsible, scale-aligned measures (mean differences, log risk ratios under a log-link SEMA), direct transport is justified; analysts can design analyses around these measures when clinically reasonable.","Network-based methods that estimate effects in a pre-specified target population can avoid the implicit second step, but only under stronger network-wide assumptions about shared prognostic and effect-modification functions."],"fun_headline_variants":["Transportability requires collapsible measures and scale alignment","Pairwise MAIC/STC results are comparator-population specific","Even shared effect modifiers don't make marginal effects portable","Conditional linear-predictor effects transport better","PAIC transportability hinges on collapsibility and effect-measure scale"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The formal results rest on the structural model in which the same baseline prognosis function applies to all treatments and the two active treatments share an identical effect-modification function on the model's scale; if those functions differ across treatments, the cancellation that produces transportable contrasts does not occur.","fun_headline_variants_meta":{"raw":{"variants":["Transportability requires collapsible measures and scale alignment","Pairwise MAIC/STC results are comparator-population specific","Even shared effect modifiers don't make marginal effects portable","Conditional linear-predictor effects transport better","PAIC transportability hinges on collapsibility and effect-measure scale"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000983,"raw_usage":{"total_tokens":4033,"prompt_tokens":795,"completion_tokens":3238,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":539,"completion_tokens_details":{"reasoning_tokens":3156}},"tokens_in":539,"tokens_out":3238,"duration_ms":18949,"temperature":1.0,"reasoning_tokens":3156,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T22:21:09.831206+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate an anchored comparison with SEMA and scale alignment holding by construction for a marginal odds ratio or hazard ratio, letting the index and comparator populations differ in the mean of an effect modifier; if the true marginal active-active effect is identical in every population, the paper's central claim is false, and if it varies, the claim is confirmed.","supporting_citations":[],"review_version":1}