{"id":"a8067e90-ac61-4e4d-b7fc-0af9cc0db01d","arxiv_id":"2505.12308","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"EQPS-rMAP, a Bayesian method that borrows external trial and real-world data through propensity score stratification and equivalence probability weighting, is reported in simulations to reduce bias and sample size versus MAP, PS-MAP, and EB-rMAP.","lead":"A new statistical method for drug trials combines overseas trial data and local real-world patient data to reduce the number of patients needed in a new region, while adjusting for differences in patient characteristics. It matters because regulators and drug developers are actively looking for ways to use real-world data to make global clinical trials faster and smaller.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Data-dependent selection of omega_r from current trial data (Eqs 2-18/2-19) before using that same data in the final posterior (Eq 2-20) is unproven to preserve type I error; the claimed robustness and sample-size savings rest on this.","rationale":"I read the paper as proposing an adaptive borrowing method whose central promise is to reduce sample size while preserving estimation accuracy and error control. The most load-bearing step is the data-dependent selection of omega_r: the current trial data are used first to choose how much to borrow (Eqs 2-18 and 2-19) and then again in the final posterior (Eq 2-20). This is a double use of the outcome data, and the paper provides neither a theoretical proof nor a systematic calibration showing that nominal type I error is maintained. The reader's weakest-assumption analysis identified exactly this issue, and I agree with it. The paper's own text acknowledges that parameter operationalization remains to be examined, and the simulation section reports only graphical summaries without Monte Carlo errors or numeric tables, so the evidence is thinner than the abstract's claim. The synthetic case study further limits external validation, but that is secondary to the operating-characteristic gap. The method is not obviously invalid; a null-scenario calibration study of the kind described would settle whether the concern actually lands. Because the conditional verdict already reflects the need for additional evidence, no verdict adjustment is needed.","tokens_in":14488,"tokens_out":5262,"duration_ms":55330,"concrete_test":"Run a null-scenario calibration study with the paper's own setup (Section 3.1): n=500 per arm in current trial, external trial, and RWD; true response rates equal under H0 (e.g., p_T=p_C=0.5), with baseline shifts alpha_RWD and alpha_ext nonzero in Eq 3-2 but beta3=beta4=0. For each lambda in {0.7,0.8,0.9} and delta in {0.05,0.1,0.15}, select omega_r via Eqs 2-18/2-19 and evaluate the final posterior Eq 2-20, declaring success if P(p_T>p_C|data)>0.95. Repeat 10,000 times per cell and report empirical type I error with Monte Carlo standard error sqrt(alpha(1-alpha)/N), plus the empirical distribution of selected omega_r. If any cell exceeds 0.05 by more than 2 SEs, the claimed type I error control and the sample-size reduction derived from it are not established; if all cells are within tolerance, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that EQPS-rMAP 'maintains estimation robustness ... while reducing sample size demands' is only valid if the data-adaptive choice of the vague-prior weight omega_r preserves frequentist operating characteristics. In Step 3, omega_r is chosen from the current data: Eq 2-18 computes p = Pr(theta_mix - delta < theta_current < theta_mix + delta), where theta_current is the response rate from the current trial, and Eq 2-19 sets omega_r to the smallest weight with p >= lambda (or 1 if p < lambda). The same current data are then used again in the final posterior, Eq 2-20. Hence the prior is a function of the outcome data used for inference. No theorem or calibration argument is provided that this double use controls type I error; Section 3 only gives graphical results for six simulated configurations, with user-chosen lambda and delta and no Monte Carlo standard errors. The paper itself says in Section 2.3 that operationalization of lambda and delta 'will be systematically examined in subsequent simulation trials' and in Section 5 that 'systematic simulations are required to identify optimal parameter combinations.' Since sample-size reduction is achieved exactly by letting omega_r permit borrowing, any type I error inflation converts claimed efficiency into overstated confidence. The abstract's general conclusion is therefore not supported by the evidence presented.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes EQPS-rMAP, a three-stage Bayesian hybrid method that combines domestic real-world data and overseas external trial data with a new-region randomized controlled trial. Stage 1 uses propensity score trimming and stratification to address baseline covariate imbalance; Stage 2 builds stratum-specific robust meta-analytic predictive (MAP) priors with heterogeneity variances scaled by propensity-score overlap; Stage 3 introduces an equivalence-probability weight omega_r, defined as the smallest vague-prior weight such that the probability that the hybrid posterior lies within a margin delta of the current-data response distribution reaches a threshold lambda. The authors claim that this adaptive borrowing preserves estimation robustness under heterogeneity and reduces required sample sizes. The method is evaluated with six simulation scenarios comparing EQPS-rMAP with MAP, PS-MAP, and EB-rMAP, and with an illustrative risankizumab psoriasis case analysis in which external and real-world data are real but the current-trial data are simulated. The central claim is that EQPS-rMAP resolves baseline-heterogeneity conflicts while maintaining type I error control and estimation accuracy.","tokens_in":14765,"tokens_out":4457,"duration_ms":45119,"significance":"If the claims were fully supported, the method would be a practically useful contribution to hybrid trial design, bridging studies, and multi-regional clinical trials, where regulatory interest in Bayesian borrowing from external data is high. The paper deserves credit for a clear conceptual structure, for combining propensity-score stratification with stratum-specific MAP priors, and for attempting a realistic case analysis; the authors also state that R code for Section 4 is available. However, the validation evidence is currently insufficient in two load-bearing respects: the borrowing weight is chosen from the current outcome data that are then used again in the final posterior, and the simulation results are reported only graphically, without numeric summaries or Monte Carlo standard errors. The significance of the method can be assessed only after these operating characteristics are quantified.","major_comments":[{"comment":"The vague-prior weight omega_r is a function of the current trial data: Eq (2-18) computes p by comparing the hybrid posterior with the Beta posterior of the current data, and Eq (2-19) selects omega_r as the smallest weight with p >= lambda. The same current data are then used again in the final posterior, Eq (2-20). No theorem, calibration argument, or extensive simulation demonstrates that this double use of the data preserves the nominal 5% type I error. This issue is load-bearing because the claimed sample-size savings are achieved precisely by allowing omega_r < 1; the paper's own statement in Section 2.3 that operationalization of lambda and delta 'will be systematically examined in subsequent simulation trials' does not resolve the concern.","section":"Section 2.3, Eqs (2-18)-(2-20)"},{"comment":"The central performance comparison is presented only graphically, with no numeric table of absolute bias, mean squared error, or type I error and no Monte Carlo standard errors. The abstract's claim that EQPS-rMAP 'maintains estimation robustness under significant heterogeneity' and the conclusion that it 'effectively manages Type I error' cannot be quantitatively assessed from the current figures. The authors should report point estimates with Monte Carlo standard errors for all six scenarios, and ideally across the full 54-scenario grid described in Section 3.1.","section":"Section 3.2, Figures 7 and 8"},{"comment":"The parameters lambda and delta are user-specified tuning parameters, and the comparisons in Section 3.1 use a single pair (lambda = 0.8, delta = 0.1) without a sensitivity analysis. The paper itself states in Section 5 that 'systematic simulations are required to identify optimal parameter combinations.' Until a calibration rule or sensitivity results are provided, the claimed superiority over EB-rMAP, MAP, and PS-MAP remains conditional on unexamined choices of lambda and delta.","section":"Section 3.1 and Section 5"},{"comment":"The 'current trial data' in the illustrative example are simulated with prespecified response rates (40% control, 65% treatment) and are generated using baseline characteristics from the external trial data. The case analysis therefore does not validate the method on observed current-trial outcomes, and the statement in the conclusion that 'case analyses confirm superior external bias control and accuracy' overstates what this example can show. The example should be described as a feasibility illustration based on a simulated current trial.","section":"Section 4"}],"minor_comments":[{"comment":"The text refers to 'EQPS-MAP (lambda = 0.8, delta = 0.1)' while the abstract and the rest of the paper define the method as EQPS-rMAP; the terminology should be standardized.","section":"Section 3.1"},{"comment":"There are repeated typographical errors, including 'External trail data' instead of 'External trial data' and 'Jefferys' prior' instead of 'Jeffreys' prior'; a careful proofread is needed.","section":"Throughout"},{"comment":"Equation (2-14) is typeset in a way that makes the integral and the density arguments difficult to parse; it should be rewritten with standard integral notation and clearly defined variables.","section":"Equation (2-14)"},{"comment":"The figures would be much easier to evaluate if they included numeric axis labels and a legend identifying the curves for the different parameter values; currently several figure descriptions refer to lines and columns that are not explicitly labeled.","section":"Figures 3-6"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the manuscript is better positioned as a methodological proposal with a feasibility illustration than as a fully validated method. The most serious risk is the data-dependent selection of omega_r followed by reuse of the same current data in the final posterior; this needs either a formal operating-characteristic argument or a very extensive simulation calibration. The paper is likely salvageable through a major revision that adds numeric simulation tables with Monte Carlo errors, sensitivity analyses for lambda and delta, and a more cautious interpretation of the illustrative example."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper does something real: it combines propensity-score stratification with stratum-specific robust MAP priors and an equivalence-probability rule for setting the vague-prior weight. That combination is not in the cited literature, and it targets a genuine regulatory need for bridging trials and MRCTs where local RWD and overseas trial data differ in baseline and heterogeneity. The method is clearly described, and the authors cite the relevant building blocks honestly.\n\nWhat the paper does well: the three-stage structure is sensible, the overlap-coefficient idea for stratum-specific similarity is reasonable, and the authors are upfront in Section 5 that lambda and delta need systematic calibration and that the framework currently only handles qualitative efficacy comparisons. The case study is transparent about using simulated current data.\n\nThe soft spots are in proportion. The big one is the circularity the stress-test flags: omega_r is chosen from the current trial data via Eqs 2-18 and 2-19, and the same data then enter the final posterior in Eq 2-20. No calibration argument or theorem shows this double use preserves frequentist type I error. The paper itself punts on operationalizing lambda and delta, saying it will be examined in subsequent simulation trials. That is an honest admission, but it undercuts the abstract's claim that the method 'reduces sample size demands' while controlling error. Sample-size savings are exactly where data-dependent borrowing can inflate type I error, and six simulated scenarios with user-chosen thresholds and no Monte Carlo standard errors do not establish the operating characteristics.\n\nSecond, the simulation reporting is too thin. Results are presented only through figures, with no numeric tables, no MCSE, and no confidence intervals for bias, MSE, or type I error. Third, the data accessibility statement says R code is available for Section 4, but no repository link appears in the text. Fourth, the abstract and conclusion call the case study a 'retrospective case analysis' when it is actually a synthetic illustration using real external data and simulated current data. That overstatement should be fixed.\n\nNone of this makes the method invalid. The central idea is plausible, and the paper identifies a real gap. But the evidence presented does not support the strength of the claims. I would not cite this as a validated method yet, and I would not use it in a regulatory setting without a calibration study.\n\nWho is this for? Statisticians in pharma or regulatory science working on Bayesian borrowing methods. They will get a clear description of a workable extension and a list of open questions worth addressing.\n\nRecommendation: send it to peer review, but require major revision: add a code repository, report numeric simulation results with Monte Carlo errors, provide a calibration rule or sensitivity analysis for lambda and delta, re-label the case study honestly, and address the data-dependent weight selection directly. As it stands, the paper is a promising methods note rather than a validated method.","headline":"A plausible incremental Bayesian borrowing method whose headline efficiency claims rest on unverified data-dependent weight selection and thin simulation reporting.","tokens_in":15319,"tokens_out":2027,"would_cite":false,"duration_ms":22947,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"EQPS-rMAP maintains accuracy while borrowing external real-world data.","keywords":["hybrid clinical trial design","real-world data","meta-analytic predictive prior","propensity score stratification","adaptive borrowing","equivalence probability weight","bridging trials","risankizumab"],"falsifier":"Simulate the full procedure repeatedly under the null hypothesis, computing $\\omega_r^*$ by Eq. (2-19) on each replicate and testing $P(\\theta_T>\\theta_C)>0.95$ as the decision rule; if the proportion of false successes exceeds 5% by more than simulation error, the claimed error control does not hold in that scenario.","tokens_in":14290,"feed_emoji":"📊","tokens_out":8010,"duration_ms":74348,"temperature":0.7,"pith_summary":"The paper aims to establish that a hybrid Bayesian prior, EQPS-rMAP, can combine overseas randomized trial data with domestic real-world data in bridging and multi-regional trials without the bias typical of fixed-borrowing methods. It does this in three steps: stratify all patients by propensity score, build meta-analytic predictive priors inside each stratum, and set the borrowing weight through an equivalence probability that measures conflict between external and current data. The method is designed so that compatible evidence is borrowed heavily, while incompatible evidence is effectively shut off. If the simulations and case analysis are right, this gives trialists an adaptive, pre-specifiable way to reduce sample size in new regions without giving up estimation accuracy.","feed_headline":"EQPS-rMAP maintains accuracy while borrowing external real-world data","feed_subtitle":"Stratifying patients before borrowing lets the prior reject incompatible data while shrinking sample size.","key_machinery":"The machinery is the data-dependent prior weight $\\omega_r^*$ inside a robust MAP mixture. The weight is computed from the equivalence-probability consistency metric $p = P(\\theta_{\\text{post}}-\\delta < \\theta_{\\text{current}} < \\theta_{\\text{post}}+\\delta)$, with $\\omega_r^*$ set to the smallest value satisfying $p \\ge \\lambda$ and to 1 when no such value exists. Before that, propensity-score stratification trims subjects whose propensity scores fall outside the current trial's range and divides the rest into strata, and stratum-specific MAP priors carry between-source heterogeneity through half-normal variance parameters whose scales are informed by the overlap coefficient of the propensity-score distributions. The weight is updated from the current trial's own data after seeing the current outcomes, which is what lets the borrowing proportion adapt, but also what makes the final operating characteristics depend on the data twice.","core_discovery":"The central claim is that baseline discrepancies and multi-source heterogeneity can be handled together by making the borrowing weight a function of measured data conflict. Within each propensity-score stratum, the external and real-world sources enter as stratum-specific robust MAP priors, and the final EQPS-rMAP posterior is a mixture of that informative component and a vague prior. The weight of the vague component, $\\omega_r^*$, is the smallest value such that the mixed posterior's response probability stays within a clinical equivalence margin $\\delta$ of the current-trial response distribution with probability at least $\\lambda$. If agreement is poor, the weight goes to one, turning off borrowing entirely; if agreement is good, the method uses the maximal safe amount of external information. The paper reports that this keeps bias and mean squared error low across six scenarios and in a risankizumab psoriasis case study, while reducing the sample size the current trial would otherwise need.","pith_inferences":["Because the same current data both selects $\\omega_r^*$ and enters the final posterior, the effective type I error could diverge from 5% in settings not covered by the reported scenarios; a calibration study sweeping sample size, endpoint type, and $\\lambda$/ $\\delta$ would be the natural next check.","The weighting scheme could be ported to non-binary endpoints by replacing the beta-binomial mixture with conjugate normal or gamma mixtures; the paper states this as future work but does not implement it.","Regulatory use would likely require pre-specifying $\\lambda$ and $\\delta$ before unblinding; the paper demonstrates the trade-off in simulations but does not supply a default combination that guarantees operating characteristics."],"forward_implications":["In bridging and multi-regional trials, investigators can pre-specify $\\lambda$ and $\\delta$ and let the data decide how much foreign or real-world information to borrow, instead of committing to a fixed proportion.","A new region that wants to run a smaller trial can quantify how much sample size it saves under EQPS-rMAP when external data are compatible, because the posterior precision rises with the borrowed information.","The propensity-score stratification means baseline differences between domestic real-world patients and overseas trial patients are removed before the prior is built, so the method does not require exchangeability across sources.","The risankizumab case shows the posterior estimate staying near the true treatment effect as the proportion of borrowed data varies, whereas standard MAP and PS-MAP estimates drift toward the external data."],"supporting_citations":[{"why":"Defines the original meta-analytic predictive prior that the method extends.","marker":"[18]"},{"why":"Provides the robust MAP prior with a vague-prior component, the base of the borrowing construction.","marker":"[19]"},{"why":"Supplies the empirical robust MAP prior with an adaptive vague weight, a key comparator and conceptual source.","marker":"[22]"},{"why":"Introduces the equivalence-probability weighting idea used to quantify prior-data conflict.","marker":"[24]"},{"why":"Gives the conjugate mixture approximation that turns the MAP prior into a beta mixture for computation.","marker":"[31]"},{"why":"Shows how propensity-score weighting incorporates real-world evidence, the basis for the baseline adjustment.","marker":"[32]"},{"why":"Combines propensity score with MAP prior for real-world and historical data, the direct predecessor of stratum-specific borrowing.","marker":"[33]"},{"why":"Defines the tail-region probability used to measure consistency between two response distributions.","marker":"[35]"},{"why":"Provides the overseas randomized trial data used in the risankizumab case study.","marker":"[36]"},{"why":"Provides the domestic retrospective real-world cohort used in the same case study.","marker":"[37]"}],"fun_headline_variants":["EQPS-rMAP: borrow external data only if it agrees","Stratify patients, borrow safely: EQPS-rMAP","EQPS-rMAP: adaptive borrowing shrinks sample sizes","Borrow RWD only when conflict is low: EQPS-rMAP"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claimed bias and sample-size advantages assume that selecting the vague-prior weight from the current trial's own data and then using that same data in the final posterior preserves the nominal frequentist type I error, even though the weight itself is random and data-dependent.","fun_headline_variants_meta":{"raw":{"variants":["EQPS-rMAP: borrow external data only if it agrees","Stratify patients, borrow safely: EQPS-rMAP","EQPS-rMAP: adaptive borrowing shrinks sample sizes","Borrow RWD only when conflict is low: EQPS-rMAP"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000213,"raw_usage":{"total_tokens":1453,"prompt_tokens":1006,"completion_tokens":447,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":622,"completion_tokens_details":{"reasoning_tokens":375}},"tokens_in":622,"tokens_out":447,"duration_ms":4447,"temperature":1.0,"reasoning_tokens":375,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:37:01.993316+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate the full procedure repeatedly under the null hypothesis, computing $\\omega_r^*$ by Eq. (2-19) on each replicate and testing $P(\\theta_T>\\theta_C)>0.95$ as the decision rule; if the proportion of false successes exceeds 5% by more than simulation error, the claimed error control does not hold in that scenario.","supporting_citations":[],"review_version":1}