{"id":"b81dea28-8f3d-4c65-b1e1-84eabafd8aec","arxiv_id":"2602.00434","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Across 50 real randomized trials, simple prespecified covariate adjustment improved precision as much as or more than machine-learning estimators, with no systematic bias.","lead":"Drawing on 29,094 participants from 50 public randomized trials, this paper tests 18 ways to adjust for baseline patient characteristics when estimating treatment effects. Simple prespecified regression methods improved precision (median 13.3% variance reduction for continuous outcomes) and performed as well as or better than machine-learning approaches, supporting routine use of transparent covariate adjustment.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Binary-outcome PVR medians for ML estimators are conditional on successful runs (DML 12% errors, TMLE 4%, ANCOVA 0%); differential failure may confound the conclusion that ML adds no precision.","rationale":"The reader's weakest assumption was external representativeness of the 50-trial corpus. That is a real limitation, but it concerns generalizability: even if the corpus were unrepresentative, the internal comparison among methods on these trials could still be valid. The concern I identify is internal: the binary-outcome PVR comparison across methods is confounded by differential computational failures. This is more load-bearing because it can invalidate the key comparative claim within the very data analyzed. The paper reports error rates transparently, but it does not account for them when comparing PVR medians. A reader who accepts the external-validity limitation can still learn which method performs better in the curated corpus; a reader concerned about selective success cannot trust even that internal ranking. The reader's rationale did mention 'binary-outcome precision estimates may be biased by dropping error cases' as one of several concerns, but it was not identified as the weakest assumption. I therefore mark agreement as 'disagree' relative to the weakest-assumption field. My recommendation is to keep the verdict CONDITIONAL: the paper's conclusions are plausible and the concern is addressable with a robustness analysis, but the current analysis has a concrete, potentially decisive gap. If the proposed reanalysis changes the rankings, the central claim would need substantial revision; if it does not, the conclusion is strengthened.","tokens_in":22479,"tokens_out":6877,"duration_ms":90012,"concrete_test":"Recompute all binary-outcome PVR summaries on the intersection of treatment–outcome pairs where every method ran successfully (complete-case analysis across methods). Also compute a sensitivity analysis that assigns plausible PVR values to failed runs (e.g., the 10th percentile of that method's observed PVR, or the ANCOVA PVR for the same pair). If the median PVR difference between ANCOVA and DML/TMLE changes sign or becomes negligible under these reanalyses, the 'ML adds no precision' conclusion is an artifact of conditioning on successful runs.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that machine-learning estimators 'did not provide additional precision' rests heavily on comparing median PVR across methods. For binary outcomes, error rates differ sharply: DML fails in 12% of All-covariate analyses, TMLE in 4%, while ANCOVA and g-logistic fail in 4% or less (Figure 2). The paper reports PVR only for runs that returned results, so the median PVR for DML/TMLE is computed on a selected subset of trials/pairs. The paper itself states that failures 'mostly arose from rare events in outcomes' (Error rate section). If those rare-event or small-sample settings have systematically different covariate-adjustment gains, then the comparison is not apples-to-apples: DML's median PVR reflects only the settings where it ran, while ANCOVA's median reflects all settings. This could bias the relative ranking in either direction, making ML look better or worse than it truly is. Because the paper's practical recommendation to prefer parsimonious regression is partly justified by the claim that ML offers no efficiency gain, this selective-success issue directly threatens the internal validity of that comparison. The concern is not external generalizability; it is a structural bias in the performance metric itself.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper assembles individual-level data from 50 publicly available randomized trials (29,094 participants; 574 treatment–outcome pairs) and benchmarks 18 covariate-adjustment strategies: six estimators (ANCOVA, ANHECOVA, IPW, g-logistic, DML, TMLE) crossed with three covariate-selection rules (All, Top-3, Baseline+). Performance is measured by proportional variance reduction (PVR), standardized point-estimate shifts, covariate adjustment gain/loss, and error rates. The headline findings are that covariate adjustment improves precision in most settings (median PVR 13.3% continuous, 4.6% binary), that parsimonious prespecified regression approaches such as ANCOVA perform as well as or better than machine-learning estimators, and that ML estimators add little precision while failing more often for binary outcomes. The authors release all curated data and code and provide practical, decision-oriented recommendations favoring prespecified adjustment with a small set of prognostic covariates.","tokens_in":22804,"tokens_out":4951,"duration_ms":67631,"significance":"If the empirical comparisons are trustworthy, this is a valuable and timely contribution. It is, to my knowledge, one of the first large-scale benchmarks of covariate-adjustment strategies on real RCT data rather than simulations, and the open release of 50 harmonized datasets is a reusable resource for future methodological work. The paper is also unusually transparent about limitations, including the representativeness of public-repository trials and the correlation induced by multiple treatment–outcome pairs from the same trial. However, the central claim that machine-learning estimators 'did not provide additional precision' is currently vulnerable to a differential-success bias in the binary-outcome analysis, and the Top-3 strategy is evaluated in-sample despite being data-driven. Those issues are fixable but load-bearing for the paper's practical recommendations.","major_comments":[{"comment":"The comparison of median PVR across estimators for binary outcomes is computed only on runs that returned results, while error rates differ sharply by method: DML fails in 12% of All-covariate analyses, 12% of Top-3, and 8% of Baseline+; TMLE fails in 4%, 3%, and 3%; ANCOVA fails in 0% (Figure 2, error-rate row). The paper states that most failures arose from rare events in outcomes. If rare-event or small-sample settings have systematically different covariate-adjustment gains, the DML/TMLE medians are estimated on a selected subset, making the comparison with ANCOVA not apples-to-apples. This directly affects the claim that ML offers no precision gain. Please provide sensitivity analyses: for example, (i) restrict all methods to the subset of trials where every estimator returns results; (ii) assign failed runs conservative PVR values (e.g., 0 or worst-case) and re-plot; and (iii) repo","section":"§Error rate; Figure 2"},{"comment":"The Top-3 strategy selects the three covariates most strongly correlated with the outcome using the same dataset in which PVR is then evaluated. As the manuscript acknowledges, this carries 'invalidity from data-driven selection.' Because no sample splitting or cross-validation is used, the Top-3 PVR values are in-sample estimates that can be optimistically biased, especially in small trials. Since Top-3 is one of the parsimonious strategies used to support the recommendation that small covariate sets perform well, the analysis needs either a cross-validated version of Top-3 or an explicit statement that these figures are exploratory and not direct estimates of real-world performance. At minimum, a sensitivity analysis using a prespecified or split-sample selection rule would clarify whether the Top-3 results are an artifact of selection on the outcome.","section":"§Results, Benchmarking pipeline; Figure 2"},{"comment":"The main performance summaries are median PVR values and boxplots over 574 treatment–outcome pairs, without confidence intervals or formal comparisons between methods. The 574 pairs are not independent: they are clustered within 50 trials, and the authors acknowledge that many pairs share overlapping covariate information. The robustness check restricted to primary outcomes has only 76 continuous and 18 binary pairs, so it is not a substitute for accounting for clustering in the full analysis. To support the wording 'consistently improves precision' and 'performed as well as or better,' please add uncertainty quantification that respects the trial-level clustering—for example, per-trial median PVR with a cluster bootstrap, or paired within-trial comparisons of estimators—and report the variability of the all-pairs medians.","section":"§Figure 2; §Discussion, first limitation paragraph"}],"minor_comments":[{"comment":"The CAG/CAL definitions contain a typo: 'the adjusted and adjusted analysis' should read 'the adjusted and unadjusted analysis.'","section":"§Methods, Performance metrics"},{"comment":"The caption says 'regression adn machine-learning-based estimators'—'adn' should be 'and.'","section":"Figure 3 caption; Supplementary Figure 3 caption"},{"comment":"The figure panel includes stray text such as 'BOX ： 00 00' and the layout of the numeric summaries is hard to read. Please clean up the annotation and ensure the boxplot/error-rate panels are legible at print resolution.","section":"Figure 2"},{"comment":"The EMA guidance URL in reference [6] appears malformed ('euincremental/docpredefinedindicative-guideline'); please correct it.","section":"References"},{"comment":"The description of missing outcome handling says variables with >40% missingness were excluded and remaining missing outcomes were 'noted as N/A and dropped.' This is complete-case analysis; please state explicitly in the limitations that complete-case exclusion may affect estimates of precision, especially for binary outcomes.","section":"§Methods, Datasets and preprocessing"}],"recommendation":"major_revision","confidential_remarks":"The data-release contribution and the breadth of the empirical corpus are substantial, and the limitations section is unusually candid. The main reason for major revision is not external generalizability but internal validity of the ML-versus-parsimonious comparison: the differential failure rates for binary outcomes, combined with in-sample Top-3 selection, mean that some of the headline precision comparisons may not reflect what would happen if these strategies were applied in a new trial. These concerns are addressable with sensitivity analyses, and I would be willing to review a revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper delivers what it promises: a large, transparent benchmark of covariate-adjustment strategies across 50 public RCTs and 574 treatment–outcome comparisons, including a first-of-its-kind comparison of machine-learning estimators against classical adjustment on real trial data. The curated datasets and code are openly released, the methods are sensible, and the authors are honest about their limitations. The central finding—that parsimonious regression approaches match or beat more complex ML estimators—is broadly consistent with the continuous-outcome results, where error rates are zero and the comparison is clean.\n\nThat is why the soft spot in the binary-outcome analysis matters. DML fails in 12% of binary analyses under the All-covariates strategy, TMLE in 4%, while ANCOVA and g-logistic fail 4% or less. The reported PVR medians for DML and TMLE are therefore computed only on the runs that succeeded, which are disproportionately the trials without rare events. If those rare-event settings have different covariate-adjustment gains—and the authors themselves say the failures mostly came from rare outcomes—then comparing the ML medians to the complete ANCOVA median is not apples-to-apples. This could bias the ranking in either direction. It does not sink the paper, because the continuous results are unaffected and the high error rates are themselves a relevant operating characteristic, but it is a real threat to the internal validity of the 'ML adds no precision' claim for binary outcomes.\n\nThe other weaknesses are more minor. The performance summaries are descriptive medians with no confidence intervals or formal between-method comparisons; the Top-3 strategy selects covariates in-sample, which the authors acknowledge; and the 50-trial corpus from public repositories may not represent the broader trial landscape (a limitation the authors flag). These are addressable with some added analysis and hedging.\n\nOverall, this is a solid empirical contribution that deserves peer review. I would recommend a revision-focused referee process, asking for the binary-outcome PVR analysis to be reported conditional on success versus complete-case, ideally with sensitivity analyses that impute failures as worst-case, plus bootstrap intervals on the medians. The paper should also soften the claim that ML 'did not provide additional precision' to 'did not provide additional precision in the settings where it ran successfully for binary outcomes.'\n\nFor anyone working in clinical trial methodology or teaching covariate adjustment, this is a useful benchmark and a good starting point for discussion. I would cite it for the scale of the evidence and the open-data resource, and I would bring it to a reading group.","headline":"A genuinely useful empirical benchmark of covariate adjustment in RCTs, with a real but fixable selective-success problem in the binary-outcome ML comparison.","tokens_in":23226,"tokens_out":1630,"would_cite":true,"duration_ms":24325,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"Across 50 real randomized trials, simple prespecified covariate adjustment improves precision consistently — 13.3% median variance reduction for continuous, 4.6% for binary — and matches or beats machine-learning estimators.","keywords":["covariate adjustment","randomized clinical trials","ANCOVA","machine learning","precision gains","baseline covariates","empirical benchmarking"],"falsifier":"Take any independent set of, say, 50 archived randomized trials (ideally including non-publicly-deposited industry trials) and recompute the median proportional variance reduction and the ANCOVA-versus-TMLE/DML gap; the central claim fails if the new median PVR is near zero or negative, or if machine-learning estimators show a consistent precision advantage in small samples with binary outcomes.","tokens_in":22417,"feed_emoji":"📊","tokens_out":7156,"duration_ms":73937,"temperature":0.7,"pith_summary":"Using individual-level data from 50 completed randomized trials (29,094 participants, 574 treatment-outcome comparisons), the paper benchmarks six covariate-adjustment estimators under three covariate-selection rules. It tries to establish that routine adjustment with a small, prespecified set of prognostic baseline covariates consistently improves precision — median variance reductions of 13.3% for continuous outcomes and 4.6% for binary outcomes — without systematically shifting point estimates or increasing the risk of losing a significant result. The paper also argues that parsimonious regression methods such as ANCOVA are at least as precise as machine-learning estimators in these real trials, and more reliable in small samples. That matters because regulatory guidance already recommends covariate adjustment, but practicing trialists lack empirical evidence on which strategy to prespecify. If correct, trial protocols can safely prespecify simple, transparent adjustment and reap power gains.","feed_headline":"Parsimonious regression outperforms machine learning across 50 trials","feed_subtitle":"Median variance drops 13.3% (continuous) and 4.6% (binary), so prespecified adjustment can shrink trial size.","key_machinery":"The object doing the analytical work is the benchmark corpus plus the two operating-characteristic metrics: proportional variance reduction (PVR = 1 − Var(adjusted)/Var(unadjusted)) and the covariate-adjustment gain/loss (CAG/CAL), the conditional probabilities that adjustment creates or removes a 5%-significant finding relative to the unadjusted analysis. PVR translates directly into the sample-size reduction needed to maintain power, and CAG/CAL translate method performance into clinical decision impact. The estimators under study are anchored by ANCOVA, which here is the reference parsimonious strategy and is guaranteed asymptotically no less precise than the unadjusted estimator under eq","core_discovery":"Across 574 treatment-outcome comparisons drawn from 50 real RCTs, covariate adjustment improves precision in most settings, with median proportional variance reductions of 13.3% (continuous) and 4.6% (binary). ANCOVA adjusting for all available covariates gives the largest precision gains and rarely loses precision; in small samples, more complex estimators — ANHECOVA with interactions, IPW, g-computation, and the machine-learning estimators TMLE and DML — are either similar or worse, and the machine-learning methods show elevated error rates for binary outcomes. Under a prespecified 'Baseline+' adjustment set (baseline outcome, stratification factors, age, sex, weight), gains remain consist","pith_inferences":["A natural next test is whether hyperparameter-tuned or stratified-cross-validation versions of TMLE/DML close the gap; the paper's own error-rate analysis suggests the two-layer sample splitting is the likely culprit, not the estimators' statistical framework.","Because the corpus is 50 trials with many treatment-arm comparisons sharing covariates, the median PVR should be treated as a distribution, not a guarantee; replicating on a per-therapeutic-area basis (or on trials with large samples and highly nonlinear outcomes) would show where flexible models start paying off.","If these results generalize, the practice of reporting only unadjusted analyses in trial registries and primary papers carries a quantifiable efficiency cost; one could compute the expected power loss directly from the observed PVR distribution and use it to argue for prespecified adjustment in protocols.","The CAG/CAL asymmetry suggests a screening metric: trials whose unadjusted analysis is just at the significance boundary are the ones where covariate adjustment most plausibly flips the conclusion, so sensitivity analyses should be reported there."],"forward_implications":["Prespecifying a small prognostic set (baseline outcome, stratification factors, age, sex, weight) and fitting ANCOVA for continuous endpoints or g-computation from logistic regression for binary endpoints is a sufficient, transparent primary-analysis strategy in most trials.","Trials can be sized using the expected PVR: a 13.3% median variance reduction for continuous outcomes corresponds to roughly a 13% sample-size saving at equal power, and the paper's guidance gives a way to anticipate this at the planning stage.","Default-hyperparameter machine-learning estimators should not be the primary analysis in small-to-moderate samples: they did not beat linear regression and had error rates up to 12% for binary outcomes when all covariates were used.","Covariate adjustment in real trial data is more likely to convert a non-significant unadjusted result into a significant one (CAG) than the reverse (CAL), so routine adjustment will rarely 'lose' an effect that was genuinely there.","When sample size is roughly above 100, estimator choice matters less; below that, simple regression is the safer default."],"fun_headline_variants":["Simple covariate adjustment beats ML in 50 RCTs","ANCOVA wins over ML for trial precision","Prespecified covariates: 13% variance cut, ML flops","Covariate adjustment: parsimony pays in 50 trials","RCTs: small model, big precision gains"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the 50 publicly deposited trials, which the authors themselves say 'may not be fully representative of the broader landscape of modern clinical trials' (Discussion), represent the range of real trials where these recommendations will be applied; if public-repository trials differ systematically from commercial or preprint-only trials in sample size, outcome type, or covariate quality, the median gains and the ANCOVA-over-machine-learning ranki","fun_headline_variants_meta":{"raw":{"variants":["Simple covariate adjustment beats ML in 50 RCTs","ANCOVA wins over ML for trial precision","Prespecified covariates: 13% variance cut, ML flops","Covariate adjustment: parsimony pays in 50 trials","RCTs: small model, big precision gains"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000154,"raw_usage":{"total_tokens":1089,"prompt_tokens":827,"completion_tokens":262,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":571,"completion_tokens_details":{"reasoning_tokens":180}},"tokens_in":571,"tokens_out":262,"duration_ms":3890,"temperature":1.0,"reasoning_tokens":180,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T06:01:15.344760+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take any independent set of, say, 50 archived randomized trials (ideally including non-publicly-deposited industry trials) and recompute the median proportional variance reduction and the ANCOVA-versus-TMLE/DML gap; the central claim fails if the new median PVR is near zero or negative, or if machine-learning estimators show a consistent precision advantage in small samples with binary outcomes.","supporting_citations":[],"review_version":1}