{"id":"7f4437be-2c85-4314-a48d-4f7bf75f77a2","arxiv_id":"2502.03030","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"RISE uses rank-based nonparametric screening plus sample splitting to build a weighted gene-expression signature that approximates the vaccine's effect on antibody response.","lead":"A new statistical method, RISE, screens thousands of genes to find a small set whose early changes after vaccination can stand in for the later antibody response. It is designed for small vaccine trials with many candidate markers and was tested on a flu vaccine dataset.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The data-adaptive margin in Eq. (2) cancels UY from the screening statistic, so RISE's surrogacy test reduces to testing whether each gene responds to vaccination; with a fixed epsilon=0.10 the signature disappears (Table S2) and the evaluation TOST p-value rises to ~0.051.","rationale":"The reader's weakest assumption correctly targets the data-dependent margin. My analysis strengthens it: because epsilon is constructed from bUY, the estimated outcome effect cancels in the numerator of the Step-1 statistic whenever bUY exceeds u*. The 'surrogacy' test then asks whether the gene's own treatment effect is detectable at the specified power, rather than whether it matches the outcome effect. This is not merely a calibration issue; it changes the null hypothesis being tested. The application is a single-arm pre/post study, so UY is a within-person change score, not a causal treatment effect, and a single trial cannot support a trial-level surrogacy claim. These problems do not erase the paper's useful engineering: the sample-splitting design, the paired variance derivation, the reproducible R package, and simulation studies over multiple DGPs and correlation structures are real contributions. But the central application conclusion—that gamma_S is a reasonable trial-level surrogate—is an artifact of the chosen margin, as shown by Table S2 and by the fixed-margin evaluation p-value of about 0.051. The conditional verdict is appropriate: the authors should fix or externally justify the margin, rerun the application under pre-specified margins, and reframe the single-trial result as internal consistency rather than validation.","tokens_in":35,"tokens_out":15767,"duration_ms":266275,"concrete_test":"Recompute the application with a fixed, pre-specified epsilon=0.10 in both the screening and evaluation stages, using the same paired two one-sided test; report the number of selected genes and the evaluation p-value for gamma_S. If zero genes are selected or the evaluation p-value is >= 0.05, the 'reasonable surrogate' conclusion is an artifact of the data-adaptive margin rather than an externally validated property.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Equation (2) sets epsilon = max(0, bUY - u*). In the application bUY > u*, so for any candidate gene the screening contrast is bdelta_j - epsilon = (bUY - bUS_j) - (bUY - u*) = u* - bUS_j. The outcome Y enters the test statistic only through the estimated variance of bdelta_j, not through the effect comparison; the test is effectively a one-sample test that the gene's own pre/post change exceeds the power-based threshold u*. The same cancellation occurs in the evaluation stage, where the reported p=0.003 for gamma_S is driven by bUS_gamma exceeding u*, not by bUS_gamma matching bUY. Table S2 shows that a fixed screening margin of epsilon=0.10 selects zero genes; for the evaluation data, fixing epsilon=0.10 gives TOST p = max(Phi((-0.038-0.10)/0.038), 1-Phi((-0.038+0.10)/0.038)) ≈ 0.051, so the 'reasonable surrogate' claim fails. The single-arm paired design also cannot separate vaccination effects from time trends, so trial-level surrogacy is not identifiable even with a fixed margin.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces RISE, a two-stage rank-based procedure for screening and evaluating high-dimensional surrogate markers in small-sample settings. Stage one applies a univariate non-inferiority test to each candidate surrogate, with multiple testing correction, and stage two combines the selected markers into a weighted composite that is evaluated on an independent data split using a two one-sided test (TOST) procedure. The treatment effect measures are U-statistics (UY and US) based on comparing treated and control observations. The authors report simulation studies of false positive rate, false discovery proportion, and power under two data-generating processes, and apply RISE to a single-arm influenza vaccination study, claiming a 222-gene expression signature is a reasonable trial-level surrogate for the neutralising antibody response. The RISE package is available on CRAN and the paper includes reproducible R markdown files.","tokens_in":28641,"tokens_out":10965,"duration_ms":99037,"significance":"The paper targets a real gap: high-dimensional surrogate evaluation in small samples, where parametric surrogacy methods and kernel-based nonparametric alternatives struggle. The use of U-statistics, sample splitting, and multiple testing correction is sensible and the code/package availability is a strength. If the method worked as advertised, it would be a practical addition to the vaccinology toolkit. However, the current formulation of the non-inferiority margin is data-adaptive in a way that makes the surrogate test nearly tautological: it largely tests whether each candidate marker has its own detectable treatment effect rather than whether it approximates the effect on the primary outcome. The application additionally lacks a placebo arm, so the estimated 'treatment effects' are pre-post contrasts that cannot identify trial-level surrogacy. These issues are load-bearing for the paper's central claims, although they are in principle addressable through a fixed, pre-specified margin and a properly controlled study design.","major_comments":[{"comment":"The adaptive margin epsilon = max(0, bUY - u*) causes the primary outcome to cancel from the screening statistic. For any candidate j, the one-sided contrast is bdelta_j - epsilon = (bUY - bUS_j) - (bUY - u*) = u* - bUS_j, so the screening test reduces to testing whether the surrogate's own treatment effect exceeds the power-based threshold u*. The outcome Y enters only through the variance estimate, not through the effect comparison. This means the screening step does not evaluate whether US_j approximates UY. In the application, bUY = 0.97 yields epsilon = 0.29, and any gene with a sufficiently strong pre-post change passes; Table S2 shows that with a fixed epsilon = 0.10, zero genes are selected, and the evaluation TOST p-value for the reported bdelta = -0.038 and bsigma = 0.038 would be approximately 0.051. The 'reasonable surrogate' conclusion is therefore an artifact of the data-adaptive margin rather than evidence of closeness between US and UY.","section":"Section 2.2, Eq. (2)"},{"comment":"The simulation definition of a valid surrogate is circular. A variable is classified as valid if US > UY - epsilon, and epsilon is set to bUY - 0.5. Since UY* = 0.985 in DGP 1, this means any variable with US > 0.5 is valid. Consequently, the reported type I error, false discovery proportion, and power characterise the test of whether a surrogate has any positive treatment effect, not whether it captures the treatment effect on Y. The false positive rate simulation at the boundary US = 0.5 is a calibration check for a null hypothesis that is unrelated to the surrogacy question. The authors acknowledge in Section 2.5 that this choice of epsilon is 'unlikely to be useful', but the simulation results are still presented as supporting RISE as a surrogate identification tool.","section":"Section 2.5"},{"comment":"The application is a single-arm study without a placebo control. Pre-vaccination day 0 measurements are treated as the 'control' observations (Y0, S0), while post-vaccination measurements are taken on day 1 for the surrogates and day 28 for the outcome. Under the counterfactual notation of Section 2.1, Y0 is the outcome under no vaccine at the same follow-up time. Comparing day 0 to day 28 post-vaccine estimates the probability that a post-vaccine value exceeds a pre-vaccine value, which conflates vaccine effects with time trends, assay drift, and regression to the mean. As a result, UY and US are not identified treatment effects, and trial-level surrogacy is not identifiable in this application, even with a fixed and clinically meaningful margin.","section":"Section 3.2"},{"comment":"The evaluation stage uses the same adaptive margin, epsilon = bUY - u* = 0.14. In the TOST, p(1) = Phi((u* - bUS)/sigma) and p(2) = 1 - Phi((2bUY - bUS - u*)/sigma). With bUY = 0.96, p(2) is essentially zero for any plausible value of bUS, so the combined test is effectively a one-sided test that bUS > u*. The reported p = 0.003 is therefore not evidence that delta is small in any fixed sense; it is evidence that the composite marker shows a strong pre-post change. The Spearman correlation of 0.77 in Figure 5 is a marginal association and does not establish trial-level surrogacy. This reinforces that the evaluation step, like the screening step, is not testing whether the surrogate's treatment effect approximates that of the primary outcome.","section":"Section 3.2, evaluation stage"}],"minor_comments":[{"comment":"The text uses 'empirical FDR' and 'false discovery proportion' interchangeably (e.g., Figure 3 caption and the paragraph discussing Scenario 2). The paper should use 'FDP' consistently, since FDR traditionally denotes the expected value of the FDP under a multiple testing procedure.","section":"Section 3.1"},{"comment":"The column header says 'Bonferroni Adjusted p-value' but the table also includes unadjusted p-values; the header should distinguish the columns more clearly. Also, several gene symbols contain typographical spacing errors (e.g., 'V AMP5' and 'TNF AIP6' should be 'VAMP5' and 'TNFAIP6').","section":"Table 1 and Table S1"},{"comment":"The p-value formula 'pj = P(Z < bδj) where Z ∼ N(ϵ, bσδj )' is unconventional; it is clearer to write pj = Phi((bδj - epsilon)/bσδj ) or to define the test statistic explicitly as a standard normal variate under the null.","section":"Section 2.2"},{"comment":"The caption contains a duplicated string 'ρ = 0.77ρ = 0.77ρ = 0.77'; this should be a single value.","section":"Figure 5 caption"}],"recommendation":"major_revision","confidential_remarks":"The paper is borderline. The methodological idea of using rank-based U-statistics for high-dimensional surrogate screening is reasonable and the authors have made their code and data reproducible. However, the data-adaptive margin in Eq. (2) is not a harmless tuning choice: it changes the scientific null hypothesis being tested. The screening and evaluation effectively test whether each marker has a strong pre-post change, and the application's conclusion disappears when a fixed margin of 0.10 is used. The lack of a placebo arm in the application is a separate identifiability problem. These issues are fixable in principle by pre-specifying epsilon and re-analyzing a properly controlled dataset, but as it stands the paper's central applied claim is not supported. I would not reject outright because the framework could be salvaged, but the revision needs to be substantive, not cosmetic."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know before you read it. The two-step screening-plus-evaluation architecture is a genuine extension of Parast et al. (2024), and the simulations are clean. But the adaptive non-inferiority margin in Eq. (2) guts the surrogacy test: when bUY > u*, the screening p-value is a function of u* - bUS, not of bdelta, so screening is effectively testing whether each gene's own pre/post change exceeds a power threshold. The outcome Y enters only through the variance. The same cancellation affects the evaluation TOST; the reported p=0.003 is mostly a statement about the combined marker's own effect, not about its proximity to UY. A fixed margin of 0.10 selects zero genes in screening, and the evaluation p rises to ~0.051. That is a load-bearing flaw for the application claim.\n\nThe paper does several things well. It ships code and a CRAN package; the paired-sample variance derivation is straightforward and useful; the sensitivity analysis is reported honestly; the discussion of limitations (ad hoc epsilon, univariate screening, surrogate-of-a-surrogate) is candid. The simulations, while simple, show controlled FPR and reasonable power under their DGPs.\n\nThe application also lacks a control arm; pre/post measurements in one trial cannot identify a trial-level treatment effect, so calling gamma_S a 'reasonable trial-level surrogate' overreaches. The authors acknowledge the design limitation but still frame the result as a demonstration of RISE's practicality.\n\nFor a methods reader, the two-stage idea is worth discussing, and the code is a nice resource. But as written, the default adaptive margin undermines the method's claim to screen for surrogates. The fix is to require a prespecified or externally justified margin, and to add paired-design simulations that vary the true delta. The application should be reframed as a case study in sensitivity, not a validated surrogate.\n\nI would send it to peer review — a serious referee can push on the margin issue and the design — but my own verdict is that the application does not support the abstract's conclusion.","headline":"RISE has a useful two-stage skeleton and honest simulations, but the adaptive margin cancels the outcome from the screening test, so the application's surrogate claim rests on a permissive, data-dependent threshold.","tokens_in":29151,"tokens_out":4343,"would_cite":false,"duration_ms":38229,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G10","62P10","62G20"],"pacs":[],"model":"deepseek-v4-flash","headline":"A new rank-based procedure, RISE, claims that a 222-gene early-blood-expression signature is a trial-level surrogate for the neutralizing antibody response to influenza vaccination.","keywords":["surrogate marker","high-dimensional screening","rank-based inference","nonparametric testing","gene expression signature","vaccine trial","transcriptomics","non-inferiority test"],"falsifier":"The applied claim would be settled by re-running the screening with a pre-specified ε fixed at 0.10 rather than the data-derived 0.29; the paper's own sensitivity table shows that no genes pass, so the conclusion that a reasonable surrogate exists would fail. A complementary check is external validation: applying the 222-gene combination to independent data from another trivalent influenza vaccine cohort and observing a δ confidence interval that does not fall within the pre-specified margin would falsify the trial-level surrogacy claim.","tokens_in":28132,"feed_emoji":"🧬","tokens_out":9573,"duration_ms":88462,"temperature":0.7,"pith_summary":"The paper proposes RISE, a two-stage procedure for finding, among thousands of candidate variables, a small combined marker that can stand in for a later clinical outcome, in settings where sample sizes are small and the number of candidate variables is large. RISE first tests each candidate individually with a non-parametric rank-based comparison of treatment effects, applying multiple-testing correction, then combines the candidates that pass into a weighted score and evaluates that score's treatment effect on independent data. Simulation results across two data-generating processes indicate controlled false positive rates for samples above about 30 and high power when surrogate strength is high. Applied to gene expression from an inactivated influenza vaccine study, RISE selects a 222-gene signature whose early day-1 expression tracks the day-28 neutralising antibody response among females, with a rank correlation of 0.77 and an evaluation p-value of 0.003. The signature is enriched in innate antiviral and interferon-stimulated genes, giving it a coherent biological interpretation.","feed_headline":"222-gene signature stands in for flu vaccine antibody response","feed_subtitle":"Early gene activity, screened and validated by RISE, predicts day-28 antibody response in small trials.","key_machinery":"The load-bearing object is the rank-based effect measure U = P(value under vaccine > value under control) + 1/2 P(tie), estimated by a U-statistic, and the difference δ = UY - US, which places surrogacy on a common probability scale. RISE tests the non-inferiority hypotheses H0: δ ≥ ε and H0: δ ≤ -ε, so a surrogate is accepted only if both one-sided tests reject and the confidence interval for δ lies within [-ε, ε]. Screening uses these tests per candidate with a false-discovery-rate correction; evaluation combines selected candidates into γ_S = Σ |δ_j|^{-1} S̄_j, a standardized weighted sum in which stronger surrogates receive more weight, and then applies the same rank-based test to γ_S on split-sample data. The paired-data extension, closed-form variance from correlated U-statistic theory, and adaptive choice of ε from the desired power complete the procedure.","core_discovery":"In the paper's own terms, RISE establishes that trial-level surrogacy can be assessed by comparing two rank-based effect measures: UY, the probability that a randomly chosen vaccinated individual has a higher outcome than a randomly chosen control, and USj, the same quantity for a candidate marker. A candidate is deemed valid when the difference δj = UY - USj lies within a non-inferiority margin ε, tested with unpaired or paired U-statistics and a two one-sided procedure that requires δj to fall inside [-ε, ε]. The screening step applies this test to every variable with multiplicity correction; the evaluation step combines the survivors into a standardized weighted sum γ_S whose weights are the inverse of the estimated |δj|, splits the data so evaluation uses independent observations, and re-tests the combined marker. In the application, the screening stage with ε = 0.29 and the strictest multiple-testing correction retained 222 genes; on the held-out evaluation data, the combined γ_S had δ = -0.038 (95% CI [-0.10, 0.025], p = 0.003). The authors conclude that γ_S is a reasonable trial-level surrogate for the neutralising antibody response to the 2008-2009 trivalent inactivated influenza vaccine in females.","pith_inferences":["The data-dependent choice of ε makes 'reasonable surrogate' a relative statement: an external standard with ε fixed at 0.10 would eliminate the signature, so RISE users should pre-register ε before screening.","RISE addresses trial-level surrogacy, not individual-level prediction; combining it with principal-stratification or meta-analytic surrogacy methods could connect the selected marker to individual treatment effects and cross-trial generalization.","A testable extension is to screen gene sets rather than single genes, aggregating co-expressed or pathway-defined modules before applying RISE; this could recover multivariate surrogates that univariate screening misses.","If the signature generalizes, the same two-stage rank pipeline could be repurposed to other vaccines or to correlates-of-protection searches, where the endpoint is infection rather than antibody titre."],"forward_implications":["In trials with 30-200 participants, RISE can screen tens of thousands of markers while keeping the false positive rate near nominal and achieving good power for strong surrogates.","A day-1 post-vaccination gene expression score could serve as a trial-level surrogate for the day-28 neutralizing antibody response, allowing vaccine candidates to be down-selected weeks earlier.","The identified 222-gene signature points to innate antiviral and interferon pathways as the biological bridge between early transcriptomic changes and later antibody production.","Because surrogate evaluation uses independent split-sample data, the reported strength is not an artifact of selecting genes on the same observations that evaluate them.","With paired designs and ordinal outcomes supported, the method can be applied beyond this vaccine study to other settings where early molecular readouts are candidates for later clinical endpoints."],"supporting_citations":[{"why":"Supplies the surrogate definition motivating the criterion: a valid marker is one on which a test of treatment effect mirrors the test on the outcome.","marker":"[13]"},{"why":"Provides the single-marker rank-based non-inferiority procedure that RISE extends to multiple markers, including the adaptive epsilon calibration from desired power.","marker":"[23]"},{"why":"Provides the variance theory for differences of correlated U-statistics, giving standard errors and confidence intervals for δ.","marker":"[26]"},{"why":"Supplies the two one-sided tests used to require δ to fall within [-ε, ε].","marker":"[27]"},{"why":"False-discovery-rate adjustment used in screening; one of the three corrections compared in simulations.","marker":"[30]"},{"why":"Conservative multiple-testing correction applied in the vaccine data application to select the 222 genes.","marker":"[31]"},{"why":"Source of the influenza vaccine trial transcriptomic and antibody data used to construct and evaluate the gene signature.","marker":"[33]"},{"why":"Documents sex differences in the response to influenza vaccination, motivating the female-only analysis.","marker":"[35]"},{"why":"Functional enrichment method that identifies innate antiviral and interferon pathways in the selected gene set.","marker":"[42]"}],"fun_headline_variants":["RISE pinpoints 222-gene surrogate for flu vaccine response","222-gene signature predicts flu vaccine antibody response via RISE","New method RISE identifies surrogate genes for flu vaccine immunity","Two-stage RISE method finds 222 genes as flu vaccine surrogate"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything rests on the non-inferiority margin ε being meaningful: RISE computes ε from the observed treatment effect and the desired power in the same data, so the set of genes declared surrogates changes with this internal choice, and a smaller externally fixed ε would leave no signature at all.","fun_headline_variants_meta":{"raw":{"variants":["RISE pinpoints 222-gene surrogate for flu vaccine response","222-gene signature predicts flu vaccine antibody response via RISE","New method RISE identifies surrogate genes for flu vaccine immunity","Two-stage RISE method finds 222 genes as flu vaccine surrogate"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000431,"raw_usage":{"total_tokens":2248,"prompt_tokens":1042,"completion_tokens":1206,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":658,"completion_tokens_details":{"reasoning_tokens":1134}},"tokens_in":658,"tokens_out":1206,"duration_ms":10740,"temperature":1.0,"reasoning_tokens":1134,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T10:07:38.526887+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"The applied claim would be settled by re-running the screening with a pre-specified ε fixed at 0.10 rather than the data-derived 0.29; the paper's own sensitivity table shows that no genes pass, so the conclusion that a reasonable surrogate exists would fail. A complementary check is external validation: applying the 222-gene combination to independent data from another trivalent influenza vaccine cohort and observing a δ confidence interval that does not fall within the pre-specified margin would falsify the trial-level surrogacy claim.","supporting_citations":[{"cited_title":"Prentice","cited_arxiv_id":null,"evidence_quote":"Supplies the surrogate definition motivating the criterion: a valid marker is one on which a test of treatment effect mirrors the test on the outcome."},{"cited_title":"A rank-based approach to evaluate a surrogate marker in a small sample setting","cited_arxiv_id":null,"evidence_quote":"Provides the single-marker rank-based non-inferiority procedure that RISE extends to multiple markers, including the adaptive epsilon calibration from desired power."},{"cited_title":"DeLong, David M","cited_arxiv_id":null,"evidence_quote":"Provides the variance theory for differences of correlated U-statistics, giving standard errors and confidence intervals for δ."},{"cited_title":"Schuirmann","cited_arxiv_id":null,"evidence_quote":"Supplies the two one-sided tests used to require δ to fall within [-ε, ε]."},{"cited_title":"Controlling the false discovery rate: A practical and pow- erful approach to multiple testing","cited_arxiv_id":null,"evidence_quote":"False-discovery-rate adjustment used in screening; one of the three corrections compared in simulations."},{"cited_title":"Multiple comparisons among means","cited_arxiv_id":null,"evidence_quote":"Conservative multiple-testing correction applied in the vaccine data application to select the 222 genes."},{"cited_title":"Thomas, Barry Smith, Henry Schaefer, Jieming Chen, Zicheng Hu, Kelly A","cited_arxiv_id":null,"evidence_quote":"Source of the influenza vaccine trial transcriptomic and antibody data used to construct and evaluate the gene signature."},{"cited_title":"Hejblum, Noah Simon, Vladimir Jojic, Cornelia L","cited_arxiv_id":null,"evidence_quote":"Documents sex differences in the response to influenza vaccination, motivating the female-only analysis."},{"cited_title":"Systematic and integrative analysis of large gene lists using david bioinformatics resources","cited_arxiv_id":null,"evidence_quote":"Functional enrichment method that identifies innate antiviral and interferon pathways in the selected gene set."}],"review_version":1}