{"id":"5f3b032c-1644-430a-8554-970ab8500095","arxiv_id":"2506.07387","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"A Bayesian utility-based endpoint combining tumor burden AUC with survival time to estimate treatment effects in oncology trials, evaluated only on simulations with an internal data-generation mismatch.","lead":"This paper proposes a single per-patient score that combines tumor burden changes with progression-free survival time, and defines the treatment effect as the difference in average scores between arms. The Bayesian method is tested only on simulations that contain an internal inconsistency, so its value in real early-phase trials is not yet established.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The simulation evidence for controlled Type I error is invalidated by a data-generation/fitted-model mismatch: outcomes are simulated with ψ=t/S_i but modeled with ψ=t/T_i; re-running the simulation under the stated model is required before the central claim can be assessed.","rationale":"The reader's overall REJECT verdict is appropriate, and my concern reinforces it rather than changing it. The reader flagged the ψ=t/S_i versus ψ=t/T_i mismatch in the rationale but did not make it the formal weakest assumption; that is why I mark agreement as partial. My analysis sharpens the empirical objection: the mismatch means the simulation never evaluates the proposed estimator under its own model, so the claimed controlled Type I error is unsupported. I also note that the Scenario 3 TBAUC 'null' is not a null in the estimand actually defined, since the integration bound min(L_i,T_i) depends on survival; this makes the 0.062 rejection uninterpretable as a Type I error. Both issues are checkable by re-simulation because the authors provide code. I do not see a separate basis to accept the paper; the central claim remains unverified even though the proposed utility-based endpoint is a coherent extension of quality-adjusted survival and joint modeling ideas.","tokens_in":14034,"tokens_out":6287,"duration_ms":78253,"concrete_test":"Re-run the simulation with Section 6.1.3 changed to ψ_ij = t_ij/T_i (T_i generated from the exponential in 6.1.2), all other settings unchanged, and reproduce Table 3. If the Scenario 4 total-AUC rejection rate moves materially away from 0.025 or Scenario 3 TBAUC rejection changes, the reported Type I error control is an artifact of the mismatch; if unchanged, the mismatch is not the operative cause and a separate misspecification analysis is needed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 6.1.3 generates tumor burden with ψ_ij = t_ij/S_i, where S_i = min(T_i,C_i), but the fitted model in Section 4.2 and the likelihood in (11) use ψ_ij = t_ij/T_i with T_i the true event time. The two agree only for uncensored subjects (δ_i=1). For censored subjects, S_i=C_i<T_i, so the observed Y_ij are drawn from a quadratic in t/C_i while the model treats them as a quadratic in t/T_i and imputes T_i>C_i. The longitudinal sub-model is therefore misspecified for every censored patient in every replication, and the posterior TBAUC draws used in the Wald tests are computed under a likelihood that did not generate the data. Hence Table 3's Scenario 4 rejection rates cannot be read as a valid Type I error check of the proposed procedure. Separately, the Scenario 3 test of H0:θ_TB=0 is not a null test: because TBAUC integrates up to min(L_i,T_i), a survival effect alone changes E[TBAUC|A] even with β_A=0; the 0.062 rejection is a real effect on that estimand, not the claimed propagation artifact. Both issues undermine the paper's central empirical claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a Bayesian utility-based composite endpoint for early-phase oncology trials. For each patient, a utility function u_i(t) equals a modeled tumor burden trajectory m_i(t) before the true event time T_i and a constant Gamma after T_i; the endpoint U_i is the integral of u_i over a follow-up window [0,L_i]. The joint model factorizes into an exponential survival sub-model and a longitudinal sub-model in which tumor burden is quadratic in scaled time psi_ij = t_ij/T_i, with censoring handled by data augmentation. The average treatment effect ATE = E[U|A=1] - E[U|A=0] is estimated from posterior draws, and Wald-type tests are proposed for the TBAUC, SAUC, and total AUC components. A simulation study with four scenarios is used to claim controlled Type I error rates for the proposed tests.","tokens_in":14283,"tokens_out":13050,"duration_ms":137681,"significance":"The proposed estimand is an interesting attempt to operationalize a composite endpoint in the spirit of ICH E9(R1), and the analytic integration for TBAUC in Eq. (16) is clean and computationally convenient. The use of the Bayesian bootstrap to account for posterior variance is also reasonable. However, the manuscript's central empirical claim rests on a simulation study that is internally inconsistent with the stated model, and the component-wise tests do not isolate the effects their labels imply. The paper would make a useful contribution if these issues were resolved, but in its current form the evidence for controlled Type I error is not valid.","major_comments":[{"comment":"Section 6.1.3 simulates Y_ij using psi_ij = t_ij/S_i with S_i = min(T_i,C_i), whereas the model in Section 4.2 and the likelihood in Eq. (11) use psi_ij = t_ij/T_i with T_i the true event time. For censored patients, S_i = C_i < T_i, so the generated trajectory is quadratic in t/C_i while the fitted model treats it as quadratic in t/T_i and imputes T_i > C_i. In addition, the baseline measurement Y_i0 is generated as N(15,0.5) independently of X_i and A_i, while Eq. (9) implies E[Y_i0 | X_i,A_i] = beta_X X_i + beta_A A_i, which is approximately 9 under the simulation settings. Thus the simulated data are not a draw from the fitted model even at baseline, and the rejection rates in Table 3 cannot be interpreted as frequentist operating characteristics of the proposed procedure. The simulation must be re-run with the longitudinal model used for inference, or the model must be changed to match the data generation, and the paper should clarify whether baseline measurements are included in the likelihood.","section":"6.1.3, 4.2, Eq. (11)"},{"comment":"In Scenario 3, the null hypothesis H_TB: theta_TB = 0 is false by construction even though beta_A = 0. The TBAUC estimand in Eq. (16) integrates the trajectory up to min(L,T_i) and scales time by T_i, so E[TBAUC|A] depends on the survival distribution through T_i. With gamma_A = -0.75, the treatment arm has a different T_i distribution, and therefore theta_TB is nonzero. The observed rejection rate of 0.062 is thus the power of the test against a true alternative, not a Type I error rate. More generally, the labels 'longitudinal AUC only' and 'survival AUC only' in Table 1 are misleading: TBAUC is not a pure longitudinal effect, and SAUC is not a pure survival effect, because the posterior distribution of censored T_i is influenced by the longitudinal model. The paper should either redefine these hypotheses in terms of the structural coefficients (beta_A, gamma_A) or clearly state that the tests are for composite estimands whose null values are not implied by the absence of a coefficient-level effect.","section":"6.2, 6.3, Eq. (16), Table 3"}],"minor_comments":[{"comment":"Table 1 states the alternatives as H_A: theta > 0 for all three tests, but Section 6.3 describes the tests as one-sided left-tailed. With the sign convention that beneficial treatment reduces tumor burden and prolongs survival, the relevant alternative is theta < 0. Please reconcile the table and the text so that the reported rejection rates are interpretable.","section":"Table 1 and Section 6.3"},{"comment":"The phrase 'relatively controlled Type I error rates' is vague; in Scenario 4 the one-sided rejection rates are 0.006, 0.020, and 0.010, which are far below the nominal 0.025. The authors should discuss whether this conservatism is a feature of the method and what its implications are for power.","section":"Abstract and Section 6.3"},{"comment":"The statement that the survival AUC contribution is 'Gamma times one minus the restricted mean survival time' is only exact when L_i is a common constant; otherwise the relevant quantity is Gamma * E[(L_i - T_i) I(T_i < L_i)]. Please clarify this point.","section":"Section 3"},{"comment":"The variance estimators are defined as variances across posterior iterations of Bayesian bootstrap weighted averages, but the text does not explicitly state that a fresh Dirichlet(1_{n_a}) weight vector is drawn at each iteration q. Please make this explicit for reproducibility.","section":"Section 5.2"},{"comment":"The reference 'Group, I.E.W. et al.' should be replaced with the proper citation of the ICH E9(R1) addendum, and the figure captions describing 'administrative censoring' should be harmonized with Section 6.1.2, which generates exponential censoring times.","section":"References and figure captions"}],"recommendation":"major_revision","confidential_remarks":"I am not persuaded that the current simulation evidence can be salvaged without re-running all four scenarios under the model in Section 4.2. The Scenario 3 issue is conceptual rather than computational: the paper needs to reframe what the component-wise tests are testing. If the authors are willing to make these changes, the paper could be publishable; in its present form, the central claims are not supported."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear Colleague,\n\nThe short version: the paper proposes a utility-based composite endpoint for early-phase oncology trials that integrates tumor burden and PFS. The idea is worth taking seriously, but the simulation evidence that carries the paper is not valid as it stands. The data generation uses ψ=t/S_i while the model uses ψ=t/T_i, so the longitudinal sub-model is misspecified for every censored patient in every replication. That undercuts the Type I error claims in Table 3.\n\nWhat is actually new: the specific combination of a conditional-on-survival longitudinal model with scaled time by the true event time, plus the post-event penalty in the AUC, is a reasonable extension of Q-TWiST and RMST-style endpoints. The analytic integrals in Equations 16–17 are correct, the Bayesian implementation is transparent, and the code is available. The paper also honestly acknowledges the parametric assumptions and that the functional form is not tied to the method.\n\nThe main problems are the ones above. First, Section 6.1.3 simulates Y_ij using S_i = min(T_i,C_i) in the denominator of ψ, but Section 4.2 and the likelihood in (11) use T_i. For uncensored patients these coincide; for censored patients they do not. The model imputes T_i>C_i while the data have generated the trajectory as a quadratic in t/C_i. So the model is fitting the wrong covariance structure, and the rejection rates in Scenario 4 cannot be read as a valid Type I error check. This needs to be re-run with data generated from the actual model, or the model changed to match the simulation.\n\nSecond, Scenario 3's test of θ_TB=0 is not a null test. TBAUC integrates up to min(L_i,T_i), so a survival-only effect changes the estimand even when β_A=0. The 0.062 rejection rate is a real effect on that estimand, not a propagation artifact. The paper's framing is misleading.\n\nThird, there is no real-data application and no comparison to simpler alternatives like RMST or Q-TWiST. The quadratic-in-scaled-time assumption is load-bearing and untested.\n\nIf the simulation is corrected and the tests are interpreted properly, the approach could be useful for early-phase decision-making. As it stands, the central empirical claim is not established. I'd send it to a referee who understands joint modeling and clinical trials, but the referee should be asked to verify the simulation design carefully. It's not a desk-reject in my view, because the underlying proposal is coherent and potentially valuable.","headline":"A promising composite endpoint undermined by a simulation that does not follow the fitted model; needs correction before the claims can be assessed.","tokens_in":14837,"tokens_out":3047,"would_cite":false,"duration_ms":33831,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62N01","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"One utility score captures tumor burden and survival effects.","keywords":["utility-based endpoint","tumor burden","progression-free survival","joint modeling","Bayesian inference","average treatment effect","area under the curve","early-phase oncology trials"],"falsifier":"Simulate trials where the true tumor burden trajectory is not quadratic in scaled time — for example, a steep early drop followed by a plateau — fit the proposed quadratic model, and check whether the total-AUC test keeps Type I error near the nominal level and the ATE estimate stays unbiased. If bias or error inflation appears, the estimand's validity under misspecification is refuted; the paper currently provides no such robustness check or real-data validation.","tokens_in":13773,"feed_emoji":"🎯","tokens_out":6895,"duration_ms":60638,"temperature":0.7,"pith_summary":"This paper proposes a utility-based Bayesian endpoint that combines longitudinal tumor burden measurements with progression-free survival into a single area-under-the-curve score, and a joint model for estimating the average treatment effect on that score in early-phase oncology trials. The central claim is that this scalar estimand, $U_i = \\int_0^{L_i} u_i(t)\\,dt$ with $u_i(t)=m_i(t)$ before the event and $u_i(t)=\\Gamma$ after, captures treatment efficacy signals from both outcomes efficiently even when data are scarce, and could be developed into a Phase 3 confirmatory endpoint. Simulations across four scenarios show low bias in parameter estimates and controlled Type I error for the total-AUC test under the no-effect scenario, with high power when both tumor-burden and survival effects are present. The approach matters because it turns two outcomes measured in different units into one interpretable composite estimand without requiring large confirmatory studies.","feed_headline":"One utility score captures tumor burden and survival effects","feed_subtitle":"Bayesian joint model estimates this composite AUC with controlled Type I error in simulations.","key_machinery":"The central object is the utility function $u_i(t)=m_i(t)\\mathbf{1}(t\\le T_i)+\\Gamma\\mathbf{1}(t>T_i)$, whose definite integral $U_i=\\int_0^{L_i}u_i(t)\\,dt$ forms the composite endpoint. The tumor burden trajectory $m_i(\\psi)$ is modeled as a quadratic in scaled time $\\psi_{ij}=t_{ij}/T_i$ with random effects, and the survival sub-model is exponential with hazard $\\lambda_i=\\exp(\\gamma_A A_i+\\gamma_X^\\top X_i)$. Data augmentation draws censored event times $T_i$ from their posterior, allowing closed-form computation of $U_i$ as the sum of a tumor-burden AUC and a survival AUC $\\Gamma(L_i-T_i)\\mathbf{1}(T_i<L_i)$; Bayesian bootstrap weights estimate the variance for Wald tests.","core_discovery":"The paper's discovery is a new estimand and estimation procedure: define each patient's utility as the tumor burden trajectory before death or progression and a fixed penalty $\\Gamma$ afterward, integrate it over the study window, and take the average treatment effect as the difference in these integrals between arms. The joint model factorizes the likelihood into an exponential survival sub-model and a longitudinal sub-model where tumor burden is quadratic in time scaled by the patient's event time, with censored event times imputed by data augmentation. The endpoint decomposes exactly into a tumor-burden AUC and a survival AUC equal to $\\Gamma$ times the post-event time, linking the composite to restricted mean survival time. The authors show by simulation that Wald tests on the total AUC control Type I error and that the component tests behave as expected, apart from a propagation of survival effects into the tumor-burden test when only survival is affected.","pith_inferences":["The paper's reliance on a quadratic-in-scaled-time trajectory is the main fragility; a natural stress test is to simulate from non-quadratic trajectories (e.g., a two-phase or spline curve) and check whether the total-AUC test keeps its Type I error and low bias.","The Scenario 3 inflation in the tumor-burden test suggests the component tests are not independent when the joint factorization conditions on survival; users should interpret them as descriptive rather than as separate confirmatory tests.","Sensitivity to the choice of $\\Gamma$ is unexplored; for large $\\Gamma$ the endpoint approaches a survival-only estimand, so reporting results across a clinically motivated range of $\\Gamma$ would strengthen practical use.","A real-data application with RECIST measurements would be the decisive next step, since simulations are generated under the same functional form the model assumes."],"forward_implications":["Early-phase oncology trials can summarize both tumor response and progression-free survival in one pre-specified estimand, reducing the need to choose between biomarker and survival endpoints.","The decomposition into TBAUC and SAUC lets analysts attribute a treatment effect to tumor shrinkage, survival benefit, or both, which is helpful for go/no-go decisions.","The closed-form integrals and posterior data augmentation keep computation tractable, so the method is feasible in the small samples typical of phase II studies.","The estimand aligns with the ICH E9 (R1) composite-strategy framework, giving a regulatory path toward using it as a secondary or confirmatory endpoint."],"supporting_citations":[{"why":"Supplies the factorization of the joint distribution into survival and longitudinal-conditional-on-event components that the proposed model uses.","marker":"Hogan and Laird (1997)"},{"why":"Provides the standard joint modeling framework for longitudinal and survival data that the paper adapts to a utility-based endpoint in small samples.","marker":"Rizopoulos (2012)"},{"why":"The ICH E9 (R1) addendum justifies composite-strategy estimands, the regulatory basis for combining tumor burden and survival into one utility.","marker":"Group et al. (2020)"},{"why":"Defines RECIST progression criteria for tumor burden, connecting the longitudinal outcome to the clinical definition of progression.","marker":"Eisenhauer et al. (2009)"},{"why":"Introduces the Bayesian bootstrap used to compute variance estimates for the AUC-based treatment effect tests.","marker":"Rubin (1981)"},{"why":"Provides the restricted mean survival time quantity that links the survival component of the utility endpoint to a standard survival estimand.","marker":"Uno et al. (2014)"},{"why":"Gives the sample-size formula used to size the simulation's event count and power.","marker":"Schoenfeld (1983)"}],"fun_headline_variants":["Composite AUC integrates tumor burden and survival","Bayesian utility score merges tumor and survival data","One estimand fuses tumor burden and survival effects"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes each patient's true tumor burden follows a quadratic curve when plotted against time scaled by their true event time, and that censored event times can be imputed from this model; if the real trajectory is not quadratic, the estimated treatment effect is biased.","fun_headline_variants_meta":{"raw":{"variants":["Composite AUC integrates tumor burden and survival","Bayesian utility score merges tumor and survival data","One estimand fuses tumor burden and survival effects"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000254,"raw_usage":{"total_tokens":1523,"prompt_tokens":852,"completion_tokens":671,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":468,"completion_tokens_details":{"reasoning_tokens":634}},"tokens_in":468,"tokens_out":671,"duration_ms":6669,"temperature":1.0,"reasoning_tokens":634,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:36:35.192806+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate trials where the true tumor burden trajectory is not quadratic in scaled time — for example, a steep early drop followed by a plateau — fit the proposed quadratic model, and check whether the total-AUC test keeps Type I error near the nominal level and the ATE estimate stays unbiased. If bias or error inflation appears, the estimand's validity under misspecification is refuted; the paper currently provides no such robustness check or real-data validation.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the factorization of the joint distribution into survival and longitudinal-conditional-on-event components that the proposed model uses."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the standard joint modeling framework for longitudinal and survival data that the paper adapts to a utility-based endpoint in small samples."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The ICH E9 (R1) addendum justifies composite-strategy estimands, the regulatory basis for combining tumor burden and survival into one utility."},{"cited_title":"A., Therasse, P., Bogaerts, J., Schwartz, L","cited_arxiv_id":null,"evidence_quote":"Defines RECIST progression criteria for tumor burden, connecting the longitudinal outcome to the clinical definition of progression."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the Bayesian bootstrap used to compute variance estimates for the AUC-based treatment effect tests."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the restricted mean survival time quantity that links the survival component of the utility endpoint to a standard survival estimand."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the sample-size formula used to size the simulation's event count and power."}],"review_version":1}