{"id":"285e7d85-73cc-4237-9df6-693a051d1b4d","arxiv_id":"2507.15985","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A response to Chiba (2025) showing the claimed contradiction in the Kaplan-Meier plug-in estimator for average hazard rests on undefined quantities, with a simulation confirming the estimator's reliability.","lead":"This commentary argues that a recent published critique of the Kaplan-Meier plug-in estimator for average hazard is based on a false mathematical contradiction. It uses logic and a simulation study to show the estimator stays approximately unbiased whether the truncation time hits an observed event time or not.","discovery_kind":"replication","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The rebuttal of Chiba's proof is sound, but the affirmative claim of general small-sample reliability rests on one constant-hazard simulation; non-constant hazards and heavy censoring remain untested.","rationale":"The paper's core negative claim—that Chiba's proof by contradiction fails—does not depend on the simulation and is well supported: the disputed term is either undefined or 0/0, so no contradiction follows. The paper also gives a correct didactic example showing why average hazard is not flat when the survival curve is flat. The weakest point is the positive generalization: the simulation covers only one parametric family (exponential), one censoring pattern, and no hazard changes within intervals, yet the abstract and conclusion state that the estimator provides reliable estimates across a range of truncation times even in small samples. The reader's weakest_assumption identifies exactly this gap, and I agree. The appropriate verdict remains CONDITIONAL: accept the rebuttal, but require either broader simulations or a more qualified claim about small-sample reliability. Since my concern matches the reader's and does not change the verdict, the verdict should be UNCHANGED.","tokens_in":4402,"tokens_out":5571,"duration_ms":64948,"concrete_test":"Re-run the simulation with a non-constant hazard, e.g., h(t)=0.01 for t<50 and h(t)=0.05 for t≥50 (or a decreasing hazard), with censoring at 120, n ∈ {10, 30, 50, 100}, 1000 replications, and compute the Monte Carlo bias E[\\hat AH(τ)] − AH(τ) on a fine τ grid, especially just before and after 50. If the maximum absolute relative bias stays below, say, 5% for all τ, the generality claim is supported; if the bias spikes at τ slightly above 50 for small n, the conclusion should be qualified to constant or slowly varying hazards and the verdict stays conditional.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The logical rebuttal in the section 'Logical Gaps in Chiba's Proof by Contradiction' is convincing: Chiba's alleged contradiction depends on the term \\hat f(τ)/\\hat h(τ), which is undefined (or the indeterminate form 0/0) when τ is not an observed event time, so the proof does not establish that Formula (4) is incorrect. However, the paper's affirmative conclusion—'the Kaplan-Meier plug-in estimator (1) remains approximately unbiased regardless of whether τ falls between observed event times, as our simulation study demonstrates'—is supported only by the exponential/hazard-0.01/censor-at-120 simulation in 'Sampling Variability Should Be Taken into Account' (Figure A). That scenario has a constant true hazard, so AH(τ) is constant and the estimator's decline between event times is sampling noise that averages out across replications. The simulation does not exercise cases where the true hazard changes within a censoring interval, for example a step increase in h(t) at a time t0 that is not an observed event time. In such a case \\hat S(t) is flat through t0 until the next observed event, so Formula (4)'s denominator accumulates person-time with an overestimated survival while the numerator is unchanged, potentially producing a nontrivial bias for τ just after t0, especially in small samples. This does not contradict the asymptotic consistency of the plug-in estimator (already established in the literature), but it undermines the paper's general claim of reliability 'even in small samples.' The conclusion that investigators can continue to apply the estimator with confidence therefore goes beyond the evidence provided.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a commentary on Chiba (2025), which claimed that the Kaplan–Meier plug-in estimator of the average hazard AH(τ) is incorrect when τ is not an observed event time and proposed a harmonic-mean reinterpretation. The authors argue that Chiba's proof-by-contradiction fails because the ratio f_hat(τ)/h_hat(τ) is 0/0 or undefined at non-event times, that Chiba's single-sample illustration confuses sampling variability with systematic bias, and that a flat-survival example shows AH(τ) itself declines between events. A simulation study with exponential hazards and administrative censoring at 120 indicates that the plug-in estimator is approximately unbiased across τ for sample sizes 10–100. The paper concludes that Formula (4) is not incorrect and that investigators can continue using it.","tokens_in":4736,"tokens_out":8948,"duration_ms":95645,"significance":"The paper's logical rebuttal is persuasive and useful: it correctly identifies the undefined quantity in Chiba's contradiction, and the flat-survival example is a nice illustration of why the average hazard is not a step function. The simulation, while limited to a single distribution, supports the claim that the specific example in Chiba (2025) is a small-sample artifact. If the overbroad generality of the affirmative conclusion is addressed, this commentary would be a valuable correction to the literature. The authors are also to be credited for making the software (survAH) available, which aids reproducibility.","major_comments":[{"comment":"The simulation study uses only one data-generating process—exponential with constant hazard 0.01 and administrative censoring at 120—so the conclusion that the estimator 'remains approximately unbiased regardless of whether τ falls between observed event times' is not supported for general survival distributions. In particular, if the true hazard changes at a time t0 that is not an observed event time, the Kaplan–Meier survival estimate S_hat(t) is flat through t0 until the next observed event, while the true S(t) declines; the plug-in estimator then uses an overestimated survival in the denominator (and an underestimated cumulative incidence in the numerator) for τ just after t0, which can produce non-negligible finite-sample bias in small samples. The authors should either add simulations with non-constant hazards (e.g., Weibull or step-hazard models) and heavier censoring, or qualify the conclusion to the scenarios actually studied.","section":"Sampling Variability Should Be Taken into Account (Figure A) and Conclusion"}],"minor_comments":[{"comment":"The caption says 'Average deviation from the true average hazard' while the y-axis is labeled 'Average of the AH estimates'; please align the caption with the axis labels.","section":"Figure A caption"},{"comment":"The simulation section reports results only graphically; adding a table with average bias, Monte Carlo standard error, and coverage probability would make the 'approximately unbiased' claim easier to assess.","section":"Sampling Variability Should Be Taken into Account"},{"comment":"The displayed harmonic-mean rewrite of Formula (4) is hard to read in the manuscript; please ensure the numerator and denominator are clearly typeset as a ratio.","section":"Logical Gaps in Chiba's Proof by Contradiction"},{"comment":"The flat-survival example uses a hazard that is exactly zero on [2,5]; while mathematically valid, it would be more persuasive to also show a nonzero but time-varying hazard example, since zero hazard over an interval is a special limiting case.","section":"Why dAH(τ) Based on Formula (4) Is Not Flat Between Events"},{"comment":"The conclusion says 'applying discrete-time logic to a continuous-time estimator'; earlier the paper notes Chiba never states a distribution assumption, so it may be clearer to say 'applying discrete-time reasoning to a continuous-time estimand and its plug-in estimator.'","section":"Conclusion"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within scope for a statistical methods journal. The main concern is that the affirmative claim of general finite-sample reliability goes beyond the simulation evidence; this can be fixed with additional simulations or by scoping the conclusion. No concerns about novelty or citation pattern."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper before the next round of the survival-analysis argument: the logical rebuttal is solid, and the stress-test note's main concern is fair but shouldn't block the paper. Chiba's proof-by-contradiction fails because the ratio f-hat(τ)/h-hat(τ) is 0/0 or undefined when τ is not an observed event time. That's a clean diagnosis, and the paper shows it clearly. The flat-survival example is also well chosen: it demonstrates that a flat survival curve doesn't imply a flat average hazard, because the denominator keeps accumulating person-time. That part is genuinely new and useful.\n\nThe paper does not introduce a new estimator or theory; its novelty is the targeted refutation of a published claim. That's fine for a comment. The simulation is simple—exponential hazard 0.01, censoring at 120, sample sizes 10–100—and it shows the plug-in estimator is approximately unbiased across truncation times in that scenario. The stress-test note is right that this doesn't establish general small-sample reliability for non-constant hazards or heavy censoring. The paper's conclusion that investigators can continue to apply the estimator 'with confidence' goes a bit beyond the evidence shown. But this is a minor overreach, not a flaw in the central argument. The central claim—that Chiba's contradiction doesn't hold—is independent of the simulation and stands on the logic alone.\n\nI'd also note that the paper is honest about relying on earlier asymptotic results (ref [1]), and it doesn't hide the limitation of the simulation. The writing is clear, and the tone is measured. This is a paper for applied statisticians who might have been confused by Chiba's critique, and for anyone citing the average hazard literature. It deserves serious peer review—the rebuttal is important enough to get into the permanent record, even if the positive evidence is narrower than the conclusion. I'd accept it with a request to soften the general reliability claim or add one non-constant hazard scenario.\n\nFor your own use: I'd cite it if I were writing on average hazards, and I'd bring it to a reading group as a good example of a focused methodological correction. Recommend sending it to review.","headline":"A correct, narrowly scoped logical rebuttal of Chiba's proof, though the affirmative small-sample claim rests on a single constant-hazard simulation.","tokens_in":5232,"tokens_out":972,"would_cite":true,"duration_ms":12814,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62N01","62N02","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"The Kaplan-Meier plug-in estimator of the average hazard is valid even when the truncation time falls between observed event times.","keywords":["average hazard","Kaplan-Meier plug-in estimator","harmonic mean","survival analysis","continuous-time estimator","discrete-time logic","indeterminate form","finite-sample bias"],"falsifier":"Simulate event times from a non-constant hazard, such as a Weibull with decreasing hazard, apply heavy censoring, place the truncation time $\\tau$ inside a long event-free gap, and average the plug-in estimates over many replications; if the average deviates systematically from the true $AH(\\tau)$, the claim that the estimator is reliable for any $\\tau$ would be contradicted.","tokens_in":4253,"feed_emoji":"📊","tokens_out":10178,"duration_ms":89945,"temperature":0.7,"pith_summary":"This commentary defends the Kaplan-Meier plug-in estimator of the average hazard against a recently published claim that the estimator is 'incorrect' whenever the truncation time τ falls between observed event times. The authors argue that the alleged contradiction in that critique comes from an indeterminate 0/0 ratio and from applying discrete-time reasoning to a continuous-time estimator. They demonstrate with simulations, including sample sizes as small as 10, that the estimator's finite-sample bias is negligible for all truncation times in the range considered. The practical consequence is that investigators can keep using the plug-in estimator without restricting τ to observed event times.","feed_headline":"Average-hazard estimator shows negligible bias between observed events","feed_subtitle":"A rebuttal dissolves the claimed contradiction into a 0/0 form and shows negligible bias in small-sample simulations.","key_machinery":"The carrying object is the ratio defining the average hazard, $AH(\\tau)=\\{1-S(\\tau)\\}/\\{\\int_0^\\tau S(u)\\,du\\}$, with the Kaplan-Meier estimate $\\hat S(u)$ plugged in for the unknown survival curve. The argument turns on two mechanisms: first, Chiba's algebraic rewrite of the estimator into harmonic-mean form requires evaluating the hazard at $\\tau$ when no event was observed, where the hazard estimate is undefined and the ratio $\\hat f(\\tau)/\\hat h(\\tau)$ becomes $0/0$; second, even when the survival curve is exactly flat between event times, the denominator $\\int_0^\\tau S(u)\\,du$ continues to grow, so the true average hazard declines on open intervals between events. This second mechanism shows that a declining estimate between events is a feature of the estimand, not evidence of bias.","core_discovery":"The central claim is that Formula (4), the Kaplan-Meier plug-in estimator of the average hazard $AH(\\tau)=\\{1-S(\\tau)\\}/\\{\\int_0^\\tau S(u)\\,du\\}$, is not incorrect when $\\tau$ lies between observed event times. Chiba's proof by contradiction fails because the quantity $\\hat f(\\tau)/\\hat h(\\tau)$ is the indeterminate form $0/0$: $\\hat h(\\tau)$ is undefined at a non-event time, so asserting that $\\hat f(\\tau)=0$ contradicts $\\hat f(\\tau)/\\hat h(\\tau)>0$ only if that ratio were well defined. The paper also shows that the flatness of the Kaplan-Meier and Nelson-Aalen step functions between events does not imply the average-hazard estimate should be flat, since the integral in the denominator keeps accumulating person-time and pushes $AH(\\tau)$ downward during a gap. Simulation results with a constant hazard of 0.01 and censoring at 120 show that, averaged over 1000 replicates, the plug-in estimate tracks the true value for sample sizes 10 to 100 across all tested truncation times.","pith_inferences":["A natural stress test the authors leave untried is a simulation under a non-constant hazard (for example, a Weibull with shape below 1) with heavier censoring; if the plug-in estimator shows systematic bias there, the paper's general reassurance would need to be narrowed.","The paper evaluates bias by averaging estimates; a practitioner examining a single realized curve will still see the step-like declines between event times, so the visual pattern Chiba pointed to remains a real feature even though it is not bias.","The same continuous-versus-discrete distinction should apply to other estimators formed by plugging step functions into integrals, suggesting that similar critiques based on step-function intuition may fail for related survival estimands."],"forward_implications":["The Kaplan-Meier plug-in estimator can be used for the average hazard at any truncation time $\\tau$, including times that do not coincide with observed events, without introducing material bias.","A single-sample scattered pattern in the estimate across $\\tau$ is expected finite-sample variation, not evidence that the estimator is invalid.","The apparent contradiction in the harmonic-mean reinterpretation disappears once the undefined hazard at non-event times is acknowledged.","Investigators do not need to restrict their choice of $\\tau$ to the set of observed event times when reporting average-hazard estimates."],"supporting_citations":[{"why":"Defines the average hazard, introduces the Kaplan-Meier plug-in estimator, and reports its asymptotic properties—the method whose validity is defended here.","marker":"[1]"},{"why":"Alternative treatment-effect measures under nonproportional hazards; cited to situate the average hazard as a generalized hazard measure.","marker":"[2]"},{"why":"Another source for generalized hazard treatment-effect measures, supporting the estimand's definition in the introduction.","marker":"[3]"},{"why":"The paper criticized; its harmonic-mean reinterpretation and proof-by-contradiction supply the target of the rebuttal.","marker":"[4]"}],"fun_headline_variants":["Average hazard estimator survives Chiba's critique","Chiba's average-hazard paradox reduced to 0/0","Simulations vindicate plug-in estimator for average hazard","Average hazard: Chiba's objection rests on 0/0"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The rebuttal's positive claim that the estimator is reliable regardless of where $\\tau$ falls relies on a single simulation scenario—an exponential hazard with censoring at 120 and sample sizes 10 to 100—so its generality to other survival distributions is assumed rather than demonstrated.","fun_headline_variants_meta":{"raw":{"variants":["Average hazard estimator survives Chiba's critique","Chiba's average-hazard paradox reduced to 0/0","Simulations vindicate plug-in estimator for average hazard","Average hazard: Chiba's objection rests on 0/0"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000844,"raw_usage":{"total_tokens":3641,"prompt_tokens":875,"completion_tokens":2766,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":491,"completion_tokens_details":{"reasoning_tokens":2699}},"tokens_in":491,"tokens_out":2766,"duration_ms":18325,"temperature":1.0,"reasoning_tokens":2699,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T15:20:21.499032+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate event times from a non-constant hazard, such as a Weibull with decreasing hazard, apply heavy censoring, place the truncation time $\\tau$ inside a long event-free gap, and average the plug-in estimates over many replications; if the average deviates systematically from the true $AH(\\tau)$, the claim that the estimator is reliable for any $\\tau$ would be contradicted.","supporting_citations":[{"cited_title":"Ratio and difference of average hazard with survival weight: New measures to quantify survival benefit of new therapy","cited_arxiv_id":null,"evidence_quote":"Defines the average hazard, introduces the Kaplan-Meier plug-in estimator, and reports its asymptotic properties—the method whose validity is defended here."},{"cited_title":"Treatment effect measures under nonproportional hazards","cited_arxiv_id":null,"evidence_quote":"Alternative treatment-effect measures under nonproportional hazards; cited to situate the average hazard as a generalized hazard measure."},{"cited_title":"Treatment effect measures under nonproportional hazards","cited_arxiv_id":null,"evidence_quote":"Another source for generalized hazard treatment-effect measures, supporting the estimand's definition in the introduction."},{"cited_title":"Average hazard as harmonic mean","cited_arxiv_id":null,"evidence_quote":"The paper criticized; its harmonic-mean reinterpretation and proof-by-contradiction supply the target of the rebuttal."}],"review_version":1}