Pith. sign in

REVIEW 1 major objections 5 minor 6 references

Comment on "Average Hazard as Harmonic Mean" by Chiba

T0 review · 1 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The Kaplan-Meier plug-in estimator of the average hazard is valid even when the truncation time falls between observed event times.

desk verdict A correct, narrowly scoped logical rebuttal of Chiba's proof, though the affirmative small-sample claim rests on a single constant-hazard simulation. read the letter →

arxiv 2507.15985 v1 pith:37QK3VNY submitted 2025-07-21 stat.ME stat.CO

classification stat.MEstat.CO MSC 62N0162N0262P10
keywords averagehazardKaplan-Meierplug-inestimatorharmonicmeansurvivalanalysiscontinuous-timediscrete-timelogicindeterminateformfinite-samplebias
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This commentary defends the Kaplan-Meier plug-in estimator of the average hazard against a recently published claim that the estimator is 'incorrect' whenever the truncation time τ falls between observed event times. The authors argue that the alleged contradiction in that critique comes from an indeterminate 0/0 ratio and from applying discrete-time reasoning to a continuous-time estimator. They demonstrate with simulations, including sample sizes as small as 10, that the estimator's finite-sample bias is negligible for all truncation times in the range considered. The practical consequence is that investigators can keep using the plug-in estimator without restricting τ to observed event times.

What carries the argument

The carrying object is the ratio defining the average hazard, $AH(\tau)=\{1-S(\tau)\}/\{\int_0^\tau S(u)\,du\}$, with the Kaplan-Meier estimate $\hat S(u)$ plugged in for the unknown survival curve. The argument turns on two mechanisms: first, Chiba's algebraic rewrite of the estimator into harmonic-mean form requires evaluating the hazard at $\tau$ when no event was observed, where the hazard estimate is undefined and the ratio $\hat f(\tau)/\hat h(\tau)$ becomes $0/0$; second, even when the survival curve is exactly flat between event times, the denominator $\int_0^\tau S(u)\,du$ continues to grow, so the true average hazard declines on open intervals between events. This second mechanism shows that a declining estimate between events is a feature of the estimand, not evidence of bias.

What would settle it

Simulate event times from a non-constant hazard, such as a Weibull with decreasing hazard, apply heavy censoring, place the truncation time $\tau$ inside a long event-free gap, and average the plug-in estimates over many replications; if the average deviates systematically from the true $AH(\tau)$, the claim that the estimator is reliable for any $\tau$ would be contradicted.

Watch

Extended reading notes

Core claim

The central claim is that Formula (4), the Kaplan-Meier plug-in estimator of the average hazard $AH(\tau)=\{1-S(\tau)\}/\{\int_0^\tau S(u)\,du\}$, is not incorrect when $\tau$ lies between observed event times. Chiba's proof by contradiction fails because the quantity $\hat f(\tau)/\hat h(\tau)$ is the indeterminate form $0/0$: $\hat h(\tau)$ is undefined at a non-event time, so asserting that $\hat f(\tau)=0$ contradicts $\hat f(\tau)/\hat h(\tau)>0$ only if that ratio were well defined. The paper also shows that the flatness of the Kaplan-Meier and Nelson-Aalen step functions between events does not imply the average-hazard estimate should be flat, since the integral in the denominator keeps accumulating person-time and pushes $AH(\tau)$ downward during a gap. Simulation results with a constant hazard of 0.01 and censoring at 120 show that, averaged over 1000 replicates, the plug-in estimate tracks the true value for sample sizes 10 to 100 across all tested truncation times.

Load-bearing premise

The rebuttal's positive claim that the estimator is reliable regardless of where $\tau$ falls relies on a single simulation scenario—an exponential hazard with censoring at 120 and sample sizes 10 to 100—so its generality to other survival distributions is assumed rather than demonstrated.

Editorial extensions

If this is right

  • The Kaplan-Meier plug-in estimator can be used for the average hazard at any truncation time $\tau$, including times that do not coincide with observed events, without introducing material bias.
  • A single-sample scattered pattern in the estimate across $\tau$ is expected finite-sample variation, not evidence that the estimator is invalid.
  • The apparent contradiction in the harmonic-mean reinterpretation disappears once the undefined hazard at non-event times is acknowledged.
  • Investigators do not need to restrict their choice of $\tau$ to the set of observed event times when reporting average-hazard estimates.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural stress test the authors leave untried is a simulation under a non-constant hazard (for example, a Weibull with shape below 1) with heavier censoring; if the plug-in estimator shows systematic bias there, the paper's general reassurance would need to be narrowed.
  • The paper evaluates bias by averaging estimates; a practitioner examining a single realized curve will still see the step-like declines between event times, so the visual pattern Chiba pointed to remains a real feature even though it is not bias.
  • The same continuous-versus-discrete distinction should apply to other estimators formed by plugging step functions into integrals, suggesting that similar critiques based on step-function intuition may fail for related survival estimands.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 5 minor

Summary. This manuscript is a commentary on Chiba (2025), which claimed that the Kaplan–Meier plug-in estimator of the average hazard AH(τ) is incorrect when τ is not an observed event time and proposed a harmonic-mean reinterpretation. The authors argue that Chiba's proof-by-contradiction fails because the ratio f_hat(τ)/h_hat(τ) is 0/0 or undefined at non-event times, that Chiba's single-sample illustration confuses sampling variability with systematic bias, and that a flat-survival example shows AH(τ) itself declines between events. A simulation study with exponential hazards and administrative censoring at 120 indicates that the plug-in estimator is approximately unbiased across τ for sample sizes 10–100. The paper concludes that Formula (4) is not incorrect and that investigators can continue using it.

Significance. The paper's logical rebuttal is persuasive and useful: it correctly identifies the undefined quantity in Chiba's contradiction, and the flat-survival example is a nice illustration of why the average hazard is not a step function. The simulation, while limited to a single distribution, supports the claim that the specific example in Chiba (2025) is a small-sample artifact. If the overbroad generality of the affirmative conclusion is addressed, this commentary would be a valuable correction to the literature. The authors are also to be credited for making the software (survAH) available, which aids reproducibility.

major comments (1)
  1. [Sampling Variability Should Be Taken into Account (Figure A) and Conclusion] The simulation study uses only one data-generating process—exponential with constant hazard 0.01 and administrative censoring at 120—so the conclusion that the estimator 'remains approximately unbiased regardless of whether τ falls between observed event times' is not supported for general survival distributions. In particular, if the true hazard changes at a time t0 that is not an observed event time, the Kaplan–Meier survival estimate S_hat(t) is flat through t0 until the next observed event, while the true S(t) declines; the plug-in estimator then uses an overestimated survival in the denominator (and an underestimated cumulative incidence in the numerator) for τ just after t0, which can produce non-negligible finite-sample bias in small samples. The authors should either add simulations with non-constant hazards (e.g., Weibull or step-hazard models) and heavier censoring, or qualify the conclusion to the scenarios actually studied.
minor comments (5)
  1. [Figure A caption] The caption says 'Average deviation from the true average hazard' while the y-axis is labeled 'Average of the AH estimates'; please align the caption with the axis labels.
  2. [Sampling Variability Should Be Taken into Account] The simulation section reports results only graphically; adding a table with average bias, Monte Carlo standard error, and coverage probability would make the 'approximately unbiased' claim easier to assess.
  3. [Logical Gaps in Chiba's Proof by Contradiction] The displayed harmonic-mean rewrite of Formula (4) is hard to read in the manuscript; please ensure the numerator and denominator are clearly typeset as a ratio.
  4. [Why dAH(τ) Based on Formula (4) Is Not Flat Between Events] The flat-survival example uses a hazard that is exactly zero on [2,5]; while mathematically valid, it would be more persuasive to also show a nonzero but time-varying hazard example, since zero hazard over an interval is a special limiting case.
  5. [Conclusion] The conclusion says 'applying discrete-time logic to a continuous-time estimator'; earlier the paper notes Chiba never states a distribution assumption, so it may be clearer to say 'applying discrete-time reasoning to a continuous-time estimand and its plug-in estimator.'

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the rebuttal's 0/0 argument and its simulation against a known true average hazard are self-contained; the paper's self-citations are peripheral background, not load-bearing.

full rationale

The paper rebuts Chiba's proof-by-contradiction with a purely logical argument: Chiba's alleged contradiction depends on the ratio \(f̂(\ tau)/\ĥ(\ tau)\), which is the indeterminate form 0/0 when \(\ĥ(\ tau)=0\), and Chiba himself concedes \(\ĥ(t)\) is undefined at non-event times. This step is self-contained and does not rest on any fitted parameter, prior result, or citation. The simulation study (Figure A) is an independent Monte Carlo check against an analytically known truth: for an exponential hazard of 0.01, AH(\ tau) equals 0.01 exactly for every \ tau, so the estimator is evaluated against an external benchmark rather than against a quantity derived from its own outputs; no parameter is fitted and renamed a prediction. The 'concrete scenario' with a flat survival curve is a direct computation from the definition of AH (AH(2)=1 versus AH(5)≈0.68), illustrating that the true average hazard itself declines when survival is flat; it explains why the estimator's between-event decline is expected, rather than deriving the estimator's validity from itself. Self-citations are present (refs [1], [5], [6] are by the present authors), but none is load-bearing: the 0/0 rebuttal and the simulation stand alone, and ref [1] merely supplies background asymptotic properties of the standard KM plug-in estimator, while refs [5]-[6] point to software. No equation in the paper reduces to its own input, no uniqueness theorem or ansatz is imported from the authors' prior work, and no known result is renamed. The affirmative generalization beyond the one simulated scenario (constant hazard, censoring at 120) is broader than the evidence shown, but that is an evidentiary scope concern, not circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper introduces no free parameters because the simulation settings (hazard 0.01, censoring time, sample sizes) are illustrative, not fitted to data. The main burden falls on interpreting Chiba's notation and on the adequacy of one simulation scenario.

assumptions (3)
  • domain assumption Kaplan-Meier plug-in estimator for AH is consistent and asymptotically normal (ref [1])
    The paper relies on previously established asymptotic properties to frame the simulation check, citing Uno and Horiguchi (2023).
  • domain assumption Chiba's harmonic-mean rewrite defines f-hat(τ)=0 and h-hat(τ)=0 (or leaves h-hat(τ) undefined) at non-event times
    The refutation hinges on reading Chiba's inserted terms at τ as zero or undefined; this is a reasonable reading given Chiba's own statement that h-hat cannot be defined at non-event times.
  • domain assumption The continuous-time framework is the appropriate setting for defining AH and its plug-in estimator
    Chiba's discrete-time view is the target of the critique; the paper assumes AH is a continuous-time quantity as originally defined.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Comment on "Average Hazard as Harmonic Mean" by Chiba." pith.science (2026). https://pith.science/paper/37QK3VNY

@misc{pith2026250715985,
  author       = {Pith},
  title        = {Pith review of: Comment on "Average Hazard as Harmonic Mean" by Chiba},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/37QK3VNY}},
  note         = {Machine review of arXiv:2507.15985}
}
read the original abstract

In a recent article published in Pharmaceutical Statistics, Chiba proposed a reinterpretation of the average hazard as a harmonic mean of the hazard function and questioned the validity of the Kaplan-Meier plug-in estimator when the truncation time does not coincide with an observed event time. In this commentary, we examine the arguments presented and highlight several points that warrant clarification. Through simulation studies, we further show that the plug-in estimator provides reliable estimates across a range of truncation times, even in small samples. These support the continued utilization of the Kaplan-Meier plug-in estimator for the average hazard and help clarify its proper interpretation and implementation.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

6 extracted references · 6 canonical work pages

  1. [1]

    Ratio and difference of average hazard with survival weight: New measures to quantify survival benefit of new therapy

    Uno H, Horiguchi M. Ratio and difference of average hazard with survival weight: New measures to quantify survival benefit of new therapy. Stat. Med. 2023; 42(7): 936–952

  2. [2]

    Treatment effect measures under nonproportional hazards

    Snapinn S, Jiang Q, Ke C. Treatment effect measures under nonproportional hazards. Pharm. Stat. 2023; 22(1): 181-193

  3. [3]

    Treatment effect measures under nonproportional hazards

    Jackson D, Sweeting M, Baker R. Treatment effect measures under nonproportional hazards. Pharm. Stat. 2025; 24(2): e2449

  4. [4]

    Average hazard as harmonic mean

    Chiba Y. Average hazard as harmonic mean. Pharm. Stat. 2025; 24(2): e70009

  5. [5]

    survAH: Survival Data Analysis using Average Hazard

    Uno H, Horiguchi M. survAH: Survival Data Analysis using Average Hazard. https://cran.r-project. org/web/packages/survAH/index.html

  6. [6]

    uno1lab/survAH

    Uno H, Horiguchi M, Qian Z. uno1lab/survAH. https://doi.org/10.5281/zenodo.15312214. 5

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.