REVIEW 3 major objections 4 minor 6 references
A revisit to maximum likelihood estimation of Weibull model parameters
T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read The paper proves Weibull maximum-likelihood estimates are unique and approximately Weibull-distributed.
desk verdict Clean uniqueness theorem, promising empirical idea, but the main validation table is internally inconsistent and the claim is not supported. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The paper's engine is the likelihood-equation function $h(\delta|\mathbf{X}) = \frac{1}{\delta} + \ln \tilde{X} - \frac{\sum_i X_i^\delta \ln X_i}{\sum_i X_i^\delta}$, where $\tilde{X}$ is the sample geometric mean. The proof that the MLE exists and is unique runs through this function: it goes to $+\infty$ as $\delta\to 0^+$, to $\ln(\tilde{X}/X_{(n)})<0$ as $\delta\to\infty$, and its derivative is non-positive by the Cauchy–Schwarz inequality, forcing exactly one crossing. The second mechanism is the moment-matching approximation: simulated $\hat\delta$ and $\hat\beta/\beta$ values are fitted to Weibull distributions with parameters $a,b,c,d$, which are then summarized by the four regression equations $a=a_0+a_1\ln(n+a_2)$, $b=b_0\delta/(1+b_1\ln(\ln n))$, $c=c_0\delta\ln(n+c_1)$, and $d=d_0+d_1/n+d_2/\delta+d_3/(n\delta)$. These equations turn the simulation output into a closed-form recipe for quantiles.
What would settle it
Run a fresh simulation on off-grid values such as $n=15$, $\delta=0.75$ and $\delta=7.3$, with $10^5$ replications, and compare the empirical 10th and 90th percentiles of $\hat\delta$ and $\hat\beta/\beta$ against the quantiles predicted by equations (3.2)–(3.5); a discrepancy larger than the reported simulation variation would show the Weibull approximation fits its training grid but does not predict.
Extended reading notes
Core claim
The paper's central claims are two. The first is Theorem 2.1: the maximum likelihood equation $h(\delta|\mathbf{X})=0$ for the Weibull shape parameter always has a solution, and that solution is unique. The proof decomposes $h$ into a term that diverges to $+\infty$ as $\delta\to 0^+$, a limit $\ln(\tilde{X}/X_{(n)})<0$ as $\delta\to\infty$, and a derivative that is non-positive by the Cauchy–Schwarz inequality, so the continuous function crosses zero exactly once. The second claim is empirical: based on $10^5$ simulated replications for $n=10,20,\ldots,100$ and $\delta=0.5,1.0,\ldots,10$, the paper hypothesizes that $\hat\delta \approx W(a,b)$ and $\hat\beta/\beta \approx W(c,d)$, with $a,b,c,d$ obtained by moment matching and then summarized by the regression equations (3.2)–(3.5). Quantile comparisons in Section 4 indicate the fit is good for moderate and large $p$ and improves with $n$, and the implied biases and variances approach the previously reported first-order asymptotic expressions for $n\ge 50$ for $\delta$ and $n\ge 10$ for $\hat\beta/\beta$.
Load-bearing premise
The load-bearing premise is that the sampling distributions of the maximum-likelihood estimators are themselves Weibull, which the paper adopts after visual inspection and moment-matching on the same simulated data used later for validation rather than proving it or testing it out of sample.
Editorial extensions
If this is right
- For a practitioner, the regression equations turn the MLE's sampling uncertainty into closed-form Weibull quantiles, so confidence intervals for the shape and scale can be built without re-simulation.
- The reported fit improves with sample size, and the implied bias and variance approach the known first-order asymptotic values at $n$ around 50 for the shape and around 10 for the scaled scale estimator, giving a concrete rule for when large-sample formulas are safe.
- Because the distributions of $\hat\delta$ and of $\hat\beta/\beta$ depend only on $\delta$ and $n$, a common set of tables covers all scale values, making the approximation scale-invariant in practice.
- The authors state that the approximation is intended for small to moderately large samples ($n$ up to 100), so it fills the region where asymptotic normality is not yet comfortable.
- The same quantile-comparison procedure can be extended to censored data, though the censoring mechanism would enter the approximation.
Reading between the lines
- Because the fitted shape parameter $a$ is nearly constant in $\delta$ for fixed $n$, a testable prediction is that the coefficient of variation of $\hat\delta$ depends mainly on $n$; an independent simulation could check whether this collapses to a simple sample-size law.
- The same moment-matching-plus-regression recipe could be tried on other two-parameter families such as gamma or lognormal to see whether MLE sampling distributions are generically self-similar; the paper does not claim this.
- A concrete next check is to reproduce Table 4.1 from equations (3.2)–(3.5) and the coefficients in Table 3.5; if the entries do not match, the regression summaries are not internally consistent with the tabulated sampling-distribution parameters.
- An out-of-sample check on off-grid $(n,\delta)$ pairs would separate the regression's predictive value from its in-sample fit, since the current validation uses the same simulation grid that produced the fitted coefficients.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper revisits maximum likelihood estimation for the two-parameter Weibull distribution. It proves existence and uniqueness of the solution to the profile score equation h(δ|X)=0 (Theorem 2.1), and it shows that the sampling distributions of the MLEs depend only on n and δ apart from scale (Result 2.1). The main empirical claim is that the sampling distributions of δ̂ and β̂/β can be approximated by two-parameter Weibull distributions W(a,b) and W(c,d), with parameters obtained by moment matching and then summarized by regression equations in n and δ. The paper validates this claim by comparing Weibull percentiles with empirical percentiles from M=10^5 simulations and by comparing first-order bias and variance expressions with asymptotic counterparts in Section 4.3.
Significance. If the proposed Weibull approximations to the sampling distributions were properly validated, the paper would offer a practical tool for small-sample inference on Weibull parameters and would complement existing first-order asymptotic results. The proof of Theorem 2.1 is clean and correct: the monotonicity of h(δ|X) follows from the Cauchy-Schwarz inequality and the stated limits are correct. Result 2.1 is a useful invariance observation. However, the central empirical claim is currently not supported. Table 4.1, which is the main quantitative evidence, is internally inconsistent with the stated formulas, and the validation is in-sample, so the claimed approximation is not demonstrated. No code or data are provided to allow reproduction of the simulations.
major comments (3)
- [Section 4.1, Table 4.1] Table 4.1 cannot be reproduced from Eq. (4.2) and Tables 3.1–3.2. For n=10, δ=1, p=0.5, Eq. (4.2) with a=3.236 and b=1.297 gives Q̂_p = 1.297 × (−ln 0.5)^{1/3.236} ≈ 1.16, but the table lists (4.17, 5.09, 4.18) for (Q̂_p, Q̃_p, Q̂̂_p). The same problem occurs for n=100, δ=1, p=0.5, where the formula gives approximately 1.02 but the table lists approximately 4.60. Since Remark 4.1 uses Table 4.1 to conclude that the W(a,b) approximation has a very good fit, this numerical inconsistency removes the main quantitative support for the paper's central distributional claim.
- [Section 4, Eqs. (4.3)–(4.8)] The validation in Section 4 is in-sample. The same M=10^5 replications that are used to moment-match (a,b) and (c,d) via Eq. (3.1) are also used to compute the empirical percentiles Q̃_p in Eqs. (4.3) and (4.8). Therefore Tables 4.1–4.2 measure how well a Weibull distribution with parameters estimated from a sample matches quantiles of that same sample; they provide no out-of-sample evidence for the hypothesized distributional form. The Weibull-family assumption is asserted after visual inspection of histograms, and the residual diagnostics in Table 3.5 show that model (3.5) for d fails the Shapiro-Wilk and Anderson-Darling tests with p<0.001. An independent validation, either with fresh simulations or a split-sample procedure, and a justification of the Weibull family are needed.
- [Section 3.3, Table 3.5 and Remark 4.2] The authors themselves state after Table 3.5 that model (3.5) for d is not unbiased and that its residuals fail the Shapiro-Wilk and Anderson-Darling tests with p<0.001, yet Remark 4.2 uses this model to conclude that the sampling distribution of β̂/β is approximated very well by W(c,d). This is an internal inconsistency in a load-bearing part of the paper: the scale parameter d is essential to the W(c,d) approximation, so the evidence for the β̂/β distributional claim is not just weak but explicitly contradicted by the authors' own diagnostics. A revised manuscript should either provide a corrected regression form for d with adequate residuals or restrict the W(c,d) claim to settings where the fitted model is actually supported.
minor comments (4)
- [Section 1.1] There are several typos: 'probability destiny function' should be 'probability density function', 'our objective is much brother' should be 'much broader', and Eq. (1.6) has 'form W' which should be 'from W'.
- [Section 4.3, Eqs. (4.16)–(4.23)] The displayed formula for \b\ after Eq. (4.17) has a misplaced division; it should read \b = \b_0 δ / (1 + \b_1 ln(ln n)) rather than \b_0/δ times the bracket. Also, using \B_1 both for the asymptotic bias constant in (1.10) and for the sample-size-dependent quantity in (4.17) is confusing; please rename one of them.
- [Table 4.1, n=80, δ=2, p=0.5] The entry (4.72, 4.10, 4.70) is anomalous because the empirical quantile 4.10 is far below the adjacent values in the same row and below the corresponding theoretical quantiles; this looks like a typographical error and should be checked.
- [References and notation] The in-text citation 'Collett (2015)' does not match the reference list entry 'Collett, D. (2023)'. In addition, the notation \dot\sim is never defined; please state explicitly that it means 'approximately distributed as'.
Circularity Check
In-sample moment matching and validation make the W(a,b) approximation and Section 4.3 bias/variance 'confirmations' reductions of the same simulated data; Theorem 2.1 itself is not circular.
-
fitted input called prediction
[Section 3.2, Eq (3.1); Section 4.1, Eqs (4.2)-(4.3); Remark 4.1]
"Therefore, using the simulated \hat\delta^{(m)}, 1 \leq m \leq M, the parameters a and b are found by solving b\Gamma(1 + 1/a) = \bar{\hat\delta}(.) and b^{2}[\Gamma(1 + 2/a) - {\Gamma(1 + 1/a)}^{2}] = s^{2}(\hat\delta). (3.1) ... We now compare the above \hat Q_p value with the empirical (100p)th percentile of \hat\delta based on \hat\delta^{(m)}, 1 \leq m \leq M = 10^5, defined as \tilde Q_p = (M p)th ordered value of {\hat\delta^{(m)}}."
The parameters (a,b) of the hypothesized W(a,b) distribution are fitted by matching the first two moments of the very same M=10^5 MLE samples (Eq 3.1). The validation then compares quantiles of that fitted W(a,b), Qhat, with empirical quantiles Qtilde from exactly those same samples, so Table 4.1 and Remark 4.1's 'very good fit' is an in-sample goodness-of-fit, not an independent test of the distributional claim. The distributional form is only suggested by histograms, and no held-out simulation or external benchmark is used to check it.
-
fitted input called prediction
[Section 3.2, Eqs (3.2)-(3.3) and Table 3.5; Section 4.1, Remark 4.1]
"This motivates us to fit the following regression line (after several trial and error attempts) ... (3.2) ... Based on the values of a = a(n, \delta) as shown in Table 3.1, the above regression equation (3.2) gives a staggering R-square of almost0.99 indicating a very good fit. ... Remark 4.1: Table 4.1 ... it is noted that (i) \hat Q_p and \hat{\hat Q}_p are nearly identical ... This shows a very good fit for the regression equations (3.2) and (3.3)."
The regression forms and coefficients are selected and fitted to the same tabulated a,b values that they are then used to reproduce. R^2 approximately 0.99 and the closeness of Qhat and Qhat-hat in Remark 4.1 are in-sample measures against the training data, with no held-out (n,\delta) combinations. The residual normality tests in Table 3.5 are also computed on the same residuals used to fit the models, so the regressions are not independently validated as predictive approximations.
1 more flagged steps
-
self definitional
[Section 3.2, Eq (3.1); Section 4.3, Eqs (4.16)-(4.19)]
"Therefore, using the simulated \hat\delta^{(m)}, 1 \leq m \leq M, the parameters a and b are found by solving b\Gamma(1 + 1/a) = \bar{\hat\delta}(.) and b^{2}[\Gamma(1 + 2/a) - {\Gamma(1 + 1/a)}^{2}] = s^{2}(\hat\delta). (3.1) ... Bias(\hat\delta) \approx b\Gamma(1 + 1/a) - \delta = (1/n)[n{b\Gamma(1 + 1/a) - \delta}] \approx (\hat B_1/n) (say), (4.16)"
By Eq (3.1), b\Gamma(1 + 1/a) equals the simulated sample mean of \hat\delta by construction, and the variance expression equals the simulated sample variance. Therefore the 'W-based' bias in Eq (4.16) and variance in Eq (4.18) are just the sample mean minus \delta and the sample variance. The Weibull approximation contributes no independent content to Tables 4.3-4.6; those tables are direct simulation comparisons of sample moments with Tanaka et al.'s asymptotic formulas, despite being presented as products of the hypothesized W(a,b) sampling distribution.
full rationale
The analytical part of the paper is self-contained and not circular: Theorem 2.1 proves existence and uniqueness of the MLE equation using limits and a Cauchy-Schwarz monotonicity argument, with no reliance on the later distributional approximations. The circularity is in the empirical half. In Eq (3.1), the parameters a and b are solved from the sample mean and variance of the same M=10^5 MLE replications that later provide the empirical percentiles in Eq (4.3); hence Table 4.1 and Remark 4.1 are in-sample goodness-of-fit checks rather than independent predictions. Similarly, the regression equations (3.2)-(3.5) are fitted to the very a,b,c,d values whose reproduction (R^2 about 0.99, Remark 4.1) is offered as evidence, with no held-out grid. Section 4.3's 'W-based' bias and variance reduce to the sample moments by Eq (3.1), so Tables 4.3-4.6 are essentially simulation checks of Tanaka et al.'s asymptotic formulas, not independent consequences of the Weibull approximation. A separate correctness problem, not itself circularity, is that Table 4.1 is internally inconsistent with Eq (4.2) and Tables 3.1-3.2: for n=10, delta=1, p=0.5 the computed Qhat is about 1.16 while 4.17 is reported, so the main quantitative support for the W(a,b) fit is not reproducible. Self-citations to Tanaka et al. (2018) are not the load-bearing circular issue; the load-bearing issue is fitting and validating on the same simulated data.
Assumptions & free parameters
free parameters (5)
- Regression coefficients (a0, a1, a2) for shape parameter a(n,delta) =
-26.485, 7.915, 33.125
- Regression coefficients (b0, b1) for scale parameter b(n,delta) =
1.775, 0.463
- Regression coefficients (c0, c1) for shape parameter c(n,delta) =
1.944, -5.782
- Regression coefficients (d0, d1, d2, d3) for scale parameter d(n,delta) =
0.999, -0.031, 0.048, 1.003
- Moment-matched Weibull parameters a,b,c,d on the simulation grid =
Tables 3.1-3.4, e.g., a=3.236, b=1.297 for n=10, delta=1
assumptions (5)
- standard math Cauchy-Schwarz inequality implies h'(delta) <= 0 in the proof of Theorem 2.1.
- domain assumption Observations are iid from a two-parameter Weibull W(delta,beta) with delta,beta>0.
- ad hoc to paper The two-parameter Weibull family is an adequate approximating family for the MLE sampling distributions.
- ad hoc to paper The regression forms (3.2)-(3.5) correctly capture the dependence of a,b,c,d on n and delta.
- domain assumption Monte Carlo replications of size M=100,000 accurately represent the sampling distributions.
Cite this review
Pith. "Pith review of A revisit to maximum likelihood estimation of Weibull model parameters." pith.science (2026). https://pith.science/paper/G3RZFI4A
@misc{pith2026250111604,
author = {Pith},
title = {Pith review of: A revisit to maximum likelihood estimation of Weibull model parameters},
year = {2026},
howpublished = {\url{https://pith.science/paper/G3RZFI4A}},
note = {Machine review of arXiv:2501.11604}
}
read the original abstract
In this work, we revisit the estimation of the model parameters of a Weibull distribution based on iid observations, using the maximum likelihood estimation (MLE) method which does not yield closed expressions of the estimators. Among other results, it has been shown analytically that the MLEs obtained by solving the highly non-linear equations do exist (i.e., finite), and are unique. We then proceed to study the sampling distributions of the MLEs through both theoretical as well as computational means. It has been shown that the sampling distributions of the two model parameters' MLEs can be approximated fairly well by suitable Weibull distributions too. Results of our comprehensive simulation study corroborate some recent results on the first-order bias and first-order mean squared error (MSE) expressions of the MLEs.
Figures
Figures from the paper (12 more)
Reference graph
Works this paper leans on
-
[1]
Almeida, J. B. (1999). Application of Weibull Statistics to the Failure of Coatings, Journal of Materials Processing Technology, 93, 257–263. Aljeddani, S. M. A. and Mohammed, M. A. (2023). A Novel Approach to Weibull Distribution for the Assessment of Wind Energy Speed, Alexandria Engineering Journal, 78, 56 -
work page 1999
-
[3]
Chatfield, C., & Goodhardt, G. J. (1973). A consumer purchasing model with Erlang inter- purchase times. Journal of the American Statistical Association , 68(344), 828-835. Chen, M., Zhang, Z., & Cui, C. (2017). On the bias of the maximum likelihood estimators of parameters of the Weibull distribution. Mathematical and Computational Applications , 22(1),
work page 1973
-
[5]
Estimation, Asymptotic Variances, Journal of Hydrology, 242, 157–170. Johnson, N. L., Kotz, S. and Balakrishnan, N. (1994). Continuous Univariate Distributions , V ol. 1 (Wiley Series in Probability and Statistics) 2nd Edition. Jiang, R., & Murthy, D. N. P. (2011). A study of Weibull shape parameter: Properties and significance. Reliability Engineering & ...
work page 1994
-
[19]
Collett, D. (2023). Modelling survival data in medical research . Chapman and Hall/CRC. Eliazar, I. (2017). Lindy’s law. Physica A: Statistical Mechanics and its Applications , 486, 797-805. Cox, D. R., & Snell, E. J. (1968). A general definition of residuals. Journal of the Royal Statistical Society: Series B (Methodological) , 30(2), 248-265. Fok, S. L....
work page 2023
-
[64]
Austin, L. G., Klimpel, R. R., & Luckie, P. T. (1984). Process engineering of size reduction: ball milling. Society of Mining Engineers of the AIME. Carroll, K. J. (2003). On the Use and utility of the Weibull Model in the Analysis of Survival Data, Controlled Clinical Trials, 24, 682–701. Chandrasekhar, S. (1943). Stochastic problems in physics and astro...
work page 1984
-
[69]
Matsushita, S., Hagiwara, K., Shiota, T., Shimada, H., Kuramoto, K. and Toyokura, Y . (1992). Lifetime Data Analysis of Disease and Aging by the Weibull Probability Distribution, Journal of Clinical Epidemiology, V ol. 45, No. 10, 1165-1175. McCool, J. I. (2012). Using the Weibull distribution: reliability, modeling, and inference (V ol. 950). John Wiley ...
work page 1992
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.