REVIEW 6 minor 23 references
This paper proves that, for any finite Markov-chain characteristic admitting a smooth matrix extension, the plug-in MLE's asymptotic distribution is the Fréchet derivative of that extension applied to the limiting Gaussian matrix of the est
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 18:49 UTC pith:Y5YHVM36
load-bearing objection A clean, correct matrix-level restatement of the delta method for finite Markov chains—modest novelty, solid execution, no hidden flaw.
Matrix asymptotic calculus for plug-in maximum likelihood estimators in finite Markov chains
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central discovery is Theorem 2: for an irreducible transition matrix P and a functional φ defined on stochastic matrices, if φ agrees with a Fréchet-differentiable (or C^r, or real-analytic) map Φ on an open neighborhood of P, then √n(φ(P̂)-φ(P)) converges in distribution to Φ'_P(W_P), where W_P is the centered Gaussian matrix limit of the transition-matrix MLE. The theorem also gives finite-order expansions and shows that any two smooth extensions have identical differentials on the tangent space T_P = {H : H1_s=0, h_ij=0 if p_ij=0}, making the limits representation-independent. The row-stochastic constraints are thus carried by the limiting object rather than removed by a minimal param
What carries the argument
The key machinery is the limiting Gaussian random matrix W_P, with covariance Cov(W_ij, W_kl) = δ_ik π_i^{-1} p_ij(δ_jl - p_il), which stands in for the estimated transition matrix in all asymptotic statements. Theorem 2 is the engine: it applies Fréchet (or higher-order, or analytic) differentials of a smooth matrix-space extension Φ of the functional to W_P, evaluated on the tangent space T_P that encodes the stochastic constraints. Kronecker-product and vectorization identities then convert the resulting linear operator into the covariance matrix needed for inference, which is why one differentiation at matrix level serves every functional.
Load-bearing premise
The load-bearing premise is that the functional admits a smooth (Fréchet-differentiable, C^r, or analytic) extension to an open neighborhood of the true transition matrix in the full matrix space, with irreducibility of that matrix also assumed throughout.
What would settle it
Simulate a long trajectory for a three-state irreducible chain with a zero transition probability, and compare the empirical law of √n(φ(P̂)-φ(P)) for a smooth functional (e.g., the stationary availability) against the Gaussian law Φ'_P(W_P) computed from the paper's covariance formula. Because the limiting matrix W_P has a zero component exactly where p_ij=0, any systematic mismatch would reveal that the tangent-space encoding of the constraints is incomplete, refuting Theorem 2.
If this is right
- Asymptotic distributions for matrix powers, stationary probabilities, resolvent-based characteristics, reliability indicators and entropy-type quantities all become evaluations of the same derivative operator, removing the need for per-functional derivations.
- Finite-dimensional curves such as availability curves inherit a joint Gaussian limit, enabling simultaneous confidence bands that achieve target coverage with smaller width than Bonferroni bands.
- The second-order term is the curvature correction of the functional; for the mean time to failure it reproduces the right skew of the finite-sample distribution at moderate sample sizes.
- Covariance computation through the Kronecker representation is faster than element-wise Jacobian calculation, with the gap widening as the state space grows.
- The same covariance operators directly give Wald tests and confidence regions for linear restrictions on plug-in estimates.
Where Pith is reading between the lines
- The smooth-extension assumption is checked case by case, so a natural next step is to map which common Markov characteristics break it—functions involving zero transition probabilities or singular inverses likely need a separate boundary calculus.
- The tangent-space view suggests the calculus should extend to other constrained parameter spaces, such as semi-Markov kernels, provided a suitable infinite-dimensional Gaussian limit and Fréchet differentiability can be defined.
- The second-order expansion could be turned into a practical bias correction once a closed form for the first-order bias of the transition-matrix MLE is available; the paper keeps that term explicit but does not supply it.
- The computational comparison hints that for large state spaces, deriving and computing covariances at the matrix level may be the preferred default in software, even when the final output is a coordinate-wise interval.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a unified matrix-level asymptotic calculus for plug-in nonparametric maximum likelihood estimators of the transition matrix of a finite Markov chain. After recalling the Gaussian random matrix limit W_P of sqrt(n)(hat P - P), with block-multinomial covariance (2)--(4), the central Theorem 2 states that if a stochastic functional phi agrees locally with a Frechet-differentiable (or C^r, or analytic) extension Phi on an open neighborhood of an irreducible P, then the plug-in estimator satisfies sqrt(n)(phi(hat P)-phi(P)) -> Phi'_P(W_P), together with finite-order expansions (10)--(13) and representation independence along the tangent space T_P. The subsequent sections apply this calculus to matrix powers, resolvents, stationary vectors, additive-functional variances, Kemeny's constant, entropy rates, and a list of reliability indicators (availability, reliability, failure probabilities, MTTF, MTTR, and hitting intensities). Numerical experiments illustrate confidence intervals, simultaneous bands, and second-order curvature refinements.
Significance. If the central theorem holds, the paper provides a genuinely unified source of first- and second-order asymptotics for a broad class of Markov-chain plug-in estimators, avoiding coordinate-wise minimal parametrizations. I checked the main derivations: Theorem 2 is a standard functional-delta-method/Taylor argument, the derivative formulas in the applications are internally consistent, and the representation-independence result in Theorem 2(iv) is valid because tangent-space perturbations stay stochastic. The paper is also honest about its limitations: Remark 2 explicitly distinguishes the curvature term from the complete order-n^{-1} bias correction, and the entropy example in Section 4.4 is restricted to positive transition probabilities, where the smooth-extension hypothesis is available. The numerical simulations are illustrative rather than exhaustive, but they support the claimed practical value of the matrix covariances and the second-order refinement. Overall, the contribution is significant and appropriate for the journal if the local presentation issues are fixed.
minor comments (6)
- [Section 1 (Introduction)] The citation 'Angius et al. (2021)' in the paragraph on phase-type distributions is missing from the reference list. Please add the full reference.
- [References] Buchholz and Kemper are cited as (1992), but the reference entry appears to correspond to the 2004 volume 'Validation of Stochastic Systems' (LNCS 2925). Verify and correct the year.
- [Section 5.2, Proposition 10] The symbols W_{P_UU}^k and W_{P_UD} are used in (51)--(53) without explicit definition. Define W_{P_UU}^k as the block analogue of (23) applied to P_UU, and W_{P_UD} = E_U^T W_P E_D.
- [Section 6.3] The sentence 'We write e_U = 1_{r,s} for the column vector...' conflicts with the earlier definition e_U = E_U 1_r in Section 5.1. Use one consistent notation.
- [Section 4.4, Proposition 8] The entropy statement assumes all transition probabilities are positive. It would be helpful to state explicitly that this is required because the smooth-extension hypothesis of Theorem 2 fails at zeros of the transition matrix, not merely because logarithms are undefined there.
- [Section 3, Remark 2] The bias-correction discussion invokes a matrix bias limit B_P without giving its form for the multinomial transition-matrix MLE. A brief formula or reference would make the remark self-contained.
Circularity Check
No significant circularity: Theorem 2 is a standard delta-method/Taylor expansion applied to the classical transition-matrix CLT, and the derived formulas are not equivalent to their inputs by construction.
full rationale
The paper's central result, Theorem 2, is a functional delta method and finite-order Taylor expansion for the plug-in MLE of a smooth functional of the transition matrix. Its proof invokes external standard results (van der Vaart's delta method, Taylor's formula in finite-dimensional normed spaces, and the classical CLT for the transition-matrix MLE stated as Theorem 1). The smooth-extension hypothesis is an explicit assumption, not a hidden restatement of the conclusions; the paper repeatedly discloses that boundary cases such as entropy with zero transition probabilities are excluded. The derived asymptotic formulas for powers, resolvents, stationary distributions, availability, MTTF, and related characteristics follow from ordinary differentiation rules (product rule, inverse rule) evaluated on the limiting Gaussian matrix; none of these outputs is used to define the input or fitted to the target quantity. The representation-independence argument in Theorem 2(iv) is genuine: tangent directions in T_P preserve stochasticity for small perturbations, so two extensions agreeing on stochastic matrices must have equal differentials on T_P. Self-citations (Trevezas & Limnios 2009; Votsi et al.) are contextual or auxiliary characterizations and are not load-bearing for the main theorem. The numerical experiments use Monte Carlo only to compare empirical distributions with the derived approximations; they do not fit parameters that are then called predictions. Thus the derivation chain is self-contained against external benchmarks and no circular step is present.
Axiom & Free-Parameter Ledger
axioms (5)
- domain assumption The true transition matrix P is irreducible (P∈P_ir)
- domain assumption For each functional φ, a C^r or analytic extension Φ exists on an open neighborhood U of P in R^{s×s}
- standard math Multinomial row-count CLT and strong law for finite Markov chains (Anderson-Goodman/Billingsley)
- standard math Fréchet Taylor formula and continuous mapping theorem in finite-dimensional normed spaces
- domain assumption For entropy rate, all transition probabilities are strictly positive; for reliability inverse quantities, ρ(P_UU)<1 and ρ(P_DD)<1
read the original abstract
In this work, we develop a unified matrix-level asymptotic calculus for plug-in non-parametric maximum likelihood estimators in finite Markov models. Starting from the asymptotic distribution of the estimated transition matrix, the limiting object is kept in its natural matrix form as a Gaussian random matrix, while the corresponding row-wise vector representation remains immediately available. The main point is that the stochastic constraints of the transition matrix need not be removed by a minimal parametrization: they are carried by the tangent directions and by the covariance structure of the limiting Gaussian matrix, whereas the relevant differentials are computed directly in matrix spaces. A single stochastic calculus theorem gives first-order limit distributions, finite-order developments for sufficiently differentiable functionals, and analytic expansions when the functional is analytic. This provides a common source for asymptotic formulas for matrix powers, stationary characteristics, finite-dimensional curves of Markov characteristics, additive-functional variances, entropy-type quantities and reliability indicators. The resulting covariance operators lead directly to confidence intervals, confidence regions, simultaneous finite-dimensional bands and Wald-type tests. Since the derivations are expressed through matrix products and Kronecker representations rather than coordinate-wise calculations, the method also gives substantial simplifications and, in many cases, computational gains. The second-order terms identify curvature corrections of smooth functionals and provide refined approximations whenever higher-order information is useful.
Reference graph
Works this paper leans on
-
[1]
W.andGoodman, L
Anderson, T. W.andGoodman, L. A.(1957). Statistical inference about Markov chains.The Annals of Mathematical Statistics28(1)89–110
1957
-
[2]
Statistical methods in Markov chains.The Annals of Mathemat- ical Statistics32(1)12–40
Billingsley, P.(1961). Statistical methods in Markov chains.The Annals of Mathemat- ical Statistics32(1)12–40
1961
-
[3]
Kronecker based matrix representations for large Markov models
Buchholz, P.andKemper, P.(1992). Kronecker based matrix representations for large Markov models. InValidation of Stochastic Systems: A Guide to Current Research (M. Broy and E. Denert, eds.) 256–295. Springer Berlin Heidelberg
1992
-
[4]
Maximum likelihood estimator for hidden Markov models in con- tinuous time.Statistical Inference for Stochastic Processes12(2)139–163
Chigansky, P.(2009). Maximum likelihood estimator for hidden Markov models in con- tinuous time.Statistical Inference for Stochastic Processes12(2)139–163
2009
-
[5]
Springer
Dayar, T.(2012).Analyzing Markov chains using Kronecker products. Springer
2012
-
[6]
Springer
Dayar, T.(2019).Kronecker Modeling and Analysis of Multidimensional Markovian Sys- tems. Springer
2019
-
[7]
J.(2013).Mathematics for Econometrics
Dhrymes, P. J.(2013).Mathematics for Econometrics. Springer New York, NY
2013
-
[8]
Software reliability modelling and prediction with hidden Markov chains.Statistical Modelling5(1)75–93
Durand, J.-B.andGaudoin, O.(2005). Software reliability modelling and prediction with hidden Markov chains.Statistical Modelling5(1)75–93. G´amiz, M. L.,Limnios, N.andSegovia-Garc ´ıa, M. d. C.(2023). Hidden Markov models in reliability and maintenance.European Journal of Operational Research 304(3)1242–1255
2005
-
[9]
K.andNagar, D
Gupta, A. K.andNagar, D. K.(2018).Matrix variate distributions. Chapman and Hall/CRC
2018
-
[10]
A.(2025).Hidden Markov Processes and Adaptive Filtering.Springer Series in Statistics
Kutoyants, Y. A.(2025).Hidden Markov Processes and Adaptive Filtering.Springer Series in Statistics. Springer
2025
-
[11]
A.andMotrunich, A.(2016)
Kutoyants, Y. A.andMotrunich, A.(2016). On multi-step MLE-process for Markov sequences.Metrika79(6)705–724
2016
-
[12]
V.andTsatsomeros, M
Le, H. V.andTsatsomeros, M. J.(2022). Matrix analysis for continuous-time Markov chains.Special Matrices10(1)219–233
2022
-
[13]
I.(1992).Adventures in stochastic processes
Resnick, S. I.(1992).Adventures in stochastic processes. Springer Science & Business Media
1992
-
[14]
Asymptotic properties for maximum likelihood esti- mators for reliability and failure rates of Markov chains.Communications in Statistics— Theory and Methods31(10)1837–1861
Sadek, A.andLimnios, N.(2002). Asymptotic properties for maximum likelihood esti- mators for reliability and failure rates of Markov chains.Communications in Statistics— Theory and Methods31(10)1837–1861
2002
-
[15]
RMAT: Random Matrix Analysis Toolkit R package version 0.2.0
Taqi, A.andWells, J.(2021). RMAT: Random Matrix Analysis Toolkit R package version 0.2.0
2021
-
[16]
Variance estimation in the central limit theorem for Markov chains.Journal of Statistical Planning and Inference139(7)2242–2253
Trevezas, S.andLimnios, N.(2009). Variance estimation in the central limit theorem for Markov chains.Journal of Statistical Planning and Inference139(7)2242–2253
2009
-
[17]
poly- nomial ergodic
Trevezas, S.andLimnios, N.(2011). Exact MLE and asymptotic properties for non- parametric semi-Markov models.Journal of Nonparametric Statistics23(3)719–739. van der V aart, A. W.(1994). Weak convergence of smoothed empirical processes. Scandinavian Journal of Statistics21(4)501–504. van der V aart, A. W.(1998).Asymptotic Statistics.Cambridge Series in St...
2011
-
[18]
Conditional failure occurrence rates for semi-Markov chains.Journal of Applied Statistics46(15)2722–2743
Votsi, I.(2019). Conditional failure occurrence rates for semi-Markov chains.Journal of Applied Statistics46(15)2722–2743
2019
-
[19]
Bootstrap of reliability indicators for semi-Markov processes.Methodology and Computing in Applied Probability27(1)3
Votsi, I.andBouzebda, S.(2025). Bootstrap of reliability indicators for semi-Markov processes.Methodology and Computing in Applied Probability27(1)3
2025
-
[20]
Confidence interval for the mean time to failure in semi-Markov models: an application to wind energy production.Journal of Applied Statistics46(10)1756–1773
Votsi, I.andBrouste, A.(2019). Confidence interval for the mean time to failure in semi-Markov models: an application to wind energy production.Journal of Applied Statistics46(10)1756–1773
2019
-
[21]
Estimation of the intensity of the hitting time for semi- Markov chains and hidden Markov renewal chains.Journal of Nonparametric Statistics 27(2)149–166
Votsi, I.andLimnios, N.(2015). Estimation of the intensity of the hitting time for semi- Markov chains and hidden Markov renewal chains.Journal of Nonparametric Statistics 27(2)149–166
2015
-
[22]
Hidden semi- Markov modeling for the estimation of earthquake occurrence rates.Communications in Statistics—Theory and Methods43(7)1484–1502
Votsi, I.,Limnios, N.,Tsaklidis, G.andPapadimitriou, E.(2014). Hidden semi- Markov modeling for the estimation of earthquake occurrence rates.Communications in Statistics—Theory and Methods43(7)1484–1502
2014
-
[23]
W.(1966)
Zehna, P. W.(1966). Invariance of maximum likelihood estimators.The Annals of Math- ematical Statistics37(3)744–744
1966
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.