REVIEW 3 major objections 4 minor 2 references
Evaluating financial tail risk forecasts: Testing Equal Predictive Ability
T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read At the extreme quantile levels used in financial regulation, the Diebold–Mariano test and Model Confidence Set have little power against VaR and expected shortfall models that understate tail risk, and with short evaluation windows their…
desk verdict Useful simulation study, but the headline claim about unreliable critical values needs a null simulation to fully land. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the studentized loss differential built from a strictly consistent asymmetric loss function. For standalone VaR forecasts the paper uses the generalized piecewise linear (GPL) class, whose member with $b = 1$ is the tick loss $L(x,y) = (1\{y \le x\} - p)(x - y)$; for joint VaR and ES it uses the Fissler–Ziegel class with three parametrizations (AL, NZ, FZG). Both the DM statistic $t_{12} = \bar{d}_{12}/\sqrt{\widehat{\mathrm{var}}(\bar{d}_{12})}$ and the MCS's $T_{\max}$ statistic compare a sample mean of loss differences against an estimated variance. The paper's worked example shows that when the true model is more conservative than its competitor, the loss difference is positive — the wrong sign — for every observation above the competitor's forecast, which at $p = 0.01$ is the overwhelming majority, and negative only for the rare tail observations below it; the analytical expression for the expected loss difference shows that the share of correctly signed observations is roughly $F(q_2 - p(q_2 - q_1))$. Because the rare correct-sign observations also dominate the variance estimate, short samples frequently exhibit both the wrong mean and an underestimated variance, which the paper argues explains the heavily skewed statistics, low power, and type III errors.
What would settle it
Replicate the static comparison of a true Student-t VaR model against a misspecified normal model at $p = 0.01$ with $P = 251$ out-of-sample days: the paper reports DM power of roughly 1%. If a re-estimation-based version of the same comparison — with GARCH parameters refit on a rolling window before each forecast — pushes power substantially above that level, the weakness is an artifact of the fixed-parameter design rather than of the loss functions. Separately, one can directly verify the mechanism by checking whether the observed fraction of correctly signed loss differences matches the paper's formula $F(q_2 - p(q_2 - q_1))$.
Extended reading notes
Core claim
The paper's central finding is that, for out-of-sample VaR and ES forecasts evaluated with strictly consistent asymmetric losses, the DM test and the MCS procedure show little power against models that understate tail risk at the most extreme quantile levels $p \in \{0.01, 0.025\}$, while power improves with the quantile level and the out-of-sample size. For $p \in \{0.01, 0.025\}$ and evaluation windows of up to two years of daily data, the test statistics are heavily skewed and type III errors — rejections that run opposite to the true ranking — are non-negligible, in several simulated settings comparable to or larger than the power itself. The paper demonstrates both empirically and theoretically how this follows from the combination of asymmetric loss and time-varying volatility: when the true, more conservative model is compared with a risk-understating competitor, the loss difference has the correct sign only for the rare observations below both forecasts — a fraction roughly of size $p$ — and samples that miss these observations also underestimate the variance of the loss differential, producing skewed $t$-statistics and wrong-direction rejections. It further shows that joint evaluation of VaR and ES via Fissler–Ziegel-type scores dominates standalone VaR evaluation once at least two years of data are available, and that tests against over-conservative forecasts are far more powerful than tests against risk-understating ones.
Load-bearing premise
The simulation design fixes every competing model's parameters at their true population values and never re-estimates them, and the dynamic scenarios use TGARCH parameter values drawn from one sample of Dow Jones index stocks, so the reported power and type III error rates could differ in real applications where parameters are estimated and unstable.
Editorial extensions
If this is right
- Researchers comparing 1% or 2.5% VaR forecasts with one to two years of daily data should distrust standard normal or bootstrapped critical values, since the DM and MCS tests can reject in favor of the worse model.
- Reliable ranking of extreme-quantile forecasts requires evaluation windows of several years, which collides with the assumption that the forecast models remain stable over such long periods.
- Joint evaluation of VaR and ES with Fissler–Ziegel-type scores is more powerful than standalone VaR evaluation once at least two years of data are available.
- At larger quantile levels ($p \ge 0.05$) the tests behave closer to their nominal properties; for the MCS, potency exceeds its confidence level after about a year even at $p = 0.01$, so trimming the worst models at $\alpha = 0.25$ remains feasible.
- Failure to reject should not be read as evidence that a risk model is not dangerously optimistic: under-reporting models are precisely the ones the tests find hardest to detect.
Reading between the lines
- The paper's own mechanism suggests a concrete fix it leaves untested: a finite-sample correction to the variance estimator, or a bootstrap that preserves the frequency of tail observations, could restore much of the DM test's power at $p = 0.01$.
- Read from a regulator's perspective, the type III error pattern implies that a backtest approving a model on one year of 1% VaR results can actively reward an under-capitalised institution; approval frameworks may need longer windows or joint scores by design.
- The rare-signal logic should apply a fortiori to even more extreme quantiles such as $p = 0.005$, suggesting the qualitative conclusion generalises well beyond the specific losses and TGARCH design used here.
- A direct extension would run the same simulation design under conditional predictive-ability testing, where forecasts are produced with re-estimated parameters, to see whether conditioning on the recent volatility state rescues the power that unconditional tests lose.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper studies the finite-sample properties of the Diebold-Mariano (DM) test and the Model Confidence Set (MCS) procedure when applied to out-of-sample VaR and ES forecasts evaluated with strictly consistent asymmetric loss functions. The author conducts extensive Monte Carlo simulations under static (iid Student-t) and dynamic (TGARCH with Student-t innovations) DGPs, with quantile levels p in {0.01, 0.025, 0.05, 0.1} and out-of-sample sizes from one to ten years. The central findings are that at extreme quantiles and short evaluation windows the tests have very low power against models that underestimate tail risk, the test statistics are heavily skewed, and type III errors (rejections in the wrong direction) are non-negligible. The paper also provides theoretical calculations for the loss differential under the tick loss. The main practical message is that researchers should be cautious about using standard normal or bootstrapped critical values for extreme-quantile forecast comparisons and need long evaluation samples.
Significance. If the findings are accepted, the paper serves as a useful caution for applied work in financial risk forecasting, a field where DM and MCS tests are routinely used to compare VaR/ES models. The study is unusually comprehensive, covering multiple loss functions (GPL with b=0.5,1,2; AL, NZ, FZG scores), several DGPs with parameter values estimated from Dow Jones stock data, and both pairwise and multi-model comparisons. The qualitative result that power is very low at p=0.01 and p=0.025, with type III errors around 10% for one-year samples, is consistent across tables and loss functions, which strengthens the credibility of the main descriptive finding. The theoretical derivations in Appendix A.3 are a useful contribution for understanding the mechanism. However, the paper's scope is limited by its fixed-parameter design, which the author explicitly acknowledges in Section 2.1.3, and by the absence of direct null simulations for the DM test.
major comments (3)
- [Abstract, Section 3.2.1] The abstract and conclusion claim that researchers should be cautious about using standard normal or bootstrapped critical values for DM/MCS tests at extreme quantiles, but all reported simulations are conducted under alternatives (fixed parameter values, as stated in Section 2.1.3). Skewness of t12 under H1 and type III errors do not, by themselves, establish that nominal critical values are invalid under H0; a test statistic can be heavily skewed under an alternative while size remains controlled. Please add null simulations (e.g., comparing two models with identical expected loss, or the true model to itself) and report empirical size for normal and bootstrap critical values over the same (p, P) grid. Without such simulations, the critical-value warning in the abstract is an extrapolation.
- [Section 3.2.1, Table 3.7] In the dynamic volatility-misspecification design, the tick loss power at p=0.025 is essentially constant near zero (0.004, 0.006, 0.006, 0.008) while type III errors are about 10%, and the paper does not report the expected loss differential for this design. Because the GARCH-t misspecification may yield an expected tick loss very close to the true model, the near-zero power may reflect a tiny effect size rather than test deficiency. Please report the expected loss ratios (or standardized effect sizes) for each pairwise comparison and discuss the non-monotonic power pattern in Table 3.7.
- [Section 3, all tables] No Monte Carlo standard errors are reported for any of the simulation results. With 10,000 replications in the dynamic pairwise designs, the standard error for a rejection frequency of 0.05 is about 0.002, so differences such as 0.004 vs. 0.008 in Table 3.7 are not statistically distinguishable. Please provide MCSEs or confidence intervals for the main tables (at least Tables 3.1, 3.2, 3.5-3.8 and the MCS tables), and qualify statements that rely on small differences.
minor comments (4)
- [Section 2.1.1, Eq. (2.9)] The convergence arrow in equation (2.9) is typeset incorrectly: 'd→N(0,1)' should be '→d N(0,1)'. This is a minor formatting issue but obscures the meaning.
- [Section 3.1] The sentence 'We omit the VaR results for p=0.025 as the quantile of the normal and Student's t distribution are very close such that the calculation of tij is not reliable' is vague; please state the numerical reason (e.g., near-cancellation in the loss differential) or provide the table.
- [Appendix B.5, Table B.22] The note to Table B.22 contains an unresolved placeholder '(crossref?)' that should be replaced with the proper cross-reference.
- [Throughout] There are several typos, including 'Apendix' in Section 3.2.1, 'the the' in Section 2.1.1, and 'conditional volatiltity' in Appendix A.3.1. A careful proofread is recommended.
Circularity Check
No significant circularity: the simulation design fixes DGP parameters externally, and no reported result reduces by construction to its own inputs.
full rationale
This paper is a Monte Carlo study of the Diebold-Mariano test and the Model Confidence Set for VaR/ES losses. The DGP parameters in Table 3.4 are estimated from Dow Jones index data (Appendix A.1), not fitted to the paper's conclusions. The loss functions are taken from Gneiting (2011), Fissler and Ziegel (2016), Taylor (2020), and Nolde and Ziegel (2017), and the testing procedures from Diebold and Mariano (1995) and Hansen et al. (2011); there are no self-citations and no uniqueness or ansatz results imported from the author's own prior work. The theoretical expressions in Appendix A.3 are derived algebraically from the tick loss and are used only to explain the simulated patterns, not to define the test outcomes. The paper explicitly states in Section 2.1.3 that it uses fixed parameters rather than estimating them, which limits generality but is not circular. The reviewer's concern that all DM simulations are under alternatives, so skewness and type III errors do not by themselves establish that normal or bootstrapped critical values fail under the null, is a legitimate evidentiary limitation on the strength of the practical warning, but it is not a circularity: the critical-value claim is not forced by construction from the simulated alternative distributions. No fitted parameter is renamed as a prediction, and no step in the derivation chain is equivalent to its own input by definition. The appropriate finding is therefore no significant circularity.
Assumptions & free parameters
free parameters (1)
- TGARCH DGP parameters (omega, alpha, gamma, beta, nu) =
omega=0.03, alpha=0.04, gamma=0.1, beta=0.9, nu in {3,4,7,12}
assumptions (4)
- domain assumption Strictly consistent loss functions correctly rank VaR and joint VaR/ES forecasts (Gneiting 2011b; Fissler and Ziegel 2016).
- domain assumption The TGARCH(1,1) model with scaled Student-t innovations is a realistic DGP for daily financial returns.
- standard math The DM statistic is asymptotically standard normal under mixing conditions with a consistent variance estimator, and the MCS bootstrap controls the familywise error rate.
- domain assumption Forecast model parameters are fixed at population values, excluding parameter estimation error.
Cite this review
Pith. "Pith review of Evaluating financial tail risk forecasts: Testing Equal Predictive Ability." pith.science (2026). https://pith.science/paper/ZBIUPQVT
@misc{pith2026250523333,
author = {Pith},
title = {Pith review of: Evaluating financial tail risk forecasts: Testing Equal Predictive Ability},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZBIUPQVT}},
note = {Machine review of arXiv:2505.23333}
}
read the original abstract
This paper provides comprehensive simulation results on the finite sample properties of the Diebold-Mariano (DM) test by Diebold and Mariano (1995) and the model confidence set (MCS) testing procedure by Hansen et al. (2011) applied to the asymmetric loss functions specific to financial tail risk forecasts, such as Value-at-Risk (VaR) and Expected Shortfall (ES). We focus on statistical loss functions that are strictly consistent in the sense of Gneiting (2011a). We find that the tests show little power against models that underestimate the tail risk at the most extreme quantile levels, while the finite sample properties generally improve with the quantile level and the out-of-sample size. For the small quantile levels and out-of-sample sizes of up to two years, we observe heavily skewed test statistics and non-negligible type III errors, which implies that researchers should be cautious about using standard normal or bootstrapped critical values. We demonstrate both empirically and theoretically how these unfavorable finite sample results relate to the asymmetric loss functions and the time varying volatility inherent in financial return data.
Figures
Reference graph
Works this paper leans on
-
[1]
Aparicio, D., & López de Prado, M. (2018). How hard is it to pick the right model? MCS and backtest overfitting.Algorithmic Finance, 7(1-2), 53–61. BankforInternationalSettlements.(2013). Fundamental review of the trading book: A revised market risk framework ; consultative document ; issued for comment by 31 january 2014 (Oct. 2013). Bank for Internation...
work page 2018
-
[2]
In this example, we illustrate how the tick loss function evaluates VaR forecast of two models if one model consistently makes forecasts that are closer to zero than the forecasts of the other model. To illustrate how correlations affect the MCS test, assume that we compare only two models i and j and that the variances are fixed. It holds that var(Dt,ij)...
work page 2011
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.