Pith. sign in

REVIEW 3 major objections 4 minor 2 references

Evaluating financial tail risk forecasts: Testing Equal Predictive Ability

T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read At the extreme quantile levels used in financial regulation, the Diebold–Mariano test and Model Confidence Set have little power against VaR and expected shortfall models that understate tail risk, and with short evaluation windows their…

desk verdict Useful simulation study, but the headline claim about unreliable critical values needs a null simulation to fully land. read the letter →

arxiv 2505.23333 v1 pith:ZBIUPQVT submitted 2025-05-29 econ.EM stat.AP

classification econ.EMstat.AP MSC 62F0362F4062M10
keywords Diebold–MarianotestmodelconfidencesetValue-at-RiskExpectedShortfallstrictlyconsistentlosstailriskforecastingfinite-samplepowertypeIIIerror
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether the standard toolkit for comparing forecast accuracy — the Diebold–Mariano (DM) test and the Model Confidence Set (MCS) — can tell a good tail-risk forecaster from a bad one. Using Monte Carlo simulations of daily Value-at-Risk and expected shortfall forecasts, evaluated with strictly consistent asymmetric loss functions of the kind used in backtesting, the paper finds that at the extreme quantile levels $p \in \{0.01, 0.025\}$ the tests have essentially no power against models that underestimate risk, and that with one to two years of daily data the tests frequently reject in the wrong direction, declaring the inferior model superior. The paper locates the cause in the loss functions themselves: at these quantile levels only a tiny fraction of observations carry the correct ranking, and short samples both draw the wrong sign of the average loss difference and underestimate its variance. A sympathetic reader should care because these tests are routinely used to choose between risk models in research and regulation, and the paper's message is that several years of evaluation data, not the customary one or two, are needed for reliable conclusions about extreme-quantile forecasts.

What carries the argument

The load-bearing object is the studentized loss differential built from a strictly consistent asymmetric loss function. For standalone VaR forecasts the paper uses the generalized piecewise linear (GPL) class, whose member with $b = 1$ is the tick loss $L(x,y) = (1\{y \le x\} - p)(x - y)$; for joint VaR and ES it uses the Fissler–Ziegel class with three parametrizations (AL, NZ, FZG). Both the DM statistic $t_{12} = \bar{d}_{12}/\sqrt{\widehat{\mathrm{var}}(\bar{d}_{12})}$ and the MCS's $T_{\max}$ statistic compare a sample mean of loss differences against an estimated variance. The paper's worked example shows that when the true model is more conservative than its competitor, the loss difference is positive — the wrong sign — for every observation above the competitor's forecast, which at $p = 0.01$ is the overwhelming majority, and negative only for the rare tail observations below it; the analytical expression for the expected loss difference shows that the share of correctly signed observations is roughly $F(q_2 - p(q_2 - q_1))$. Because the rare correct-sign observations also dominate the variance estimate, short samples frequently exhibit both the wrong mean and an underestimated variance, which the paper argues explains the heavily skewed statistics, low power, and type III errors.

What would settle it

Replicate the static comparison of a true Student-t VaR model against a misspecified normal model at $p = 0.01$ with $P = 251$ out-of-sample days: the paper reports DM power of roughly 1%. If a re-estimation-based version of the same comparison — with GARCH parameters refit on a rolling window before each forecast — pushes power substantially above that level, the weakness is an artifact of the fixed-parameter design rather than of the loss functions. Separately, one can directly verify the mechanism by checking whether the observed fraction of correctly signed loss differences matches the paper's formula $F(q_2 - p(q_2 - q_1))$.

Watch

Extended reading notes

Core claim

The paper's central finding is that, for out-of-sample VaR and ES forecasts evaluated with strictly consistent asymmetric losses, the DM test and the MCS procedure show little power against models that understate tail risk at the most extreme quantile levels $p \in \{0.01, 0.025\}$, while power improves with the quantile level and the out-of-sample size. For $p \in \{0.01, 0.025\}$ and evaluation windows of up to two years of daily data, the test statistics are heavily skewed and type III errors — rejections that run opposite to the true ranking — are non-negligible, in several simulated settings comparable to or larger than the power itself. The paper demonstrates both empirically and theoretically how this follows from the combination of asymmetric loss and time-varying volatility: when the true, more conservative model is compared with a risk-understating competitor, the loss difference has the correct sign only for the rare observations below both forecasts — a fraction roughly of size $p$ — and samples that miss these observations also underestimate the variance of the loss differential, producing skewed $t$-statistics and wrong-direction rejections. It further shows that joint evaluation of VaR and ES via Fissler–Ziegel-type scores dominates standalone VaR evaluation once at least two years of data are available, and that tests against over-conservative forecasts are far more powerful than tests against risk-understating ones.

Load-bearing premise

The simulation design fixes every competing model's parameters at their true population values and never re-estimates them, and the dynamic scenarios use TGARCH parameter values drawn from one sample of Dow Jones index stocks, so the reported power and type III error rates could differ in real applications where parameters are estimated and unstable.

Editorial extensions

If this is right

  • Researchers comparing 1% or 2.5% VaR forecasts with one to two years of daily data should distrust standard normal or bootstrapped critical values, since the DM and MCS tests can reject in favor of the worse model.
  • Reliable ranking of extreme-quantile forecasts requires evaluation windows of several years, which collides with the assumption that the forecast models remain stable over such long periods.
  • Joint evaluation of VaR and ES with Fissler–Ziegel-type scores is more powerful than standalone VaR evaluation once at least two years of data are available.
  • At larger quantile levels ($p \ge 0.05$) the tests behave closer to their nominal properties; for the MCS, potency exceeds its confidence level after about a year even at $p = 0.01$, so trimming the worst models at $\alpha = 0.25$ remains feasible.
  • Failure to reject should not be read as evidence that a risk model is not dangerously optimistic: under-reporting models are precisely the ones the tests find hardest to detect.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's own mechanism suggests a concrete fix it leaves untested: a finite-sample correction to the variance estimator, or a bootstrap that preserves the frequency of tail observations, could restore much of the DM test's power at $p = 0.01$.
  • Read from a regulator's perspective, the type III error pattern implies that a backtest approving a model on one year of 1% VaR results can actively reward an under-capitalised institution; approval frameworks may need longer windows or joint scores by design.
  • The rare-signal logic should apply a fortiori to even more extreme quantiles such as $p = 0.005$, suggesting the qualitative conclusion generalises well beyond the specific losses and TGARCH design used here.
  • A direct extension would run the same simulation design under conditional predictive-ability testing, where forecasts are produced with re-estimated parameters, to see whether conditioning on the recent volatility state rescues the power that unconditional tests lose.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper studies the finite-sample properties of the Diebold-Mariano (DM) test and the Model Confidence Set (MCS) procedure when applied to out-of-sample VaR and ES forecasts evaluated with strictly consistent asymmetric loss functions. The author conducts extensive Monte Carlo simulations under static (iid Student-t) and dynamic (TGARCH with Student-t innovations) DGPs, with quantile levels p in {0.01, 0.025, 0.05, 0.1} and out-of-sample sizes from one to ten years. The central findings are that at extreme quantiles and short evaluation windows the tests have very low power against models that underestimate tail risk, the test statistics are heavily skewed, and type III errors (rejections in the wrong direction) are non-negligible. The paper also provides theoretical calculations for the loss differential under the tick loss. The main practical message is that researchers should be cautious about using standard normal or bootstrapped critical values for extreme-quantile forecast comparisons and need long evaluation samples.

Significance. If the findings are accepted, the paper serves as a useful caution for applied work in financial risk forecasting, a field where DM and MCS tests are routinely used to compare VaR/ES models. The study is unusually comprehensive, covering multiple loss functions (GPL with b=0.5,1,2; AL, NZ, FZG scores), several DGPs with parameter values estimated from Dow Jones stock data, and both pairwise and multi-model comparisons. The qualitative result that power is very low at p=0.01 and p=0.025, with type III errors around 10% for one-year samples, is consistent across tables and loss functions, which strengthens the credibility of the main descriptive finding. The theoretical derivations in Appendix A.3 are a useful contribution for understanding the mechanism. However, the paper's scope is limited by its fixed-parameter design, which the author explicitly acknowledges in Section 2.1.3, and by the absence of direct null simulations for the DM test.

major comments (3)
  1. [Abstract, Section 3.2.1] The abstract and conclusion claim that researchers should be cautious about using standard normal or bootstrapped critical values for DM/MCS tests at extreme quantiles, but all reported simulations are conducted under alternatives (fixed parameter values, as stated in Section 2.1.3). Skewness of t12 under H1 and type III errors do not, by themselves, establish that nominal critical values are invalid under H0; a test statistic can be heavily skewed under an alternative while size remains controlled. Please add null simulations (e.g., comparing two models with identical expected loss, or the true model to itself) and report empirical size for normal and bootstrap critical values over the same (p, P) grid. Without such simulations, the critical-value warning in the abstract is an extrapolation.
  2. [Section 3.2.1, Table 3.7] In the dynamic volatility-misspecification design, the tick loss power at p=0.025 is essentially constant near zero (0.004, 0.006, 0.006, 0.008) while type III errors are about 10%, and the paper does not report the expected loss differential for this design. Because the GARCH-t misspecification may yield an expected tick loss very close to the true model, the near-zero power may reflect a tiny effect size rather than test deficiency. Please report the expected loss ratios (or standardized effect sizes) for each pairwise comparison and discuss the non-monotonic power pattern in Table 3.7.
  3. [Section 3, all tables] No Monte Carlo standard errors are reported for any of the simulation results. With 10,000 replications in the dynamic pairwise designs, the standard error for a rejection frequency of 0.05 is about 0.002, so differences such as 0.004 vs. 0.008 in Table 3.7 are not statistically distinguishable. Please provide MCSEs or confidence intervals for the main tables (at least Tables 3.1, 3.2, 3.5-3.8 and the MCS tables), and qualify statements that rely on small differences.
minor comments (4)
  1. [Section 2.1.1, Eq. (2.9)] The convergence arrow in equation (2.9) is typeset incorrectly: 'd→N(0,1)' should be '→d N(0,1)'. This is a minor formatting issue but obscures the meaning.
  2. [Section 3.1] The sentence 'We omit the VaR results for p=0.025 as the quantile of the normal and Student's t distribution are very close such that the calculation of tij is not reliable' is vague; please state the numerical reason (e.g., near-cancellation in the loss differential) or provide the table.
  3. [Appendix B.5, Table B.22] The note to Table B.22 contains an unresolved placeholder '(crossref?)' that should be replaced with the proper cross-reference.
  4. [Throughout] There are several typos, including 'Apendix' in Section 3.2.1, 'the the' in Section 2.1.1, and 'conditional volatiltity' in Appendix A.3.1. A careful proofread is recommended.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the simulation design fixes DGP parameters externally, and no reported result reduces by construction to its own inputs.

full rationale

This paper is a Monte Carlo study of the Diebold-Mariano test and the Model Confidence Set for VaR/ES losses. The DGP parameters in Table 3.4 are estimated from Dow Jones index data (Appendix A.1), not fitted to the paper's conclusions. The loss functions are taken from Gneiting (2011), Fissler and Ziegel (2016), Taylor (2020), and Nolde and Ziegel (2017), and the testing procedures from Diebold and Mariano (1995) and Hansen et al. (2011); there are no self-citations and no uniqueness or ansatz results imported from the author's own prior work. The theoretical expressions in Appendix A.3 are derived algebraically from the tick loss and are used only to explain the simulated patterns, not to define the test outcomes. The paper explicitly states in Section 2.1.3 that it uses fixed parameters rather than estimating them, which limits generality but is not circular. The reviewer's concern that all DM simulations are under alternatives, so skewness and type III errors do not by themselves establish that normal or bootstrapped critical values fail under the null, is a legitimate evidentiary limitation on the strength of the practical warning, but it is not a circularity: the critical-value claim is not forced by construction from the simulated alternative distributions. No fitted parameter is renamed as a prediction, and no step in the derivation chain is equivalent to its own input by definition. The appropriate finding is therefore no significant circularity.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The paper introduces no new entities or fitted-to-conclusion parameters. The DGP parameters are estimated on external data, and the loss functions are taken from prior literature. The main load-bearing assumptions are that the loss functions are strictly consistent, that the TGARCH-t DGP is representative, and that fixing forecast parameters does not change the qualitative finite-sample conclusions.

free parameters (1)
  • TGARCH DGP parameters (omega, alpha, gamma, beta, nu) = omega=0.03, alpha=0.04, gamma=0.1, beta=0.9, nu in {3,4,7,12}
    Estimated on Dow Jones index returns in Appendix A.1 and treated as fixed simulation inputs. The finite-sample results depend on these values, but they are not fitted to the paper's conclusions.
assumptions (4)
  • domain assumption Strictly consistent loss functions correctly rank VaR and joint VaR/ES forecasts (Gneiting 2011b; Fissler and Ziegel 2016).
    The whole evaluation framework relies on consistency of the GPL and FZ-family loss functions, cited in Section 2.
  • domain assumption The TGARCH(1,1) model with scaled Student-t innovations is a realistic DGP for daily financial returns.
    Used throughout Section 3.2 with parameters estimated on Dow Jones indices; the simulation conclusions inherit this assumption.
  • standard math The DM statistic is asymptotically standard normal under mixing conditions with a consistent variance estimator, and the MCS bootstrap controls the familywise error rate.
    Invoked in Section 2.1.1 and Section 2.1.2; the paper is testing finite-sample deviations from these asymptotic properties.
  • domain assumption Forecast model parameters are fixed at population values, excluding parameter estimation error.
    Stated explicitly in Section 2.1.3; this isolates the effect of loss functions but limits direct applicability to real forecasting settings where parameters are estimated.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Evaluating financial tail risk forecasts: Testing Equal Predictive Ability." pith.science (2026). https://pith.science/paper/ZBIUPQVT

@misc{pith2026250523333,
  author       = {Pith},
  title        = {Pith review of: Evaluating financial tail risk forecasts: Testing Equal Predictive Ability},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZBIUPQVT}},
  note         = {Machine review of arXiv:2505.23333}
}
read the original abstract

This paper provides comprehensive simulation results on the finite sample properties of the Diebold-Mariano (DM) test by Diebold and Mariano (1995) and the model confidence set (MCS) testing procedure by Hansen et al. (2011) applied to the asymmetric loss functions specific to financial tail risk forecasts, such as Value-at-Risk (VaR) and Expected Shortfall (ES). We focus on statistical loss functions that are strictly consistent in the sense of Gneiting (2011a). We find that the tests show little power against models that underestimate the tail risk at the most extreme quantile levels, while the finite sample properties generally improve with the quantile level and the out-of-sample size. For the small quantile levels and out-of-sample sizes of up to two years, we observe heavily skewed test statistics and non-negligible type III errors, which implies that researchers should be cautious about using standard normal or bootstrapped critical values. We demonstrate both empirically and theoretically how these unfavorable finite sample results relate to the asymmetric loss functions and the time varying volatility inherent in financial return data.

Figures

Figures reproduced from arXiv: 2505.23333 by the authors.

Figure 3.1
Figure 3.1. Density histogram of t12 for the tick loss This figure displays density histograms of tij for one-step-ahead forecasts evaluated using the tick loss, where model i is the true Student’s t model while model j is a misspecified model that assumes a normal distribution. Furthermore, [PITH_FULL_IMAGE:figures/full_fig_p012_3_1.png] view at source ↗
Figure 3.2
Figure 3.2. Finite sample properties for the tick loss [PITH_FULL_IMAGE:figures/full_fig_p018_3_2.png] view at source ↗
Figure 3.3
Figure 3.3. Finite sample properties for the FZG score [PITH_FULL_IMAGE:figures/full_fig_p020_3_3.png] view at source ↗
Figures from the paper (1 more)
Figure 3.4
Figure 3.4. Figure 3.4: Finite sample properties for the tick loss [PITH_FULL_IMAGE:figures/full_fig_p023_3_4.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

2 extracted references · 2 canonical work pages

  1. [1]

    Aparicio, D., & López de Prado, M. (2018). How hard is it to pick the right model? MCS and backtest overfitting.Algorithmic Finance, 7(1-2), 53–61. BankforInternationalSettlements.(2013). Fundamental review of the trading book: A revised market risk framework ; consultative document ; issued for comment by 31 january 2014 (Oct. 2013). Bank for Internation...

  2. [2]

    To illustrate how correlations affect the MCS test, assume that we compare only two models i and j and that the variances are fixed

    In this example, we illustrate how the tick loss function evaluates VaR forecast of two models if one model consistently makes forecasts that are closer to zero than the forecasts of the other model. To illustrate how correlations affect the MCS test, assume that we compare only two models i and j and that the variances are fixed. It holds that var(Dt,ij)...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.