REVIEW 4 major objections 7 minor 3 references
Conditional Method Confidence Set
T0 review · 4 major / 7 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Running the Model Confidence Set state by state yields per-regime forecast-model sets with an asymptotic coverage guarantee.
desk verdict A sensible statewise extension of Hansen-Lunde-Nason with an informative simulation study, but the central CLT and bootstrap arguments for non-contiguous statewise subsamples are asserted rather than proved, so the asymptotic claims need careful refereeing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing device is the statewise application of the MCS closure principle. The paper splits the full out-of-sample loss series into state-specific subsequences $L^l_{i,\tau}$ by conditioning on a categorical state variable $S_t$ at the forecast origin, computes per-state loss differentials $d^l_{ij,\tau}$ and standardized excess losses $t^l_{i\cdot}$, and applies the elimination rule $e^l_{\max,M}=\arg\max_{i\in M} t^l_{i\cdot}$ whenever $T^l_{\max,M}$ exceeds a bootstrap critical value. A coherency condition ties the test to the elimination rule so that no conditionally superior method is removed with probability exceeding $\alpha$. The circular block bootstrap supplies the critical values, implicitly accounting for the correlation among loss differentials and for the temporal dependence in the state subsamples.
What would settle it
Take a state with at least two competing methods and simulate a DGP in which the conditional expected loss differential changes sign within that state, for example with $E[d]=+\delta$ in the first half of the state subsample and $-\delta$ in the second half. If the CMCS is applied and its coverage is computed over many repeated samples, coverage below $1-\alpha$ for the true statewise superior set would directly contradict Theorem A.1, and the failure should become visible when the sign flip occurs near the middle of the subsample with $n_l$ large.
Extended reading notes
Core claim
The central claim is that forecast evaluation can be made conditional on a discrete state variable without losing the set-based guarantees of the Model Confidence Set. For each state $l$, the procedure tests $H^l_{0,M}$: all methods in $M$ have equal conditional expected loss, using statewise $t$-statistics $t^l_{i\cdot}$ and the maximum statistic $T^l_{\max,M}=\max_i t^l_{i\cdot}$. Theorem 2.1 states that under the null $T^l_{\max,M}$ converges in distribution to the maximum of a normal vector with the statewise correlation matrix, while under the alternative it diverges and the elimination rule removes an inferior method. Theorem A.1 then delivers the headline guarantee: the conditional confidence set $\widehat{M}^{l,*}_{1-\alpha}$ contains the set of conditionally superior methods with asymptotic probability at least $1-\alpha$, and eliminates each inferior method with probability tending to one. Thus the CMCS controls the familywise error rate state by state.
Load-bearing premise
The procedure assumes that, within each economic regime, the ranking of the forecasting methods never changes over time; if the best method switches within a regime, the set the test is trying to cover is no longer well defined and the coverage guarantee does not apply.
Editorial extensions
If this is right
- For each regime, the analyst obtains a set of models whose conditional predictive ability is indistinguishable at level $\alpha$, rather than a single forced winner.
- A model that looks best unconditionally but is weak in one state is still eliminated in that state, so forecast combinations adapt to the regime.
- Because the procedure is run separately on disjoint states, false eliminations are controlled within each state and do not contaminate other regimes.
- In the Expected Shortfall application, the CMCS selects different models across stocks, liquidity horizons, and forecasting horizons, yielding lower and more stable average ES forecasts than the Wald-type dynamic forecast combination benchmark.
Reading between the lines
- The proof relies on discrete states with positive probability; conditioning on continuous or high-dimensional states would require local smoothing and is not covered by the present asymptotics.
- If within-state rankings can switch, the set of conditionally superior methods may cease to be well defined; an alternative target would be the best method inside rolling sub-periods, replacing Assumption 2.3(i) with piecewise stability.
- The same statewise closure idea could transfer to other set-valued selection tasks, such as variable selection or nowcast averaging, whenever the loss can be indexed by a regime.
- One can test the weakest assumption directly: estimate statewise loss differentials on rolling windows and look for sign changes within a state; frequent sign flips would signal that the asymptotic coverage guarantee should not be relied on.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a Conditional Method Confidence Set (CMCS) that applies the MCS algorithm of Hansen et al. (2011) separately to statewise loss subsequences selected by a discrete state variable S_t. It defines conditionally superior sets M^{l,*}, statewise hypotheses H^l_{0,M}, and a T^l_max test statistic. Theorem 2.1 claims that T^l_max converges to the distribution of the maximum of a normal vector under the null and diverges under the alternative, and Theorem A.1 claims asymptotic coverage at least 1−α for the statewise MCS. The paper also provides simulation evidence comparing CMCS with the unconditional MCS and with the Wald-based dynamic forecast combination of Borup et al. (2024), and an empirical application to Expected Shortfall stress testing across liquidity horizons.
Significance. If the central claims were fully established, the paper would make a useful contribution: a statewise MCS is a natural and practically relevant extension, especially for regulatory stress testing where model rankings may differ across regimes. The simulation design in Section 3.2 is informative about finite-sample power differences between statewise t-tests and the Wald test. The paper is not circular: no fitted parameters are used to make predictions, and the procedure is derived from stated assumptions. However, the proof of the main asymptotic result is incomplete, the bootstrap validity is not established, and the empirical evidence is descriptive and lacks backtesting. The contribution is therefore conditional on closing these gaps.
major comments (4)
- [§2.4 and Appendix A.1, Lemma A.1] The statewise CLT is asserted, not proved. Corollary A.2 shows that z_t = h_t d_{ij,t} is α-mixing of the stated size when Assumption 2.1 holds, but the statewise sequence d^l_{ij,n_lτ} is obtained by selecting times {t: S_t = s_l}. Since S_t is F_t-measurable and in the application is a stress indicator (often a 252-day window selected by a risk factor, Section 4.1 and Figure 3), the selected times are neither independent of the losses nor regularly spaced; the statewise sample is a concatenation of possibly separated stress episodes. Mixing of a subsequence selected by a dependent indicator does not follow from mixing of the original sequence. Lemma A.1 simply cites "the CLT for α-mixing processes" after defining X^l_τ from the statewise losses. Without a proof that {X^l_τ} is itself strong mixing, or an alternative CLT for the statewise sample mean, the convergence statement in Theorem 2.1 is unsupported. Please provide the missing argument or add assumptions that make the statewise sequence mixing by construction.
- [Appendix A.2, Steps l.1–l.3] The circular block bootstrap is applied to the concatenated statewise loss sequence, but no consistency result is proved for this object. Even if Assumption 2.3(i) holds, blocks drawn from the concatenated sequence can bridge distinct stress episodes and thereby create resampled samples that do not reflect the dependence structure of the original process. Gonçalves and White (2002) is cited for the block bootstrap variance estimator under dependence, but that result concerns the mean of an observed array, not a bootstrap applied to a subsequence defined by a data-dependent state indicator. The bootstrap p-value in Step l.3(c) is therefore not shown to approximate the distribution F^l_{ρ_{n_l}} needed in Theorem 2.1, and the coverage claim in Theorem A.1 is not closed.
- [Theorem 2.1 and Lemma A.2] The limit distribution is stated as F^l_{ρ_{n_l}}, which depends on n_l through the correlation matrix ρ_{n_l} implied by Ω_{n_l}. Lemma A.2 states (n_l)^{1/2}(V̄^l − ψ^l) →^d N(0, Ω_{n_l}) without assuming convergence of Ω_{n_l} to a fixed limit. Consequently "T^l_{max,M} →^d F^l_{ρ_{n_l}}" is not a standard convergence-in-distribution statement: the limiting object changes with the sample size. If Ω_{n_l} does not converge, the bootstrap critical values cannot be expected to converge either. Please either impose convergence of Ω_{n_l} (or of ρ_{n_l}) to a fixed limit and prove it under the maintained assumptions, or formulate the result as a uniform approximation over the sequence.
- [Section 2.3, Assumption 2.3(i) and Definition 2.1] The paper assumes that sgn(E[d^l_{ij,n_lτ}]) is constant across the within-state index τ, and states that this "implies that the null and the alternative are exhaustive." If this sign varies within a state, the set M^{l,*} defined in Definition 2.1 is not a well-defined target: the conditionally superior set could differ across episodes within the same state. The authors acknowledge this restriction in the text, but it is central to the coverage claim in Theorem A.1 and is not tested or discussed in the application, where stress states are defined by 252-day windows of a risk factor and may contain heterogeneous episodes. Please add a formal discussion of what M^{l,*} means when Assumption 2.3(i) is violated, or provide a diagnostic test for the assumption.
minor comments (7)
- [Abstract] The term "CMC S" appears with a spacing typo; it should be "CMCS".
- [Section 4.1 and Appendix A.5] There are typos: "defied" should be "defined" in Section 4.1, and "Appedix" should be "Appendix" in the heading of Appendix A.5.
- [Table 3 caption] The caption duplicates the sentence "The level of all tests is α = 0.05." Please remove the repetition.
- [Appendix A.2, Step l.1(b)(i)] The notation "(v_{b,1},...,v_{b,p}) = (ξ_{b1}, ξ_{b1}+1, ξ_{b1}+p−1)" does not describe a sequence of length p; it should read "(ξ_{b1}, ξ_{b1}+1, ..., ξ_{b1}+p−1) modulo n_l."
- [Appendix A.2, Step l.3(a)] The quantity \(\hat{var}(\bar d_{i·})\) is defined as (1/B)Σ_b(ξ*_{b,i}−ξ*_{b,·})^2, where ξ*_{b,·} is the cross-method average for bootstrap b. This is not the usual bootstrap variance estimator for \(\bar d_{i·}\); please clarify whether this is intentional and justify it, since Theorem 2.1 uses t^l_{i·} with \(\hat{var}(\bar d^l_{i·,n_l})\).
- [Theorem A.1 proof] The proof refers to "Assumption 4.1" instead of Assumption A.1, and part (ii) states lim inf P(i ∈ \hat M^l_{1−α}) = 0 where the stronger statement lim P = 0 is what is actually proved.
- [Section 4] The empirical comparison of CMCS and DFC is descriptive; no backtesting of the resulting ES forecasts or inference on the differences is provided, so the concluding claim that CMCS is "a robust tool" is stronger than what the tables demonstrate.
Circularity Check
No circularity: the CMCS derivation is self-contained and no claim reduces to its own inputs by construction.
full rationale
The paper does not fit any parameter and then rename it a prediction. The CMCS test statistic is built from statewise loss differentials under explicit Assumptions 2.1–2.4, and Theorem 2.1 follows from Lemma A.2, which invokes a CLT for alpha-mixing processes. The bootstrap implementation is explicitly the circular block bootstrap of Politis and Romano (1991) and the consistency argument cites Gonçalves and White (2002), both external to the authors. Hansen et al. (2011) is used as a benchmark and template, not as a self-citation, and no uniqueness theorem from the authors' own prior work is imported. The empirical application compares CMCS with the DFC procedure of Borup et al. (2024) on out-of-sample ES forecasts, but the comparison is descriptive and not a fitted-input-called-prediction loop. There are open technical questions about the step from mixing of z_t to mixing of the non-contiguous statewise subsequence {d^l_ij,nlτ}, and about the sample-size-dependent limit F^l_{rho_nl} in Theorem 2.1; however, those are concerns about proof correctness and asymptotic statement rigor, not about circularity. No load-bearing step reduces, by the paper's own equations, to its own inputs.
Assumptions & free parameters
free parameters (3)
- Bootstrap block length p =
Not reported
- Number of bootstrap replications B =
100
- MCS significance level alpha =
0.05
assumptions (8)
- domain assumption The process {W_t} and test functions {h_t} are alpha-mixing of order -r/(r-2) for some r>2 and gamma>0 (Assumption 2.1).
- domain assumption Finite (r+gamma)-th moments of conditional loss differentials (Assumption 2.2).
- ad hoc to paper The sign of E[d^l_{ij,nlτ}] is constant in τ, with positive variance and uniformly positive variance of normalized partial sums (Assumption 2.3).
- domain assumption Each state l has probability at least epsilon>0, so n_l goes to infinity (Assumption 2.4).
- ad hoc to paper The statewise loss subsequences are alpha-mixing of at most order -r/(r-2) (Corollary A.2).
- ad hoc to paper The circular block bootstrap applied to statewise losses consistently estimates the distribution of the statewise mean (Section 2.5).
- standard math A CLT for heterogeneous mixing arrays applies to the statewise means (Lemma A.1).
- domain assumption The loss function is strictly consistent for the target functional (Gneiting-Raftery).
Cite this review
Pith. "Pith review of Conditional Method Confidence Set." pith.science (2026). https://pith.science/paper/WQCSQB53
@misc{pith2026250521278,
author = {Pith},
title = {Pith review of: Conditional Method Confidence Set},
year = {2026},
howpublished = {\url{https://pith.science/paper/WQCSQB53}},
note = {Machine review of arXiv:2505.21278}
}
read the original abstract
This paper proposes a Conditional Method Confidence Set (CMCS) which allows to select the best subset of forecasting methods with equal predictive ability conditional on a specific economic regime. The test resembles the Model Confidence Set by Hansen et al. (2011) and is adapted for conditional forecast evaluation. We show the asymptotic validity of the proposed test and illustrate its properties in a simulation study. The proposed testing procedure is particularly suitable for stress-testing of financial risk models required by the regulators. We showcase the empirical relevance of the CMCS using the stress-testing scenario of Expected Shortfall. The empirical evidence suggests that the proposed CMCS procedure can be used as a robust tool for forecast evaluation of market risk models for different economic regimes.
Figures
Figures from the paper (9 more)
Reference graph
Works this paper leans on
-
[1]
For δM and ǫM, we suppress ( ·)l for ease of notation, though the pairs of equivalence test and elimination rule could differ per state. Regarding notation, δM = 0 refers to the case when H l 0,M cannot be rejected, and δM = 1 refers to the case when it is rejected. Note that the former als o implies that the testing procedure is halted. In the latter case...
work page 1984
-
[2]
− (p∆ 1 + (1 −p)∆ 2)p∆ 1 =pσ2 +p(1 −p)(∆ 2 1 − ∆ 1∆ 2) Σ 22 =Var (Dt · /BD St=1) =E[(Dt · /BD St=1)2] −E[Dt · /BD St=1]2 =E[D2 t |St = 1] ·E[/BD St=1] −E[Dt|St = 1] 2 ·E[/BD St=1]2 =p(σ2 + ∆ 2
-
[3]
−p2∆ 2 1 =pσ2 +p(1 −p)∆ 2 1 Given that it exists, the inverse of the 2 × 2 matrix Σ is given by Σ −1 = 1 det(Σ) Σ 22 −Σ 12 −Σ 12 Σ 11 , 42 with det(Σ) = (1 −p)pσ2( (1 −p)∆ 2 1 +p∆ 2 2 +σ2) . (13) Plugging into T h Thus, as factors in D,E,F in the expression of T h in Equation 12 above we obtain Σ −1 11 + 2Σ −1 12 + Σ −1 22 = 1 det(Σ) × ( pσ2 +p(1 −...
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.