REVIEW 4 major objections 4 minor 60 references
Predicting Realized Variance Out of Sample: Can Anything Beat The Benchmark?
T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper claims that volatility forecasts which do not clearly beat the HAR benchmark on standard error metrics can still produce economically large gains in daily straddle-portfolio returns, with a nested PCA-HAR model reaching a daily…
desk verdict A large empirical study with a novel evaluation angle, but the headline economic claim is not yet supported by the tables, inference, or a transaction-cost analysis. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the daily, firm-level volatility-risk-premium sorting variable $\ln(\hat{RV}_{i,t+1}/IV_{i,t})$, the log gap between a model's forecast of next-day realized variance and the market's at-the-money implied variance. Portfolios are formed each day by ranking S&P 500 straddles into quintiles on this variable and taking the long-short high-minus-low portfolio, with all variables measured before 3:55pm and trades placed before the close to avoid look-ahead bias. The model that carries the result is the nested PCA-HAR specification, which uses a rolling one-year eigendecomposition of the firm-level realized-variance panel to extract common factors, forecasts factor levels with daily/weekly/monthly lags, forecasts the residual with the same HAR structure, and adds the firm's own HAR predictors. The walk-forward pseudo out-of-sample protocol with a 250-day rolling window, nested cross-validation for hyperparameters, and reconstruction of forecasts each day for every firm is what gives the portfolio returns their claim to being out-of-sample.
What would settle it
Lock the PCA-HAR specification (number of factors, inclusion of residual forecasts, estimation window) before any analysis of the option sample, then evaluate it on a genuinely untouched later period such as 2020–2024; if the long-short straddle Sharpe ratio no longer beats HAR's, the paper's economic-significance claim collapses.
Extended reading notes
Core claim
The central discovery is that the ranking of volatility forecasts changes when they are judged by the returns of portfolios built from the forecasts rather than by forecast-error statistics, and that a model combining common factors with firm-specific predictors can exploit this. The PCA-HAR model extracts principal components from the cross-section of firm-level realized variances over a rolling year, forecasts those factors with HAR-style daily, weekly, and monthly lags, and adds the firm's own HAR predictors; it nests both the HAR model and a pure factor model. On pooled and cross-sectional error measures the model does not consistently beat HAR, yet the long-short at-the-money straddle portfolio sorted on $\ln(\hat{RV}/IV)$ earns 2.06% per day with a Sharpe ratio of 0.80, against 1.39% and 0.63 for the HAR-based portfolio. The paper interprets this as evidence that standard forecast evaluation is an incomplete guide to economic value, and that the realized-minus-implied volatility spread remains a strong daily predictor of option returns when the realized side is forecast more effectively.
Load-bearing premise
The headline improvement rests on the assumption that the PCA-HAR model was chosen before peeking at the 1996–2019 option-return sample, rather than selected after trying many forecasting variants on that same sample; if that assumption fails, the measured Sharpe and return gains could be luck.
Editorial extensions
If this is right
- If the paper's claim is right, forecast-error rankings by RMSE, MAE, or QLIKE are not reliable proxies for the economic value of a volatility forecast, and models that tie or lose on those metrics can still win in portfolio construction.
- The realized-minus-implied volatility spread, already documented at the monthly horizon, appears to work at the daily frequency for individual equity options when the realized-volatility forecast is improved.
- The nested PCA-HAR model's success suggests common-factor information in the cross-section of firm variances contains predictive content for option returns that firm-specific HAR dynamics alone miss.
- Portfolio-based evaluation could become a standard complement to statistical loss functions for comparing volatility forecasts, and model training could be re-targeted at portfolio objectives rather than squared error.
Reading between the lines
- A natural extension would be to train the volatility model directly on the sorting signal or on option-portfolio returns, rather than on squared forecast error; the paper's own evidence implies this could sharpen the spread further, but it does not test this.
- The Sharpe improvement could be checked against multiple-testing corrections, since the paper concedes that pseudo-out-of-sample evaluation can still be mined; applying a correction for the number of model variants tried may shrink the gap between PCA-HAR and HAR.
- The same forecast-to-portfolio pipeline could be transferred to other asset classes or to volatility products such as VIX futures, where the realized-minus-implied spread is also observable; this is an extension the paper leaves implicit.
- Transaction costs are deliberately set aside; with realistic bid-ask costs and margin on short straddles, the 2.06% daily gross return would shrink, and the ranking across models could change.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper forecasts daily realized variance for a large panel of S&P 500 firms from 1993 to 2019 using several models: the HAR benchmark, LASSO regressions on a high-dimensional cross-section, a nested PCA-HAR factor model, and equal-weighted forecast ensembles. Forecasts are evaluated both by standard loss functions (RMSE, MAE, QLIKE, MZ-R2) and by the returns of daily at-the-money delta-neutral straddle portfolios sorted on the forecast-minus-implied volatility spread. The paper reports that, while no model unambiguously beats HAR on forecast-error metrics, the PCA-HAR model and the ensemble produce long-short straddle portfolios with higher average returns and Sharpe ratios than HAR (e.g., 2.06% per day and Sharpe 0.80 for PCA-HAR versus 1.39% and 0.63 for HAR, Table 4), and interprets this as evidence that portfolio-based evaluation can reveal value not captured by standard forecast-loss rankings.
Significance. If the results hold, the paper makes a useful empirical contribution by scaling up daily firm-level realized-variance forecasting to the full S&P 500 cross-section and by drawing attention to the disconnect between standard forecast-error rankings and end-user portfolio performance. The design has notable strengths: a walk-forward out-of-sample procedure, a clear timing convention that uses information only up to 3:55pm to allow trades before the close, careful filtering of option data, and explicit discussion of look-ahead bias and data-snooping concerns. However, the economic-significance conclusion currently rests on gross midpoint-based returns with no transaction costs, on differences in Sharpe ratios without any inference, and on a model that was selected after trying several variants on the same sample. The significance of the paper is therefore conditional on these load-bearing issues being addressed.
major comments (4)
- [Section 6.2, Table 4; Section 5.2] The headline economic-significance claim is computed from OptionMetrics end-of-day bid-ask midpoints, with no transaction costs or margin requirements, as the paper itself concedes in Section 6.2. Because the long-short portfolios are rebalanced daily and equal-weighted, each rebalance crosses the bid-ask spread on both legs of each straddle; the PCA-HAR advantage over HAR is 0.67% per day (0.0206 vs 0.0139), which is within a plausible one-way ATM-straddle round-trip cost. The manuscript should either report returns net of realistic transaction costs or compute the break-even cost per rebalance that would erase the advantage; without this, the 'economically significant' conclusion is not established.
- [Table 4 and Table 5] No standard errors, confidence intervals, or significance tests are reported for the differences in mean returns or Sharpe ratios across forecasts. Given the daily long-short portfolio return kurtosis of 18.64 for PCA-HAR (Table 5), the difference between the PCA-HAR Sharpe (0.8044) and the HAR Sharpe (0.6317) may be within sampling noise. Please provide HAC or bootstrap inference for the average return and Sharpe difference, and ideally a test of equal performance across the two portfolios.
- [Section 3.1, Section 4.3, footnotes 20 and 24] The paper's main model, PCA-HAR, is presented in Section 4.3 only after describing several alternatives, and footnotes 20 and 24 report that various other specifications were tried with little difference. Section 3.1 explicitly concedes that pseudo-OOS tests can still suffer from p-hacking, multiple testing, and data mining, yet no correction (e.g., a multiple-testing adjustment) or a true holdout period is applied. Since the central claim is that PCA-HAR improves option-portfolio performance, the possibility of selection on the same sample is load-bearing; a pre-specified specification, a final holdout, or a multiple-testing assessment is needed.
- [Abstract, Section 6.1, Table 2] The abstract's claim that 'marginal improvements to standard forecast error measurements can lead to economically significant gains' is not supported by Table 2: PCA-HAR has larger RMSE, MAE, and QLIKE than HAR in Panels A and B, and only in the pooled MZ-R2 (Panel C) does it beat HAR. If anything, the model looks worse on most loss functions, so the paper should either identify which specific forecast-error improvement drives the portfolio gains, or reframe the conclusion to say that forecast-error rankings are ambiguous and do not predict portfolio performance. The current phrasing mischaracterizes the evidence in the paper's own tables.
minor comments (4)
- [Table 1] Table 1 reports kurtosis values without stating whether they are excess kurtosis; Section 6.2 later refers to excess kurtosis, so define this consistently in the table note.
- [Throughout] There are numerous typographical and encoding issues, such as 'A ¨ ıt-Sahalia' in references, 'varainces' in Section 4.3, 'for for' in Section 2.3, and 'of of' in Section 4.2, which should be corrected.
- [Equations (9) and (10)] Equations (9) and (10) use notation like 'Et [Rstocks t+1 ]' and 'Et [Roptions t+1 ]' without definitions; clarify what these expectations represent.
- [Section 6.2] The paper states that the results 'confirm the so-called volatility risk premium is large and economically significant' in daily option returns, but Table 3 Panel A shows unconditional ATM straddle excess returns are slightly negative; the confirmation pertains to long-short sorted portfolios, so the wording should be qualified.
Circularity Check
No significant circularity: the option-portfolio results are genuine out-of-sample outcomes and the PCA-HAR advantage is not forced by construction.
full rationale
The paper's derivation chain is linear and data-driven. Realized variances are measured from TAQ intraday prices; each forecast model (HAR, LASSO, PCA-HAR, equal-weighted ensemble) is re-estimated on rolling 250-day windows ending before 3:55pm ET; the sorting variable is VRP_{t-1} = ln(RV_hat_t / IV_{t-1}) (Eqs. 16-17 and 40), using only information available when portfolios are formed; and straddle returns are computed from OptionMetrics midpoints over the subsequent day. Option returns never enter the estimation of any volatility model or the selection of hyperparameters, and the portfolio weights are equal weights, so no parameter is fitted to the outcome being predicted. The paper's own warning that pseudo-OOS tests may suffer from p-hacking and multiple testing (Section 3.1) is a model-selection and data-mining caveat, not a circular reduction: no equation in the paper equates the reported Sharpe ratios or average returns to a fitted input, and the PCA-HAR specification is not defined in terms of the option returns it is used to predict. Citations to Jones et al. (2021a, 2021b) and Jones (2001, 2006) are supporting references for data construction, option-return properties, and factor extraction; they do not function as a uniqueness theorem or as the sole justification for the central claim. The central economic-significance finding is therefore an independent out-of-sample empirical result, subject to the usual robustness concerns about multiple testing and midpoint pricing, but not circular.
Assumptions & free parameters
free parameters (4)
- LASSO penalty lambda =
not reported
- Number of PCA factors K =
3
- Rolling window length W =
250 days
- Ensemble weights =
equal weights
assumptions (5)
- standard math Realized variance converges to integrated variance as sampling frequency increases (RV to IV under a continuous-time semimartingale)
- domain assumption Returns evolve as a continuous-time semimartingale with finite quadratic variation
- ad hoc to paper The 3:55pm partial-day realized variance is a valid, consistently scaled predictor for next-day full-day realized variance
- domain assumption At-the-money straddle returns computed from bid-ask midpoints approximate realizable returns
- domain assumption The realized-minus-implied volatility spread predicts the cross-section of option returns
Cite this review
Pith. "Pith review of Predicting Realized Variance Out of Sample: Can Anything Beat The Benchmark?." pith.science (2026). https://pith.science/paper/REP22CJ2
@misc{pith2026250607928,
author = {Pith},
title = {Pith review of: Predicting Realized Variance Out of Sample: Can Anything Beat The Benchmark?},
year = {2026},
howpublished = {\url{https://pith.science/paper/REP22CJ2}},
note = {Machine review of arXiv:2506.07928}
}
read the original abstract
The discrepancy between realized volatility and the market's view of volatility has been known to predict individual equity options at the monthly horizon. It is not clear how this predictability depends on a forecast's ability to predict firm-level volatility. We consider this phenomenon at the daily frequency using high-dimensional machine learning models, as well as low-dimensional factor models. We find that marginal improvements to standard forecast error measurements can lead to economically significant gains in portfolio performance. This makes a case for re-imagining the way we train models that are used to construct portfolios.
Figures
Reference graph
Works this paper leans on
-
[1]
Y. A\"it-Sahalia and J. Jacod. High F requency F inancial E conometrics . Princeton University Press, first edition, 2014
work page 2014
-
[2]
Y. A\"it-Sahalia, Y. Wang, and F. Yared. Do Option Markets Correctly Price the Probabilities of Movement of the Underlying Asset? Journal of Econometrics, 102 0 (1): 0 67--110, 2001
work page 2001
-
[3]
Y. A\"it-Sahalia, C. Li, and C. X. Li. Implied Stochastic Volatility Models . The Review of Financial Studies, 34 0 (1): 0 394--450, 03 2020
work page 2020
-
[4]
M. Ammann and M. Moerke. Credit Variance Risk Premiums . Working paper , 2021
work page 2021
-
[5]
T. Andersen and T. Bollerslev. Answering the Skeptics: Yes, Standard Volatility Models Do Provide Accurate Forecasts . International Economic Review, 39 0 (4): 0 885--905, 1998
work page 1998
-
[6]
T. Andersen, T. Bollerslev, F. X. Diebold, and H. Ebens. The distribution of realized stock return volatility . Journal of Financial Economics, 61 0 (1): 0 43--76, 2001
work page 2001
-
[7]
T. Andersen, T. Bollerslev, P. F. Christoffersen, and F. X. Diebold. Volatility and Correlation Forecasting , pages 777--878. Handbook of Economic Forecasting. 2006
work page 2006
-
[8]
G. Bakshi and N. Kapadia. Delta-Hedged Gains and the Negative Market Volatility Risk Premium . The Review of Financial Studies, 16 0 (2): 0 527--566, 06 2015
work page 2015
Show all 60 references
-
[9]
T. G. Bali, R. F. Engle, and S. Murray. Empirical Asset Pricing: The Cross Section of Stock Returns . Wiley, 2016
2016
-
[10]
F. Black. Studies of Stock Price Volatility Changes . Proceedings of the Business and Economics Section of the American Statistical Association, page 177–181, 1976
1976
-
[11]
Black and M
F. Black and M. Scholes. The Pricing of Options and Corporate Liabilities . Journal of Political Economy, 81 0 (3): 0 637--654, 1973
1973
-
[12]
Bollerslev
T. Bollerslev. Generalized autoregressive conditional heteroskedasticity. Journal of Econometrics, 31 0 (3): 0 307--327, 1986
1986
-
[13]
Bollerslev, G
T. Bollerslev, G. Tauchen, and H. Zhou. Expected Stock Returns and Variance Risk Premia . The Review of Financial Studies, 22 0 (11): 0 4463--4492, 02 2009
2009
-
[14]
Bollerslev, N
T. Bollerslev, N. Sizova, and G. Tauchen. Volatility in Equilibrium: Asymmetries and Dynamic Dependencies . Review of Finance, 16 0 (1): 0 31--80, 03 2011
2011
-
[15]
Bollerslev, D
T. Bollerslev, D. Osterrieder, N. Sizova, and G. Tauchen. Risk and return: Long-run relations, fractional cointegration, and return predictability . Journal of Financial Economics, 108 0 (2): 0 409--424, 2013
2013
-
[16]
M. W. Brandt and C. S. Jones. Volatility Forecasting With Range-Based EGARCH Models . Journal of Business & Economic Statistics, 24: 0 470 -- 486, 2006
2006
-
[17]
Broadie, M
M. Broadie, M. Chernov, and M. Johannes. Understanding Index Option Returns . The Review of Financial Studies, 22 0 (11): 0 4493--4529, 05 2009
2009
-
[18]
J. Y. Campbell and S. B. Thompson. Predicting Excess Stock Returns Out of Sample: Can Anything Beat the Historical Average? The Review of Financial Studies, 21 0 (4): 0 1509--1531, 11 2007
2007
-
[19]
Cao and B
J. Cao and B. Han. Cross section of option returns and idiosyncratic stock volatility . Journal of Financial Economics, 108 0 (1): 0 231--249, 2013
2013
-
[20]
Carr and R
P. Carr and R. Lee. Volatility Derivatives . Annual Review of Financial Economics, 1 0 (1): 0 319--339, 2009
2009
-
[21]
Carr and L
P. Carr and L. Wu. Vol, Skew, and Smile Trading
-
[22]
Carr and L
P. Carr and L. Wu. Variance Risk Premiums . The Review of Financial Studies, 22 0 (3): 0 1311--1341, 04 2008
2008
-
[23]
Carr and L
P. Carr and L. Wu. Leverage Effect, Volatility Feedback, and Self-Exciting Market Disruptions . Journal of Financial and Quantitative Analysis, 52 0 (5): 0 2119–2156, 2017
2017
-
[24]
P. Carr, L. Wu, and Z. bai Zhang. Using Machine Learning to Predict Realized Variance . Journal of Investment Management, 18: 0 1--16, 2020
2020
-
[25]
M. D. Cattaneo, R. K. Crump, M. H. Farrell, and E. Schaumburg. Characteristic-Sorted Portfolios: Estimation and Inference . The Review of Economics and Statistics, 102 0 (3): 0 531--551, 07 2020
2020
-
[26]
Chinco, A
A. Chinco, A. D. Clark-Joseph, and M. Ye. Sparse Signals in the Cross-Section of Returns . The Journal of Finance, 74 0 (1): 0 449--492, 2019
2019
-
[27]
Chordia, A
T. Chordia, A. Goyal, and A. Saretto. Anomalies and False Rejections . The Review of Financial Studies, 33 0 (5): 0 2134--2179, 02 2020
2020
-
[28]
R. Cont. Empirical properties of asset returns: stylized facts and statistical issues. Quantitative Finance, 1 0 (2): 0 223--236, 2001
2001
-
[29]
F. Corsi. A Simple Approximate Long-Memory Model of Realized Volatility . Journal of Financial Econometrics, 7 0 (2): 0 174--196, 02 2009
2009
-
[30]
J. D. Coval and T. Shumway. Expected Option Returns . The Journal of Finance, 56 0 (3): 0 983--1009, 2001
2001
-
[31]
J. C. Cox and M. Rubinstein. Options Markets. Englewood Cliffs, N.J, Prentice-Hall, 1985
1985
-
[32]
F. X. Diebold and M. Shin. Machine learning for regularized survey forecast combination: Partially-egalitarian lasso and its derivatives. International Journal of Forecasting, 35 0 (4): 0 1679--1691, 2019
2019
-
[33]
R. F. Engle. Autoregressive Conditional Heteroscedasticity with Estimates of the Variance of United Kingdom Inflation . Econometrica, 50 0 (4): 0 987--1007, 1982
1982
-
[34]
R. F. Engle and A. J. Patton. What good is a volatility model? Quantitative Finance, 1, 02 2001
2001
-
[35]
W. Ferson. Empirical Asset Pricing: Models and Methods . MIT Press, 2019
2019
-
[36]
M. B. Garman and M. J. Klass. On the Estimation of Security Price Volatilities from Historical Data . The Journal of Business, 53 0 (1): 0 67--78, 1980
1980
-
[37]
Goyal and A
A. Goyal and A. Saretto. Cross-Section of Option Returns and Volatility . Journal of Financial Economics, 94 0 (2): 0 310 -- 326, 2009
2009
-
[38]
C. R. Harvey, Y. Liu, and H. Zhu. … and the Cross-Section of Expected Returns . The Review of Financial Studies, 29 0 (1): 0 5--68, 10 2015
2015
-
[39]
Hasanhodzic and A
J. Hasanhodzic and A. W. Lo. On Black's Leverage Effect in Firms with No Leverage . The Journal of Portfolio Management, 46 0 (1): 0 106--122, 2019
2019
-
[40]
Hastie, R
T. Hastie, R. Tibshirani, and J. Friedman. The Elements of Statistical Learning . Springer Series in Statistics. Springer, New York, second edition, 2009. Data mining, inference, and prediction
2009
-
[41]
J. C. Jackwerth. Recovering Risk Aversion from Option Prices and Realized Returns . The Review of Financial Studies, 13 0 (2): 0 433--451, 06 2015
2015
-
[42]
C. S. Jones. Extracting factors from heteroskedastic asset returns. Journal of Financial Economics, 62 0 (2): 0 293--325, 2001
2001
-
[43]
C. S. Jones. A Nonlinear Factor Analysis of S&P 500 Index Option Returns . Journal of Finance, 61 0 (5): 0 2325--2363, 2006
2006
-
[44]
C. S. Jones, J. Duarte, and J. Wang. Very Noisy Option Prices and Inferences Regarding Option Returns . Working Paper, 2021 a
2021
-
[45]
C. S. Jones, S. L. Heston, M. Khorram, S. Li, and H. Mo. Option Momentum . Working paper , 2021 b
2021
-
[46]
L. Y. Liu, A. J. Patton, and K. Sheppard. Does anything beat 5-minute RV? A comparison of realized measures across multiple asset classes . Journal of Econometrics, 187 0 (1): 0 293--311, 2015
2015
-
[47]
Mandelbrot
B. Mandelbrot. The Variation of Certain Speculative Prices . The Journal of Business, 1963
1963
-
[48]
N. Meddahi. A Theoretical Comparison between Integrated and Realized Volatility . Journal of Applied Econometrics, 17 0 (5): 0 479--508, 2002
2002
-
[49]
R. C. Merton. Theory of Rational Option Pricing . The Bell Journal of Economics and Management Science, 4 0 (1): 0 141--183, 1973
1973
-
[50]
Mincer and V
J. Mincer and V. Zarnowitz. The Evaluation of Economic Forecasts . In Economic Forecasts and Expectations: Analysis of Forecasting Behavior and Performance , pages 3--46. National Bureau of Economic Research, Inc, 1969
1969
-
[51]
Muravyev
D. Muravyev. Order Flow and Expected Option Returns . The Journal of Finance, 71 0 (2): 0 673--708, 2016
2016
-
[52]
S. Nagel. Machine Learning in Asset Pricing , volume 8. Princeton University Press, 2021
2021
-
[53]
Parkinson
M. Parkinson. The Extreme Value Method for Estimating the Variance of the Rate of Return . The Journal of Business, 53 0 (1): 0 61--65, 1980
1980
-
[54]
A. J. Patton. Data-Based Ranking of Realised Volatility Estimators . Journal of Econometrics, 161 0 (2): 0 284 -- 303, 2011 a
2011
-
[55]
A. J. Patton. Volatility Forecast Comparison Using Imperfect Volatility Proxies . Journal of Econometrics, 160 0 (1): 0 246--256, 2011 b
2011
-
[56]
L. C. G. Rogers and S. E. Satchell. Estimating Variance From High, Low and Closing Prices . The Annals of Applied Probability, 1 0 (4): 0 504 -- 512, 1991
1991
-
[57]
W. G. Schwert. Why Does Stock Market Volatility Change Over Time? The Journal of Finance, 44 0 (5): 0 1115--1153, 1989
1989
-
[58]
P. V. Tassel. The Law of One Price in Equity Volatility Markets . Working Paper, 2020
2020
-
[59]
Tibshirani
R. Tibshirani. Regression shrinkage and selection via the lasso. Journal of the Royal Statistical Society. Series B (Methodological), 58 0 (1): 0 267--288, 1996
1996
-
[60]
Yang and Q
D. Yang and Q. Zhang. Drift‐Independent Volatility Estimation Based on High, Low, Open, and Close Prices . The Journal of Business, 73 0 (3): 0 477--492, 2000
2000
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.