{"id":"dff1831b-076e-4a51-995a-5e352369c017","arxiv_id":"2505.11163","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"Incremental fine-tuning of the TimesFM foundation model improves one-day-ahead realized volatility forecasts and beats HAR, ARFIMA, CHAR, and RGARCH benchmarks on average losses across 21 global equity indices.","lead":"This paper tests whether a general-purpose AI time-series model, Google's TimesFM, can forecast stock market volatility better than standard financial models. It finds that fine-tuned versions improve accuracy over traditional benchmarks, but the statistical evidence is undermined by the way the significance tests were applied.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"DM/GW tests are applied to 21 cross-sectional average losses rather than to the time series of daily loss differentials, so the claimed statistical outperformance is not established.","rationale":"The reader identified the same load-bearing concern: DM/GW tests applied to cross-sectional averages of per-stock losses rather than to time-series loss differentials. This is indeed the most serious flaw because the abstract's central claim is explicitly framed around those tests. If the tests are misapplied, the statistical outperformance claim is unsupported regardless of the point estimates. The RGARCH inconsistency is also serious and independently justifies caution, but the DM/GW misapplication is more central to the advertised contribution. I agree with the reader's REJECT verdict: the descriptive results may be useful, but the paper as submitted does not establish statistical superiority. The requested check—recomputing DM p-values from daily loss differentials—would settle whether the concern lands, and if it does, the core claim fails. The reader's confidence of MODERATE seems appropriate because a corrected analysis could still favor TimesFM on average losses, but the current evidence is insufficient.","tokens_in":54788,"tokens_out":4717,"duration_ms":45446,"concrete_test":"Reconstruct the daily out-of-sample QLIKE loss differential series for TFM512PT vs HAR and TFM512IL vs HAR for each of the 21 stocks. Compute the DM statistic on the pooled daily loss differentials (or per-stock DM statistics combined with a stock-level cluster-robust variance estimator) using Newey-West HAC. If the corrected p-values exceed 0.10, the conclusion that these models statistically outperform under QLIKE collapses. Separately, recompute a paired t-test on the 21 stock-level average QLIKEs to verify whether the cross-sectional comparison even reaches significance, and compare the resulting p-values with Table 12's entries to check the reported direction.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is statistical superiority of TimesFM variants over HAR, ARFIMA, CHAR, and RGARCH, based on Diebold-Mariano (DM) and Giacomini-White (GW) tests. The DM test is asymptotically valid for a stationary time series of loss differentials d_t = L(y_t, yhat_i,t) - L(y_t, yhat_j,t), with a HAC variance estimator. Table 7's caption, however, states: 'We applied these tests to cross-section (21 values) of average MSE obtained from different models for 21 stocks.' This replaces the time-series observations with 21 per-stock average losses. The resulting test is a cross-sectional comparison of stock-level averages, not the DM/GW test whose null and asymptotic theory are described in Section 4. The p-values in Tables 7-12 therefore do not have the claimed standard normal or chi-square null distributions. This is the load-bearing flaw because the abstract and conclusion assert 'statistically outperform' specifically through these tests. A secondary issue compounds it: for the QLIKE table (Table 12), the text says TFM512PT and TFM512IL achieve p-values below 0.10 against econometric models, but the displayed row p-values for those models are all above 0.9; only the reverse comparisons (econometric row vs TimesFM column) show small p-values, contradicting the caption's row/column convention. Both issues undermine the headline statistical claim, even though the descriptive average-loss comparisons may still be informative.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript evaluates Google's TimesFM, a decoder-only time-series foundation model, for one-day-ahead realized variance forecasting on 21 global equity indices from the Oxford-Man Institute Realized Library over 2000-2021. Six TimesFM configurations are compared with HAR, ARFIMA, CHAR, and RGARCH benchmarks (plus log-transformed variants): three context lengths (64, 128, 512) in zero-shot (pretrained) form and after an incremental-learning fine-tuning scheme that iteratively refines the model on expanding windows. Performance is assessed under six loss functions (MSE, MAE, MDA, MAPE, sMAPE, QLIKE), with Diebold-Mariano and Giacomini-White tests, Model Confidence Sets, and skill-score tables. The headline claims are that fine-tuned variants improve forecast accuracy and that TimesFM models 'statistically outperform' the traditional benchmarks, with TFM512PT and TFM512IL highlighted under QLIKE and TFM64logIL identified as the overall best model.","tokens_in":55038,"tokens_out":23896,"duration_ms":206693,"significance":"If the statistical claims could be substantiated, this would be a useful contribution to the emerging literature on foundation models in financial forecasting. The paper's strengths are genuine: the incremental fine-tuning protocol is described in enough detail to reproduce (data splits, hyperparameters, checkpointing), the walk-forward evaluation is applied uniformly to the neural and econometric models, and the descriptive layer is concrete and falsifiable, notably the finding that TFM64logIL attains the lowest average MSE and MAPE and the highest MDA in Table 3. The credibility of the headline result, however, depends entirely on the DM/GW significance tests, and that layer is, as presented, methodologically invalid and numerically inconsistent; the claimed 'statistical outperformance' cannot be taken as established from the manuscript as written.","major_comments":[{"comment":"The DM and GW tests are applied to the wrong sampling unit for the asymptotic theory the paper itself spells out. Section 4 defines the loss differential d_t,i,j over the out-of-sample time series t = 1,...,P and states that the DM statistic is asymptotically standard normal under a HAC variance estimator, and the GW statistic chi-square. Every DM/GW table (Tables 7-12) instead reports, in the caption's words, tests applied to the 'cross-section (21 values) of average MSE' (resp. MAE, MDA, MAPE, sMAPE, QLIKE) obtained for the 21 stocks. Replacing the time series of daily loss differentials with 21 per-stock average losses removes the temporal structure on which the null distributions, the HAC estimator, and the P going to infinity asymptotics all rely; with 21 observations the normal and chi-square approximations and the Newey-West variance estimator are not meaningful. Since the abstract and conclusion base the 'statistically outperform' claim precisely on these p-values, the central statistical claim is not established by the reported tests. The same concern applies to the Model Confidence Set summarized in Figure 4, whose implementation (cross-sectional versus time-series loss differentials, bootstrap scheme) is not described.","section":"§4 (Evaluation Framework) and Tables 7-12 (captions)"},{"comment":"The p-value tables contain patterns that cannot arise from any pairwise comparison of the type described. In Table 7, the RGARCH column is 0.8379 for all 18 non-diagonal rows and the RGARCH row is 0.1621 or 0.1622 for all columns; Table 8 shows the RGARCH row essentially constant at 0.1439-0.1440 and Table 10 at 0.1553-0.1577. These near-constant entries across 19 distinct model pairs are not compatible with pairwise tests on 21 per-stock average losses, which would necessarily differ from one row or column to the next. Moreover, Table 3 reports RGARCH's average MSE (0.00804) as roughly six to seven times larger than every other model's, so a paired cross-sectional test of H1: MSE_RGARCH > MSE_other would be rejected with p-values near zero, not p of about 0.16 (RGARCH row) or p of about 0.84 (RGARCH column). The authors must trace the computation or the data feeding it and regenerate the p-value tables before the statistical claims can be evaluated.","section":"Tables 7-12 (p-value patterns) and Table 3"},{"comment":"The text reading of the QLIKE results is reversed relative to the tables' stated convention. Table 12's caption defines H0: Qlike_i = Qlike_j versus H1: Qlike_i > Qlike_j with row model i and column model j, so a low p-value means the row model is significantly worse. Under this convention, the evidence that TFM512PT and TFM512IL 'consistently outperform traditional models under the Q-like loss function' (Conclusion, and the analogous passage in Section 5) lives in the econometric rows versus the TimesFM columns, for example Table 12 ARFIMA row versus TFM512PT column (DM p = 0.0493) and HAR row versus TFM512IL column (DM p = 0.0219); the TFM512PT and TFM512IL rows themselves show p-values above 0.94 against every econometric model. As written, the sentence 'TFM512PT and TFM512IL consistently achieve p-values below 0.10 in all pairwise comparisons against econometric models' points the reader to the wrong entries. The MDA tables (Table 9) need a separate sign check for the same reason, since higher MDA indicates better directional accuracy rather than a larger loss, which is the opposite of the 'row is worse when p is low' interpretation used in the captions.","section":"§5 'Is there any statistically better model?' and Conclusion (Tables 7-12 orientation)"},{"comment":"The Section 5 narrative reverses the direction of the relative-error tables. Tables 4-6 define each entry as the error of the column model divided by the error of the row benchmark, so values below 1 mean the column model is better. The text instead says 'A model is considered strong if its skill scores exceed 1 in its respective row, meaning it outperforms the benchmark' and later asserts that models 'achieve skill scores exceeding 1 in MSE, MAD, and Qlike, reinforcing their robustness'; both statements are backward relative to the table definitions, and specific claims are contradicted by the tables themselves, for example Table 5 Panel B shows TFM64IL with 0.836 against the ARFIMA row, below 1, and the text's claim that 'TFM64IL achieves skill scores greater than 1 in all six loss functions' is false under either reading of Tables 4-6. This reversal runs through the 'Identifying the Best Performing Model' discussion and needs to be corrected throughout. In addition, the abstract's blanket statement that 'Fine-tuned variants not only improve forecast accuracy' is not supported by Table 3 for the linear IL models, whose average MSE (0.00124-0.00128) exceeds ARFIMA's (0.00120), CHAR's (0.00117), and HAR's (0.00118); the improvement over econometric benchmarks in MSE is specific to the log-transformed variants.","section":"§5 'Comparative Performance' and 'Identifying the Best Performing Model' (Tables 4-6, Table 3)"}],"minor_comments":[{"comment":"Metric terminology is inconsistent: Section 4 defines MAE, Table 3 reports 'MAD' (Mean Absolute Deviation), Tables 4-6 Panel B and Table 8 switch back to MAE, and Table 11's caption refers to 'average MAPE' while the panel reports sMAPE. Standardize the names and the captions.","section":"Metrics and captions"},{"comment":"Small errors in figures and text: Figure 5's caption says models are evaluated 'across all remaining 18 models' although 19 models are compared; Figure 3's caption says the central line represents the median MSE even though the panels show relative errors for six different loss functions; Section 1 contains the typo 'dicusses'; Section 3 has 'a incremental fine-tuning procedure'; and Section 2's heading 'Literature Review Realized Volatility Forecasting' is missing punctuation.","section":"Figures and text"},{"comment":"The fine-tuning section says the authors 'adopted the code fine-tuning in the TimesFM library'; for reproducibility they should cite the repository version or commit and state how the hyperparameters and the three context lengths were chosen relative to the out-of-sample period, since the same test window appears to drive both the configuration comparison and the headline results.","section":"§3 Incremental Fine-Tuning"},{"comment":"Tables 7-12 report only p-values; the per-stock average losses or the loss-differential series that feed the tests should be made available, or the code provided, so that the computations can be verified and the cross-sectional dispersion assessed.","section":"Tables 7-12"},{"comment":"The novelty statement 'to the best of our knowledge, this study is the first to extensively explore the application of time series foundation models for volatility forecasting' should be reconciled with the authors' own reference [27] applying foundation models to VaR forecasting and with the active concurrent literature; the claim should be scoped more carefully.","section":"Introduction"},{"comment":"Table 2 labels the series 'realized volatility' although the data are 5-minute sub-sampled realized variance from the Oxford-Man library; the labels should be aligned with the definitions in the dataset section.","section":"Table 2"}],"recommendation":"major_revision","confidential_remarks":"The near-constant RGARCH p-values in Tables 7-12 (for instance 0.8379 in all 18 rows of the RGARCH column of Table 7, and 0.1621 across the RGARCH row) are what one would expect from a coding or bookkeeping error rather than from the pairwise testing procedure described; before the next round the authors should be required to supply the code reproducing Tables 7-12 and the underlying per-stock loss values. The manuscript is not circular, but the post hoc selection of TFM64logIL and TFM512PT as the favored variants, without any multiple-comparison control, should be acknowledged. The descriptive benchmark result may still be publishable in an empirical finance or ML venue once the statistical layer is redone; whether that reworked analysis meets this journal's bar is for the editor to judge."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague — quick read of arXiv:2505.11163. The genuinely new piece is the application of a time-series foundation model (TimesFM) to one-day-ahead realized variance forecasting, with an incremental fine-tuning scheme tested on 21 global equity indices against HAR, CHAR, ARFIMA, and RGARCH. That is a reasonable empirical benchmark, and the descriptive results are worth having: the fine-tuned models, especially TFM64log IL, post lower average MSE/MAE/MAPE than the econometric benchmarks across the cross-section. The finding that short context plus incremental updating beats longer zero-shot context is concrete and plausible.\n\nThe soft spot is the significance testing, and it is load-bearing. The DM and GW tests are applied to 21 per-stock average losses, not to the time series of daily loss differentials the tests are designed for. The paper's own caption says exactly that. So the p-values in Tables 7–12 do not have the standard normal or chi-square null distributions claimed. On top of that, the QLIKE interpretation in the text runs opposite to the table caption: a low p-value in the caption means the row model is worse, not better, so the claim that TFM512PT/TFM512IL are 'statistically superior' with p<0.10 is reading the numbers backwards. And the RGARCH rows show nearly constant p-values (~0.16) across all columns, which suggests a mechanical error in the computation. These are not cosmetic; the abstract's 'statistically outperform' is exactly what the tests are supposed to support.\n\nThe average-loss comparisons may still hold up under a corrected testing procedure (e.g., proper time-series DM tests with HAC errors, or a panel approach), but that work is not in the paper.\n\nReadership: people working on ML volatility forecasting will want to know about this application, and the descriptive benchmark section is a useful reference point. But as submitted, the statistical claims overreach.\n\nRecommendation: send it to peer review with a clear request to redo the significance tests and fix the interpretation. It deserves a serious referee even though the current version needs heavy revision.","headline":"A useful first benchmark of TimesFM for realized volatility, but the claimed statistical outperformance rests on a misapplied DM/GW test and needs to be redone.","tokens_in":55568,"tokens_out":2220,"would_cite":false,"duration_ms":23244,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62M10","91B84","68T07"],"pacs":[],"model":"deepseek-v4-flash","headline":"A time-series foundation model, fine-tuned on rolling chunks of recent data, delivers one-day-ahead volatility forecasts that statistically beat HAR, ARFIMA, CHAR, and RGARCH benchmarks.","keywords":["volatility forecasting","realized variance","foundation models","time series analysis","incremental learning","transfer learning","TimesFM","Diebold-Mariano test"],"falsifier":"Re-run the comparison at the individual-stock level: for each of the 21 indices, form the daily QLIKE loss differential between $\\mathrm{TFM512}_{IL}$ and HAR over the out-of-sample period and apply the Diebold-Mariano test with Newey-West standard errors to that daily series. If few stocks show significance, or if a pooled panel of daily loss differentials yields p-values above 0.10, then the claimed statistical outperformance is an artifact of testing on 21 averaged numbers rather than genuine evidence of superior predictive ability.","tokens_in":54545,"feed_emoji":"📈","tokens_out":15424,"duration_ms":123499,"temperature":0.7,"pith_summary":"This paper is trying to establish that a general-purpose time-series foundation model, the TimesFM decoder-only transformer, can be turned into a superior one-day-ahead volatility forecaster by fine-tuning it incrementally on each index's own history. Tested on 21 major global equity indices between 2000 and 2021, the fine-tuned variants produce lower realized-variance forecast errors than the standard econometric benchmarks HAR, ARFIMA, CHAR, and RGARCH across most of six loss functions. The sharpest claim is statistical: under the asymmetric QLIKE loss, the 512-context pretrained and incrementally fine-tuned variants beat every econometric benchmark with Diebold-Mariano and Giacomini-White p-values below 0.10, while the log-transformed 64-context fine-tuned model is never statistically worse than any competitor. A sympathetic reader would care because the result suggests a single pretrained model, updated with lightweight retraining on recent data, could replace bespoke econometric machinery for a core risk-management task, and because the paper finds that zero-shot performance alone is only a reasonable baseline, with the adaptation step carrying the result.","feed_headline":"Fine-tuned AI beats classic models at volatility forecasting","feed_subtitle":"On 21 global indices, a foundation model fine-tuned on recent data wins under the risk-relevant QLIKE loss.","key_machinery":"The load-bearing object is TimesFM v2.0, a decoder-only transformer — a neural network that reads a sequence of past observations and directly emits the next value — pretrained on over 6 billion real-world and synthetic time-series points, which the paper uses to produce one-day-ahead point forecasts of realized variance from a look-back context of 64, 128, or 512 daily observations, with no covariates and no feature engineering. The mechanism that carries the argument is the incremental fine-tuning protocol: the pretrained model is first adapted to the first 50% of an index's data with linear probing, keeping the transformer layers frozen and updating only the core layer, then the resulting checkpoint is fine-tuned again on the next 20% segment, and again on the segment after that, so the model inherits all previous learning while tracking the recent regime; the econometric benchmarks are re-estimated on the same rolling splits. The evidence is evaluated with six loss functions — MSE, MAE, MAPE, MDA, QLIKE, and sMAPE — and the statistical claims ride on Diebold-Mariano and Giacomini-White tests together with Model Confidence Set inclusion rates. The QLIKE loss is the hinge of the strongest claim, since it penalizes under-prediction of variance asymmetrically and is the loss relevant to risk management.","core_discovery":"The paper claims that TimesFM, a decoder-only transformer pretrained on billions of diverse time-series points, can match and often beat the standard econometric models of volatility forecasting after a modest adaptation step. The adaptation is an incremental fine-tuning scheme: the pretrained model is fine-tuned on the first half of each index's realized variance series, then repeatedly re-fine-tuned, checkpoint onward, on each successive block of around 20% of the data, so the model always chases the most recent regime. Out of sample, one day ahead, on 21 global equity indices from 2000 to 2021, the fine-tuned models achieve lower average losses than HAR, CHAR, ARFIMA, and RGARCH under most of the six loss functions examined. The headline statistical result is that the 512-context variants $\\mathrm{TFM512}_{PT}$ and $\\mathrm{TFM512}_{IL}$ reject equal predictive accuracy against every econometric benchmark under the QLIKE loss at the 10% significance level in the Diebold-Mariano and Giacomini-White tests, which the authors interpret as evidence that the foundation model captures tail-sensitive volatility dynamics better than the classical parametric rivals; the log-transformed, 64-point-context fine-tuned variant $\\mathrm{TFM64}^{\\log}_{IL}$ is reported as never statistically worse than any other model on any loss function.","pith_inferences":["Because the paper's significance tests run on the cross-section of 21 average losses per stock, the defensible reading of the outperformance claim, until a proper time-series implementation is carried out, is that the fine-tuned models hold lower average losses — which the skill-score tables establish directly — rather than proven forecast superiority.","The incremental fine-tuning scheme is a moving-window retraining protocol, which suggests the gains come from regime-adaptivity rather than from pretraining per se; a testable extension would refit the simple HAR model on the identical rolling windows to see how much of the gap survives when econometric models are updated just as often.","The QLIKE advantage of the 512-context variants may reflect better calibration of the conditional variance, meaning fewer under-predictions, rather than sharper point prediction; checking whether the gain translates into out-of-sample Value-at-Risk coverage rates would separate those mechanisms.","A natural extension of the same protocol points at other volatility targets the paper's own benchmarks can distinguish, such as bipower variation, jump components, or implied volatility, since the continuous and discontinuous parts of quadratic variation are known to follow different dynamics."],"forward_implications":["A financial institution could deploy one pretrained foundation model, refreshed by incremental fine-tuning on recent realized variance, and obtain one-day-ahead forecasts that are statistically better than HAR, ARFIMA, CHAR, and RGARCH under the risk-relevant QLIKE loss at the 10% significance level.","Shorter context lengths, around 64 daily observations, are the better choice after fine-tuning: longer contexts improve the zero-shot model but appear to add noise once the model is adapted, so deployment guidance is to keep the look-back short for fine-tuned variants.","Zero-shot foundation-model forecasting alone is not enough for volatility: the paper's comparisons show that static pretraining is only a reasonable baseline, and that incremental adaptation is what produces the outperformance.","The log-transformed 64-context fine-tuned variant is the safest pick of the study, being never statistically worse than any competitor under any of the six loss functions, which the paper recommends as the preferred configuration.","Because fine-tuning updates only the core layers, keeping the model current is computationally cheap, arguing for rolling foundation models forward in production rather than periodically retraining from scratch."],"supporting_citations":[{"why":"Supplies TimesFM, the pretrained decoder-only foundation model whose zero-shot and fine-tuned variants are the objects of the entire evaluation.","marker":"[18]"},{"why":"Defines the HAR benchmark, the standard realized-variance forecaster that the TimesFM variants must beat.","marker":"[16]"},{"why":"Defines the CHAR benchmark with bipower variation, the preferred econometric reference model in the paper's boxplots and decile analyses.","marker":"[1]"},{"why":"Provides the Realized-GARCH benchmark that extends GARCH with realized measures.","marker":"[31]"},{"why":"Supplies the Diebold-Mariano test behind the claim that fine-tuned TimesFM variants statistically outperform the econometric benchmarks.","marker":"[19]"},{"why":"Supplies the Giacomini-White test, the second significance test behind the outperformance claim.","marker":"[25]"},{"why":"Provides the Model Confidence Set procedure used to identify the robust models across losses and confidence levels.","marker":"[30]"},{"why":"Motivates the choice of MSE and QLIKE as robust loss functions for comparing volatility forecasts under an imperfect proxy.","marker":"[45]"}],"fun_headline_variants":["Fine-tuned foundation model tops volatility benchmarks","Foundation model + incremental tuning wins volatility forecasts","TimesFM fine-tuned beats HAR, CHAR, ARFIMA, RGARCH","Adaptive fine-tuning makes AI volatility forecaster better","Incremental learning lifts time-series AI over econometrics"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim of statistical superiority rests on running the Diebold-Mariano and Giacomini-White tests on only 21 numbers, one average forecast loss per stock, even though the tests' theory is built on a long time series of daily loss differentials, so if that cross-sectional application is invalid the statistical outperformance collapses into a statement about average losses only.","fun_headline_variants_meta":{"raw":{"variants":["Fine-tuned foundation model tops volatility benchmarks","Foundation model + incremental tuning wins volatility forecasts","TimesFM fine-tuned beats HAR, CHAR, ARFIMA, RGARCH","Adaptive fine-tuning makes AI volatility forecaster better","Incremental learning lifts time-series AI over econometrics"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000743,"raw_usage":{"total_tokens":3348,"prompt_tokens":1016,"completion_tokens":2332,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":632,"completion_tokens_details":{"reasoning_tokens":2253}},"tokens_in":632,"tokens_out":2332,"duration_ms":15591,"temperature":1.0,"reasoning_tokens":2253,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:56:22.085501+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the comparison at the individual-stock level: for each of the 21 indices, form the daily QLIKE loss differential between $\\mathrm{TFM512}_{IL}$ and HAR over the out-of-sample period and apply the Diebold-Mariano test with Newey-West standard errors to that daily series. If few stocks show significance, or if a pooled panel of daily loss differentials yields p-values above 0.10, then the claimed statistical outperformance is an artifact of testing on 21 averaged numbers rather than genuine evidence of superior predictive ability.","supporting_citations":[{"cited_title":"A decoder- only foundation model for time-series forecasting","cited_arxiv_id":null,"evidence_quote":"Supplies TimesFM, the pretrained decoder-only foundation model whose zero-shot and fine-tuned variants are the objects of the entire evaluation."},{"cited_title":"A simple approximate long-memory model of realized volatility","cited_arxiv_id":null,"evidence_quote":"Defines the HAR benchmark, the standard realized-variance forecaster that the TimesFM variants must beat."},{"cited_title":"Andersen, Tim Bollerslev, and Francis X","cited_arxiv_id":null,"evidence_quote":"Defines the CHAR benchmark with bipower variation, the preferred econometric reference model in the paper's boxplots and decile analyses."},{"cited_title":"Realized garch: a joint model for returns and realized measures of volatility","cited_arxiv_id":null,"evidence_quote":"Provides the Realized-GARCH benchmark that extends GARCH with realized measures."},{"cited_title":"Diebold and Robert S","cited_arxiv_id":null,"evidence_quote":"Supplies the Diebold-Mariano test behind the claim that fine-tuned TimesFM variants statistically outperform the econometric benchmarks."},{"cited_title":"Tests of conditional predictive ability","cited_arxiv_id":null,"evidence_quote":"Supplies the Giacomini-White test, the second significance test behind the outperformance claim."},{"cited_title":"The model confidence set","cited_arxiv_id":null,"evidence_quote":"Provides the Model Confidence Set procedure used to identify the robust models across losses and confidence levels."},{"cited_title":"Volatility forecast comparison using imperfect volatility proxies","cited_arxiv_id":null,"evidence_quote":"Motivates the choice of MSE and QLIKE as robust loss functions for comparing volatility forecasts under an imperfect proxy."}],"review_version":1}