{"id":"94f0063d-221a-4d08-a4ed-f840b35ee405","arxiv_id":"2607.05291","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Zero-shot time series foundation models largely fail to beat econometric benchmarks for realized volatility forecasting, with only TTM achieving a narrow, calibration-driven edge.","lead":"This paper compares nine zero-shot time series foundation models against eight econometric benchmarks for forecasting realized volatility across 50 assets. It finds that most foundation models do not beat a simple Log-HAR benchmark, with only one small model (TTM) edging it out narrowly.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"The 1.3–1.8% average QLIKE advantage of TTM over Log-HAR lacks a panel-level significance test; per-asset MCS results show both models co-included for most assets, suggesting the cross-sectional average may not be statistically distinguishable from parity.","rationale":"The reader correctly identified that the ranking depends on QLIKE and the loss-ratio aggregation, and the paper does provide some robustness checks (geometric mean, MSE, averaged target). However, the more load-bearing issue is the absence of a formal panel-level significance test for the specific TTM-vs-Log-HAR comparison. The paper reports DM win rates pooled across all 16 opponents and MCS inclusion rates that show both models co-included for most assets, but never isolates the pairwise TTM-vs-Log-HAR significance count. This matters because the 1.3–1.8% improvement is small enough that it could easily be within cross-sectional sampling variation, and the MCS evidence suggests the two models are statistically indistinguishable on most individual assets. The paper is otherwise thorough — the MZ recalibration, GR tests, COVID subsample, context-length sensitivity, and combination analysis are all well-executed. The claim that most TSFMs do not beat Log-HAR is well-supported. But the specific claim that TTM 'beats' Log-HAR 'at every horizon' should be conditioned on a panel-level test that the paper does not provide. If such a test shows significance, the ACCEPT verdict stands; if not, the claim should be softened to 'marginally improves on average, statistically indistinguishable on most individual assets.' The paper's own framing ('thin edge,' 'not best on every asset') is already cautious, so this is a refinement rather than a rejection.","tokens_in":39946,"tokens_out":4782,"duration_ms":94658,"concrete_test":"For each of the 50 assets, compute the time-series mean QLIKE loss differential d_i = QLIKE_TTM,i − QLIKE_LogHAR,i and its Newey–West standard error (h−1 lags). Then: (a) report how many of the 50 per-asset DM tests (TTM vs. Log-HAR only, not pooled across 16 opponents) reject at 5% in TTM's favor; (b) compute the cross-sectional t-statistic for H0: mean(d_i) = 0. If fewer than ~15 of 50 per-asset tests are significant, or if the cross-sectional t-statistic is below ~2.0, replace 'beats Log-HAR' with 'matches or marginally improves on Log-HAR on average.'","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim — that TTM is the only TSFM to 'beat' Log-HAR — rests on average QLIKE loss ratios of 0.982, 0.986, and 0.987 (Tab. 6). These are cross-sectional means of per-asset ratios, but the paper does not report a formal test of whether this average is significantly below 1.0. The statistical evidence presented does not directly address this: (1) The DM win rates in Tab. 8 are pooled across all 16 opponents (800 tests), not isolated to TTM vs. Log-HAR, so we cannot tell how many of the 50 specific TTM-vs-Log-HAR comparisons are significant. (2) The MCS includes TTM for 98% of assets and Log-HAR for 86% at h=1 (Tab. 8), meaning both are frequently retained together — i.e., statistically indistinguishable on most individual assets. (3) The GR fluctuation tests (Fig. 3) show TTM's rolling DM statistic against Log-HAR hovering near zero for much of the sample at h=1, rarely breaching the 5% band. Taken together, the 1.3–1.8% edge could reflect cross-sectional sampling noise rather than a genuine panel-level advantage. The paper acknowledges the edge is 'thin' but still uses the language of 'beats' and 'beats at every horizon,' which implies statistical rather than merely point-estimate superiority. Without a cross-sectional t-test on the per-asset loss differentials (or at minimum the isolated TTM-vs-Log-HAR DM win count), the headline overstates what the statistics show.","agreement_with_reader":"partial"},"referee_report":{"model":"glm-5.2","summary":"This paper conducts a systematic zero-shot comparison of nine time series foundation models (TSFMs) against eight econometric benchmarks for forecasting realized volatility, using the VOLARE dataset across 50 assets (equities, FX, futures) and three horizons (h=1, 5, 22). The evaluation employs a walk-forward scheme with a 1,000-day rolling window, Diebold-Mariano (DM) tests, Model Confidence Set (MCS) analysis, Mincer-Zarnowitz (MZ) regressions, and Giacomini-Rossi (GR) fluctuation tests. The central finding is that most TSFMs do not improve on a well-specified Log-HAR benchmark under equal-weighted loss ratios. Only Tiny Time Mixers (TTM) beats Log-HAR at every horizon, by a narrow margin of 1.3-1.8% in average QLIKE loss ratios. An MZ recalibration shows that TTM's short-horizon edge is largely a calibration effect, while a genuine informational advantage remains at the monthly horizon. A simple equal-weight combination of TTM and Log-HAR enters the MCS for 98-100% of assets.","tokens_in":40225,"tokens_out":1383,"duration_ms":85596,"significance":"The paper addresses a timely and well-defined question: whether pretrained TSFMs can compete with established econometric models for realized volatility forecasting. The empirical design is thorough, the first multi-model, multi-asset evaluation of its kind for this target. Strengths include the use of formal forecast comparison tests (DM, MCS, GR), a symmetric MZ recalibration applied to all models, an honest assessment of contamination risk, and publicly available replication code. The finding that architecture choice matters more than the foundation-vs-econometric distinction is a durable and practically useful contribution. The MZ decomposition separating calibration from informational content is a particularly insightful analytical choice.","major_comments":[{"comment":"§5.1, Tab. 6 and §6, Tab. 8: The paper's headline claim that TTM 'beats' Log-HAR at every horizon rests on average QLIKE loss ratios of 0.982, 0.986, and 0.987. However, the paper does not report a formal panel-level (cross-sectional) significance test for whether these averages are significantly below 1.0. The MCS results in Tab. 8 show that both TTM (98% inclusion at h=1) and Log-HAR (86% inclusion at h=1) are frequently co-included, suggesting they are statistically indistinguishable on most individual assets. The DM win rates in Tab. 8 are pooled across 16 opponents (800 tests), not isolated to TTM vs. Log-HAR. Without either a cross-sectional t-test on the per-asset loss differentials or the isolated TTM-vs-Log-HAR DM win count, the language of 'beats' and 'beats at every horizon' overstates what the statistics demonstrate. The authors should either add a panel-level test or soften ","section":null},{"comment":"§4.4, Eq. (9) and surrounding text: The QLIKE loss is evaluated on the variance scale by squaring the volatility forecast (dRV_t = sigma_hat^2_t), while the conditional mean is extracted as the point forecast. The authors acknowledge a Jensen term (footnote 6) but state that the conditional mean 'nonetheless remains preferable to the conditional median.' This is a non-trivial approximation. Since QLIKE is the primary evaluation metric and the basis for the central claim, a brief sensitivity check or at least a more formal justification that this Jensen gap does not systematically bias the TSFM-vs-econometric comparison (e.g., do TSFMs with heavier-tailed predictive distributions suffer larger Jensen gaps?) would strengthen the analysis.","section":null}],"minor_comments":[{"comment":"§4.1.1, Eq. (1): The HAR model is defined on the volatility scale (sigma), while the augmented variants (HAR-J, HAR-RS, HARQ) in Eqs. (2)-(4) are defined on the variance scale (RV). The text explains this transition, but a brief clarifying note in the equation captions or at the start of §4.1.1 would improve readability.","section":null},{"comment":"Tab. 1: The equity panel reports cross-sectional averages but does not specify whether the kurtosis and skewness are also averaged. Given the extreme kurtosis values in the FX and futures panels (e.g., 3,550 for Crude Oil), clarifying the aggregation method for the equity panel would be helpful.","section":null},{"comment":"Fig. 3: The y-axis labels are difficult to read. The critical-value bands (±2.80) are mentioned in the caption but not clearly marked in some panels. Improving the annotation would aid interpretation.","section":null},{"comment":"§7, Tab. 11: The MZ-corrected QLIKE values differ from the original values in Tab. 5 due to the 252-day warm-up window. While footnote 11 explains this, adding a column in Tab. 11 showing the number of observations used for the corrected evaluation would make the comparison more transparent.","section":null},{"comment":"§4.2.2: The context window for TTM and Moirai-MoE is 512 days, while other TSFMs use 1,000 days. Tab. 14 shows TTM performs best at 512. While this is an architectural constraint, a brief discussion of whether the 512-day window could be an advantage (e.g., less overfitting to stale regimes) rather than a limitation would add nuance.","section":null},{"comment":"Typo in §5.1: 'The level HAR variants HAR-RS and HARQ produce inflated QLIKE on futures' — 'level' should likely be 'level-scale' or simply 'The HAR variants'.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The stress-test concern about the lack of a panel-level significance test for TTM vs. Log-HAR is well-founded and is the primary issue to address. The paper is otherwise a solid empirical contribution with honest reporting. The authors' own acknowledgment that the edge is 'thin' suggests they are aware of this fragility, making the fix (adding a test or softening language) straightforward. The paper fits well within the journal's scope."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for a careful and constructive report. Both major comments identify genuine gaps in the statistical evidence supporting our headline claim. We address each below and indicate the revisions we will make.","responses":[{"response":"The referee is correct. Our headline claim that TTM 'beats' Log-HAR at every horizon is supported by the average loss ratios in Tab. 6 and the MCS inclusion rates in Tab. 8, but we do not report a formal panel-level test for whether the per-asset loss ratios are jointly significantly below one, nor do we isolate the TTM-vs-Log-HAR pairwise DM comparison. The DM win rates in Tab. 8 are pooled across 16 opponents (800 tests per model), so they do not directly answer whether TTM significantly outperforms Log-HAR on a meaningful share of individual assets. The MCS co-inclusion rates the referee cites (TTM 98%, Log-HAR 86% at h=1) do suggest that the two models are frequently statistically indistinguishable at the asset level, which is consistent with the narrow margin (1.3–1.8%) we report. We agree that the language 'beats' and 'beats at every horizon' overstates what the existing statistics demonstrate without the isolated comparison. We will make two changes in the revision. First, we will add the isolated TTM-vs-Log-HAR pairwise DM results: for each of the 50 assets, we will report the DM test statistic and p-value for the QLIKE loss differential between TTM and Log-HAR, along with the count of assets where TTM achieves significantly lower QLIKE at the 5% level (and the reverse count). Second, we will add a cross-sectional t-test on the per-asset loss-ratio differentials (i.e., testing whether the mean of the 50 per-asset QLIKE ratios is significantly below 1.0), with a Newey–West variance to account for any cross-sectional dependence. We expect these tests to confirm that the edge is statistically detectable but narrow—consistent with the MCS co-inclusion pattern—and we will adjust the language accordingly. Specifically, we will replace 'beats Log-HAR at every horizon''","revision_made":"no","referee_comment":"§5.1, Tab. 6 and §6, Tab. 8: The paper's headline claim that TTM 'beats' Log-HAR at every horizon rests on average QLIKE loss ratios of 0.982, 0.986, and 0.987. However, the paper does not report a formal panel-level (cross-sectional) significance test for whether these averages are significantly below 1.0. The MCS results show that both TTM and Log-HAR are frequently co-included, suggesting they are statistically indistinguishable on most individual assets. The DM win rates are pooled across 16 opponents, not isolated to TTM vs. Log-HAR. Without either a cross-sectional t-test on the per-asset loss differentials or the isolated TTM-vs-Log-HAR DM win count, the language of 'beats' and 'beats at every horizon' overstates what the statistics demonstrate."},{"response":"The referee raises a legitimate concern. The QLIKE-optimal forecast is E[RV_{t+h} | F_t], the conditional mean of the variance, but we construct the variance forecast by squaring the conditional mean of the volatility, sigma_hat^2, which differs from E[sigma^2] by a Jensen term equal to the conditional variance of the volatility forecast. This term is model-specific: models whose predictive distributions have higher conditional variance (heavier tails) will have larger Jensen gaps, so the approximation could in principle bias the comparison if TSFMs and econometric models differ systematically in their predictive dispersion. Our defense in the manuscript is limited to two points: (i) the conditional mean is still preferable to the conditional median, which carries an additional bias, and (ii) we apply the same point-forecast construction to all 17 models. These points show the treatment is symmetric but do not establish that the Jensen gap is negligible or uniform across models. We agree this needs to be addressed more rigorously. In the revision we will add a brief sensitivity analysis. Specifically, for the subset of TSFMs that expose full predictive distributions (Lag-Llama, Toto, Sundial, Moirai-MoE, and the quantile-output models), we will compute E[sigma^2] directly from the predictive distribution—by integrating the quantile function or computing the second moment of the sampled trajectories—and compare the resulting QLIKE against the squared-mean construction. For models that output only a point forecast (TTM) or a mean plus quantiles without a full distribution (TimesFM 2.5, Chronos-Bolt), the squared-mean construction is the only available option, so the sensitivity check will cover the models where the Jensen gap is potentially largest. We will also add a few","revision_made":"no","referee_comment":"§4.4, Eq. (9) and surrounding text: The QLIKE loss is evaluated on the variance scale by squaring the volatility forecast (dRV_t = sigma_hat^2_t), while the conditional mean is extracted as the point forecast. The authors acknowledge a Jensen term (footnote 6) but state that the conditional mean 'nonetheless remains preferable to the conditional median.' This is a non-trivial approximation. Since QLIKE is the primary evaluation metric and the basis for the central claim, a brief sensitivity check or at least a more formal justification that this Jensen gap does not systematically bias the TSFM-vs-econometric comparison would strengthen the analysis."},{"response":"We thank the referee for the recommendation and for the constructive tone of the report. The suggested revisions are well-targeted and will strengthen the statistical foundations of the paper without changing its central message. We will implement both sets of changes in the revised manuscript.","revision_made":"no","referee_comment":"REFEREE RECOMMENDATION: minor_revision"}],"tokens_in":39745,"tokens_out":2154,"duration_ms":60400,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"The paper's main contribution is clear: it is the first to systematically compare nine zero-shot time series foundation models against eight econometric benchmarks for realized volatility forecasting, across 50 assets and three horizons, with formal statistical testing. The central finding — that most TSFMs do not beat a well-specified Log-HAR, and only TTM does so consistently but by a thin 1.3–1.8% margin — is useful and honestly reported. The MZ recalibration decomposition, showing that TTM's short-horizon edge is largely calibration while a genuine informational gain remains only at the monthly horizon, is a genuinely interesting analytical step that elevates the paper above a pure horse-race exercise. The walk-forward design, DM tests, MCS analysis, GR fluctuation tests, and robustness checks (pre/post-COVID, context length, averaged target) are thorough and appropriate. Replication code is available on GitHub, which is good practice. The equal-weight combination result (TTM + Log-HAR entering the MCS for 98–100% of assets) is a practical and well-motivated finding. The stress-test concern about the lack of a panel-level significance test on the per-asset loss ratios is valid but moderate in severity. The paper does report DM win rates pooled across all 16 opponents (800 tests), and the MCS results show TTM and Log-HAR frequently co-included, which the paper honestly acknowledges. The language of 'beats' is slightly stronger than what the isolated TTM-vs-Log-HAR evidence supports, but the paper does qualify this with 'thin' and 'narrow margin' throughout. A cross-sectional t-test on the per-asset loss differentials, or at minimum the isolated TTM-vs-Log-HAR DM win count, would strengthen the headline claim. This is a fixable gap, not a structural problem. The reader's assessment is fair. The significance score of 5.0 is about right — the paper clarifies an open question for practitioners and does so carefully, but the findings are confirmatory rather than paradigm-shifting. The soundness score of 7.0 is reasonable given the missing panel-level test. The paper is for financial econometricians and practitioners evaluating whether to deploy TSFMs for volatility forecasting. It deserves a serious referee who can push the author to add the isolated pairwise significance evidence and soften the 'beats' language where the statistics do not directly support it.","headline":"First systematic multi-TSFM vs. econometric benchmark comparison for realized volatility; finds only TTM edges out Log-HAR, narrowly and partly via calibration","tokens_in":40948,"tokens_out":575,"would_cite":true,"duration_ms":70392,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"Only one foundation model beats Log-HAR for volatility forecasting","keywords":[],"falsifier":"Re-run the full evaluation under MSE loss ratios (instead of QLIKE) with the same equal-weight per-asset aggregation. If TTM no longer beats Log-HAR at every horizon under MSE, the claim that it is the sole consistent winner is metric-specific. Alternatively, if a different aggregation scheme (e.g., weighted by liquidity or trading volume) produces a different sole winner, the ranking is an artifact of the equal-weight choice.","tokens_in":40205,"feed_emoji":"📉","tokens_out":1426,"duration_ms":106574,"temperature":0.7,"pith_summary":"This paper asks whether pretrained time series foundation models (TSFMs) — large neural networks trained on diverse time series from weather, energy, retail, and other domains — can outperform established econometric benchmarks for forecasting realized volatility, the standard measure of ex-post price variation built from intraday returns. Using the VOLARE dataset, the author evaluates nine zero-shot TSFMs spanning eight architectures against eight econometric specifications, including the Heterogeneous Autoregressive (HAR) family, across 50 assets in equities, foreign exchange, and futures, and three forecast horizons (one day, one week, one month). The central finding is that foundation models do not deliver a uniform gain. Pooled loss averages initially appear to favor several TSFMs, but this advantage is concentrated in a few high-volatility outlier assets. When each asset is weighted equally through per-asset loss ratios relative to a well-specified Log-HAR benchmark, only one model — Tiny Time Mixers (TTM), the smallest in the evaluation at under one million parameters — beats Log-HAR at every horizon, and only by a margin of roughly 1.3 to 1.8 percent. The other eight TSFMs do not improve on Log-HAR on the typical asset. A Mincer-Zarnowitz recalibration, which strips out level and scale bias from every forecast, then shows that TTM's short-horizon edge is largely a calibration effect — its forecasts already sit at the right level and scale — that several other foundation models also exhibit once the same correction is applied. A genuine informational advantage survives only at the monthly horizon. Because the edge is thin and TTM is not best on every asset, a simple equal-weight average of TTM and Log-HAR matches the best single model and enters the Model Confidence Set for 98 to 100 percent of assets, more often than either component alone. The paper's most durable conclusion is that performance varies so widely across TSFM architectures that choosing the right architecture matters more than the broader choice between foundation and econometric models.","feed_headline":"Only one AI foundation model edges out econometric volatility benchmark","feed_subtitle":"Nine pretrained time series models tested on 50 assets: just TTM beats Log-HAR, and only by 1-2%. A simple average of both is the safest bet","key_machinery":"The paper's argument turns on three formal tools. First, per-asset QLIKE loss ratios relative to Log-HAR, averaged equally across assets, replace pooled-mean losses that are distorted by outlier series. Second, the Model Confidence Set (MCS) of Hansen et al. (2011), using a Tmax statistic with a moving-block bootstrap, identifies the subset of models that cannot be statistically distinguished from the best at each asset. Third, recursive Mincer-Zarnowitz regressions decompose forecast accuracy into calibration (level and scale alignment) and information (genuine predictive content about future volatility dynamics), applied symmetrically to all 17 models. The TSFMs themselves are evaluated in","core_discovery":"The paper's central object is the QLIKE loss ratio: each model's per-asset forecast loss divided by Log-HAR's loss on the same asset, then averaged equally across all 50 assets. This aggregation, unlike pooled-mean losses that are dominated by a few high-volatility series, reveals that only TTM consistently beats the Log-HAR benchmark, and narrowly. The Mincer-Zarnowitz recalibration then decomposes that edge into a calibration component (correct forecast level and scale, shared by several models) and an information component (genuine predictive content about volatility dynamics), with the informational advantage surviving only at the 22-day horizon. The mechanism carrying the argument is tw","pith_inferences":["If TTM's advantage at the monthly horizon reflects genuine long-memory capture rather than calibration, fine-tuning TTM specifically on realized volatility series could widen the gap with Log-HAR at all horizons, an extension the author flags but does not test.","The wide dispersion across TSFM architectures suggests that the pretraining corpus composition (what types of time series the model has seen) may be more predictive of financial-volatility performance than model size or architectural family, which could be tested by pretraining identical architectures on different domain mixtures.","The result that a sub-1M-parameter model outperforms models with 100-200M parameters on volatility forecasting may indicate that realized volatility's statistical properties (positive, mean-reverting, long-memory) are sufficiently close to common pretraining domains that a lightweight architecture can capture them, while larger models overfit or misallocate capacity.","The MZ recalibration result — that several TSFMs carry hidden predictive information masked by scale bias — implies that a systematic post-hoc calibration layer applied to any zero-shot TSFM could substantially narrow the gap with tuned benchmarks, a practical pipeline improvement the paper does not explicitly propose."],"forward_implications":["Practitioners should not treat foundation models as a uniform class that automatically improves on econometric benchmarks; architecture selection within the TSFM family is the first-order decision.","A simple equal-weight average of the best foundation model (TTM) and the best econometric benchmark (Log-HAR) provides a robust, no-tuning forecast that enters the Model Confidence Set for nearly all assets, sidestepping the model-selection problem.","Mincer-Zarnowitz recalibration can recover predictive signal hidden by level and scale bias in underperforming foundation models, suggesting that raw zero-shot outputs from several TSFMs carry useful information that affine correction exposes.","The finding that the smallest model (under 1M parameters) is the only consistent winner raises questions about whether pretraining scale and parameter count are the right objectives for time series foundation models, or whether architecture and training-data composition matter more."],"fun_headline_variants":["Foundation models barely beat econometric volatility benchmarks","Architecture choice beats model class for volatility forecasting","Only TTM edges Log-HAR for volatility forecasting—and narrowly","Ensemble of TTM and Log-HAR is safest volatility forecast","Zero-shot TSFMs offer thin, uneven edge over HAR volatility models"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The conclusion that TTM is the sole foundation model to consistently beat Log-HAR depends on the choice of the QLIKE loss function and the equal-weighted per-asset loss-ratio aggregation. Under MSE, TTM's dominance is less pronounced, and a different loss function or aggregation scheme could reorder the rankings.","fun_headline_variants_meta":{"raw":{"variants":["Foundation models barely beat econometric volatility benchmarks","Architecture choice beats model class for volatility forecasting","Only TTM edges Log-HAR for volatility forecasting—and narrowly","Ensemble of TTM and Log-HAR is safest volatility forecast","Zero-shot TSFMs offer thin, uneven edge over HAR volatility models","Recalibration erases most foundation-model volatility edge","Econometric benchmarks stay competitive against AI volatility forecasts","Foundation vs econometric matters less than picking the right architecture","TTM wins on average but no model dominates every asset","Informational edge for AI volatility models survives only at monthly horizon"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":828,"prompt_tokens":665,"completion_tokens":163,"prompt_tokens_details":null},"tokens_in":665,"tokens_out":163,"duration_ms":3209,"temperature":1.0,"reasoning_tokens":45,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-07T19:24:11.162649+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"Re-run the full evaluation under MSE loss ratios (instead of QLIKE) with the same equal-weight per-asset aggregation. If TTM no longer beats Log-HAR at every horizon under MSE, the claim that it is the sole consistent winner is metric-specific. Alternatively, if a different aggregation scheme (e.g., weighted by liquidity or trading volume) produces a different sole winner, the ranking is an artifact of the equal-weight choice.","supporting_citations":[],"review_version":1}