{"id":"9e64ff82-d525-438b-8a16-4616ad9cd9b2","arxiv_id":"2411.13792","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"A multiscale extension of Markowitz optimization, averaging covariance matrices across time scales, is claimed to improve out-of-sample Sharpe and drawdown versus single-scale Markowitz in a 2019-2024 US ETF backtest.","lead":"This paper proposes a portfolio optimization method that controls risk not only at daily frequency but across multiple time scales, using a Hurst-exponent scaling law. The authors report that their multiscale version beats traditional Markowitz on Sharpe ratio and drawdown in a 2019-2024 backtest on US sector and factor ETFs.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The superiority claim rests on a single 5-year backtest whose multiscale estimator includes a monthly covariance block with only ~6 independent observations; without error bars or scale-set sensitivity, the reported Sharpe gap may be a finite-sample artifact.","rationale":"Read in good faith: the paper's central claim is empirical, and the estimator plus backtest are the evidence for it. The weakest point is the finite-sample reliability of the multiscale covariance estimator. The text defines Σ_MS in Section 4.2 and Section 7 but never reports effective sample sizes per scale, the exact scale grid, or the rebalancing schedule. Simple arithmetic from the stated 125-day window shows the monthly block has about 6 observations, making that covariance matrix singular. The normalization Σ(Δt)/Δt is not motivated; under the paper's scaling law it is only scale-independent if 2H=1. Because the backtest summary lacks any measure of uncertainty, the reported superiority cannot be distinguished from estimation noise. A single robustness run (dropping the monthly block) would directly test whether the headline depends on the noisiest input. This is a fixable empirical gap, not a logical contradiction, so the reader's conditional verdict remains appropriate.","tokens_in":8335,"tokens_out":10426,"duration_ms":93816,"concrete_test":"Re-run the Section 7 backtest after removing the lowest-frequency scale (the monthly block, Δt≈21, ~6 non-overlapping observations) from the average in Σ_MS = ⟨Σ(Δt)/Δt⟩, keeping all other parameters unchanged. If the Multiscale Markowitz Sharpe ratio in Tables 1–2 drops to or below the Traditional Markowitz values (0.35 sector, 0.43 factor), the claimed superiority is an artifact of a statistically unreliable low-frequency covariance block; if it persists, the low-frequency-noise concern is not the driver.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim (Introduction; Tables 1–2) is that the multiscale estimator Σ_MS = ⟨Σ_ij(Δt)/Δt⟩_Δt, computed on a rolling 125-day window (Section 7), beats daily-covariance Markowitz. The load-bearing condition is that this estimator is a reliable risk measure and that the reported Sharpe/Sortino gaps are not finite-sample artifacts. That condition is not met. With 125 trading days, the weekly block (Δt=5) has 25 non-overlapping observations and a monthly block (Δt≈21) has about 6. For 11 sector ETFs or 9 factors, the monthly covariance matrix is rank-deficient and extremely noisy; no shrinkage is described. The normalization Σ(Δt)/Δt is also scale-dependent under the paper's own scaling law: if σ(Δt)∝(Δt)^H, then Σ(Δt)/Δt∝(Δt)^{2H-1}, so for H≠1/2 the average depends on the arbitrary choice of scale grid and weights. The paper provides no error bars, no bootstrap, no subperiod checks, and no sensitivity to the scale set. Thus the headline 0.53 vs 0.35 Sharpe difference may be driven by noisy low-frequency blocks or by the ad-hoc normalization rather than by a genuine multiscale effect.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a 'multiscale Markowitz' framework in which portfolio risk is measured with a covariance estimator averaged over multiple return horizons, Sigma^MS_ij = <Sigma_ij(Delta t)/Delta t>, motivated by anomalous scaling sigma(Delta t) proportional to (Delta t)^H. It presents a sensitivity analysis for minimum-variance weights and an out-of-sample backtest on 11 SPDR sector ETFs and 9 factor ETFs over 2019-2024, reporting higher Sharpe and Sortino ratios and lower drawdowns than daily-covariance Markowitz (Tables 1-2). The abstract and introduction also promise a toy example and a target-Hurst optimization, but the delivered backtest uses only the averaged-covariance option (Section 3), and no toy example appears in the text.","tokens_in":8607,"tokens_out":8807,"duration_ms":81256,"significance":"If the empirical claim were robust, the multiscale covariance estimator would be a simple and potentially useful alternative to single-scale covariance estimation for long-only portfolio optimization. The paper's out-of-sample design has strengths: it compares against equal-weight and traditional Markowitz in two asset universes, and it does not fit parameters to the Sharpe outcome, so the central comparison is not circular. However, the main contribution is empirical, and the evidence is not statistically established: there are no confidence intervals, no sensitivity analysis for the scale grid, and the scale normalization has a load-bearing arbitrariness. The reported Sharpe gaps (0.53 vs 0.35 and 0.53 vs 0.43) are therefore not yet convincing evidence of superiority.","major_comments":[{"comment":"The paper's central claim, stated in the introduction as 'we evidence on US sector index tracking ETFs the superiority of multifrequency optimization over traditional Markowitz,' rests on a single five-year backtest with no measure of statistical uncertainty. The estimator Sigma^MS_ij = <Sigma_ij(Delta t)/Delta t> is computed on a 125-day lookback with non-overlapping lower-frequency blocks; at Delta t = 5 this gives about 25 weekly observations and at Delta t ≈ 21 about 6 monthly observations. For 11 sector ETFs or 9 factors, the monthly covariance block is rank-deficient, and no shrinkage or regularization is described. The Sharpe differences in Tables 1–2 therefore have no confidence intervals, bootstrap, or subperiod checks, so finite-sample noise cannot be ruled out. Section 7 also states that transaction costs are 'assumed to be negligible' without reporting turnover or rebalancing frequency, which can bias a Sharpe comparison in favor of whichever method trades more. Please add error bars or bootstrap/subperiod splits, report the exact scale grid and the number of independent observations per block, and quantify turnover and transaction-cost sensitivity.","section":"Section 7, Tables 1–2"},{"comment":"The definition Sigma^MS_ij = <Sigma_ij(Delta t)/Delta t> is not scale-invariant under the paper's own scaling law. If sigma(Delta t) is proportional to (Delta t)^H, then Sigma_ij(Delta t) is proportional to (Delta t)^{2H}, so Sigma_ij(Delta t)/Delta t is proportional to (Delta t)^{2H-1}; unless H = 1/2 for all pairs, the average over scales depends on the choice of scale grid and on the weights in the average. The manuscript does not specify the set of scales used, the weights, or whether the grid is varied in the backtest. Consequently, the reported improvement in Tables 1–2 could be driven by the arbitrary normalization rather than by a genuine multiscale effect. Please report the scale grid and demonstrate robustness to alternative grids and weights, for example by excluding the monthly block or using equal weights across scales.","section":"Section 4.2"},{"comment":"The abstract promises both a target-Hurst formulation and a toy example ('We illustrate this concept with a toy example'), and the introduction repeats the toy-example plan. However, Section 3 explicitly drops the target-Hurst implementation ('We consider only the last case since the first does not sufficiently constrain the portfolio'), and no toy example or numerical illustration of the multifractal optimization problem appears in Sections 5–7. The delivered backtest uses only option 3, the averaged covariance matrix. Either add the promised toy example and target-Hurst experiments, or revise the abstract and introduction to state that the delivered contribution is the averaged-covariance estimator and its backtest.","section":"Abstract and Sections 1, 3, 5"}],"minor_comments":[{"comment":"The notation for H is overloaded: the abstract uses H in sigma(Delta t) proportional to (Delta t)^H, while Section 2.1 redefines H as beta/alpha via the fractional PDE. Please align the notation throughout.","section":"Section 2.1"},{"comment":"The Lagrangian solution w = Sigma^{-1}1/S does not enforce the non-negative weight constraint stated in Section 4.2, so the sensitivity conclusions do not directly apply to the long-only optimization used in the backtest.","section":"Section 6"},{"comment":"The statement 'Empirically from studies of the Epps effect, we find that H_rho ≈ 0.3' is given without a citation or derivation; please provide a reference or a supporting calculation.","section":"Section 5.2"},{"comment":"The file 'epps1.png' is listed at the end of the manuscript but is never referenced in the text; add a figure with a caption or remove the dangling file.","section":"End of manuscript"},{"comment":"There are scattered typographical errors, e.g., 'constra int' in the abstract and 'T V olatility' in Section 1.2; a careful proofread is needed.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The manuscript currently resembles an extended abstract: the advertised toy example is missing, and the empirical evidence is not yet at the standard expected for a journal article. The core idea is defensible, and the backtest is simple and transparent, so a focused revision adding bootstrap/confidence intervals, scale-grid sensitivity, turnover reporting, and the promised toy example could make the paper publishable. I do not recommend rejection at this stage."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a working-paper-level contribution with a genuinely testable idea and a backtest that doesn't yet support the headline. The new bit is the estimator Sigma_MS = <Sigma_ij(dt)/dt> and the standardized Hurst H = beta/alpha; the out-of-sample sector/factor ETF results are new. The paper does a decent job laying out the motivation and the special elliptical case, and it cites the right literature (Mandelbrot, Epps, Gatheral, Berman-Hochberg, Bianchi et al.).\n\nThe soft spots are real and several. The abstract promises a toy example and a target-Hurst optimization; Section 3 explicitly drops the target-Hurst approach ('We consider only the last case') and no toy example appears anywhere. That mismatch alone would need fixing. More importantly, the central backtest is one five-year window, 2019-2024, with no error bars, no bootstrap, no subperiod checks, no transaction costs, and no comparison against the multiscale baselines the paper cites. The stress-test math is right: with a 125-day lookback, the monthly block has about six non-overlapping observations; for 11 assets that covariance is rank-deficient and extremely noisy without shrinkage. The normalization Sigma(dt)/dt is also not scale-invariant under the paper's own sigma ~ dt^H law — the average depends on the arbitrary choice of scale grid unless H = 1/2. So the 0.53 vs 0.35 Sharpe gap could easily be finite-sample noise or a normalization artifact. The sensitivity analysis has a derivative slip (they state dsigma/dH when the chain rule needs dsigma^2/dH), though the sign conclusion happens to be correct.\n\nNone of this is fatal to the idea. The estimator is simple, computable, and the question — does averaging covariances across scales improve out-of-sample risk control — is worth answering properly. A clean version with code, error bars, multiple periods, and a direct comparison to Berman-Hochberg or wavelet methods could be a solid paper.\n\nRecommendation: send to peer review. A serious referee can force the robustness analysis that the current draft lacks. I would not cite it in its present form, but I'd bring it to a reading group to discuss what a proper test of this estimator would look like.","headline":"A clean but under-supported empirical idea: the multiscale covariance estimator is worth testing rigorously, but the paper's own implementation doesn't yet establish the claimed edge.","tokens_in":9135,"tokens_out":2697,"would_cite":false,"duration_ms":25656,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91G10"],"pacs":[],"model":"deepseek-v4-flash","headline":"A scale-averaged covariance matrix gives higher out-of-sample Sharpe and Sortino ratios and lower drawdowns than daily Markowitz, according to the paper's US sector and factor ETF backtests.","keywords":["multiscale portfolio optimization","Markowitz","Hurst exponent","multifractality","Epps effect","minimum variance portfolio","Sharpe ratio","drawdown"],"falsifier":"Rerun the same 2019-2024 backtest on sector ETF returns that have been randomized within each asset to destroy time-series scaling while preserving cross-sectional correlations and volatility; if the multiscale portfolio still beats daily Markowitz, the apparent advantage is not due to genuine scale-dependent risk.","tokens_in":8088,"feed_emoji":"📈","tokens_out":7707,"duration_ms":63629,"temperature":0.7,"pith_summary":"This paper proposes replacing the single daily covariance matrix in Markowitz optimization with a multiscale covariance estimator averaged over time scales. The authors argue that real return series show scale-dependent variance and correlations, so single-frequency risk targets miss how risk compounds across holding periods. In backtests on S&P 500 sector ETFs and factor ETFs from 2019 to 2024, the multiscale minimum-variance portfolio achieves higher Sharpe and Sortino ratios and lower maximum drawdowns than traditional daily Markowitz. The central claim is that tailoring risk at multiple frequencies, optionally through a target Hurst exponent, improves out-of-sample portfolio performance.","feed_headline":"Scale-averaged risk beats daily Markowitz in backtests","feed_subtitle":"Multiscale Markowitz raised Sharpe to 0.53 from 0.35 on US sector ETFs.","key_machinery":"The central object is the multiscale covariance estimator $\\Sigma^{MS}_{ij} = \\langle \\Sigma_{ij}(\\Delta t)/\\Delta t \\rangle_{\\Delta t}$, the average over a range of time scales $\\Delta t$ of covariances estimated from returns aggregated to those scales. The optimization uses this estimator in place of the daily covariance matrix, with the usual minimum-variance or maximum-Sharpe objective; a target Hurst exponent $H_{\\rm target}$ can be imposed so that portfolio variance scales as $\\sigma^2_{\\rm target}(\\Delta t) \\propto (\\Delta t)^{H_{\\rm target}}$ across frequencies. This estimator carries the argument because it is what makes the optimization sensitive to how risk builds up from daily to monthly horizons, and the paper derives sensitivity results showing that weights fall as volatility, Hurst exponent, or correlation rise.","core_discovery":"The paper's central claim is that optimizing a portfolio with the multiscale covariance matrix $\\Sigma^{MS}_{ij} = \\langle \\Sigma_{ij}(\\Delta t)/\\Delta t \\rangle_{\\Delta t}$ outperforms traditional Markowitz based on daily variances and covariances. This estimator averages covariances computed at several return aggregation intervals, capturing how volatility and correlation scale with time; the paper connects this scaling to a standardized Hurst exponent $H := \\beta/\\alpha$ from a fractional diffusion equation. On 11 S&P 500 sector ETFs and 9 factor ETFs, the multiscale minimum-variance portfolio reports Sharpe ratios around 0.53 in both universes, versus 0.35 and 0.43 for traditional Markowitz, with Sortino ratios and maximum drawdowns also improving. The authors interpret this as evidence that scale-dependent risk structure, including the Epps effect and rough volatility, is economically exploitable by choosing a target Hurst exponent.","pith_inferences":["The reported Sharpe improvement is measured over a single five-year window with two major drawdowns; testing over multiple decades or many rolling sub-periods would show how much of the gain is regime-specific.","The low-frequency part of the estimator relies on a small number of non-overlapping sub-samples inside a 125-day window; if those terms are noisy, a different choice of scale set could weaken or strengthen the advantage.","A direct test would use simulated multifractal data with known Hurst spectra to verify that the multiscale estimator recovers the true scale-dependent risk and that the optimization gains are not an artifact of the averaging scheme.","The framework could be extended to cross-asset portfolios with bonds and commodities, where the Epps effect and rough volatility are more pronounced, to see whether the performance gap widens."],"forward_implications":["If the claim is correct, investors can choose a target Hurst exponent to express preferences for short-horizon versus long-horizon risk, rather than fixing only daily variance.","Multiscale optimization should reduce exposure during crashes, since the scale-averaged covariance penalizes the low-frequency volatility spikes that daily Markowitz misses.","The same estimator applies to factor rotation and other long-only allocation problems, not just sector rotation.","The results imply that single-scale Markowitz misallocates to assets with Hurst exponents far from one half, such as illiquid or rough-volatility assets.","The multiscale covariance can be replaced by an L1-modified version to preserve convergence under fat-tailed return distributions."],"supporting_citations":[{"why":"Supplies the scaling law $\\sigma(\\Delta t) \\propto (\\Delta t)^H$ from fractional Brownian motion that motivates measuring variance at multiple scales.","marker":"[1]"},{"why":"Documents rough volatility with Hurst exponents below one half, the phenomenon the authors say causes single-scale Markowitz to over-allocate.","marker":"[2]"},{"why":"Provides evidence that asset returns are multifractal, justifying a scale-dependent variance structure in the optimization.","marker":"[3]"},{"why":"Prior use of multiscale analysis in financial performance evaluation that the authors build on for scale-averaged portfolio construction.","marker":"[10]"},{"why":"Establishes the Epps effect, the increase of correlations with time scale, which the scale-averaged covariance is designed to capture.","marker":"[15]"}],"fun_headline_variants":["Multiscale Markowitz lifts Sharpe from 0.35 to 0.53","Targeting Hurst exponent beats daily covariance in ETFs","Scale-aware risk model outshines classic Markowitz in tests","Sharpe 0.53 vs 0.35: multiscale portfolio optimization wins"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the scale-averaged covariance matrix estimated from a rolling 125-day window, with lower frequencies built from non-overlapping sub-samples, remains a reliable estimate of true risk at those horizons; if those low-frequency estimates are noisy or the normalization across scales is arbitrary, the reported improvement could be an artifact.","fun_headline_variants_meta":{"raw":{"variants":["Multiscale Markowitz lifts Sharpe from 0.35 to 0.53","Targeting Hurst exponent beats daily covariance in ETFs","Scale-aware risk model outshines classic Markowitz in tests","Sharpe 0.53 vs 0.35: multiscale portfolio optimization wins"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00022,"raw_usage":{"total_tokens":1441,"prompt_tokens":934,"completion_tokens":507,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":550,"completion_tokens_details":{"reasoning_tokens":427}},"tokens_in":550,"tokens_out":507,"duration_ms":4911,"temperature":1.0,"reasoning_tokens":427,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:52:51.601388+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the same 2019-2024 backtest on sector ETF returns that have been randomized within each asset to destroy time-series scaling while preserving cross-sectional correlations and volatility; if the multiscale portfolio still beats daily Markowitz, the apparent advantage is not due to genuine scale-dependent risk.","supporting_citations":[{"cited_title":"Mandelbrot and John W","cited_arxiv_id":null,"evidence_quote":"Supplies the scaling law $\\sigma(\\Delta t) \\propto (\\Delta t)^H$ from fractional Brownian motion that motivates measuring variance at multiple scales."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents rough volatility with Hurst exponents below one half, the phenomenon the authors say causes single-scale Markowitz to over-allocate."},{"cited_title":"Calvet and Adlai J","cited_arxiv_id":null,"evidence_quote":"Provides evidence that asset returns are multifractal, justifying a scale-dependent variance structure in the optimization."},{"cited_title":"Bianchi, Michael E","cited_arxiv_id":null,"evidence_quote":"Prior use of multiscale analysis in financial performance evaluation that the authors build on for scale-averaged portfolio construction."},{"cited_title":"Multiscale Markowitz","cited_arxiv_id":"2411.13792","evidence_quote":"Establishes the Epps effect, the increase of correlations with time scale, which the scale-averaged covariance is designed to capture."}],"review_version":1}