{"id":"c9e54116-240a-4994-8bf6-9565c1522318","arxiv_id":"2411.19649","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":8,"one_line_summary":"Transformer-based covariance and semi-covariance forecasts are claimed to improve ETF portfolio returns, but the supporting backtest is short, leaky, and unreproducible.","lead":"This paper applies Transformer-based models to predict covariance and semi-covariance matrices for ETF portfolios, claiming better risk-adjusted returns than a sample-based baseline. The evidence is a one-month backtest with undisclosed parameters and possible look-ahead bias, so the headline claim is not yet supported.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's central claim rests on a one-month backtest whose asset-selection rule may use returns from inside the test window; no cutoff date is given, so the reported outperformance is not a valid out-of-sample result.","rationale":"The most load-bearing condition for the central claim is that the reported performance difference is generated without using future information. Section 4.1 selects assets by three-year expected returns but never states the cutoff date; the two-year data span means the ranking must come from undisclosed external history, so it is impossible to verify that the test month (February 12 to March 12, 2024) was excluded. If the ranking includes that month, the entire comparison is contaminated. This concern is not merely hypothetical: the paper's own Section 5.2 concedes that gains may be due to concentration or the selected period, and the numerical record is internally inconsistent (Section 5.4 vs Table 2), so the results cannot be independently checked. The Transformer/semi-covariance idea has some merit, and the PSD post-processing step is a reasonable engineering choice, but a valid empirical demonstration would require a disclosed selection protocol, a longer out-of-sample period, and multiple rebalancing windows. These missing pieces mean the central claim is unsupported, sustaining the reader's REJECT verdict.","tokens_in":15847,"tokens_out":6917,"duration_ms":54958,"concrete_test":"Ask the authors to disclose the exact cutoff date used for the three-year expected-return ranking in Section 4.1 and the date the ETF pool in Table 3 was fixed. Then re-run the selection with the ranking window ending February 11, 2024 (the day before the test period) and compare the resulting pool to Table 3. If any ETF is replaced, re-run the full backtest on the clean pool; if no Transformer variant then beats the sample-covariance baseline over the same month, the reported gains are an artifact of selection lookahead.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that Transformer-based covariance and semi-covariance predictions 'significantly enhance' portfolio performance cannot be evaluated because the empirical protocol in Section 4.1 contains a potential lookahead bias in the asset-selection step. The paper selects, within each asset class, the ETF with the highest expected return over the past three years, but it never states the date at which that three-year lookback ends. The data window is stated as March 21, 2022 to March 21, 2024 (with a typo in the following sentence reading 'March 12'), the training set ends February 12, 2024, and the test set is February 12 to March 12, 2024. If the three-year ranking is computed as of any date on or after February 12, 2024, the test month is inside the selection window, so assets are chosen with knowledge of their returns during the test period. Because the observed data span is only two years, the three-year ranking must rely on external data that is not described, so the cutoff cannot be inferred from the text. The resulting ETF pool in Table 3 is therefore not demonstrably free of test-period information. Since the reported outperformance of all Transformer models over the sample baseline is based on a single month of returns, even a small selection bias could produce the results. The authors themselves note in Section 5.2 that the return gains 'may be partly attributed to the greater concentration of the portfolio ... or the selected period,' and no alternative explanation is ruled out. Internal inconsistencies in the reported numbers further prevent the central claim from being checked: Section 5.4 gives Informer's covariance-return as 3.31% while Table 2 reports 1.17%, and the Transformer Sortino with semi-covariance is reported as 8.25 in the text versus 8.38 in Table 2.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes an end-to-end pipeline in which the lower-triangular entries of rolling covariance and semi-covariance matrices of an ETF pool are compressed into vectors, fed into Transformer variants (vanilla Transformer, Autoformer, Informer, Reformer), and reconstructed as forecast matrices, with regularization and eigenvalue clipping used to enforce positive semi-definiteness. The forecasts are then used to construct minimum-variance portfolios. The paper reports MSE improvements in Table 1 and return/Sortino improvements in Table 2 relative to a sample-covariance baseline, and concludes that Transformer-based dynamic risk-matrix forecasts, especially with the semi-covariance matrix, significantly improve portfolio performance.","tokens_in":16137,"tokens_out":4662,"duration_ms":39687,"significance":"Forecasting a semi-covariance matrix with attention-based models and using it directly in minimum-variance optimization is a reasonable research direction, and the paper is careful to motivate the downside-risk objective and to compare against a standard baseline. The MSE table (Table 1) provides a concrete quantitative target, and the use of the Sortino ratio is appropriate given the downside-risk focus. However, the empirical evidence is not sufficient to support the abstract's strong claims: the test period is a single month, the ETF selection rule may overlap the test window, and key rolling-window parameters are withheld. As it stands, the contribution is an architectural proposal with an illustrative backtest rather than a validated result.","major_comments":[{"comment":"The asset-selection rule chooses, within each asset class, the ETF with the highest expected return over the past three years, but the text never states the cutoff date for that three-year lookback. The observed data are stated to span March 21, 2022 to March 21, 2024, and the test set is February 12 to March 12, 2024. If the ranking is computed as of any date on or after February 12, 2024, the test month is inside the selection window, so the ETF pool in Table 3 may be selected using returns from the test period. Because the stated data span is only two years, the three-year history must come from an unstated external source, so the cutoff cannot be inferred. This potential lookahead bias undermines the out-of-sample interpretation of Table 2 and Figures 2 and 3.","section":"Section 4.1 (Data Description and Experimental Setup)"},{"comment":"The entire performance comparison rests on a single one-month test window (February 12 to March 12, 2024). No standard errors, confidence intervals, or significance tests are reported for the return or Sortino differences in Table 2, and the MSE values in Table 1 are given as averages over five runs without dispersion. The abstract's claim that the predictions 'significantly enhance portfolio performance' is therefore not supported by the evidence presented; the observed differences could easily be within noise for one month and a small pool of assets.","section":"Section 5.1 and Table 2"},{"comment":"The paper states that the precise rolling-window parameters, including window size and rebalancing frequency, 'are not disclosed here due to confidentiality restrictions.' Without these parameters, the experiment cannot be reproduced, the number of effective rebalancing periods is unknown, and it is unclear whether the MSE and return results are based on overlapping or independent predictions. This lack of disclosure is a load-bearing gap for any empirical claim in the paper.","section":"Section 4.1 (rolling window)"},{"comment":"The Informer results are internally inconsistent. Table 2 reports Return using Covariance = 1.17% and Sortino using Covariance = 3.31, but Section 5.4 states that Informer's 'return increase from 3,31% to 5.69%, and its Sortino ratio improves from 3.31% to 10.05.' The text appears to conflate the return and the Sortino ratio and contradicts the table, making the numerical results unreliable for at least one model.","section":"Section 5.4 and Table 2"},{"comment":"The post-processing step clips negative eigenvalues to zero (Section 3.5) to enforce positive semi-definiteness, while the regularized loss in Section 3.4 penalizes negative eigenvalues only softly. For the semi-covariance matrix, which is not positive semi-definite by construction (Section 2.4), this projection can materially alter the risk measure before portfolio optimization. The paper does not quantify the distance between the projected and unprojected matrices or test whether the reported portfolio improvements survive without the projection.","section":"Sections 3.4-3.5"}],"minor_comments":[{"comment":"The data window contains a typo: the text first states data span from March 21, 2022 to March 21, 2024, but the next sentence gives the training set as March 12, 2022 to February 12, 2024; the training-set start date should presumably be March 21, 2022.","section":"Section 4.1"},{"comment":"The bullet labeled 'Actual Result' is unclear; it is not a model and appears to be a leftover label rather than a description of a baseline.","section":"Section 4.2"},{"comment":"The text uses a European comma decimal ('3,31%') inconsistently with the rest of the paper, and the bullet about Informer mixes return and Sortino units.","section":"Section 5.4"},{"comment":"The Data Availability Statement points to an 'EarningsCall Dataset' repository, which does not match the ETF data used in this study.","section":"Data Availability Statement"},{"comment":"Some sentences are incomplete or run-on, for example 'To further investigate the model's behavior.' appears without a continuation.","section":"Section 5.2"}],"recommendation":"reject","confidential_remarks":"The paper is explicitly a working version, and the empirical section is not yet at journal standard. The most serious issue is the potential lookahead bias in the ETF selection rule combined with the single-month test window; even if the authors clarify the cutoff date, the central claim of 'significant' outperformance would require a much longer out-of-sample evaluation. The withheld rolling-window parameters also conflict with reproducibility expectations. I would encourage the authors to resubmit after a substantial revision that discloses all experimental parameters, uses a longer test period with multiple rebalancing windows, and reports dispersion or significance measures."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper does about what it says: it flattens the lower triangle of covariance and semi-covariance matrices, feeds them to Autoformer, Informer, and Reformer, reconstructs a PSD matrix, and optimizes a minimum-variance ETF portfolio. That pipeline is reasonable and the target—semi-covariance forecasting with Transformers—is genuinely less explored than LSTM-based covariance work. The MSE improvements over the sample baseline are small (0.00071 to 0.00064 for covariance, 0.00019 to 0.00015 for semi-covariance) but consistent across all models, and the Sortino ratio is the right metric for a downside-risk-focused objective. The authors also honestly note in Section 5.2 that the return gains may come from portfolio concentration or the selected period, not from the model. Give credit for that.\n\nThe empirical core, however, cannot support the abstract's claims. The test set is a single month, February 12 to March 12, 2024. There are no error bars, no significance tests, and no multiple rebalancing periods. The asset-selection rule in Section 4.1 chooses, within each asset class, the ETF with the highest expected return over the past three years, but the paper never states the date at which that lookback ends. The observed data span is only two years, so the three-year ranking must rely on external data, and if the ranking window includes the test month, the portfolio is selected with knowledge of the test period. The paper does not rule this out. It also withholds the rolling window size and rebalancing frequency for 'confidentiality,' which is not acceptable in a paper whose only evidence is a backtest.\n\nThe internal inconsistencies are concrete: Section 5.4 gives Informer's covariance return as 3.31% while Table 2 reports 1.17%, and the Transformer Sortino with semi-covariance appears as 8.25 in the text versus 8.38 in Table 2. The data availability statement links to an unrelated EarningsCall dataset repository, and no code is provided. These are load-bearing flaws, not cosmetic ones. The paper is a work in progress and reads like one.\n\nWho gets value from this? Someone thinking about sequence models for covariance forecasting might read it for the architectural choices, but not for the empirical results. I would not send this to peer review in its current form: the central claim is not reproducible from the information given. If the authors extend the backtest to multiple years, state the selection cutoff, disclose all parameters, fix the number inconsistencies, and provide code or a real data link, the idea could warrant a serious referee. Right now it is a desk reject.","headline":"A plausible covariance-forecasting pipeline undermined by a one-month backtest, an undisclosed selection cutoff that may leak the test window, and internally inconsistent reported numbers.","tokens_in":16737,"tokens_out":2324,"would_cite":false,"duration_ms":21260,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91G10","91G80","68T07"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that Transformer forecasts of covariance and semi-covariance matrices outperform the historical sample as inputs to minimum-variance ETF portfolios, with semi-covariance generally improving downside-risk-adjusted…","keywords":["Financial Forecasting","Deep Learning","Transformer","Covariance","Semi-Covariance","Portfolio Optimization","ETF","Sortino ratio"],"falsifier":"A concrete check is to re-run the experiment with the ETF pool fixed using only data up to February 11, 2024 (the day before the test set starts), or with a randomly selected pool, and to extend the test to several non-overlapping months; if the MSE and Sortino improvements for the Transformer and semi-covariance approaches vanish under this out-of-sample selection, the reported edge is a look-ahead artifact rather than a forecasting gain.","tokens_in":15597,"feed_emoji":"📉","tokens_out":6985,"duration_ms":53445,"temperature":0.7,"pith_summary":"The paper claims that Transformer-based models—Autoformer, Informer, Reformer, and the vanilla Transformer—can forecast the full covariance and semi-covariance matrices of a diversified ETF pool more accurately than the sample historical method, and that plugging these forecasts into minimum-variance optimization improves portfolio performance. The central selling point is the semi-covariance matrix, which measures only downside co-movement and better matches risk-averse investors' concern with losses; the paper reports that portfolios optimized on predicted semi-covariance generally beat those optimized on predicted covariance in return and Sortino ratio. A sympathetic reader would care because this is a concrete template for replacing static risk estimates with a data-driven, attention-based forecaster while keeping the optimization layer unchanged. The paper presents this as a working version with a single one-month test window, so the reported magnitudes are provisional.","feed_headline":"Downside-risk AI forecasts beat plain covariance in ETF tests","feed_subtitle":"Transformer models forecasting semi-covariance lifted ETF returns and Sortino ratios in the paper's backtest.","key_machinery":"The central machinery is a compression–Transformer–reconstruction pipeline. The unique entries of each day's covariance or semi-covariance matrix are flattened into a vector, fed as a time series into an attention-based sequence model (Autoformer, Informer, Reformer, or vanilla Transformer), and the forecast vector is reshaped back into a matrix. Symmetry and positive semi-definiteness are enforced by a regularization loss (penalizing asymmetry and negative eigenvalues) and by a post-hoc eigenvalue-clipping projection to the nearest PSD matrix. The resulting matrix is then substituted directly into the closed-form minimum-variance weight formula $w = \\frac{\\Sigma^{-1}\\mathbf{1}}{\\mathbf{1}^{\\top}\\Sigma^{-1}\\mathbf{1}}$, so the only change relative to the baseline is the source of $\\Sigma$.","core_discovery":"On the paper's own terms, the central discovery is that the attention mechanism can learn time-varying cross-asset dependence directly from the compressed triangular part of a covariance matrix, and that the resulting forecasts are better inputs to minimum-variance optimization than the trailing historical matrix. The paper further claims that the semi-covariance matrix, because it isolates downside co-movement, is the more suitable risk input for risk-averse investors: across its tests, semi-covariance-based portfolios mostly produced higher returns and higher Sortino ratios than covariance-based portfolios, with the clearest gains for the vanilla Transformer and Informer. The reported numbers support the direction of the claim for prediction accuracy—all Transformer variants lowered MSE for both matrices—while the performance edge varies by model, and the paper reads the overall pattern as evidence that downside-focused dynamic forecasts adapt better to volatile conditions.","pith_inferences":["The reported test is a single month with undisclosed rolling-window settings, so the specific return and Sortino numbers are likely optimistic; a multi-year out-of-sample backtest would be needed to confirm the edge is stable.","Because the ETF selection rule looks at three-year returns ending after the test month, the fairest reading is that this paper demonstrates a promising methodology, not a validated investment edge.","The mixed results for Autoformer and Reformer—where covariance actually produced higher returns than semi-covariance—suggest the downside-risk advantage may be model-specific rather than a universal property of semi-covariance optimization.","The same compression–transformer template could be applied to other non-PSD risk measures, such as expected-shortfall contributions or higher-moment co-skewness matrices."],"forward_implications":["Asset managers could adopt the pipeline as a drop-in replacement for the covariance input in minimum-variance optimization, without changing the optimizer.","If the semi-covariance effect holds, risk-averse portfolios can be built to target downside co-movement directly instead of total volatility.","The compression–reconstruction trick makes Transformer forecasting of full covariance matrices feasible even when the number of assets makes the raw matrix high-dimensional.","The PSD-enforcement scheme offers a practical solution to the non-positive-definiteness problem that has historically blocked semi-covariance optimization."],"supporting_citations":[{"why":"Establishes the mean-variance portfolio framework and the portfolio variance formula $w^T\\Sigma w$ that the optimization uses.","marker":"(Markowitz, 1952)"},{"why":"Defines the semi-covariance matrix as the downside-risk measure substituted for the covariance matrix.","marker":"(Estrada, 2002)"},{"why":"Introduces the self-attention and multi-head attention mechanisms that form the Transformer backbones.","marker":"(Vaswani et al., 2017b)"},{"why":"Supplies the Autoformer model with trend-seasonality decomposition, the paper's best-performing variant.","marker":"(Wu et al., 2021)"},{"why":"Supplies the Reformer model using locality-sensitive hashing attention.","marker":"(Kitaev et al., 2020)"},{"why":"Supplies the Informer model with Prob-Sparse attention and generative-style decoding.","marker":"(Zhou et al., 2020)"},{"why":"Provides the shrinkage and positive-semi-definiteness context motivating the regularized loss and eigenvalue-clipping projection.","marker":"(Ledoit and Wolf, 2004)"}],"fun_headline_variants":["Transformer forecasting lifts ETF portfolios via downside-risk semi-covariance","AI downside-risk forecasts beat plain covariance for ETF optimization","Semi-covariance AI models outdo plain covariance in ETF backtests","ETF portfolios get a boost from Transformer-predicted semi-covariance"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the ETF selection rule—choosing, within each asset class, the funds with the highest three-year historical returns—does not leak information about the test month, which falls inside that same three-year window; if the top pick was already a winner during the test month, the backtest advantage could come from selection rather than from Transformer forecasting.","fun_headline_variants_meta":{"raw":{"variants":["Transformer forecasting lifts ETF portfolios via downside-risk semi-covariance","AI downside-risk forecasts beat plain covariance for ETF optimization","Semi-covariance AI models outdo plain covariance in ETF backtests","ETF portfolios get a boost from Transformer-predicted semi-covariance"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001198,"raw_usage":{"total_tokens":4952,"prompt_tokens":971,"completion_tokens":3981,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":587,"completion_tokens_details":{"reasoning_tokens":3907}},"tokens_in":587,"tokens_out":3981,"duration_ms":25068,"temperature":1.0,"reasoning_tokens":3907,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T05:59:01.787280+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check is to re-run the experiment with the ETF pool fixed using only data up to February 11, 2024 (the day before the test set starts), or with a randomly selected pool, and to extend the test to several non-overlapping months; if the MSE and Sortino improvements for the Transformer and semi-covariance approaches vanish under this out-of-sample selection, the reported edge is a look-ahead artifact rather than a forecasting gain.","supporting_citations":[{"cited_title":": The semi-covariance matrix: A more intuitive alternative to the covariance matrix in portfolio optimization","cited_arxiv_id":null,"evidence_quote":"Defines the semi-covariance matrix as the downside-risk measure substituted for the covariance matrix."},{"cited_title":", Kaiser , L","cited_arxiv_id":null,"evidence_quote":"Supplies the Reformer model using locality-sensitive hashing attention."}],"review_version":1}