{"id":"3769a6d5-225f-4507-b5fe-133778d020f4","arxiv_id":"2505.01575","paper_version":3,"verdict":"REJECT","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":6,"one_line_summary":"A new encoder-only Transformer variant with autoencoder pre-training is reported to achieve high out-of-sample R2 for US stock returns, but the headline numbers come from test-set model selection.","lead":"This paper tests modified 'Transformer' AI models, including a new version called SERT, for predicting US large-cap stock returns before, during, and after COVID-19. It reports that the best model beats buy-and-hold on downside-risk-adjusted returns, but the results depend on picking the best model after seeing the test period.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Survivorship bias in stock universe selection invalidates out-of-sample R2 and Sortino claims.","rationale":"The reader's weakest assumption is exactly the load-bearing concern: the stock universe is selected using full testing-period survival, which is a look-ahead bias. This is not a minor methodological quibble—it directly affects the validity of every reported out-of-sample performance metric and the strategy evaluation. The paper's own text in Section 3 is explicit, so no inference is required. The central claim (SERT achieves highest OOS R2 and materially higher Sortino during extreme market fluctuations) would only be supported if the evaluation universe were ex-ante investable; with this filter, the numbers describe a backtest on stocks known to survive, which is not the same as a tradable strategy. I considered whether the internal abstract discrepancy or the selection of the best model across attention heads is more fundamental, but those affect interpretation or statistical significance; the survivorship filter changes what the data represent. A concrete computational test—rebuilding the universe point-in-time with delisting returns—would settle the question. The reader's REJECT verdict is appropriate and my analysis does not change it.","tokens_in":39513,"tokens_out":3449,"duration_ms":32091,"concrete_test":"Re-run the analysis with a real-time investable universe: at each month t, define the universe as stocks in the top 15% of market capitalization using only data available up to t (including CRSP delisting returns), re-estimate the same SERT and pre-trained Transformer models, and recompute OOS R2 and Sortino for the 2112 and 2212 periods. If SERT's R2 advantage over pre-trained Transformers and its Sortino advantage over buy-and-hold shrink or disappear, the central claim fails; if the qualitative advantage persists, the survivorship concern is mitigated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3 states the 420 stocks 'satisfy the conditions of having full available data in the testing period.' This is an explicit look-ahead filter: the universe is selected using information from the entire testing period (through 2022), so any stock delisted during 2013–2022 is excluded. The out-of-sample R2 and the Sortino ratios in Tables 5, 10, and 12 are therefore computed on a survivor-only universe, which is not investable in real time and removes exactly the extreme downside events (bankruptcies, distress) that the paper claims SERT hedges. The buy-and-hold benchmarks in Table 7 are computed on the same survivor universe, so the comparison is also biased; a real investor holding a point-in-time top-15% market-cap portfolio would have experienced delisting losses that are absent here. The paper's 'too-big-to-fail' and 'going concern' justification does not remove the look-ahead: a practitioner in 2013 cannot know which stocks will survive to 2022. This structural defect alone undermines the central claim of superior crisis-period R2 and downside-risk hedging. The additional issues of test-set best-model selection and the abstract/body number discrepancy (11.94%/11.47% vs 11.2%/10.91%) reinforce the concern, but the survivorship filter is the primary, load-bearing flaw.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a causal encoder-only Transformer variant called SERT, together with MLP-autoencoder-pretrained Transformers, for monthly excess-return prediction on 420 US large-cap stocks using 182 sorted-portfolio factors. It reports out-of-sample R² and trading-strategy performance across three test windows: a pre-COVID period ('1911'), a period containing the pandemic ('2112'), and a period containing one post-pandemic year ('2212'). The headline claim is that the best SERT model achieves the highest OOS R² (11.2% and 10.91%) during the extreme market periods and that its trend-following strategy achieves Sortino ratios about 47% (equal-weighted) and 28% (value-weighted) above buy-and-hold. The paper also examines attention-head counts, layer normalization first, and a softmax signal filter.","tokens_in":39816,"tokens_out":4649,"duration_ms":47889,"significance":"The architecture proposal is nontrivial and the application domain is relevant: adapting pretraining to numerical asset-pricing factors is a reasonable research direction, and the manuscript provides unusually detailed gradient derivations in Appendix B as well as a broad model comparison. If the empirical claims were supported, the paper would add useful evidence on whether simplified, causally masked encoder-only Transformers are competitive in crisis periods. However, the empirical protocol contains flaws that directly affect the central claims: a look-ahead survivorship filter in the stock universe, model selection performed on the test data, ambiguous/conflicting headline numbers, and overlapping test windows. These issues make the reported performance margins and crisis-period conclusions unsupported by the evidence as presented.","major_comments":[{"comment":"Section 3 states that the 420 stocks 'satisfy the conditions of having full available data in the testing period.' This is an explicit look-ahead survivorship filter: the universe is selected using information from the entire 2013–2022 test period, so stocks that were delisted or experienced data interruptions during that interval are excluded. All out-of-sample R² values and all strategy performance statistics, including the Sortino ratios in Tables 5, 10, and 12, are therefore computed on a survivor-only universe that is not investable in real time. This removes precisely the extreme downside outcomes that the paper claims SERT hedges. The 'too-big-to-fail' and 'going concern' arguments do not remove the information-timing problem: a practitioner forming a portfolio in 2013 cannot know which stocks will have full data through 2022. This flaw is load-bearing because the central crisis-period performance claim rests on these measurements.","section":"Section 3, Tables 5, 10, 12"},{"comment":"The 'best' model is selected using the test data itself. In Section 5.2, SERT7 is identified as the best model for '2112' and SERT5 for '2212' based on the OOS R² column of Table 5. In Section 5.3, strategy models such as SERT5, SERT2, Trans6, and Trans3 are selected as best based on their Sortino ratios in Tables 8–11 and then compared with benchmarks in Tables 12–13. Since the same test periods are used for model selection and for reporting the headline margins, the reported 'best SERT' performance is inflated by test-set selection. The claims that SERT outperforms standard encoder-only Transformers by the reported margins are not supported without a nested validation or pre-specified model-selection rule.","section":"Section 5.2, Section 5.3, Tables 5, 12, 13"},{"comment":"The abstract reports headline R² values that do not match the body of the paper. The supplied abstract text states '11.94% and 11.47%' for SERT and '11.13% and 9.72%' for pre-trained Transformers, while the paper's own abstract in the full text states '11.2% and 10.91%' and '10.38% and 9.15%'. Neither pair matches the body tables: Table 5 shows 0.1120 and 0.1091 as the best SERT values, and Table 3 shows 0.1038 and 0.0915 as the best pre-trained Transformer values. The reader cannot tell which numbers are the actual reported results, and this discrepancy directly affects the paper's principal quantitative claim.","section":"Abstract"},{"comment":"The three testing periods are nested: '2112' contains '1911' plus the pandemic, and '2212' contains '2112' plus the post-pandemic year. The OOS R² and Diebold–Mariano statistics for these periods are therefore not independent, and the paper treats them as three separate sets of evidence for the model's crisis-period advantage. The DM tests should account for the overlapping structure of the longer windows (for example, with heteroskedasticity- and autocorrelation-robust standard errors). Without this, the reported significance levels for '2112' and '2212' are overstated.","section":"Table 1, Section 5.1, Table 4"}],"minor_comments":[{"comment":"The trading rule says positions are opened 'at the point that both predicted return and actual return have a positive sign.' If the 'actual return' is the realized return of the same month for which the prediction is made, the signal would not be available at the time of trade execution; please clarify the timing and confirm the strategy is implementable in real time.","section":"Section 5.3"},{"comment":"The captions of Figure 19 and Figure 20 say 'Sign equal weighted accumulative return plots,' but the surrounding text and the tables they illustrate refer to value-weighted portfolios; the captions should be corrected.","section":"Figures 19 and 20"},{"comment":"The gradient derivations are a useful addition, but several equations contain apparent typesetting or transposition errors (e.g., Eq. B12), and the multi-head attention section is abbreviated; a careful proofread would improve reliability.","section":"Appendix B"},{"comment":"The paper does not provide code or a data-availability statement. Given that the empirical claims are extensive and the implementation details (e.g., exact hyperparameters per model) are only partially reported, a repository with code and data would be important for reproducibility.","section":"General"},{"comment":"The text says 'evaluating the Sharp ratio' where 'Sharpe ratio' is intended; there are also several typographical errors (e.g., 'gradatim') and a stray '?' in reference [17].","section":"Section 3"}],"recommendation":"reject","confidential_remarks":"The survivorship filter in Section 3 is a fundamental empirical flaw: it makes the out-of-sample R² and Sortino claims uninterpretable as evidence about real-time performance. The additional problems of test-set model selection and the mismatched abstract values reinforce the conclusion that the central claims are not supported by the reported analysis. A resubmission could potentially address these issues by rebuilding the universe point-in-time, implementing a nested validation protocol, and harmonizing the reported statistics, but that would require substantial new empirical work rather than local revisions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline first: the paper's central empirical claim—that a causal encoder-only Transformer (SERT) with MLP autoencoder pretraining gets 11% OOS R2 in crisis periods—does not survive contact with the paper's own data construction. The universe is filtered on having full data through 2022, which is textbook look-ahead survivorship bias. On top of that, the best model is selected per period using test-set R2, and the abstract and body report different numbers. So the headline result is not trustworthy as stated.\n\nWhat is actually new: SERT is a specific combination of known ingredients—causal masking for an encoder-only Transformer plus an MLP autoencoder pretraining step—and I don't know of a published test of exactly this setup on monthly US large-cap factor data. The paper also runs a broad configuration grid (number of heads, LNF vs standard, pretrained vs not) and reports DM tests and portfolio statistics. Two negative findings are genuinely informative: layer-normalization-first hurts, and adding attention heads does little. That is real, useful negative evidence for people building small Transformers for asset pricing.\n\nThe soft spots, in order of severity. First, the survivorship filter is load-bearing. The paper states that the 420 stocks 'satisfy the conditions of having full available data in the testing period.' That means any stock delisted during 2013–2022 is gone, which removes the exact downside events the paper claims to hedge. The buy-and-hold benchmarks are computed on the same survivor set, so the comparison is also biased. The 'too-big-to-fail' rationale does not fix this: a practitioner in 2013 cannot know which stocks survive to 2022. Second, choosing the best of ten configurations by OOS R2 and reporting that maximum as 'the model's' performance overstates expected performance. It is an upper bound, not an unbiased estimate. Third, the abstract says 11.94%/11.47% while the body tables show 11.2%/10.91%; that inconsistency, combined with no code or data release, makes independent verification impossible. The appendix math is standard and there is no circular derivation. Citation practice looks fine.\n\nBottom line: this paper is for readers interested in Transformer variants for return prediction, but only as a cautionary example of look-ahead bias. The flaws are correctable in principle—re-run on a point-in-time universe, pre-register model selection—and the negative results might survive a proper test. I would not desk-reject the topic, but I would send it to a referee who is explicitly asked to verify the data construction. The current version's headline claims should not be accepted.","headline":"Survivorship bias and test-set selection sink the headline R2 claims, but the paper's negative results on Transformer tweaks are worth a second look.","tokens_in":40298,"tokens_out":4042,"would_cite":false,"duration_ms":38987,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that its SERT model, an encoder-only Transformer with autoencoder pretraining and causal masks, achieves the highest out-of-sample $R^2$ in the COVID-19 period, and that its trend-following strategy beats buy-and-hold on…","keywords":["Stock pricing","Transformer","Attention mechanism","Temporal dependency","SERT","Pretrained Transformer","Factor investing","Out-of-sample forecast"],"falsifier":"Re-run the same SERT training and sign-signal backtest on a point-in-time universe: at each month from 2013 to 2022, trade every stock that was in the top market-cap bucket that month, and carry delisted stocks at their final return. If SERT's out-of-sample $R^2$ no longer exceeds the pre-trained Transformer benchmarks, or its downside-risk-adjusted advantage over buy-and-hold shrinks below the reported 47% and 28% margins, the look-ahead universe selection is the cause.","tokens_in":39328,"feed_emoji":"📈","tokens_out":9314,"duration_ms":87655,"temperature":0.7,"pith_summary":"This paper tries to establish that Transformer models, built for language, can be adapted to predict stock returns in sparse economic data, and that a specifically modified version does this better than the standard architecture exactly when markets are most turbulent. The proposed model, SERT, replaces BERT-style random masking with an MLP autoencoder pretraining step and enforces causal masks so future information cannot leak into the attention computation. Comparing ten Transformer variants and ten encoder-only variants over the pre-COVID, COVID, and one-year-post-COVID periods, the paper reports that the best SERT reaches 11.2% and 10.91% out-of-sample $R^2$ in the two turbulent periods, ahead of pre-trained Transformers at 10.38% and 9.15%. A sign-signal strategy built on the SERT predictions achieves a downside-risk-adjusted return about 47% higher than the equal-weighted buy-and-hold benchmark and 28% higher than the value-weighted benchmark during the pandemic. If these numbers hold, the paper's contribution is practical: Transformer-based return forecasts can hedge downside risk in crises rather than merely fitting calm markets.","feed_headline":"Transformer variant hits 11.2% out-of-sample R² in crashes","feed_subtitle":"An encoder-only Transformer with autoencoder pretraining beats buy-and-hold on downside risk during the pandemic.","key_machinery":"The load-bearing object is the causal, autoencoder-pretrained attention block. SERT modifies the standard Transformer's encoder self-attention by imposing an upper-triangular causal mask, and replaces BERT's random-word masking pretraining with an MLP autoencoder that maps the 182 factor inputs to 420 latent features before the Transformer body sees them. This pretraining step does double duty: it enlarges the factor dimension in the spirit of large-factor models and fills missing values, which the paper argues lets the attention mechanism capture temporal dependencies in sparse, noisy monthly return data. The causal mask is what stops the model from peeking at future months, and the pretrained dimension expansion is what the paper credits for the model's advantage during high-volatility periods.","core_discovery":"SERT is a single-directional encoder-only Transformer: it keeps the encoder block of the standard Transformer, removes the decoder and cross-attention, adds causal masks to the self-attention layer so that each monthly forecast uses only past months, and prepends an MLP autoencoder that pretrains the 182 portfolio-sorted factors, projecting them to the 420-dimensional output space and simultaneously denoising missing values. The paper's central empirical discovery is that this architecture produces the top out-of-sample model fit in the COVID-19 period (11.2% $R^2$) and the year after (10.91% $R^2$), outperforming standard Transformers, standard encoder-only Transformers, and the pre-trained Transformer models, with pairwise error-difference tests marking the gap as significant only when volatility is extreme. On the strategy side, the discovery is that the simplest trading rule, going long when predicted and realized returns share the same sign, turns the model's crisis-period forecasts into downside-risk-adjusted returns above buy-and-hold, with the advantage concentrated in the pandemic window. The paper also reports negative findings about common Transformer enhancements: layer normalization first does not help, adding attention heads helps only marginally, and a softmax signal filter erases differences between models without improving risk-adjusted performance.","pith_inferences":["Editorial inference: because the 420-stock universe is selected using full data through 2022, the crisis-period $R^2$ and downside-risk-adjusted advantage are upper bounds for a real-time strategy; a point-in-time replication that includes delisted stocks is needed before the edge can be traded.","Editorial inference: the paper's own finding that performance improves with volatility suggests the architecture should be tested on higher-volatility cross-sections such as small caps or cryptocurrencies; if the mechanism is right, the gap over benchmarks should widen there.","Editorial inference: the flat response to head count implies the model's capacity can be shrunk well below current configurations, which would make the approach feasible for larger universes where per-stock computation matters."],"forward_implications":["If the reported out-of-sample $R^2$ is real, Transformer-based factor models are most informative when volatility is high, the regime where traditional linear factor models typically degrade.","The trend-following sign-signal strategy built from SERT forecasts offers a crisis-period hedge: during the pandemic it beats the equal-weighted and value-weighted buy-and-hold benchmarks on downside-risk-adjusted return by roughly 47% and 28%.","The softmax signal filter's failure to improve risk-adjusted returns implies that the model's edge is in the sign of the forecast, not in the ranking confidence; filtering signals only makes different architectures look alike.","Since attention head count has negligible effect on fit, practitioners can use low-head-count models for speed without losing forecast accuracy in this data regime."],"supporting_citations":[{"why":"Supplies the machine-learning asset-pricing setup, factor definitions, and rolling-window evaluation conventions the paper adapts.","marker":"[1]"},{"why":"Provides the autoencoder asset-pricing rationale that motivates the MLP autoencoder pretraining module.","marker":"[2]"},{"why":"Defines the standard Transformer whose encoder, self-attention, and positional encoding SERT modifies.","marker":"[5]"},{"why":"Argues for large factor models over sparse ones, supporting the paper's pretraining dimension expansion from 182 to 420.","marker":"[13]"},{"why":"Supplies the 182 sorted-portfolio factors used as observable inputs to the models.","marker":"[52]"},{"why":"Introduces the encoder-only BERT architecture and its masking pretraining, which SERT converts to a causal single-directional form.","marker":"[57]"},{"why":"Defines the standard encoder-only Transformer baseline that SERT is compared against.","marker":"[58]"}],"fun_headline_variants":["Encoder-only Transformer tops asset pricing in COVID crash","SERT model beats buy-and-hold on downside risk in pandemic","Pre-trained Transformer for stocks: best in high volatility","Single-directional Transformer wins stock pricing during COVID","Transformer strategy hedges crash risk, beats buy-and-hold"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The stock universe used for the out-of-sample test is built from stocks that are known today to have complete data through 2022, so the test never sees stocks that were delisted or went bankrupt during the evaluation period; the reported $R^2$ and downside-risk-adjusted returns would not be the ones an investor could actually have earned.","fun_headline_variants_meta":{"raw":{"variants":["Encoder-only Transformer tops asset pricing in COVID crash","SERT model beats buy-and-hold on downside risk in pandemic","Pre-trained Transformer for stocks: best in high volatility","Single-directional Transformer wins stock pricing during COVID","Transformer strategy hedges crash risk, beats buy-and-hold"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000616,"raw_usage":{"total_tokens":2934,"prompt_tokens":1090,"completion_tokens":1844,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":706,"completion_tokens_details":{"reasoning_tokens":1765}},"tokens_in":706,"tokens_out":1844,"duration_ms":12840,"temperature":1.0,"reasoning_tokens":1765,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:15:19.696913+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same SERT training and sign-signal backtest on a point-in-time universe: at each month from 2013 to 2022, trade every stock that was in the top market-cap bucket that month, and carry delisted stocks at their final return. If SERT's out-of-sample $R^2$ no longer exceeds the pre-trained Transformer benchmarks, or its downside-risk-adjusted advantage over buy-and-hold shrinks below the reported 47% and 28% margins, the look-ahead universe selection is the cause.","supporting_citations":[{"cited_title":"National Bureau of Economic Research (2024)","cited_arxiv_id":null,"evidence_quote":"Argues for large factor models over sparse ones, supporting the paper's pretraining dimension expansion from 182 to 420."},{"cited_title":"In: Proceedings of naacL-HLT, vol","cited_arxiv_id":null,"evidence_quote":"Introduces the encoder-only BERT architecture and its masking pretraining, which SERT converts to a causal single-directional form."}],"review_version":1}