{"id":"cc885c9e-4ee3-42bd-8483-b7532495a6fd","arxiv_id":"2502.05210","paper_version":3,"verdict":"REJECT","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"high","formal_verification":"none","parameter_count":7,"one_line_summary":"The authors fit standard factor models and an LSTM to U.S. sector returns and report that the five-factor model and LSTM each look best in different sectors.","lead":"This preprint fits three standard factor models and a neural network to monthly returns of three U.S. stock sectors and claims the five-factor model fits best, with LSTM best for the high-technology sector. It is a routine empirical comparison with too little methodological detail to support the forecasting claim.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Comparing in-sample factor-model R² to test-set LSTM R² invalidates the claimed model-superiority comparison.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the factor-model R² values and the LSTM R² values are not measured under a common evaluation protocol. This is not a stylistic objection; it directly undermines the paper's central comparative claim. A full-sample in-sample fit and a 70/30 held-out test fit answer different questions, and no conclusion about which model predicts better can be drawn from their side-by-side comparison. I also note a second operational gap: even if the R² comparison were made fair, the paper's forecasting recipe requires future values of SMB, HML, RMW, and CMA, and no forecasting method for these factors is given. The text explicitly tells investors to substitute predicted factor values, but such predictions are never defined. This makes the proposed use of equations (5)–(7) non-operational. The paper has no code, no data, and no LSTM architecture or hyperparameter details, so the LSTM result cannot currently be reproduced. These issues reinforce the reader's REJECT verdict rather than changing it. I agree with the reader's assessment and recommend no change to the verdict.","tokens_in":8778,"tokens_out":2279,"duration_ms":21508,"concrete_test":"Replicate the analysis with a single chronological split: use the first 70% of monthly observations for training and the last 30% for testing. Estimate F-F3, Carhart4, and F-F5 on the training sample only, then compute out-of-sample R², RMSE, and MAE on the test sample using the same definitions as Table 8. If the test-set F-F5 R² in Hitec is still below the LSTM's 0.929, and if test-set F-F5 R² in Manuf and Other remains comparable to the LSTM results, the paper's recommendation survives. If the out-of-sample F-F5 R² in Hitec meets or exceeds 0.929, the claimed LSTM advantage is an artifact of the protocol mismatch. Also report the LSTM architecture and hyperparameters so the comparison is reproducible.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the Fama-French five-factor model has better validity and that LSTM better predicts certain industries rests on comparing the R² values in Tables 5–7 with those in Table 8. The factor-model regressions in §3.1 are fit to the full sample; the reported R², RMSE, and MAE are in-sample fit statistics. In contrast, §3.2 states that the LSTM uses a 7:3 train/test split, so the Table 8 R² is (or should be) an out-of-sample test-set statistic. These numbers are not comparable: a full-sample in-sample R² will typically overstate predictive accuracy relative to a held-out test-set R², especially for a 240-month sample with potential regime shifts. The paper then recommends using the full-sample F-F5 equations (5)–(7) for forecasting, but provides no method for forecasting SMB, HML, RMW, and CMA. The text even instructs investors to substitute predicted future factor values without saying how those predictions are obtained. Thus the conclusion that F-F5 is the best model in Manuf and Other, and that LSTM adds value only in Hitec, is not supported by the presented evidence. A matched evaluation protocol is required before any comparative claim can be assessed.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript analyzes monthly U.S. stock returns for three sectors (Manuf, Hitec, Other) over January 2004 to January 2024, fitting the Fama-French three-factor, Carhart four-factor, and Fama-French five-factor regressions and an LSTM regression model. It reports R-squared, RMSE, and MAE for each model and concludes that the Fama-French five-factor model has better validity for all three sectors, while the LSTM model better predicts returns in the Hitec sector. The paper recommends using the five-factor model for Manuf and Other, and combining LSTM for Hitec.","tokens_in":9030,"tokens_out":6488,"duration_ms":59528,"significance":"If the comparison were valid, the paper would offer evidence on sector-specific performance differences between classical factor models and an LSTM for U.S. stock returns. The topic is relevant and the use of standard factor definitions is sensible. However, the execution prevents the results from being interpretable: the factor-model and LSTM metrics are not computed under a common evaluation protocol, the LSTM is not described sufficiently for reproduction, and the proposed forecasting procedure for the factor model is incomplete. The paper does not provide code, data, or enough methodological detail to assess whether the central claims are supported.","major_comments":[{"comment":"The R-squared values in Tables 5–7 are in-sample fits from OLS regressions on the full 2004–2024 sample, whereas the LSTM in §3.2 is trained on 70% of the data and evaluated on the remaining 30%. If Table 8 reports test-set R-squared, then the two sets of R-squared values are not directly comparable, and the comparison is biased against the LSTM (or in favor of the factor models, depending on the direction of overfitting). The paper must compute factor-model R-squared, RMSE, and MAE on the same 30% hold-out, or report LSTM training-set performance, before any claim such as 'the Fama-French five-factor model has better validity' or 'LSTM better predicts Hitec' can be assessed.","section":"§3.1, Tables 5–7 vs. §3.2, Table 8"},{"comment":"The sentence instructing an investor to substitute predicted future values of SMB, HML, RMW, and CMA into the fitted regression presupposes that these factor values can be forecast. No method, model, or validation for forecasting these factors is provided, and no ex-ante availability argument is given. Without such a method, Equations (5)–(7) are in-sample fitted relationships rather than forecasting equations, so the investment-strategy recommendation in §4 is not supported.","section":"§3.1.3, after Eq. (7)"},{"comment":"The LSTM model is unspecified: the text does not report the input features, sequence length or lookback, number of layers, hidden units, activation functions, optimizer, learning rate, batch size, epochs, regularization, or whether the R-squared is computed on the training set or the test set. Without this information, the reported R-squared values, including the high Hitec value of 0.929, cannot be reproduced or audited for data leakage or overfitting. This is a load-bearing gap because the paper's only evidence for the LSTM claim is Table 8.","section":"§3.2, Table 8"},{"comment":"The conclusion that the Fama-French five-factor model is 'best' rests on small in-sample R-squared differences (e.g., Manuf: 0.901, 0.904, 0.909; Hitec: 0.864, 0.864, 0.871; Other: 0.936, 0.940, 0.946) with no statistical test of whether the increments are significant, and no adjusted R-squared or out-of-sample comparison. In the Manuf regression, RMW and CMA are insignificant (p = 0.420 and p = 0.859 in Table 4), yet Equation (5) still includes them and the model is recommended. The paper should report incremental F-tests, adjusted R-squared, and cross-validated predictive comparisons before making the superiority claim.","section":"§3.1, Tables 5–7"}],"minor_comments":[{"comment":"Equations (2)–(4) omit the intercept term even though the general regression form in Equation (1) includes β0, and Equation (4) and Equation (5) use 'MB' where 'SMB' is intended; this notation inconsistency should be corrected.","section":"§2.2 and §3.1.1"},{"comment":"The abstract contains a duplicated phrase: 'French five-factor model for the three sectors of the market' appears twice in consecutive sentences.","section":"Abstract"},{"comment":"The 'Investment Strategy' paragraph is repeated verbatim twice in the conclusion; one copy should be removed.","section":"§4"},{"comment":"The description of data preprocessing is vague: the paper states that vacant values are filled using Lagrange interpolation and outliers are eliminated 'in a similar way,' but it does not define the outlier criterion, the interpolation window, or the source and construction of the sector return series. More detail is needed for reproducibility.","section":"§2.4"},{"comment":"Several references appear unrelated to the claims they are attached to (e.g., references [12], [13], [14], and [17] concern generalized linear models, regression modeling strategies, weighted log-rank tests, and local regression, none of which is actually used in the paper). The citation list should be trimmed to relevant sources.","section":"References"}],"recommendation":"reject","confidential_remarks":"The paper appears to be a preliminary draft: the central empirical comparison is invalid because the factor-model and LSTM metrics are not produced under the same evaluation protocol, the LSTM section has no architectural or hyperparameter details, and the proposed factor-forecasting procedure is not defined. Correcting these problems would require redoing the empirical analysis and adding substantial reproducibility material, which goes beyond a routine revision. I therefore recommend rejection rather than major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a routine empirical exercise: Fama-French three-, four-, and five-factor regressions on three U.S. sector return series, plus a generic LSTM, with a claim that the five-factor model fits best and LSTM adds value in the high-tech sector. There is no new factor, no new architecture, no new evaluation protocol. What the paper does do is assemble standard data, run textbook OLS regressions, and report R², RMSE, and MAE in tables. Those tables are internally consistent for the factor models, and the authors correctly note that RMW and CMA are insignificant in the Manuf sector. That part is unremarkable but not wrong.\n\nThe central problem is that the comparison between Tables 5–7 and Table 8 is apples-to-oranges. The factor-model R² values are full-sample in-sample fits, while the LSTM R² comes from a 70/30 train/test split. Those numbers are not comparable, and the paper's conclusion that the five-factor model is best in Manuf and Other, and that LSTM is better in Hitec, is not supported by the presented evidence. A matched out-of-sample protocol is required before any comparative claim can be assessed. This is not a minor quibble; it is the load-bearing comparison of the paper.\n\nThere are also smaller but symptomatic issues. The LSTM section gives no architecture, no hyperparameters, no training details, and no code. The text instructs investors to substitute forecasted SMB, HML, RMW, and CMA values but gives no method for forecasting those factors. Equation (4) has an obvious typo, \"0.08MB\" instead of \"0.08SMB.\" The reference list contains many entries unrelated to the topic, and the abstract and conclusion contain duplicated fragments. No data or code are shipped, so nothing is independently reproducible.\n\nThis paper does not deserve a serious referee. It is a draft-quality manuscript with a fatal evaluation mismatch and missing artifacts. The only path to usefulness would be to redo the comparison under a single evaluation protocol, report the LSTM architecture, and make the data and code available. As is, it is a reject.\n\nWho might get something from it? A reader looking for a template for standard factor regressions on sector data, or a cautionary example of how not to compare in-sample and out-of-sample results. But that is not enough to recommend engaging with it as a research contribution.","headline":"Routine factor-model comparison undercut by comparing in-sample fits to out-of-sample LSTM and by vanishing LSTM implementation detail.","tokens_in":9596,"tokens_out":1295,"would_cite":false,"duration_ms":12822,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that the Fama-French five-factor model is the best linear specification for all three U.S. stock sectors studied, and that an LSTM neural network predicts high-technology sector returns better than any factor model.","keywords":["Fama-French five-factor model","LSTM stock prediction","factor models","sector stock returns","US stock market","machine learning finance","R-squared comparison"],"falsifier":"Re-estimate the five-factor model on the same 70 percent training window used for the LSTM and compute its R-squared on the 30 percent test window for the Hitec sector. If that out-of-sample R-squared reaches or exceeds the LSTM's 0.929, the paper's central claim fails. A supporting check is to run a walk-forward comparison across subperiods to see whether the LSTM advantage is stable or an artifact of the single split.","tokens_in":8532,"feed_emoji":"📈","tokens_out":4912,"duration_ms":40012,"temperature":0.7,"pith_summary":"The paper sets out to test whether three standard factor models—the Fama-French three-factor, Carhart four-factor, and Fama-French five-factor models—explain monthly returns in three U.S. stock sectors (Manuf, Hitec, Other), and whether a long short-term memory (LSTM) network adds predictive power beyond them. It claims the five-factor model is the strongest linear specification in every sector, with R-squared values of 0.909, 0.871, and 0.946 respectively. The LSTM achieves a higher R-squared than the five-factor model only in the high-technology sector (0.929 vs 0.871). The paper's conclusion is a sector-specific recommendation: rely on the five-factor model for Manuf and Other, and combine it with LSTM for Hitec. The broader point is that linear factor models and deep learning can be complementary in forecasting stock returns.","feed_headline":"LSTM beats five-factor model for high-tech stocks","feed_subtitle":"Sector-by-sector test shows deep learning wins where linear factors lag","key_machinery":"The comparison runs on two kinds of machinery: ordinary least squares regressions of excess returns on the standard factor sets (market, size, value, and for the five-factor model also profitability and investment), and an LSTM recurrent neural network with gated memory units. The factor regressions produce the estimating equations (5)-(7) and the R-squared/RMSE/MAE benchmarks in Tables 5-7. The LSTM, trained on 70 percent of the monthly data and tested on 30 percent, produces Table 8. The paper's inference about LSTM's edge in Hitec rests on reading the two sets of R-squared values side by side.","core_discovery":"The central discovery is that the Fama-French five-factor model is the most valid of the three linear factor models for all three sectors, and that an LSTM network can capture sector-specific, nonlinear return drivers that the five-factor model misses, most clearly in high technology. For Manuf and Other, the five-factor model already explains over 90 percent of return variation, so the paper argues that more complexity buys little. For Hitec, the LSTM's R-squared of 0.929 versus the five-factor model's 0.871 is presented as evidence that neural networks can improve prediction when linear factors fall short, due to LSTM's ability to model long-term dependencies and nonlinear patterns.","pith_inferences":["The R-squared comparison mixes evaluation protocols: factor regressions use the full sample while LSTM uses a holdout, so the reported gap in Hitec may overstate LSTM's true advantage; an equal-protocol test is a natural extension.","The paper recommends substituting future SMB, HML, RMW, and CMA values into its equations, but provides no way to forecast these factors; adding a factor-forecasting module would make the recommendation actionable.","A walk-forward validation that retrains the LSTM and re-estimates factor betas on rolling windows would test whether the Hitec edge persists out of sample or reflects memorization of the 2004-2024 period."],"forward_implications":["For the Manuf and Other sectors, the five-factor model should remain the default tool, since it explains more than 90 percent of return variation.","In the Hitec sector, predictions from an LSTM can exceed the five-factor model's accuracy, supporting a hybrid approach there.","The insignificance of RMW and CMA in the Manuf sector suggests that a leaner model may suffice for manufacturing stocks.","Combining factor models with LSTM offers a practical path that balances interpretability with nonlinear predictive power.","The results imply that the value of deep learning in return forecasting is sector-dependent, not universal."],"supporting_citations":[{"why":"Supplies the Fama-French five-factor model, the central linear benchmark the paper tests across all three sectors.","marker":"[1]"},{"why":"Provides evidence that LSTM networks are effective for financial market prediction, grounding the paper's choice of LSTM as the nonlinear comparator.","marker":"[9]"},{"why":"Demonstrates stock market price prediction with LSTM RNN, supporting the claim that LSTM handles time-series dependencies.","marker":"[6]"},{"why":"Shows a neural-network approach (BP-GA) applied to index volatility and returns, serving as inspiration for using neural networks after traditional models.","marker":"[4]"},{"why":"Confirms the effectiveness of modern generative models in return prediction, motivating the exploration of LSTM in this paper.","marker":"[5]"}],"fun_headline_variants":["LSTM tops five-factor model for high-tech stock returns","LSTM beats five-factor model in tech sector only","Deep learning wins for high-tech, linear model for rest","LSTM captures nonlinear returns in high-tech stocks","For tech stocks, LSTM beats five-factor model"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the full-sample R-squared of the factor models and the holdout R-squared of the LSTM are directly comparable; if they are not, the claimed LSTM superiority in high technology is not established.","fun_headline_variants_meta":{"raw":{"variants":["LSTM tops five-factor model for high-tech stock returns","LSTM beats five-factor model in tech sector only","Deep learning wins for high-tech, linear model for rest","LSTM captures nonlinear returns in high-tech stocks","For tech stocks, LSTM beats five-factor model"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000368,"raw_usage":{"total_tokens":1926,"prompt_tokens":848,"completion_tokens":1078,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":464,"completion_tokens_details":{"reasoning_tokens":1001}},"tokens_in":464,"tokens_out":1078,"duration_ms":8294,"temperature":1.0,"reasoning_tokens":1001,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T14:31:11.521304+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-estimate the five-factor model on the same 70 percent training window used for the LSTM and compute its R-squared on the 30 percent test window for the Hitec sector. If that out-of-sample R-squared reaches or exceeds the LSTM's 0.929, the paper's central claim fails. A supporting check is to run a walk-forward comparison across subperiods to see whether the LSTM advantage is stable or an artifact of the single split.","supporting_citations":[{"cited_title":"F., & French, K","cited_arxiv_id":null,"evidence_quote":"Supplies the Fama-French five-factor model, the central linear benchmark the paper tests across all three sectors."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides evidence that LSTM networks are effective for financial market prediction, grounding the paper's choice of LSTM as the nonlinear comparator."},{"cited_title":"S., & Tiwari, V","cited_arxiv_id":null,"evidence_quote":"Demonstrates stock market price prediction with LSTM RNN, supporting the claim that LSTM handles time-series dependencies."},{"cited_title":"A Consolidated Volatility Prediction with Back Propagation Neural Network and Genetic Algorithm","cited_arxiv_id":"2412.07223","evidence_quote":"Shows a neural-network approach (BP-GA) applied to index volatility and returns, serving as inspiration for using neural networks after traditional models."},{"cited_title":"Developing Cryptocurrency Trading Strategy Based on Autoencoder-CNN-GANs Algorithms","cited_arxiv_id":"2412.18202","evidence_quote":"Confirms the effectiveness of modern generative models in return prediction, motivating the exploration of LSTM in this paper."}],"review_version":1}