{"id":"4108559a-6890-47dc-8f3e-da9f3d355e8d","arxiv_id":"1908.08168","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Two basic machine learning trading strategies earned reliable intra-day profits before 2009 but not after, which the authors interpret as evidence of rising market efficiency.","lead":"This paper runs simple neural network and logistic regression trading strategies on US equities from 2003 to 2017 and finds they earned steady intra-day profits until 2008, then went flat. It proposes that the profitability of such algorithms is a measure of market efficiency and links the decline to the rise of high-frequency trading.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Zero-cost execution is the load-bearing premise; early-period returns (4.6-4.9 bps/day) may be smaller than realistic intra-day round-trip costs, so the profitability evidence for weak-form inefficiency before 2009 is not yet established.","rationale":"The paper's central claim is that two simple ML strategies earn positive intra-day returns before 2009 and not after, and this is used as an objective measure of relative weak-form market efficiency. The load-bearing condition is that the strategies are actually profitable pre-2009 in a tradable sense. The paper's own numbers (Table 1: 4.6 and 4.9 bps/day, cumulative ~94-102% over about six years) are gross returns from a backtest that makes no mention of transaction costs. In intra-day trading, the bid-ask spread alone is a first-order cost. A 5 bps round-trip cost is conservative for large caps in the early 2000s; with 500-stock daily rebalancing, market impact adds more. Thus the gross return is of the same order as plausible costs, so the profitability evidence is not robust. This directly threatens the central claim, not a peripheral detail.\n\nThe reader's verdict CONDITIONAL is appropriate: the paper's methodology is otherwise careful (rolling retraining, validation-based hyperparameter selection, random classifier control, no look-ahead), but the central quantitative result needs to be re-evaluated net of trading frictions. We considered the post hoc break date as an alternative concern; it affects the significance of the pre/post difference, but the existence of profitability is prior. If costs eliminate early-period profits, the break-date issue is moot. Hence we select transaction costs as the single most load-bearing concern.\n\nThe proposed concrete test is straightforward and uses the same TAQ data, so it should be feasible for the authors. If the net returns remain positive, the paper's claim is substantially strengthened; if not, the inference to market efficiency cannot be sustained. Given that the reader already asked for transaction-cost modeling, our read agrees with the reader's weakest_assumption; no change in verdict is needed.","tokens_in":8586,"tokens_out":5208,"duration_ms":52300,"concrete_test":"Re-run the early-period backtest (2003-2008) with a realistic execution model: fill buys at the ask plus a $0.005/share commission, sells at the bid minus commission, using NBBO quotes from TAQ at the entry minute and at the close; add a market-impact term linear in the fraction of the stock's daily volume traded. If the net daily return is not significantly above zero, the profitability claim fails. Alternatively, compute the break-even round-trip cost per trade (the cost that makes early-period mean daily return zero) and compare it with the average quoted spread for the same stocks and dates.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that intra-day prices were predictable and markets weak-form inefficient before 2009 because two ML strategies were 'profitable'—rests on the backtest's implicit assumption in the Evaluation section that orders fill at recorded one-minute close prices with zero transaction costs, slippage, and market impact. The reported early-period average daily returns are 4.6 bps (Neural) and 4.9 bps (Logistic) (Table 1). The strategy is a daily balanced long/short portfolio over the 500 most-traded stocks; entry occurs one minute after the prediction minute and exit at the market close, so each position incurs at least one round trip per day. For US equities in 2003-2008, quoted bid-ask spreads on large caps were commonly several cents, or 5-20 bps round trip, plus commissions and impact; spreads were wider in exactly the early period where the strategy shows positive gross returns. If realistic costs are only 5 bps round trip, the early-period net return is near zero or negative; if they are 10 bps, the strategy loses money. The paper neither estimates costs nor provides a break-even analysis. Without this, the profitability evidence, and hence the inference to relative market efficiency, does not survive contact with actual trading. The post hoc choice of the Sep/Oct 2008 break (Experimental Results) is a separate concern about inference on timing, but it is secondary: if net-of-cost returns are non-positive in both periods, there is no profitability to attribute to changing efficiency.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes using the profitability of two simple machine learning trading strategies—a neural network and a logistic regression—as an objective measure of relative weak-form market efficiency. Using NYSE TAQ data from 2003 to 2017 and a monthly roll-forward training/validation/test design on the 500 most actively traded US stocks, the authors report average daily gross returns of 4.6–4.9 basis points in the period before October 2008 and near-zero or negative returns afterward, while a random classifier control earns zero. They interpret the decline as evidence that US equity markets became more weak-form efficient over time and present a candidate explanation based on the rise of high-frequency trading volume, while acknowledging the single-event nature of that comparison.","tokens_in":8897,"tokens_out":3065,"duration_ms":33072,"significance":"If the empirical claims were robust, the paper would make a useful contribution by operationalizing market efficiency as the out-of-sample profitability of flexible ML learners and by documenting a historical decline in a specific class of intra-day predictability. Strengths include a clear experimental protocol with separate training, validation, and test periods, the use of a random classifier as a control, universe-relative features to reduce market beta, and explicit discussion of the N=1 limitation of the HFT correlation. These design choices are commendable. However, the central evidence for pre-2009 inefficiency is gross of transaction costs, and the break date is selected after seeing the data; both issues are load-bearing for the paper's main conclusion.","major_comments":[{"comment":"All reported returns are gross of transaction costs, slippage, and market impact. The strategy takes a daily balanced long/short portfolio in the 500 most-traded stocks, entering one minute after the prediction minute and exiting at the close, so each position incurs at least one round trip per day. Table 1 reports average early-period daily returns of 4.6 bps (Neural) and 4.9 bps (Logistic). For US equities in 2003–2008, realistic round-trip costs—quoted spreads, commissions, and impact—were commonly in the 5–20 bps range, especially in the earlier part of the sample. The paper provides no cost model, no break-even analysis, and no sensitivity check. If net-of-cost returns are non-positive in both periods, the profitability evidence for weak-form inefficiency before 2009, and hence the relative-efficiency inference built on it, collapses. This is the central load-bearing issue.","section":"Evaluation (Methodology) and Table 1"},{"comment":"The division into 'early' and 'late' periods is made after observing the data: the text states 'we observed an abrupt change in market efficiency after October 10, 2008 ... Accordingly, we selected the September/October 2008 boundary.' A post hoc break date inflates the apparent difference between periods and makes the reported early/late statistics, correlations, and figures uninterpretable as confirmatory evidence. The authors should either pre-specify the break (for example, using a known regulatory or structural event) or apply a formal structural-break test (e.g., Bai-Perron or a Chow test on the monthly return series) and report how the conclusions depend on the chosen breakpoint.","section":"Experimental Results (break date selection)"},{"comment":"No measures of uncertainty are provided for the key quantities. The claim that early-period returns are positive and late-period returns are zero rests on average daily returns of 4.6–4.9 bps versus -0.4 to -0.9 bps, but the paper reports no standard errors, t-statistics, bootstrap intervals, or number of daily observations underlying these means. Figure 3 suggests overlapping return distributions across periods. The authors should report significance tests (e.g., Newey-West adjusted t-tests for daily return series with autocorrelation) and show that the early-late difference is not attributable to a few volatile months around the 2008 crisis.","section":"Table 1 and Figures 1–4"}],"minor_comments":[{"comment":"The exact break date is inconsistent: the text says 'after October 10, 2008' and 'at September 30, 2008' in different places; this should be made consistent and stated precisely.","section":"Experimental Results (text and Figure 5)"},{"comment":"The Pearson correlation of -0.552 between HFT ratio and model monthly return is computed on two highly autocorrelated series; the effective sample size is much smaller than the number of months, so the correlation should be accompanied by a test that accounts for serial correlation (e.g., HAC standard errors or a test on differenced series).","section":"Figure 6"},{"comment":"Minor language issues: 'can traced' should be 'can be traced'; 'Avaramovic' appears misspelled in the Discussion; 'neither of them are profitable' should be 'neither of them is profitable.'","section":"Abstract and text"},{"comment":"The description of the long/short allocation says the strategy will not trade at all when there are no short predictions; it would be helpful to state how often the random classifier and the ML classifiers produced non-trading days, since this affects the interpretation of average daily returns.","section":"Approach"}],"recommendation":"major_revision","confidential_remarks":"The paper's central empirical claim is plausible and the experimental design is thoughtful, but the missing cost analysis and post hoc break selection are serious enough that the current version does not establish the stated conclusion. The authors should be given the opportunity to add a cost/break-even analysis, pre-specify or formally test the break date, and provide significance tests. The paper is within scope for the journal, though the efficiency interpretation will likely attract lively debate regardless of the statistical revisions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a well-written empirical paper with a genuinely useful idea - using the profitability of simple ML intraday strategies as a relative measure of weak-form market efficiency. The observed pattern, a decline in gross predictability around 2008-2009, is visible in their figures and is worth taking seriously. The problem is that the profitability evidence is computed before transaction costs, and the early-period edge (4.6-4.9 bps/day) is in the same range as realistic round-trip costs for US large caps in that era. As a result, the abstract's claim that the strategies were 'profitable' before 2009 is not yet established. What is actually new: the temporal profitability series itself. The authors cite standard EMH, liquidity, and HFT work, but none of it reports a fifteen-year rolling evaluation of simple neural and logistic classifiers on intraday data. The methodological care is real: monthly roll-forward retraining, universe-relative feature normalization, a one-minute lag between observation and order entry, and a random classifier control that correctly shows zero precision and zero return. The precision numbers (0.577 early vs. 0.519 late for both learners) support the qualitative claim that these models extracted predictive content before 2008, independent of the cost accounting. That part of the evidence is solid. Where it goes soft: the zero-cost assumption is load-bearing. The paper evaluates fills at recorded prices with no commissions, spread crossing, or impact, and never offers a break-even analysis. Given early-period average daily returns of under 5 bps and realistic costs that are often larger, the 'profitable' framing may collapse even if the raw predictability is real. The post hoc choice of the September/October 2008 break is a separate concern - they say plainly that they observed the abrupt change and then selected the boundary - so the timing inference is not robust. No significance tests accompany the early/late comparisons. The HFT correlation is interesting but, as they acknowledge, rests on a single historical event. The citation pattern looks appropriate and the limitations are honestly stated, including the N=1 nature of the HFT rise and the possibility that other algorithms could profit where theirs cannot. This is careful empirical work with a real empirical observation, but the central profitability claim needs cost modeling and robustness checks before it can support the efficiency interpretation. I would not cite it as evidence for the HFT-efficiency link, but I would cite it as a useful cautionary example of why gross backtest returns are not enough. Worth a serious referee: yes. The paper deserves peer review, not desk rejection, because the methodology is reproducible in principle and the empirical pattern is likely to be of interest even if the current conclusions are conditional. A referee can reasonably request a cost model, a break-even analysis, and robustness to alternative break dates. For a reading group: maybe. It would provoke a good discussion about backtest realism and what 'profitability' can and cannot tell us about market efficiency.","headline":"A clear, honest study that proposes ML strategy profitability as a weak-form efficiency meter and documents a real decline in gross predictability around 2008, but the zero-cost backtest leaves the central 'profitable before 2009' claim unestablished.","tokens_in":855,"tokens_out":906,"would_cite":false,"duration_ms":29570,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that the profitability of simple machine-learning trading strategies is a valid objective measure of weak-form market efficiency, and that this profitability disappeared around October 2008, coinciding with the peak of…","keywords":["market efficiency","weak-form EMH","intraday price prediction","neural networks","logistic regression","high-frequency trading","backtesting","machine learning trading"],"falsifier":"Re-run the same rolling monthly training and evaluation procedure on minute-bar data while charging a realistic round-trip transaction cost of at least five basis points per trade and a one-minute execution delay; if the early-period daily returns drop to zero or below, the profitability-based efficiency measure loses its empirical support.","tokens_in":8374,"feed_emoji":"📈","tokens_out":5534,"duration_ms":48396,"temperature":0.7,"pith_summary":"The paper proposes that the profitability of simple machine-learning trading strategies is an objective, quantitative measure of weak-form market efficiency. It shows that two such strategies — a neural network and a logistic regressor — earned steady daily returns on US equities from 2003 until late 2008, then produced no profits after 2009, while a random classifier earned nothing throughout. The authors interpret the disappearance of profits as evidence that the US equity market became more efficient, and they point to the rise of high-frequency trading volume as a candidate explanation. This matters because it offers a direct, algorithm-based test of a hypothesis that is usually evaluated indirectly through spreads and volatility.","feed_headline":"Machine-learning trading profits ended at the 2008 market peak","feed_subtitle":"A 15-year rolling backtest on minute-bar data shows both predictors fading to zero as high-frequency volume peaks.","key_machinery":"The key objects are the two reference trading systems — a fully connected feedforward neural network with two hidden layers (180 and 20 units, ReLU, Adam, early stopping) and an L2-regularized logistic regressor — each split into a long model and a short model trained on universe-relative cumulative returns ending one minute before the trade, with a monthly hyperparameter grid search over entry minute (end x ∈ {−5, −10, −30}) and movement threshold (bps ∈ {2, 5, 10, 25}) optimizing trade precision on a validation month. A random classifier trained and evaluated identically serves as the control. The mechanism these systems carry is the measurement: their daily returns, evaluated in a long/short balanced portfolio on the top-500-by-dollar-volume universe, are read as a direct probe of weak-form efficiency, with higher profitability indicating lower efficiency.","core_discovery":"On the paper's own terms, the central claim is that future intra-day stock prices could be predicted effectively, and profitably traded, until about October 2008 — that is, the weak-form efficient market hypothesis did not hold in US equities during 2003–2008 — and that the same two simple learning algorithms could no longer extract profit after 2009. The evidence is a monthly rolling backtest on one-minute bars for the top 500 stocks by dollar volume, with a long/short balanced portfolio evaluated on daily returns, cumulative returns, and trade precision. The neural-network strategy returned 4.6 basis points per day (93.5% cumulative) before the break and −0.4 basis points after; the logistic-regression strategy returned 4.9 basis points (101.6%) before and −0.9 after; the random control stayed at zero throughout. The authors propose that the time-varying profitability of such flexible learners constitutes an objective measure of relative market efficiency, and they observe that the profitability decline coincides with the rise of high-frequency trading volume, with a Pearson correlation of −0.552 between HFT share and model returns before the 2008 peak and 0.038 afterward.","pith_inferences":["The same measurement recipe could be applied to other markets (e.g., European or emerging-market equities, cryptocurrency pairs) where high-frequency trading grew at different times, turning the single historical event into a panel of natural experiments.","Because the reference learners are deliberately simple, the method could be strengthened by stress-testing the efficiency inference against more flexible learners (e.g., gradient-boosted trees or LSTMs) to see whether the post-2009 flatness is specific to the chosen models.","The paper's proposed agent-based simulation follow-up could be run with multiple HFT introductions to assess whether the negative correlation between HFT share and model returns is causal or coincidental."],"forward_implications":["If the central claim is correct, weak-form market efficiency in US equities increased sharply around October 2008, and the period 2003–2008 contained exploitable intraday inefficiencies.","The same methodology can be applied as a relative efficiency gauge: any market, asset class, or time period in which these reference learners earn positive returns is comparatively less efficient than one in which they do not.","The near-zero returns of the random classifier throughout the study confirm that the early profitability is attributable to real predictive information in the price data rather than to the long/short portfolio construction.","The strong negative correlation between HFT volume share and model returns before the peak, and its absence afterward, supports the candidate explanation that the rise of high-frequency trading drove the market toward efficiency, though the authors note the evidence is a single historical event."],"supporting_citations":[{"why":"Supplies the HFT volume ratio series used in the correlation analysis and the premise that HFT changed US market structure.","marker":"[Avramovic 2017]"},{"why":"Source of the TAQ minute-bar price data the entire backtest runs on.","marker":"[New York Stock Exchange 2018]"},{"why":"Defines the weak-form efficient market hypothesis that the paper tests.","marker":"[Malkiel and Fama 1970]"},{"why":"Updates the weak-form test classification to return predictability, which the paper adopts.","marker":"[Fama 1991]"},{"why":"Provides the simulation framework the authors propose for repeated HFT experiments.","marker":"[Wellman 2006]"}],"fun_headline_variants":["Machine-learning trading profits died after 2008","Predicting intraday prices worked only until 2009","ML market edge vanished as HFT volume rose","Deep learning lost its trading alpha by 2009","Backtest shows ML profits zero after 2008 peak"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The backtest assumes that every order is filled at the recorded one-minute close price with zero transaction costs, slippage, and market impact, so if realistic intra-day round-trip costs exceed the reported four to five basis points of daily return, the claimed profitability and the efficiency inference collapse.","fun_headline_variants_meta":{"raw":{"variants":["Machine-learning trading profits died after 2008","Predicting intraday prices worked only until 2009","ML market edge vanished as HFT volume rose","Deep learning lost its trading alpha by 2009","Backtest shows ML profits zero after 2008 peak"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000342,"raw_usage":{"total_tokens":1867,"prompt_tokens":914,"completion_tokens":953,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":530,"completion_tokens_details":{"reasoning_tokens":877}},"tokens_in":530,"tokens_out":953,"duration_ms":10257,"temperature":1.0,"reasoning_tokens":877,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:47:28.836140+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same rolling monthly training and evaluation procedure on minute-bar data while charging a realistic round-trip transaction cost of at least five basis points per trade and a one-minute execution delay; if the early-period daily returns drop to zero or below, the profitability-based efficiency measure loses its empirical support.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the HFT volume ratio series used in the correlation analysis and the premise that HFT changed US market structure."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source of the TAQ minute-bar price data the entire backtest runs on."},{"cited_title":"G., and Fama, E","cited_arxiv_id":null,"evidence_quote":"Defines the weak-form efficient market hypothesis that the paper tests."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Updates the weak-form test classification to return predictability, which the paper adopts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the simulation framework the authors propose for repeated HFT experiments."}],"review_version":1}