REVIEW 1 major objections 5 minor 26 references
The paper argues that candle-based Binance Spot timing models, despite strong event-ranking scores, did not produce positive net returns in any tested protocol, and every operational decision remains NO_TRADE.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 13:23 UTC pith:JZ4FD6C3
load-bearing objection Careful, candid negative audit showing high AUC coexists with negative policy returns; the main soft spot is the unquantified same-bar tie rule, but it's conservative and unlikely to flip the verdict. the 1 major comments →
Predictive Extrema, Unprofitable Policies: An AI-Assisted Audit of Candle-Based Binance Spot Timing Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central discovery is a systematic prediction-policy disconnect: models that rank rare extrema well (ROC AUC up to 0.896, average precision only 0.116-0.134) still produce negative net policy returns. The unchanged ten-pair mandatory-daily selector lost 6.72% over 19 July cycles at 31-bps cost (3 wins, 16 losses); local-extrema policies had gross mean advantages of 11.11-12.21 bps, below even the 21-bps lower cost stress; and a daily paired-extrema adaptation lost 44.30% versus -41.20% for buy-and-hold over seven cycles. The paper treats these as evidence that the tested protocols do not justify deployment.
What carries the argument
Key machinery: the paper couples candle-based (OHLCV) predictive models (ExtraTrees regressors, attention CNN-LSTMs, HGB proxies, and others) with deterministic paper-trading simulators that execute at next opens, apply target/stop barriers, and deduct flat completed-cycle costs (20/31/51 bps) with an adverse stop-first tie rule for same-bar target/stop touches. A two-dimensional evidence taxonomy classifies each result by artifact support (raw records vs. narrative vs. missing) and data status (frozen prospective, model-specific later period, consumed diagnostic, descriptive control). This machinery lets the paper separate ranking ability (AUC, average precision) from policy value (net retu
Load-bearing premise
The simulation assumes that when both the profit target and stop-loss are touched within the same candle, the stop is hit first; if real intrabar prices usually touch the target first, the paper's measured net returns would be systematically too negative.
What would settle it
Recompute all 19 July cycles using actual tick-by-tick Binance trade data to determine, for each same-bar target/stop touch, which price level was reached first. If applying the observed order flips the 19-cycle compounded return from -6.72% to positive after the same costs, the paper's central 'no positive policy value' conclusion for the selector would be refuted. Alternatively, a pre-registered forward test with the unchanged frozen model on a new month that yields positive net returns would also falsify the claim.
If this is right
- If the paper's results hold, a model can rank rare extrema well (AUC ≈ 0.89) and still lose money after 31-bps costs, so policy evaluation must include costs, execution timing, and abstention.
- A mandatory daily selector that always trades is worse than no-trade on this 19-cycle evidence; a selective policy with a cash action would be the next design to test.
- The evidence taxonomy introduced here—separating artifact-backed results from narratives and prospective from consumed data—provides a standard for reporting backtest credibility.
- Because even the gross edge (11–12 bps per cycle) is below the 21-bps cost stress, adding more models or thresholds is unlikely to rescue these specific protocols.
Where Pith is reading between the lines
- A natural extension of this audit would be to re-run the same models with intrabar tick data to measure how often target-before-stop ordering actually occurs; even a favorable reassignment is unlikely to flip the primary negative result because gross edge is already below the cost floor.
- The evidence-taxonomy approach could be applied to other published crypto ML trading results; many positive backtests likely lose this status once artifact provenance, search budget, and cost assumptions are made explicit.
- The paper implicitly argues for abstention as a policy action; one could design a selective-risk system that predeclares a cash position and calibrates acceptance on past-only data, which is a testable improvement.
- The dominant loss concentration in the selector (TRXUSDT contributed -334 bps over 9 cycles) suggests that a pair-specific or regime-based filter might be more fruitful than further global model tuning.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper audits whether several candle-based machine-learning models on Binance Spot data translate predictive performance into positive executable policy value after assumed costs. It reports a negative outcome across four artifact-backed campaigns: an unchanged ten-pair mandatory-daily selector lost 6.72% over 19 July cycles at 31 bps; validation-selected local-minimum and local-maximum policies lost -1.79% and -2.80%, respectively; a Gurgul-inspired daily paired-extrema adaptation lost -44.30% versus -41.20% for buy-and-hold; and a slow-rotation control was positive but concentrated and consumed. The paper also documents a forensic audit that downgraded an earlier One4All holdout due to data leakage, optimistic entry timing, and missing artifacts. The conclusion is that none of the tested policies establishes positive executable value, so every operational decision remains NO_TRADE.
Significance. If the results hold, the paper provides a carefully qualified negative result and a methodological template. It demonstrates a clear prediction-policy disconnect: post-hoc ROC AUC up to 0.896 and AP up to 0.360 coexist with negative net policy returns. The paper explicitly reports search budgets (946+ mandatory-daily candidates, 238 local-extrema validation comparisons, 15 paired-daily profiles, 240 rotation policies), applies multi-level cost stress (20/21/31/51 bps), and substantiates its audit with artifact provenance tables. The self-audit that removed a leaky holdout is a credible scientific contribution. The evidence is limited by small support (19/9/15/7 cycles) and the lack of preregistration, but the authors consistently qualify their claims and avoid overgeneralization.
major comments (1)
- [§4.2, Tables 3, Fig. 3D] The adverse stop-first tie rule is a worst-case assumption whose quantitative impact is unmeasured. The paper states that OHLC bars cannot reveal intra-bar target/stop order and then uses stop-first, but it never reports how often both barriers are touched within a single bar, nor does it provide a sensitivity analysis under a target-first or stochastic tie rule. This matters because the local-extrema gross mean advantages (11.11 and 12.21 bps) are stated to be 'below even the 21-bps stress'; if real paths are frequently target-first, these gross edges could exceed 21 bps, weakening the claim that the failure is not cost-driven. The mandatory-daily and paired-daily results are sufficiently negative that the tie rule is unlikely to reverse them, but the reader cannot verify this. Please report tie-bar frequencies per campaign and add a sensitivity analysis (e.g., target-first and randomiz
minor comments (5)
- [Table 3] The table rows appear as unbroken text without visible column separation in the provided manuscript. Use a proper tabular environment with clear columns for period, cost, support, result, and data status.
- [Eq. (1)] The multiplier '10,000' is clear but could be written as 10^4 for readability. Also, define all symbols immediately around the equation, especially p_H and p_e, even if they appear in the text.
- [Figure 3] The legend does not fully map marker shapes to tabular/sequence/hybrid families. Adding an explicit legend or a caption sentence describing each marker class would improve interpretability.
- [§6.1] For the mandatory-daily selector, the Clopper-Pearson interval is given for the win rate, but no interval is provided for the mean net cycle return (-36.40 bps). A bootstrap or t-interval would help readers understand the sampling uncertainty of the point estimate, though it is not essential to the main conclusion.
- [Figure 1] Panel B is visually dense; consider larger fonts and a clearer separation of the invalidated One4All span from the artifact-backed spans.
Circularity Check
No significant circularity: the negative conclusion is an audited empirical report, not a derivation that reduces to its inputs.
full rationale
I walked the claimed derivation chain and found no load-bearing step in which a prediction or first-principles result is equivalent, by construction, to a fitted input or to an unverified self-citation. (1) The mandatory-daily 'frozen prospective' campaign selected a candidate using April–June data and then evaluated an unchanged model on July 1–19; the paper explicitly says 'April–June served both as consumed cross-family selection data and as part of the final fit' and conditions the July result on that predecessor search. The July evaluation is temporally disjoint, so it is not the same data used for selection. (2) The local-extrema campaigns chose models on April–June validation and evaluated July 1–12; the paper explicitly labels these 'model-specific' slices, notes July 1–7 exposure, and reports that the validation gate failed. The realized-sample break-even costs (10.9616/12.0690 bps) are computed from the observed policy paths and are explicitly called 'realized-sample summaries of 9 and 15 cycles, not estimates of attainable future costs or edge'; they are descriptive, not parameters used to predict the negative result. (3) The 21/31/51-bps cost assumptions are stated as assumptions, not fitted to the outcome, and the cost sensitivity analysis is an evaluation rather than a derivation. (4) The stop-first tie rule is a fragility/assumption issue, not a circularity: it does not define the target or the policy outcome as equal to the input. (5) The only author self-citation [18] points to a mutable repository and the paper explicitly avoids relying on it for exact reconstruction, so it is not load-bearing. No equation in the paper sets a 'prediction' equal to a fitted quantity, and no central claim rests on a self-citation chain. The paper is best understood as a self-contained, conservative empirical audit; score 0.
Axiom & Free-Parameter Ledger
free parameters (8)
- Primary completed-cycle cost (flat) =
31 bps (stresses 20/21/51)
- Local-extrema label radius b =
24 bars (120 min each side)
- Local-extrema gross target/stop barriers =
+60/-50 bps, 24-bar horizon
- Mandatory-daily target/stop =
+100/-150 bps, max 4h
- Paired-daily common threshold =
0.80
- Paired-daily window w and model choice =
w=14; HGB proxy
- Deep-model seed count =
1 seed per model
- Paired-daily sleeve size =
2 sleeves of 50%
axioms (4)
- domain assumption OHLCV bars do not reveal intrabar order; when target and stop are both touched in a bar, the adverse (stop-first) outcome is assumed.
- domain assumption Transaction-cost stress levels (21–51 bps) bracket unmeasured real all-in costs.
- domain assumption Binance kline snapshots are complete and reliable, and the mechanically volume-screened universes do not introduce look-ahead or survivorship that would reverse the negative results.
- domain assumption The disclosed search ledger (946+ candidates, 238 validation comparisons per direction, 15 profiles × 7 folds, 3,996 One4All profiles) is exhaustive enough that the July evaluations count as out-of-sample for the unchanged models.
read the original abstract
We audit whether candle-based machine-learning models can turn predictions of cryptocurrency extrema or short-horizon outcomes into positive Binance Spot paper policies after assumed costs. Numerical results come from scripted fixed-seed model runs and deterministic simulators; human-supervised AI agents supported the July 20 evidence-integrity revision through literature retrieval, separately tasked critique, artifact reconciliation, documentation, and source packaging, not trading decisions. The strongest later-period evidence, conditional on extensive predecessor search, is negative: an unchanged ten-pair mandatory-daily selector lost 6.72\% over 19 July cycles at an assumed 31-bps completed-cycle cost, with 3 wins and 16 losses. In short model-specific July evaluations, the validation-selected local-minimum policy returned -1.79\%, while the local-maximum sell-to-cash/re-entry policy underperformed continuous holding by 2.80\%; their gross mean advantages of 11.11 and 12.21 bps were below even the 21-bps stress. A Gurgul-inspired, OHLCV-only daily adaptation attained minimum/maximum ROC AUC of 0.874/0.896 but average precision of only 0.134/0.116 and lost 44.30\% over seven cycles, versus -41.20\% for buy-and-hold. A forensic audit also downgraded an earlier One4All "30-day holdout": its dates had influenced prior architecture work, its four-hour outcome horizon was not purged at split boundaries, it used same-close entry, and its raw result directories were absent. Across the tested, mostly exploratory protocols, event-ranking performance did not establish positive executable policy value. Every operational decision remains NO\_TRADE.
Figures
Reference graph
Works this paper leans on
-
[1]
Anticipating cryptocurrency prices using machine learning.Complexity, 2018:1–16, 2018
LauraAlessandretti, AbeerElBahrawy, LucaMariaAiello, andAndreaBaronchelli. Anticipating cryptocurrency prices using machine learning.Complexity, 2018:1–16, 2018. doi: 10.1155/2018 /8983590. Article ID 8983590
-
[2]
Robert D. Arnott, Campbell R. Harvey, and Harry Markowitz. A backtesting protocol in the era of machine learning.The Journal of Financial Data Science, 1(1):64–74, 2019. doi: 10.3905/jfds.2019.1.064
-
[3]
Artificial intelligence risk management framework: Generative artificial intelligence profile
Chloe Autio, Reva Schwartz, Jesse Dunietz, Shomik Jain, Martin Stanley, Elham Tabassi, Patrick Hall, and Kamie Roberts. Artificial intelligence risk management framework: Generative artificial intelligence profile. NIST AI 600-1, National Institute of Standards and Technology, Gaithersburg, MD, 2024.https://doi.org/10.6028/NIST.AI.600-1
-
[4]
Bailey and Marcos López de Prado
David H. Bailey and Marcos López de Prado. The deflated sharpe ratio: Correcting for selection bias, backtest overfitting, and non-normality.The Journal of Portfolio Management, 40(5): 94–107, 2014. doi: 10.3905/jpm.2014.40.5.094
-
[5]
David H. Bailey, Jonathan M. Borwein, Marcos López de Prado, and Qiji Jim Zhu. The probability of backtest overfitting.Journal of Computational Finance, 20(4):39–69, 2017. doi: 10.21314/JCF.2016.322
-
[6]
Christoph Bergmeir, Rob J. Hyndman, and Bonsoo Koo. A note on the validity of cross- validation for evaluating autoregressive time series prediction.Computational Statistics & Data Analysis, 120:70–83, 2018. doi: 10.1016/j.csda.2017.11.003
-
[7]
Commission rates
Binance. Commission rates. https://developers.binance.com/en/docs/products/spot/f aqs/commission_faq, 2026. Accessed 2026-07-20
2026
-
[8]
Spot trading fee rate.https://www.binance.com/en/fee/trading, 2026
Binance. Spot trading fee rate.https://www.binance.com/en/fee/trading, 2026. Mutable fee schedule. Accessed 2026-07-20
2026
-
[9]
Spot REST API: Kline/candlestick data
Binance. Spot REST API: Kline/candlestick data. https://developers.binance.com/e n/docs/catalog/core-trading-spot-trading/api/rest-api/market , 2026. Accessed 2026-07-20
2026
-
[10]
Andrei Bysik and Robert Ślepaczuk. Machine learning-based bitcoin trading under transaction costs: Evidence from walk-forward forecasting. arXiv:2606.00060,https://arxiv.org/abs/ 2606.00060, 2026. Version 1
Pith/arXiv arXiv 2026
-
[11]
Evaluating time series forecasting models: An empirical study on performance estimation methods.Machine Learning, 109(11):1997–2028,
Vitor Cerqueira, Luis Torgo, and Igor Mozetič. Evaluating time series forecasting models: An empirical study on performance estimation methods.Machine Learning, 109(11):1997–2028,
1997
-
[12]
Adam N. Elmachtoub and Paul Grigas. Smart “predict, then optimize”.Management Science, 68(1):9–26, 2022. doi: 10.1287/mnsc.2020.3922
arXiv 2022
-
[13]
Selective classification for deep neural networks
Yonatan Geifman and Ran El-Yaniv. Selective classification for deep neural networks. In Advances in Neural Information Processing Systems, volume 30, pages 4878–4887, 2017. URL https://proceedings.neurips.cc/paper_files/paper/2017/hash/4a8423d5e91fda00b b7e46540e2b0cf1-Abstract.html. 22 Predictive Extrema, Unprofitable Policies Preprint
2017
-
[14]
Berend Jelmer Dirk Gort, Xiao-Yang Liu, Xinghang Sun, Jiechao Gao, Shuaiyu Chen, and Christina Dan Wang. Deep reinforcement learning for cryptocurrency trading: Practical approach to address backtest overfitting. arXiv:2209.05559,https://arxiv.org/abs/2209.0 5559, 2023. Version 6
Pith/arXiv arXiv 2023
-
[15]
Vincent Gurgul, Stefan Lessmann, and Wolfgang Karl Härdle. Deep learning and NLP in cryp- tocurrency forecasting: Integrating financial, blockchain, and social media data.International Journal of Forecasting, 41(4):1666–1695, 2025. doi: 10.1016/j.ijforecast.2025.02.007. The local study described in this paper is an OHLCV-only adaptation, not a replication
-
[16]
Peter R. Hansen. A test for superior predictive ability.Journal of Business & Economic Statistics, 23(4):365–380, 2005. doi: 10.1198/073500105000000063
-
[17]
Harvey, Yan Liu, and Heqing Zhu
Campbell R. Harvey, Yan Liu, and Heqing Zhu. ... and the cross-section of expected returns. The Review of Financial Studies, 29(1):5–68, 2016. doi: 10.1093/rfs/hhv059
-
[18]
Quantbot research framework: Simulation-only spot-crypto research repository
Ayoub Jadouli. Quantbot research framework: Simulation-only spot-crypto research repository. https://github.com/AyoubJadouli/Quantbot-Research-Framework , 2026. Mutable discovery repository; no archival DOI. Accessed 2026-07-20
2026
-
[19]
Lo and A
Andrew W. Lo and A. Craig MacKinlay. Data-snooping biases in tests of financial asset pricing models.The Review of Financial Studies, 3(3):431–467, 1990
1990
-
[20]
Takaya Saito and Marc Rehmsmeier. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets.PLOS ONE, 10(3): e0118432, 2015. doi: 10.1371/journal.pone.0118432
-
[21]
Forecasting and trading cryptocurrencies with machine learning under changing market conditions.Financial Innovation, 7(1):3, 2021
Helder Sebastião and Pedro Godinho. Forecasting and trading cryptocurrencies with machine learning under changing market conditions.Financial Innovation, 7(1):3, 2021. doi: 10.1186/ s40854-020-00217-x
2021
-
[22]
Data-snooping, technical trading rule performance, and the bootstrap.The Journal of Finance, 54(5):1647–1691, 1999
Ryan Sullivan, Allan Timmermann, and Halbert White. Data-snooping, technical trading rule performance, and the bootstrap.The Journal of Finance, 54(5):1647–1691, 1999
1999
-
[23]
A reality check for data snooping.Econometrica, 68(5):1097–1126, 2000
Halbert White. A reality check for data snooping.Econometrica, 68(5):1097–1126, 2000. doi: 10.1111/1468-0262.00152
arXiv 2000
-
[24]
Agentic trading: When LLM agents meet financial markets
Yihan Xia, Panpan You, Taotao Wang, Fang Liu, Han Qi, Xiaoxiao Wu, and Shengli Zhang. Agentic trading: When LLM agents meet financial markets. arXiv:2605.19337,https://arxi v.org/abs/2605.19337, 2026. Version 1
Pith/arXiv arXiv 2026
-
[25]
Winker, Rakesh Aggarwal, Lorraine E
Chris Zielinski, Margaret A. Winker, Rakesh Aggarwal, Lorraine E. Ferris, Markus Heinemann, José Florencio Lapeña, Sanjay A. Pai, Edsel Ing, Leslie Citrome, Murad Alam, Michael Voight, Farrokh Habibzadeh, and on behalf of the WAME Board. Chatbots, generative AI, and scholarly manuscripts: WAME recommendations on chatbots and generative artificial intellig...
arXiv 2024
-
[2020]
doi: 10.1007/s10994-020-05910-7
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.