Pith. sign in

REVIEW 1 major objections 5 minor 26 references

The paper argues that candle-based Binance Spot timing models, despite strong event-ranking scores, did not produce positive net returns in any tested protocol, and every operational decision remains NO_TRADE.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 13:23 UTC pith:JZ4FD6C3

load-bearing objection Careful, candid negative audit showing high AUC coexists with negative policy returns; the main soft spot is the unquantified same-bar tie rule, but it's conservative and unlikely to flip the verdict. the 1 major comments →

arxiv 2607.19453 v1 pith:JZ4FD6C3 submitted 2026-07-21 cs.LG cs.AIq-fin.STq-fin.TR

Predictive Extrema, Unprofitable Policies: An AI-Assisted Audit of Candle-Based Binance Spot Timing Models

classification cs.LG cs.AIq-fin.STq-fin.TR
keywords cryptocurrency forecastingextrema detectionpolicy valuechronological evaluationbacktest overfittingtransaction costsAI-assisted auditBinance Spot
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper tests whether machine-learning models that predict short-horizon price extrema in cryptocurrency pairs can be turned into profitable trading policies on Binance Spot after transaction costs. Across several model families, including a frozen daily selector, local-minimum and local-maximum entry/exit policies, and a daily paired-extrema adaptation, the models show strong ROC AUC but weak average precision and negative net returns. The strongest later-period evidence is a 19-cycle July campaign that lost 6.72% at an assumed 31-bps cost. The paper's central lesson is that event-ranking performance does not imply policy value; costs, execution timing, and abstention dominate. It concludes that every evaluated operational decision should remain NO_TRADE.

Core claim

The central discovery is a systematic prediction-policy disconnect: models that rank rare extrema well (ROC AUC up to 0.896, average precision only 0.116-0.134) still produce negative net policy returns. The unchanged ten-pair mandatory-daily selector lost 6.72% over 19 July cycles at 31-bps cost (3 wins, 16 losses); local-extrema policies had gross mean advantages of 11.11-12.21 bps, below even the 21-bps lower cost stress; and a daily paired-extrema adaptation lost 44.30% versus -41.20% for buy-and-hold over seven cycles. The paper treats these as evidence that the tested protocols do not justify deployment.

What carries the argument

Key machinery: the paper couples candle-based (OHLCV) predictive models (ExtraTrees regressors, attention CNN-LSTMs, HGB proxies, and others) with deterministic paper-trading simulators that execute at next opens, apply target/stop barriers, and deduct flat completed-cycle costs (20/31/51 bps) with an adverse stop-first tie rule for same-bar target/stop touches. A two-dimensional evidence taxonomy classifies each result by artifact support (raw records vs. narrative vs. missing) and data status (frozen prospective, model-specific later period, consumed diagnostic, descriptive control). This machinery lets the paper separate ranking ability (AUC, average precision) from policy value (net retu

Load-bearing premise

The simulation assumes that when both the profit target and stop-loss are touched within the same candle, the stop is hit first; if real intrabar prices usually touch the target first, the paper's measured net returns would be systematically too negative.

What would settle it

Recompute all 19 July cycles using actual tick-by-tick Binance trade data to determine, for each same-bar target/stop touch, which price level was reached first. If applying the observed order flips the 19-cycle compounded return from -6.72% to positive after the same costs, the paper's central 'no positive policy value' conclusion for the selector would be refuted. Alternatively, a pre-registered forward test with the unchanged frozen model on a new month that yields positive net returns would also falsify the claim.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If the paper's results hold, a model can rank rare extrema well (AUC ≈ 0.89) and still lose money after 31-bps costs, so policy evaluation must include costs, execution timing, and abstention.
  • A mandatory daily selector that always trades is worse than no-trade on this 19-cycle evidence; a selective policy with a cash action would be the next design to test.
  • The evidence taxonomy introduced here—separating artifact-backed results from narratives and prospective from consumed data—provides a standard for reporting backtest credibility.
  • Because even the gross edge (11–12 bps per cycle) is below the 21-bps cost stress, adding more models or thresholds is unlikely to rescue these specific protocols.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A natural extension of this audit would be to re-run the same models with intrabar tick data to measure how often target-before-stop ordering actually occurs; even a favorable reassignment is unlikely to flip the primary negative result because gross edge is already below the cost floor.
  • The evidence-taxonomy approach could be applied to other published crypto ML trading results; many positive backtests likely lose this status once artifact provenance, search budget, and cost assumptions are made explicit.
  • The paper implicitly argues for abstention as a policy action; one could design a selective-risk system that predeclares a cash position and calibrates acceptance on past-only data, which is a testable improvement.
  • The dominant loss concentration in the selector (TRXUSDT contributed -334 bps over 9 cycles) suggests that a pair-specific or regime-based filter might be more fruitful than further global model tuning.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 5 minor

Summary. The paper audits whether several candle-based machine-learning models on Binance Spot data translate predictive performance into positive executable policy value after assumed costs. It reports a negative outcome across four artifact-backed campaigns: an unchanged ten-pair mandatory-daily selector lost 6.72% over 19 July cycles at 31 bps; validation-selected local-minimum and local-maximum policies lost -1.79% and -2.80%, respectively; a Gurgul-inspired daily paired-extrema adaptation lost -44.30% versus -41.20% for buy-and-hold; and a slow-rotation control was positive but concentrated and consumed. The paper also documents a forensic audit that downgraded an earlier One4All holdout due to data leakage, optimistic entry timing, and missing artifacts. The conclusion is that none of the tested policies establishes positive executable value, so every operational decision remains NO_TRADE.

Significance. If the results hold, the paper provides a carefully qualified negative result and a methodological template. It demonstrates a clear prediction-policy disconnect: post-hoc ROC AUC up to 0.896 and AP up to 0.360 coexist with negative net policy returns. The paper explicitly reports search budgets (946+ mandatory-daily candidates, 238 local-extrema validation comparisons, 15 paired-daily profiles, 240 rotation policies), applies multi-level cost stress (20/21/31/51 bps), and substantiates its audit with artifact provenance tables. The self-audit that removed a leaky holdout is a credible scientific contribution. The evidence is limited by small support (19/9/15/7 cycles) and the lack of preregistration, but the authors consistently qualify their claims and avoid overgeneralization.

major comments (1)
  1. [§4.2, Tables 3, Fig. 3D] The adverse stop-first tie rule is a worst-case assumption whose quantitative impact is unmeasured. The paper states that OHLC bars cannot reveal intra-bar target/stop order and then uses stop-first, but it never reports how often both barriers are touched within a single bar, nor does it provide a sensitivity analysis under a target-first or stochastic tie rule. This matters because the local-extrema gross mean advantages (11.11 and 12.21 bps) are stated to be 'below even the 21-bps stress'; if real paths are frequently target-first, these gross edges could exceed 21 bps, weakening the claim that the failure is not cost-driven. The mandatory-daily and paired-daily results are sufficiently negative that the tie rule is unlikely to reverse them, but the reader cannot verify this. Please report tie-bar frequencies per campaign and add a sensitivity analysis (e.g., target-first and randomiz
minor comments (5)
  1. [Table 3] The table rows appear as unbroken text without visible column separation in the provided manuscript. Use a proper tabular environment with clear columns for period, cost, support, result, and data status.
  2. [Eq. (1)] The multiplier '10,000' is clear but could be written as 10^4 for readability. Also, define all symbols immediately around the equation, especially p_H and p_e, even if they appear in the text.
  3. [Figure 3] The legend does not fully map marker shapes to tabular/sequence/hybrid families. Adding an explicit legend or a caption sentence describing each marker class would improve interpretability.
  4. [§6.1] For the mandatory-daily selector, the Clopper-Pearson interval is given for the win rate, but no interval is provided for the mean net cycle return (-36.40 bps). A bootstrap or t-interval would help readers understand the sampling uncertainty of the point estimate, though it is not essential to the main conclusion.
  5. [Figure 1] Panel B is visually dense; consider larger fonts and a clearer separation of the invalidated One4All span from the artifact-backed spans.

Circularity Check

0 steps flagged

No significant circularity: the negative conclusion is an audited empirical report, not a derivation that reduces to its inputs.

full rationale

I walked the claimed derivation chain and found no load-bearing step in which a prediction or first-principles result is equivalent, by construction, to a fitted input or to an unverified self-citation. (1) The mandatory-daily 'frozen prospective' campaign selected a candidate using April–June data and then evaluated an unchanged model on July 1–19; the paper explicitly says 'April–June served both as consumed cross-family selection data and as part of the final fit' and conditions the July result on that predecessor search. The July evaluation is temporally disjoint, so it is not the same data used for selection. (2) The local-extrema campaigns chose models on April–June validation and evaluated July 1–12; the paper explicitly labels these 'model-specific' slices, notes July 1–7 exposure, and reports that the validation gate failed. The realized-sample break-even costs (10.9616/12.0690 bps) are computed from the observed policy paths and are explicitly called 'realized-sample summaries of 9 and 15 cycles, not estimates of attainable future costs or edge'; they are descriptive, not parameters used to predict the negative result. (3) The 21/31/51-bps cost assumptions are stated as assumptions, not fitted to the outcome, and the cost sensitivity analysis is an evaluation rather than a derivation. (4) The stop-first tie rule is a fragility/assumption issue, not a circularity: it does not define the target or the policy outcome as equal to the input. (5) The only author self-citation [18] points to a mutable repository and the paper explicitly avoids relying on it for exact reconstruction, so it is not load-bearing. No equation in the paper sets a 'prediction' equal to a fitted quantity, and no central claim rests on a self-citation chain. The paper is best understood as a self-contained, conservative empirical audit; score 0.

Axiom & Free-Parameter Ledger

8 free parameters · 4 axioms · 0 invented entities

The paper introduces no new theoretical entities or mediators. Its reliance is on hand-set cost assumptions, fixed strategy hyperparameters, label-construction choices, and the completeness of its own search ledger. These are disclosed but are not externally benchmarked; the negative result is conditioned on them.

free parameters (8)
  • Primary completed-cycle cost (flat) = 31 bps (stresses 20/21/51)
    Hand-set assumption approximating two 10-bps sides plus 1–10 bps buffer; not measured from account statements (§4.2). The central NO_TRADE conclusion is robust to using 20/21 bps for the main campaigns.
  • Local-extrema label radius b = 24 bars (120 min each side)
    Selected on validation-only window scout (§5.2, Figure 3C); all 30/60/120-min windows failed the predeclared gates.
  • Local-extrema gross target/stop barriers = +60/-50 bps, 24-bar horizon
    Fixed by protocol; the exact realized break-even costs of 10.96/12.07 bps are derived from the 9/15 observed cycles.
  • Mandatory-daily target/stop = +100/-150 bps, max 4h
    Fixed by protocol; affects the 19-cycle win/loss tally and symbol attribution.
  • Paired-daily common threshold = 0.80
    Calibrated on 2025-04-01 to 06-30 (§5.3); signals execute next daily open.
  • Paired-daily window w and model choice = w=14; HGB proxy
    Selected by absolute strategy return over 15 profiles × 7 folds; median CV excess vs buy-and-hold was -18.60 pp (§5.3).
  • Deep-model seed count = 1 seed per model
    Deep models use a single seed; no variance estimate across seeds is provided (§9).
  • Paired-daily sleeve size = 2 sleeves of 50%
    Portfolio construction choice; affects terminal liquidation path and benchmark comparison.
axioms (4)
  • domain assumption OHLCV bars do not reveal intrabar order; when target and stop are both touched in a bar, the adverse (stop-first) outcome is assumed.
    §4.2; without path data any tie rule is a model. Stop-first is conservative, but if target-first is the more common real path the net results would improve.
  • domain assumption Transaction-cost stress levels (21–51 bps) bracket unmeasured real all-in costs.
    Costs are 'assumptions, not measured realized commissions' (§4.2). The negative conclusion assumes the true cost is at least the stress values, which is plausible for retail Binance fees but not guaranteed for maker-rebate or VIP accounts.
  • domain assumption Binance kline snapshots are complete and reliable, and the mechanically volume-screened universes do not introduce look-ahead or survivorship that would reverse the negative results.
    §4.1 acknowledges that longer-history availability and mechanical volume screening may introduce survivorship or selection bias, especially in the rotation results; data integrity is not independently verified.
  • domain assumption The disclosed search ledger (946+ candidates, 238 validation comparisons per direction, 15 profiles × 7 folds, 3,996 One4All profiles) is exhaustive enough that the July evaluations count as out-of-sample for the unchanged models.
    §5.1 and Table 5. If any undisclosed comparison consumed July dates, the prospective status of the frozen selector would be invalid.

pith-pipeline@v1.3.0-alltime-deepseek · 16581 in / 15532 out tokens · 145515 ms · 2026-08-01T13:23:01.969484+00:00 · methodology

0 comments
read the original abstract

We audit whether candle-based machine-learning models can turn predictions of cryptocurrency extrema or short-horizon outcomes into positive Binance Spot paper policies after assumed costs. Numerical results come from scripted fixed-seed model runs and deterministic simulators; human-supervised AI agents supported the July 20 evidence-integrity revision through literature retrieval, separately tasked critique, artifact reconciliation, documentation, and source packaging, not trading decisions. The strongest later-period evidence, conditional on extensive predecessor search, is negative: an unchanged ten-pair mandatory-daily selector lost 6.72\% over 19 July cycles at an assumed 31-bps completed-cycle cost, with 3 wins and 16 losses. In short model-specific July evaluations, the validation-selected local-minimum policy returned -1.79\%, while the local-maximum sell-to-cash/re-entry policy underperformed continuous holding by 2.80\%; their gross mean advantages of 11.11 and 12.21 bps were below even the 21-bps stress. A Gurgul-inspired, OHLCV-only daily adaptation attained minimum/maximum ROC AUC of 0.874/0.896 but average precision of only 0.134/0.116 and lost 44.30\% over seven cycles, versus -41.20\% for buy-and-hold. A forensic audit also downgraded an earlier One4All "30-day holdout": its dates had influenced prior architecture work, its four-hour outcome horizon was not purged at split boundaries, it used same-close entry, and its raw result directories were absent. Across the tested, mostly exploratory protocols, event-ranking performance did not establish positive executable policy value. Every operational decision remains NO\_TRADE.

Figures

Figures reproduced from arXiv: 2607.19453 by Ayoub Jadouli.

Figure 1
Figure 1. Figure 1: Chronology and data-status audit, drawn on a common calendar scale. Panel A maps the [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Unchanged mandatory-daily selector on all 19 July 2026 cycles. Panel A recomputes [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Complete local-extrema model and policy audit. Panels A and B show all 14 minimum [PITH_FULL_IMAGE:figures/full_fig_p012_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Gurgul-inspired OHLCV-only daily adaptation. Panel A relates median buy-and-hold [PITH_FULL_IMAGE:figures/full_fig_p014_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Consumed slow-rotation control. The primary 31-bps round-trip label is charged as half [PITH_FULL_IMAGE:figures/full_fig_p016_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

26 extracted references · 2 canonical work pages

  1. [1]

    Anticipating cryptocurrency prices using machine learning.Complexity, 2018:1–16, 2018

    LauraAlessandretti, AbeerElBahrawy, LucaMariaAiello, andAndreaBaronchelli. Anticipating cryptocurrency prices using machine learning.Complexity, 2018:1–16, 2018. doi: 10.1155/2018 /8983590. Article ID 8983590

  2. [2]

    Arnott, Campbell R

    Robert D. Arnott, Campbell R. Harvey, and Harry Markowitz. A backtesting protocol in the era of machine learning.The Journal of Financial Data Science, 1(1):64–74, 2019. doi: 10.3905/jfds.2019.1.064

  3. [3]

    Artificial intelligence risk management framework: Generative artificial intelligence profile

    Chloe Autio, Reva Schwartz, Jesse Dunietz, Shomik Jain, Martin Stanley, Elham Tabassi, Patrick Hall, and Kamie Roberts. Artificial intelligence risk management framework: Generative artificial intelligence profile. NIST AI 600-1, National Institute of Standards and Technology, Gaithersburg, MD, 2024.https://doi.org/10.6028/NIST.AI.600-1

  4. [4]

    Bailey and Marcos López de Prado

    David H. Bailey and Marcos López de Prado. The deflated sharpe ratio: Correcting for selection bias, backtest overfitting, and non-normality.The Journal of Portfolio Management, 40(5): 94–107, 2014. doi: 10.3905/jpm.2014.40.5.094

  5. [5]

    Bailey, Jonathan M

    David H. Bailey, Jonathan M. Borwein, Marcos López de Prado, and Qiji Jim Zhu. The probability of backtest overfitting.Journal of Computational Finance, 20(4):39–69, 2017. doi: 10.21314/JCF.2016.322

  6. [6]

    Hyndman, and Bonsoo Koo

    Christoph Bergmeir, Rob J. Hyndman, and Bonsoo Koo. A note on the validity of cross- validation for evaluating autoregressive time series prediction.Computational Statistics & Data Analysis, 120:70–83, 2018. doi: 10.1016/j.csda.2017.11.003

  7. [7]

    Commission rates

    Binance. Commission rates. https://developers.binance.com/en/docs/products/spot/f aqs/commission_faq, 2026. Accessed 2026-07-20

  8. [8]

    Spot trading fee rate.https://www.binance.com/en/fee/trading, 2026

    Binance. Spot trading fee rate.https://www.binance.com/en/fee/trading, 2026. Mutable fee schedule. Accessed 2026-07-20

  9. [9]

    Spot REST API: Kline/candlestick data

    Binance. Spot REST API: Kline/candlestick data. https://developers.binance.com/e n/docs/catalog/core-trading-spot-trading/api/rest-api/market , 2026. Accessed 2026-07-20

  10. [10]

    Machine learning-based bitcoin trading under transaction costs: Evidence from walk-forward forecasting

    Andrei Bysik and Robert Ślepaczuk. Machine learning-based bitcoin trading under transaction costs: Evidence from walk-forward forecasting. arXiv:2606.00060,https://arxiv.org/abs/ 2606.00060, 2026. Version 1

  11. [11]

    Evaluating time series forecasting models: An empirical study on performance estimation methods.Machine Learning, 109(11):1997–2028,

    Vitor Cerqueira, Luis Torgo, and Igor Mozetič. Evaluating time series forecasting models: An empirical study on performance estimation methods.Machine Learning, 109(11):1997–2028,

  12. [12]

    predict, then optimize

    Adam N. Elmachtoub and Paul Grigas. Smart “predict, then optimize”.Management Science, 68(1):9–26, 2022. doi: 10.1287/mnsc.2020.3922

  13. [13]

    Selective classification for deep neural networks

    Yonatan Geifman and Ran El-Yaniv. Selective classification for deep neural networks. In Advances in Neural Information Processing Systems, volume 30, pages 4878–4887, 2017. URL https://proceedings.neurips.cc/paper_files/paper/2017/hash/4a8423d5e91fda00b b7e46540e2b0cf1-Abstract.html. 22 Predictive Extrema, Unprofitable Policies Preprint

  14. [14]

    Deep reinforcement learning for cryptocurrency trading: Practical approach to address backtest overfitting

    Berend Jelmer Dirk Gort, Xiao-Yang Liu, Xinghang Sun, Jiechao Gao, Shuaiyu Chen, and Christina Dan Wang. Deep reinforcement learning for cryptocurrency trading: Practical approach to address backtest overfitting. arXiv:2209.05559,https://arxiv.org/abs/2209.0 5559, 2023. Version 6

  15. [15]

    Deep learning and NLP in cryp- tocurrency forecasting: Integrating financial, blockchain, and social media data.International Journal of Forecasting, 41(4):1666–1695, 2025

    Vincent Gurgul, Stefan Lessmann, and Wolfgang Karl Härdle. Deep learning and NLP in cryp- tocurrency forecasting: Integrating financial, blockchain, and social media data.International Journal of Forecasting, 41(4):1666–1695, 2025. doi: 10.1016/j.ijforecast.2025.02.007. The local study described in this paper is an OHLCV-only adaptation, not a replication

  16. [16]

    Peter R. Hansen. A test for superior predictive ability.Journal of Business & Economic Statistics, 23(4):365–380, 2005. doi: 10.1198/073500105000000063

  17. [17]

    Harvey, Yan Liu, and Heqing Zhu

    Campbell R. Harvey, Yan Liu, and Heqing Zhu. ... and the cross-section of expected returns. The Review of Financial Studies, 29(1):5–68, 2016. doi: 10.1093/rfs/hhv059

  18. [18]

    Quantbot research framework: Simulation-only spot-crypto research repository

    Ayoub Jadouli. Quantbot research framework: Simulation-only spot-crypto research repository. https://github.com/AyoubJadouli/Quantbot-Research-Framework , 2026. Mutable discovery repository; no archival DOI. Accessed 2026-07-20

  19. [19]

    Lo and A

    Andrew W. Lo and A. Craig MacKinlay. Data-snooping biases in tests of financial asset pricing models.The Review of Financial Studies, 3(3):431–467, 1990

  20. [20]

    The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets.PLOS ONE, 10(3): e0118432, 2015

    Takaya Saito and Marc Rehmsmeier. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets.PLOS ONE, 10(3): e0118432, 2015. doi: 10.1371/journal.pone.0118432

  21. [21]

    Forecasting and trading cryptocurrencies with machine learning under changing market conditions.Financial Innovation, 7(1):3, 2021

    Helder Sebastião and Pedro Godinho. Forecasting and trading cryptocurrencies with machine learning under changing market conditions.Financial Innovation, 7(1):3, 2021. doi: 10.1186/ s40854-020-00217-x

  22. [22]

    Data-snooping, technical trading rule performance, and the bootstrap.The Journal of Finance, 54(5):1647–1691, 1999

    Ryan Sullivan, Allan Timmermann, and Halbert White. Data-snooping, technical trading rule performance, and the bootstrap.The Journal of Finance, 54(5):1647–1691, 1999

  23. [23]

    A reality check for data snooping.Econometrica, 68(5):1097–1126, 2000

    Halbert White. A reality check for data snooping.Econometrica, 68(5):1097–1126, 2000. doi: 10.1111/1468-0262.00152

  24. [24]

    Agentic trading: When LLM agents meet financial markets

    Yihan Xia, Panpan You, Taotao Wang, Fang Liu, Han Qi, Xiaoxiao Wu, and Shengli Zhang. Agentic trading: When LLM agents meet financial markets. arXiv:2605.19337,https://arxi v.org/abs/2605.19337, 2026. Version 1

  25. [25]

    Winker, Rakesh Aggarwal, Lorraine E

    Chris Zielinski, Margaret A. Winker, Rakesh Aggarwal, Lorraine E. Ferris, Markus Heinemann, José Florencio Lapeña, Sanjay A. Pai, Edsel Ing, Leslie Citrome, Murad Alam, Michael Voight, Farrokh Habibzadeh, and on behalf of the WAME Board. Chatbots, generative AI, and scholarly manuscripts: WAME recommendations on chatbots and generative artificial intellig...

  26. [2020]

    doi: 10.1007/s10994-020-05910-7