Pith. sign in

REVIEW 3 major objections 5 minor 88 references

CTBench: Cryptocurrency Time Series Generation Benchmark

T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read CTBench claims that synthetic crypto time series should be judged by economic utility, and finds no single TSG model dominates across forecasting and trading.

desk verdict A useful crypto TSG benchmark with a real evaluation protocol, but a survivor-biased token universe and missing artifacts make the headline rankings conditional. read the letter →

arxiv 2508.02758 v1 pith:T72WOA5F submitted 2025-08-03 q-fin.ST cs.AIcs.CEcs.DBcs.LG

classification q-fin.STcs.AIcs.CEcs.DBcs.LG
keywords cryptocurrencytimeseriesgenerationbenchmarkpredictiveutilitystatisticalarbitragegenerativemodelsmarketregimesfinancialevaluationmetrics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

CTBench is presented as the first benchmark for time series generation tailored to cryptocurrency markets. The paper argues that synthetic crypto data should be judged not just by statistical fidelity but by whether it supports forecasting and trading, and it builds an open dataset of 452 Binance USDT pairs with hourly returns from 2020 to 2024. Eight TSG models from five methodological families are scored across thirteen metrics under three trading strategies plus a statistical arbitrage task, spanning four distinct market regimes. The central empirical finding is that no single model dominates: Diffusion-TS achieves the best forecasting accuracy but lags in trading profitability, while TimeVAE and COSCI-GAN shine in specific regimes. If this is right, the benchmark gives practitioners a reusable protocol for choosing a generator based on market regime and strategy rather than fidelity alone.

What carries the argument

The load-bearing object is the dual-task evaluation protocol: Predictive Utility, where generated log-returns are featurized with Alpha101 factors and technical indicators, used to train a forecasting model, and then scored by trading a dollar-neutral long–short portfolio on real test data; and Statistical Arbitrage, where the trained TSG model reconstructs the test set and the residual time series are fitted to an Ornstein–Uhlenbeck process whose s-scores generate hourly mean-reverting positions. The benchmark also defines three canonical strategies—cross-sectional momentum, long-only top-quantile, and proportional weighting—and evaluates eleven financial metrics spanning error, rank, trading, risk, and efficiency, plus visualization. This machinery converts the question of how realistic a synthetic series is into the question of how much economic value it unlocks, which is what lets the authors compare model families on equal footing.

What would settle it

Run the same eight models under the same dual-task protocol on the full Binance USDT listing history, including tokens that listed or delisted inside the window with missing hours treated as gaps, and check whether Diffusion-TS still ranks first on forecasting metrics and whether TimeVAE and COSCI-GAN still dominate trading metrics; a material change in rankings would show that the curated universe drove the conclusions.

Watch

Extended reading notes

Core claim

The central claim is that synthetic-data quality in crypto cannot be equated with reconstruction or prediction error; the dual-task evaluation reveals a systematic gap between statistical fidelity and economic utility. On the Predictive Utility task, synthetic series are used to train an XGBoost forecaster that is traded cross-sectionally, and on the Statistical Arbitrage task, residuals between real and reconstructed returns are fit to an Ornstein–Uhlenbeck process and converted into mean-reverting trading signals. Across 2021–2024, Diffusion-TS ranks at or near the top on MSE, MAE, IC and IR but produces negative or weak CAGR under several strategies, whereas TimeVAE and COSCI-GAN generate strong risk-adjusted returns in their favored regimes, and Fourier-Flow is described as an all-weather but conservative choice. The paper therefore argues that model selection should be regime-aware and strategy-aware, matching a generator's inductive bias to the target alpha source rather than chasing fidelity.

Load-bearing premise

All results depend on the curation filter that keeps only tokens with complete hourly observations from 2020 to 2024; because late-listed and delisted coins are dropped, the benchmark measures generators on a survivorship-biased slice of the market.

Editorial extensions

If this is right

  • Practitioners choosing a TSG model for crypto should diagnose the intended market regime first, because a model that leads on fidelity can still lose money when traded.
  • Forecasting-error rankings should not be used as a proxy for trading viability; the benchmark's split between Predictive Utility and Statistical Arbitrage separates these questions explicitly.
  • Fee-sensitive deployments should prefer low-turnover generators such as TimeVAE and Diffusion-TS, since the paper shows ranking compression and Sharpe erosion for high-turnover models when a 0.03% fee is applied.
  • The dataset and rolling-window protocol provide a reusable substrate for future crypto TSG research, including tokens beyond the 452 that pass the curation filter.
  • No single generator is universally best; regime-specific recommendations, such as COSCI-GAN for trend-following, TimeVAE for mean-reverting markets, and FIDE for defensive risk control, follow directly from the results.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the survivorship-biased sample were replaced by the full listing history including delisted coins, the model rankings could plausibly change, so the benchmark's conclusions should be read as conditional on the curated 452-token universe.
  • The Statistical Arbitrage comparison excludes GAN models by design, so its rankings cover six models rather than eight; the model-family comparisons are therefore not uniform across the two tasks.
  • The same dual-task protocol could be transplanted to other 24/7 fragmented markets, such as tokenized equities or perpetual futures, with minimal changes provided an exchange feed and a mean-reverting residual process are available.
  • An immediate testable extension is to use generated series for stress testing: feed synthetic crash regimes into the risk metrics and check whether the generator's VaR and ES bracket the realized 2022 drawdown, which the paper does not report.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. CTBench proposes a benchmark for time-series generation in cryptocurrency markets. It curates hourly Binance USDT data for 452 tokens over 2020-2024, defines two evaluation tasks (Predictive Utility, where synthetic data train an XGBoost forecaster whose predictions drive portfolios, and Statistical Arbitrage, where OU-fitted residuals from reconstructed series generate mean-reversion signals), and evaluates eight TSG models across three strategies and multiple financial and statistical metrics. The paper reports walk-forward results by year and market regime and concludes that no single TSG family dominates: Diffusion-TS is best on fidelity but poor on trading, TimeVAE and COSCI-GAN are regime-dependent, and Fourier-Flow is a robust all-weather baseline.

Significance. If the experimental results are reliable, CTBench is a genuinely useful resource: it extends TSG benchmarking to a domain where 24/7 trading and fat tails matter, links generation quality to economic outcomes rather than statistical distance alone, and grounds conclusions in walk-forward splits with real-data and PCA baselines. The dual-task design is a clear step beyond pure fidelity benchmarks, and the paper is explicit about strategy, feature, and fee choices, which makes the protocol reproducible in principle. The main significance depends on the robustness of the rankings, which currently is not established.

major comments (3)
  1. [§3.1 and §4.2-§4.5] The asset-universe filter in §3.1 removes all tokens with missing observations over January 2020-December 2024, so the 452-token panel consists only of coins with continuous Binance USDT listing history throughout the window. Late-listed, delisted, and suspended tokens, which include much of the high-volatility, illiquid tail that the Introduction cites as crypto-specific, are systematically excluded. Because every ranking, regime comparison, and recommendation in §4.2-§4.5 and Table 3 is computed on this survivor universe, the paper's concluding claim of a benchmark for 'cryptocurrency markets' overstates the generalization. Please either restrict the scope explicitly to continuously listed USDT pairs or add a robustness study on the full listing history (for example, a time-varying asset universe with missing returns handled explicitly) and discuss how the rankings change.
  2. [§4.1 and §4.2-§4.5] No seed variation, confidence intervals, or error bars are reported for any experiment. Claims such as 'Diffusion-TS consistently ranks highest in forecasting metrics but lags in trading performance' and 'TimeVAE and COSCI-GAN exhibit regime-dependent strengths' are based on point estimates from what appears to be a single training/evaluation run. Since all TSG models are stochastic, a few seeds with mean and standard deviation, or rank distributions, are needed to establish that the observed trade-offs are not noise; for a benchmark meant to guide model selection, this uncertainty is load-bearing.
  3. [§3.1 and §5] The paper repeatedly calls the dataset and benchmark 'open-source' and 'publicly available,' but no repository URL, dataset link, or artifact identifier appears anywhere in the manuscript. For a benchmark paper, the artifact is the primary contribution; without a link, the selection rule, preprocessing choices, and all reported numbers are not independently verifiable. Please include a persistent link (for example, a GitHub repository and a Zenodo or figshare DOI) and, ideally, a reproducibility checklist covering data access, model configurations, and evaluation code.
minor comments (5)
  1. [Abstract, §3.4, §4.1] The number of evaluation metrics is inconsistent: the abstract and the §3 module overview say 13 metrics, §3.4 says 11 metrics and defines E1-E11, and §4.1 says 12 metrics. Please reconcile these counts so the metric list matches the abstract and the experimental setup exactly.
  2. [§3.4, Eq. (E5)] In the CAGR formula, the symbol s is used both as the test-step length in §2.1 and as the length of the equity series; this is ambiguous because CAGRs can be computed per split or across splits. Please introduce a separate notation for the backtest horizon and clarify whether the reported CAGR is averaged over splits or pooled.
  3. [§3.2.2] The mean-reversion threshold gamma=2 and the per-asset OU parameters are taken as fixed defaults. A short sensitivity analysis over gamma and over the OU estimation window would clarify whether the Statistical Arbitrage rankings are robust to these choices.
  4. [§3.5 and §4.3] The footnote states that GAN-based methods are used only in the forecasting task because they do not natively support reconstruction, so the dual-task comparison deliberately has different model sets. Please state this asymmetry explicitly in the §4.3 discussion and when comparing the two tasks, since it prevents a fully symmetric model-ranking conclusion.
  5. [§4.4 / Figure 14] In Figure 14, the time-axis labels '1 ms. 1 s. 1 min. 1 hour' are ambiguous about whether the axis is log-spaced or ordinal. Please clarify the scale, units, and how inference time is averaged over batches.

Circularity Check

0 steps flagged · score 1.0 of 10

No circular reduction: the benchmark's rankings are empirical results on held-out data, and the cited TSGBench basis is not load-bearing.

full rationale

The paper's central claims (model rankings, regime trade-offs, and deployment guidance) are produced by an actual experiment: TSG models are trained on rolling-window training returns and evaluated on held-out test returns through forecasting accuracy, rank fidelity, trading P&L, risk metrics, and efficiency. There is no equation-level reduction of a prediction to an input. The Statistical Arbitrage task fits OU parameters to training residuals and applies them to test residuals, which is a legitimate train/test split rather than a circular construction. The dataset curation filter (assets with no missing observations over 2020-2024) creates a survivorship-biased universe, but that is a scope and external-validity limitation, not a circularity. The paper does cite the authors' TSGBench [3] as the basis for the model-based evaluation paradigm and cites its own TSGAssist [2], but those citations are not load-bearing: the benchmark implementation, XGBoost forecaster, Alpha101 features, trading strategies, and all metrics are specified in the paper and run against real Binance data. No uniqueness theorem is imported, no fitted parameter is renamed as a prediction, and no known result is merely relabeled. Therefore no circular step is exhibited; the score reflects at most a minor self-citation that does not support the paper's conclusions.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The benchmark's results depend on several hand-chosen thresholds, a specific residual model, and a curated survivor-biased dataset. None of these is a fitted constant in a derivation; they are evaluation design choices. The paper does not introduce new physical or mathematical entities.

free parameters (5)
  • OU s-score threshold gamma = 2
    Chosen by hand in §3.2.2 to trigger long/short positions in the Statistical Arbitrage task; affects trade frequency and profitability.
  • OU per-asset parameters theta, mu, sigma = estimated per asset from training residuals
    Fitted on training residuals to define s-scores for test residuals; these are model parameters, not universal constants.
  • Training window w and test step s = w=500x24h, s=30x24h (Predictive Utility) or 15x24h (Statistical Arbitrage)
    Rolling-window design choices in §4.1 that determine the number of splits and evaluation length.
  • Trading fee assumption = 0% default; 0.03% for Statistical Arbitrage
    Fee setting in §4.1; the zero-fee default can inflate profitability metrics and is an explicit modeling choice.
  • XGBoost forecasting hyperparameters = not reported
    The forecaster hyperparameters are said to require minimal tuning but are not specified, so the Predictive Utility results depend on unstated settings.
assumptions (5)
  • domain assumption Reconstruction residuals from TSG models follow an Ornstein-Uhlenbeck mean-reverting process.
    The Statistical Arbitrage task models residuals as OU and uses training-residual parameters to generate test s-scores (§3.2.2). If residuals are not OU, the trading signals are misspecified.
  • domain assumption A forecasting model trained only on synthetic returns transfers to real market returns.
    The Predictive Utility task assumes XGBoost trained on generated features can predict real returns (§3.2.1). A negative result could reflect transfer failure rather than poor generation quality.
  • domain assumption Alpha101 factors and technical indicators are meaningful when computed on synthetic returns.
    Feature extraction applies quant factors to generated returns (§3.1); this assumes these features retain signal on synthetic data.
  • domain assumption The Binance USDT spot universe with complete 2020-2024 histories represents the crypto market.
    Dataset scope is restricted to Binance USDT pairs after filtering missing observations (§3.1); other venues and non-surviving tokens may behave differently.
  • domain assumption Calendar years 2021-2024 correspond to distinct market regimes.
    The paper labels 2021-2024 as bull, crash, consolidation, and mean-reverting regimes (§4.2) but does not statistically identify regimes; results may depend on this periodization.

how reviews work

0 comments
Cite this review

Pith. "Pith review of CTBench: Cryptocurrency Time Series Generation Benchmark." pith.science (2026). https://pith.science/paper/T72WOA5F

@misc{pith2026250802758,
  author       = {Pith},
  title        = {Pith review of: CTBench: Cryptocurrency Time Series Generation Benchmark},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/T72WOA5F}},
  note         = {Machine review of arXiv:2508.02758}
}
read the original abstract

Synthetic time series are essential tools for data augmentation, stress testing, and algorithmic prototyping in quantitative finance. However, in cryptocurrency markets, characterized by 24/7 trading, extreme volatility, and rapid regime shifts, existing Time Series Generation (TSG) methods and benchmarks often fall short, jeopardizing practical utility. Most prior work (1) targets non-financial or traditional financial domains, (2) focuses narrowly on classification and forecasting while neglecting crypto-specific complexities, and (3) lacks critical financial evaluations, particularly for trading applications. To address these gaps, we introduce \textsf{CTBench}, the first comprehensive TSG benchmark tailored for the cryptocurrency domain. \textsf{CTBench} curates an open-source dataset from 452 tokens and evaluates TSG models across 13 metrics spanning 5 key dimensions: forecasting accuracy, rank fidelity, trading performance, risk assessment, and computational efficiency. A key innovation is a dual-task evaluation framework: (1) the \emph{Predictive Utility} task measures how well synthetic data preserves temporal and cross-sectional patterns for forecasting, while (2) the \emph{Statistical Arbitrage} task assesses whether reconstructed series support mean-reverting signals for trading. We benchmark eight representative models from five methodological families over four distinct market regimes, uncovering trade-offs between statistical fidelity and real-world profitability. Notably, \textsf{CTBench} offers model ranking analysis and actionable guidance for selecting and deploying TSG models in crypto analytics and strategy development.

Figures

Figures reproduced from arXiv: 2508.02758 by the authors.

Figure 1
Figure 1. TSG model rankings on the Predictive Utility (left) [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Overall architecture of CTBench. resulting dataset comprises 452 unique cryptocurrencies, offering a robust foundation for TSG benchmarking in crypto markets. Formally, let 𝑛 denote the number of tradable crypto assets and (𝑙 + 1) the number of hourly observations after data filtering. We index assets by 1 ≤ 𝑖 ≤ 𝑛 and timestamps by 0 ≤ 𝑡 ≤ 𝑙. For each asset and timestamp pair (𝑖, 𝑡), we record the five standard fiel… view at source ↗
Figure 3
Figure 3. Histograms of the mean hourly log-return (%) (left) [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (10 more)
Figure 4
Figure 4. Figure 4: Line plots of closing returns for representative cryptocurrencies, with large-cap examples (top row), mid-cap examples [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 6
Figure 6. Figure 6: Architectures of dual-task benchmarks. asset 𝑖 and time 𝑡, we define training residual: 𝜌𝑖,𝑡 = 𝑟𝑖,𝑡 − 𝑟ˆ𝑖,𝑡, where 𝑟𝑖,𝑡 ∈ 𝑹train and 𝑟ˆ𝑖,𝑡 ∈ 𝑹ˆ train. For each asset 𝑖, these residuals are then fitted to an Ornstein–Uhlenbeck (OU) process [63]: 𝑑𝜌𝑖,𝑡 = 𝜃𝑖(𝜇𝑖 − 𝜌𝑖,𝑡)𝑑𝑡 …
Figure 7
Figure 7. Figure 7: Annual forecasting performance of TSG methods on the Predictive Utility task. [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 8
Figure 8. Figure 8: Annual trading performance of TSG methods on the Predictive Utility task. [PITH_FULL_IMAGE:figures/full_fig_p009_8.png]
Figure 9
Figure 9. Figure 9: Rankings of TSG models on the Predictive Utility task. [PITH_FULL_IMAGE:figures/full_fig_p010_9.png]
Figure 10
Figure 10. Figure 10: Simulated growth curves of a $10,000 investment over four years under three trading strategies. [PITH_FULL_IMAGE:figures/full_fig_p010_10.png]
Figure 11
Figure 11. Figure 11: Annual performance of TSG methods on the Statistical Arbitrage task. [PITH_FULL_IMAGE:figures/full_fig_p011_11.png]
Figure 12
Figure 12. Figure 12: Rankings of TSG models on the Statistical Arbitrage task. [PITH_FULL_IMAGE:figures/full_fig_p011_12.png]
Figure 13
Figure 13. Figure 13: Simulated growth curves of a $10,000 investment [PITH_FULL_IMAGE:figures/full_fig_p011_13.png]
Figure 14
Figure 14. Figure 14: Training and inference time of TSG methods. [PITH_FULL_IMAGE:figures/full_fig_p012_14.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

88 extracted references · 66 canonical work pages

  1. [1]

    Alaa, Alex James Chan, and Mihaela van der Schaar

    Ahmed M. Alaa, Alex James Chan, and Mihaela van der Schaar. 2021. Generative Time-series Modeling with Fourier Flows. In ICLR

  2. [2]

    Yihao Ang, Yifan Bao, Qiang Huang, Anthony KH Tung, and Zhiyong Huang

  3. [3]

    Yihao Ang, Qiang Huang, Yifan Bao, Anthony KH Tung, and Zhiyong Huang

  4. [4]

    Yihao Ang, Qiang Huang, Anthony KH Tung, and Zhiyong Huang. 2023. A Stitch in Time Saves Nine: Enabling Early Anomaly Detection with Correlation Analysis. In ICDE. 1832–1845

  5. [5]

    Yifan Bao, Yihao Ang, Qiang Huang, Anthony KH Tung, and Zhiyong Huang

  6. [6]

    Binance Exchange. 2025. Binance Exchange. https://binance.com/. Accessed: 1 March 2025

  7. [7]

    Binance Exchange. 2025. Trading Fee Schedule. https://www.binance.com/en/ fee/schedule

  8. [8]

    Towards controllable time series generation.arXiv preprint arXiv:2403.03698 (2024)

Show all 88 references
  1. [9]

    Tianqi Chen and Carlos Guestrin. 2016. Xgboost: A scalable tree boosting system. In KDD. 785–794

  2. [10]

    Tian Qi Chen, Yulia Rubanova, Jesse Bettencourt, and David Duvenaud. 2018. Neural Ordinary Differential Equations. In NeurIPS. 6572–6583

  3. [11]

    Ruichu Cai, Jiawei Chen, Zijian Li, Wei Chen, Keli Zhang, Junjian Ye, Zhuozhang Li, Xiaoyan Yang, and Zhenjie Zhang. 2021. Time Series Domain Adaptation via Sparse Associative Structure Alignment. In AAAI. 6859–6867

  4. [12]

    Andrea Coletta, Sriram Gopalakrishnan, Daniel Borrajo, and Svitlana Vyetrenko

  5. [13]

    Brubaker, Greg Mori, and Andreas M

    Ruizhi Deng, Bo Chang, Marcus A. Brubaker, Greg Mori, and Andreas M. Lehrmann. 2020. Modeling Continuous Stochastic Processes with Dynamic Normalizing Flows. In NeurIPS. 7805–7815

  6. [14]

    Zhicheng Chen, FENG SHIBO, Zhong Zhang, Xi Xiao, Xingyu Gao, and Peilin Zhao. 2024. Sdformer: Similarity-driven discrete transformer for time series generation. In NeurIPS. 132179–132207

  7. [15]

    McAuley, and Miller S

    Chris Donahue, Julian J. McAuley, and Miller S. Puckette. 2019. Adversarial Audio Synthesis. In ICLR

  8. [16]

    In NeurIPS

    On the constrained time-series generation problem. In NeurIPS. 61048– 61059

  9. [17]

    Cristóbal Esteban, Stephanie L Hyland, and Gunnar Rätsch. 2017. Real-valued (medical) time series generation with recurrent conditional gans. arXiv preprint arXiv:1706.02633 (2017)

  10. [18]

    Abhyuday Desai, Cynthia Freeman, Zuhui Wang, and Ian Beaver. 2021. TimeVAE: A Variational Auto-Encoder for Multivariate Time Series Generation. arXiv preprint arXiv:2111.08095 (2021)

  11. [19]

    Asadullah Hill Galib, Pang-Ning Tan, and Lifeng Luo. 2024. FIDE: Frequency- Inflated Conditional Diffusion Model for Extreme-Aware Time Series Generation. In NeurIPS

  12. [20]

    Vincent Dumoulin, Ishmael Belghazi, Ben Poole, Alex Lamb, Martin Arjovsky, Olivier Mastropietro, and Aaron Courville. 2017. Adversarially Learned Inference. In ICLR

  13. [21]

    Hu et al

    Y. Hu et al . 2025. FinTSB: A Comprehensive Benchmark for Financial Time Series Forecasting. In arXiv:2502.18834

  14. [22]

    Tsukasa Fujiwara and Hiroshi Kunita. 1985. Stochastic differential equations of jump type and Lévy processes in diffeomorphisms group.Journal of mathematics of Kyoto University 25, 1 (1985), 71–106

  15. [23]

    Daniel Jarrett, Ioana Bica, and Mihaela van der Schaar. 2021. Time-series Gener- ation by Contrastive Imitation. In NeurIPS. 28968–28982

  16. [24]

    Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2020. Generative adversarial networks. Commun. ACM 63, 11 (2020), 139–144

  17. [25]

    Jinsung Jeon, Jeonghak Kim, Haryong Song, Seunghyeon Cho, and Noseong Park. 2022. GT-GAN: General Purpose Time Series Synthesis with Generative Adversarial Networks. In NeurIPS. 36999–37010

  18. [26]

    Yang Hu, Xiao Wang, Lirong Wu, Huatian Zhang, Stan Z Li, Sheng Wang, and Tianlong Chen. 2024. FM-TS: Flow Matching for Time Series Generation. arXiv preprint arXiv:2411.07506 (2024)

  19. [27]

    Zura Kakushadze. 2016. 101 formulaic alphas. Wilmott 2016, 84 (2016), 72–81

  20. [28]

    Paul Jeha, Michael Bohlke-Schneider, Pedro Mercado, Shubham Kapoor, Ra- jbir Singh Nirwan, Valentin Flunkert, Jan Gasthaus, and Tim Januschowski. 2021. PSA-GAN: Progressive self attention GANs for synthetic time series. In ICLR

  21. [29]

    Patrick Kidger, James Foster, Xuechen Li, and Terry J Lyons. 2021. Neural SDEs as Infinite-Dimensional GANs. In ICML. 5453–5463

  22. [30]

    James Jordon, Jinsung Yoon, and Mihaela Van Der Schaar. 2018. PATE-GAN: Generating synthetic data with differential privacy guarantees. In ICLR

  23. [31]

    Hongming Li, Shujian Yu, and Jose Principe. 2023. Causal Recurrent Variational Autoencoder for Medical Time Series Generation. In AAAI. 8562–8570

  24. [32]

    Zura Kakushadze. 2016. 101 Formulaic Alphas. Wilmott Magazine 84 (2016), 72–

  25. [33]

    Yuening Li, Zhengzhang Chen, Daochen Zha, Mengnan Du, Jingchao Ni, Denghui Zhang, Haifeng Chen, and Xia Hu. 2022. Towards learning disentangled repre- sentations for time series. In KDD. 3270–3278

  26. [35]

    Daesoo Lee, Sara Malacarne, and Erlend Aune. 2023. Vector Quantized Time Series Generation with a Bidirectional Prior Model. In AISTATS. 7665–7693

  27. [36]

    Haksoo Lim, Minjung Kim, Sewon Park, and Noseong Park. 2023. Regular Time-series Generation using SGM. arXiv preprint arXiv:2301.08518 (2023)

  28. [37]

    Xiaomin Li, Vangelis Metsis, Huangyingrui Wang, and Anne Hee Hiong Ngu

  29. [38]

    Guang Liu, Yuzhao Mao, Qi Sun, Hailong Huang, Weiguo Gao, Xuan Li, Jianping Shen, Ruifan Li, and Xiaojie Wang. 2021. Multi-scale two-way deep neural network for stock trend prediction. In IJCAI. 4555–4561

  30. [39]

    Yuansan Liu, Sudanthi Wijewickrema, Ang Li, and James Bailey. 2022. Time- Transformer AAE: Connecting Temporal Convolutional Networks and Trans- former for Time Series Generation. (2022)

  31. [40]

    Olof Mogren. 2016. C-RNN-GAN: A continuous recurrent neural network with adversarial training. In Constructive Machine Learning Workshop (CML) at NIPS

  32. [41]

    Urnes, and Haipeng Chen

    Yang Li, Han Meng, Zhenyu Bi, Ingolv T. Urnes, and Haipeng Chen. 2025. Popu- lation Aware Diffusion for Time Series Generation. In AAAI. 18520–18529

  33. [42]

    Ilan Naiman, Nimrod Berman, Itai Pemper, Idan Arbiv, Gal Fadlon, and Omri Azencot. 2024. Utilizing image transforms and diffusion models for generative modeling of short and long time series. In NeurIPS, Vol. 37. 121699–121730

  34. [43]

    Zinan Lin, Alankar Jain, Chen Wang, Giulia Fanti, and Vyas Sekar. 2020. Using GANs for Sharing Networked Time Series Data: Challenges, Initial Promise, and Open Questions. In IMC. 464–483

  35. [44]

    Hao Ni, Lukasz Szpruch, Marc Sabate-Vidales, Baoren Xiao, Magnus Wiese, and Shujian Liao. 2021. Sig-Wasserstein GANs for time series generation. In Proceedings of the Second ACM International Conference on AI in Finance . 1–8

  36. [45]

    Hao Ni, Lukasz Szpruch, Magnus Wiese, Shujian Liao, and Baoren Xiao. 2020. Conditional Sig-Wasserstein GANs for Time Series Generation. arXiv preprint arXiv:2006.05421 (2020)

  37. [46]

    Alexander Nikitin, Letizia Iannucci, and Samuel Kaski. 2023. TSGM: A Flexible Framework for Generative Modeling of Synthetic Time Series. arXiv preprint arXiv:2305.11567 (2023)

  38. [47]

    Ilan Naiman, Nimrod Berman, Itai Pemper, Idan Arbiv, Gal Fadlon, and Omri Azencot. 2024. Utilizing image transforms and diffusion models for generative modeling of short and long time series. In NeurIPS. 121699–121730

  39. [48]

    Hengzhi Pei, Kan Ren, Yuqing Yang, Chang Liu, Tao Qin, and Dongsheng Li

  40. [49]

    Ilan Naiman, N Benjamin Erichson, Pu Ren, Michael W Mahoney, and Omri Azencot. [n.d.]. Generative Modeling of Regular and Irregular Time Series Data via Koopman VAEs. In ICLR

  41. [50]

    Xiangfei Qiu, Jilin Hu, Lekui Zhou, Xingjian Wu, Junyang Du, Buang Zhang, Chenjuan Guo, Aoying Zhou, Christian S Jensen, Zhenli Sheng, et al. 2024. TFB: Towards Comprehensive and Fair Benchmarking of Time Series Forecasting Methods. Proceedings of the VLDB Endowment 17, 9 (202...

  42. [51]

    Giorgia Ramponi, Pavlos Protopapas, Marco Brambilla, and Ryan Janssen. 2018. T-CGAN: Conditional Generative Adversarial Network for Data Augmentation in Noisy Time Series with Irregular Sampling. arXiv preprint arXiv:1811.08295 (2018)

  43. [52]

    Carl Remlinger, Joseph Mikael, and Romuald Elie. 2022. Conditional Loss and Deep Euler Scheme for Time Series Generation. In AAAI, Vol. 36. 8098–8105

  44. [53]

    YongKyung Oh, Dongyoung Lim, and Sungil Kim. 2024. Stable Neural Stochastic Differential Equations in Analyzing Irregular Time Series Data. In The Twelfth International Conference on Learning Representations

  45. [54]

    C Grinold Richard and Ronald Kahn. 2000. Active Portfolio Management: A Quantitative Approach for Producing Superior Returns and Controlling Risk

  46. [55]

    Yulia Rubanova, Ricky T. Q. Chen, and David K Duvenaud. 2019. Latent Ordinary Differential Equations for Irregularly-Sampled Time Series. In NeurIPS. 5320– 5330

  47. [56]

    Jian Qian, Bingyu Xie, Biao Wan, Minhao Li, Miao Sun, and Patrick Yin Chiang

  48. [57]

    arXiv preprint arXiv:2407.04211 (2024)

    Timeldm: Latent diffusion model for unconditional time series generation. arXiv preprint arXiv:2407.04211 (2024)

  49. [58]

    Padmanaba Srinivasan and William J Knottenbelt. 2022. Time-series Transformer Generative Adversarial Networks. arXiv preprint arXiv:2205.11164 (2022). 13

  50. [59]

    Shuo Sun, Rundong Wang, and Bo An. 2023. Reinforcement learning for quan- titative trading. ACM Transactions on Intelligent Systems and Technology 14, 3 (2023), 1–29

  51. [60]

    Muhang Tian, Bernie Chen, Allan Guo, Shiyi Jiang, and Anru R Zhang. 2024. Reliable generation of privacy-preserving synthetic electronic health record time series via diffusion models. JAMIA 31, 11 (2024), 2529–2539

  52. [61]

    Reuters. 2025. Crypto sector breaches $4 trillion in market value during pivotal week. Reuters (July 18 2025)

  53. [62]

    Chih-Fong Tsai and Yu-Chieh Hsiao. 2010. Combining multiple feature selec- tion methods for stock prediction: Union, intersection, and multi-intersection approaches. Decision support systems 50, 1 (2010), 258–269

  54. [63]

    George E Uhlenbeck and Leonard S Ornstein. 1930. On the theory of the Brownian motion. Physical review 36, 5 (1930), 823

  55. [64]

    Ali Seyfi, Jean-François Rajotte, and Raymond T. Ng. 2022. Generating multi- variate time series with COmmon Source CoordInated GAN (COSCI-GAN). In NeurIPS. 32777–32788

  56. [65]

    Kaleb E Smith and Anthony O Smith. 2020. Conditional GAN for timeseries generation. arXiv preprint arXiv:2006.16477 (2020)

  57. [66]

    Lei Wang, Liang Zeng, and Jian Li. 2023. AEC-GAN: Adversarial Error Correction GANs for Auto-Regressive Long Time-Series Generation. InAAAI. 10140–10148

  58. [67]

    Yuxuan Wang, Haixu Wu, Jiaxiang Dong, Yong Liu, Mingsheng Long, and Jianmin Wang. 2024. Deep time series models: A comprehensive survey and benchmark. arXiv preprint arXiv:2407.13278 (2024)

  59. [68]

    Yanlong Wang, Jian Xu, Tiantian Gao, Hongkang Zhang, Shao-Lun Huang, Danny Dongning Sun, and Xiao-Ping Zhang. 2025. FinTSBridge: A New Eval- uation Suite for Real-world Financial Prediction with Advanced Time Series Models. arXiv preprint arXiv:2503.06928 (2025)

  60. [69]

    Jack L Treynor and Fischer Black. 1973. How to use security analysis to improve portfolio selection. The journal of business 46, 1 (1973), 66–86

  61. [70]

    Julian Winkel and Wolfgang Karl Härdle. 2023. Pricing kernels and risk premia implied in bitcoin options. Risks 11, 5 (2023), 85

  62. [71]

    Tianlin Xu, Li Kevin Wenliang, Michael Munn, and Beatrice Acciaio. 2020. COT- GAN: Generating Sequential Data via Causal Optimal Transport. In NeurIPS. 8798–8809

  63. [72]

    László Vancsura, Tibor Tatay, and Tibor Bareith. 2025. Navigating AI-Driven Fi- nancial Forecasting: A Systematic Review of Current Status and Critical Research Gaps. Forecasting 7, 3 (2025), 36

  64. [73]

    Chengyu Wang, Kui Wu, Tongqing Zhou, Guang Yu, and Zhiping Cai. 2021. Tsagen: synthetic time series generation for kpi anomaly detection. IEEE Trans- actions on Network and Service Management 19, 1 (2021), 130–145

  65. [74]

    Xinyu Yuan and Yan Qiao. 2024. Diffusion-TS: Interpretable Diffusion for General Time Series Generation. In ICLR

  66. [75]

    Kyung Keun Yun, Sang Won Yoon, and Daehan Won. 2021. Prediction of stock price direction using a hybrid GA-XGBoost algorithm with a three-stage feature engineering process. Expert Systems with Applications 186 (2021), 115716

  67. [76]

    Chuheng Zhang, Yitong Duan, Xiaoyu Chen, Jianyu Chen, Jian Li, and Li Zhao

  68. [77]

    Magnus Wiese, Robert Knobloch, Ralf Korn, and Peter Kretschmer. 2020. Quant GANs: deep generation of financial time series. Quantitative Finance 20, 9 (2020), 1419–1440

  69. [78]

    Linqi Zhou, Michael Poli, Winnie Xu, Stefano Massaroli, and Stefano Ermon

  70. [79]

    Zhoufan Zhu and Ke Zhu. 2025. AlphaQCM: Alpha Discovery in Finance with Distributional Reinforcement Learning. In ICML. 14

  71. [80]

    https://doi.org/10.48550/arXiv.1601.00991 22 pages; no changes (excepting this line); to appear; also available as arXiv:1601.00991v3 [q-fin.PM]

  72. [81]

    Jinsung Yoon, Daniel Jarrett, and Mihaela van der Schaar. 2019. Time-series Generative Adversarial Networks. In NeurIPS. 5509–5519

  73. [82]

    Xinyu Yuan and Yan Qiao. [n.d.]. Diffusion-TS: Interpretable Diffusion for General Time Series Generation. In ICLR

  74. [85]

    Towards generalizable reinforcement learning for trade execution. InIJCAI. 4975–4983

  75. [86]

    Chuheng Zhang, Yuanqi Li, Xi Chen, Yifei Jin, Pingzhong Tang, and Jian Li. 2020. DoubleEnsemble: A new ensemble method based on sample reweighting and feature selection for financial data analysis. In ICDM. 781–790

  76. [88]

    Deep Latent State Space Models for Time-Series Generation. In ICML. 42625–42643

  77. [2021]

    Towards generating real-world time series data. In ICDM. 469–478

  78. [2022]

    Tts-gan: A transformer-based time-series generative adversarial network. In AIME. 133–143

  79. [2023]

    Proceedings of the VLDB Endowment 17, 3 (2023), 305–318

    TSGBench: Time Series Generation Benchmark. Proceedings of the VLDB Endowment 17, 3 (2023), 305–318

  80. [2024]

    Proceedings of the VLDB Endowment 17, 12 (2024), 4309–4312

    Tsgassist: An interactive assistant harnessing llms and rag for time se- ries generation recommendations and benchmarking. Proceedings of the VLDB Endowment 17, 12 (2024), 4309–4312

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.