REVIEW 3 major objections 4 minor 73 references
Model-free RL discovers profitable price-manipulation strategies from limited data more reliably than correctly specified models with noisy parameters, when volatility is intermediate.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-11 01:18 UTC pith:GFGYOKRT
load-bearing objection Clean finite-sample head-to-head showing DDPG can find dynamic arbitrage under nonlinear impact more reliably than a correctly-specified but estimated model-based optimizer, with a real but limited confound on the reward. the 3 major comments →
Can Reinforcement Learning Efficiently Discover Price Manipulation?
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
When permanent impact is concave, a Deep Deterministic Policy Gradient agent trained on a few hundred episodes of interaction discovers round-trip strategies with positive expected cash flow. At intermediate volatility those strategies outperform the classical route of estimating impact parameters from the same amount of data and then optimizing under the true functional form.
What carries the argument
The expected-cash objective under non-linear permanent impact f(v)= heta sign(v)|v|^δ plus linear temporary cost; this non-convex function admits dynamic arbitrage (round-trip strategies with E[X_T]>0) and is maximised either by SLSQP under full or estimated parameters or by DDPG that never sees the functional form.
Load-bearing premise
The true price process is exactly the discrete Almgren-Chriss model with the chosen power-law permanent impact; any real-market deviation from that functional form would change the relative performance of the two methods.
What would settle it
Re-run the finite-sample comparison on data generated by a different impact kernel (transient, stochastic, or empirically estimated from real meta-orders) and check whether DDPG still beats the correctly-specified-but-estimated SLSQP procedure at intermediate volatility.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper asks whether a model-free DDPG agent can discover and exploit dynamic-arbitrage (price-manipulation) opportunities arising from nonlinear permanent impact more effectively than a correctly specified but finite-sample estimated model-based optimizer. Prices follow discrete-time Almgren–Chriss dynamics with permanent impact f(v)=θ sign(v)|v|^δ (δ=0.1) and linear temporary impact. Existence of round-trip strategies with positive expected cash is proved for the two-velocity case (Lemmas 3.1–3.2) and the unrestricted problem is solved numerically by multi-start SLSQP under full information. Two finite-sample procedures are then compared on the same number of trajectories: (i) impact-parameter estimation from simulated TWAP meta-orders followed by SLSQP (regSLSQP), and (ii) direct DDPG training. Out-of-sample tests (Table 4) show that, for intermediate volatility, DDPG trained on 500 episodes yields statistically significant positive cash and outperforms both regSLSQP500 and regSLSQP5000; at high volatility no method succeeds, while at low volatility the model-based estimator regains superiority.
Significance. If the intermediate-volatility ranking is robust, the result is of genuine interest to both market-microstructure and algorithmic-trading communities: it shows that a model-free learner can recover profitable round-trip strategies from limited interaction data and can be more robust to sampling error than a correctly specified parametric procedure. The clean existence proofs, multi-start SLSQP benchmark, transparent Almgren-style regressions, and 1000-path out-of-sample design with t-tests are methodological strengths that make the comparison falsifiable. The regulatory warning about unsupervised RL discovering manipulation is timely and well-motivated by the literature cited.
major comments (3)
- [§4.1, Eq. (11) and Table 4] Section 4.1, Eq. (11): the instantaneous reward is defined as r_t = S_t v_t − κ v_t² and therefore injects the true temporary-impact coefficient κ into the agent’s objective. The paper repeatedly describes the DDPG agent as “model-free”, “agnostic” and possessing “no direct knowledge ho of the values for the impacts” (abstract, §4, §5). Because the model-based baseline must recover κ (together with θ and δ) from noisy TWAP regressions (Eqs. 8–10), the head-to-head comparison in Table 4 (especially the intermediate-vol row) does not isolate the pure effect of avoiding parameter estimation under correct specification. Either the reward must be rewritten without κ or the claim of complete model-agnosticism must be qualified and the experiments re-run.
- [§5, Table 4] Section 5 and Table 4: DDPG is trained exclusively on M=500 episodes while regSLSQP is evaluated at both N=500 and N=5000. The strongest claim—that RL “consistently outperforms the model-based approach when parameter estimates are affected by sampling error”—therefore rests on an unequal data budget. No experiment is reported in which DDPG is given a 5000-episode training set. Without that control it remains possible that the ranking simply reflects the relative sample efficiency of the two methods rather than an intrinsic advantage of model-free learning under correct specification.
- [Abstract and §5] The data-generating processes used by the two finite-sample methods are not identical. regSLSQP estimates parameters from fixed TWAP meta-orders (Section 3.3.2), whereas DDPG generates its own exploratory trajectories. The abstract and §5 assert that both methods are “trained ho on the same amount of data.” While the number of trajectories is matched, the information content is not; this asymmetry should be acknowledged and, if possible, controlled (e.g., by also estimating impact parameters from the same DDPG trajectories).
minor comments (4)
- [Introduction and §3.1] Several typographical errors appear in the main text: “preventsan accurate” (p. 4), “potentiallyunintendedmarket” (p. 4), “around-tripstrategy” (p. 4), “deltaincreases” (caption of Fig. 1). A careful proof-reading pass is needed.
- [Table 2] Table 2 formatting is hard to parse (mixed scientific notation, missing separators). Consider a cleaner layout that separates point estimates from standard errors.
- [Figs. 2 and 10] Figure 2 and Figure 10 captions refer to “the green (red) area” for positive/negative cash, but the colour coding is not explained in the main text; a short legend would help.
- [§4.2 and Table 3] The soft-update parameter, learning-rate schedule and exploration-noise decay are given only in footnotes or the algorithm box; collecting them in Table 3 would improve reproducibility.
Circularity Check
Empirical OOS simulation comparison under known DGP; reward, estimators and cash are independently defined with no reduction of reported outperformance to a fitted constant by construction.
full rationale
The paper's load-bearing claims are (i) existence of positive-E[X_T] round-trips under the stated discrete-time AC dynamics with concave permanent impact (Prop. 2.1, Lemmas 3.1-3.2, SLSQP benchmark) and (ii) finite-sample head-to-head of DDPG vs. regression+SLSQP on held-out paths (Table 4). Existence follows by direct expansion of the cash functional (Eq. 4) and elementary calculus under the model definition; it is not circular. Parameter estimation (Eqs. 8-10) is ordinary least-squares on independent TWAP meta-orders; the subsequent SLSQP step and OOS cash evaluation use the true DGP, so sampling error is measured rather than assumed away. DDPG trains on the same environment with an instantaneous reward that matches the temporary-impact term of the cash expansion; the permanent-impact channel is observed only through the price process. The final numbers in Table 4 are Monte-Carlo averages of true cash under the learned policies; they are not algebraically forced by any fitted constant or self-referential definition. Self-citations (Macrì-Lillo 2025, Macrì et al. 2025) supply only prior RL-execution context and are not used to justify uniqueness, existence, or the performance ranking. Minor design asymmetries (κ known inside the reward, unequal episode budgets) affect fairness of the comparison but do not create circularity of derivation. Score 1 reflects only the presence of non-load-bearing self-citations; the central empirical claim remains independent.
Axiom & Free-Parameter Ledger
free parameters (6)
- permanent-impact exponent δ =
0.1
- permanent-impact coefficient θ =
0.0001
- temporary-impact coefficient κ =
0.0003
- volatility levels σ =
0.0168 / 0.003 / 0.002
- horizon T and action bound b =
T=21, b=15
- DDPG training budget and architecture =
M=500, B=128, lr=0.1
axioms (4)
- domain assumption Price follows discrete-time Almgren-Chriss dynamics with permanent impact f(v)=θ sign(v)|v|^δ and linear temporary impact (Eqs. 1–3).
- domain assumption A risk-neutral agent maximises expected terminal cash subject to exact round-trip inventory constraint q0 = qT = 0.
- standard math For δ < 1 the expected-cash objective is continuous and coercive, hence a maximiser exists (Prop. 2.1).
- domain assumption Impact parameters can be recovered from TWAP meta-order regressions of the form given by Almgren et al. (2005) (Eqs. 8–10).
read the original abstract
In this paper, we investigate whether a model-free RL agent can identify and exploit price manipulation opportunities more effectively than a traditional model-based approach that assumes correct specification of the data-generating process but relies on noisy parameter estimates. We consider a single-asset market in which prices evolve according to an Almgren-Chriss framework with non-linear permanent impact and linear temporary impact. We first establish the existence of price-manipulative strategies in discrete time and compute the optimal benchmark strategy using Sequential Least Squares Quadratic Programming under full information. We then compare two finite-sample learning approaches: a model-based procedure that estimates impact parameters from simulated execution data and an agnostic RL approach based on Deep Deterministic Policy Gradient, trained directly on the same amount of data. For intermediate volatility, the RL agent successfully discovers profitable manipulative strategies without explicit knowledge of the underlying model, even when training data are quite limited. More importantly, RL consistently outperforms the model-based approach when parameter estimates are affected by sampling error, despite the latter benefiting from the correct model specification. For large volatility, all methods are unable to identify manipulation opportunities, while for small volatility, the model based approach outperforms RL. These findings highlight both the effectiveness of RL in complex control problems and the risks associated with deploying learning algorithms in financial markets without appropriate safeguards.
Figures
Reference graph
Works this paper leans on
-
[1]
Gianbiagio Curato and Jim Gatheral and Fabrizio Lillo , title =. Quantitative Finance , year =
-
[2]
Applied Mathematical Finance , volume =
Brian Ning and Franco Ho Ting Lin and Sebastian Jaimungal , title =. Applied Mathematical Finance , volume =
-
[3]
Applied Mathematical Finance , volume =
Andrea Macrì and Fabrizio Lillo , title =. Applied Mathematical Finance , volume =
-
[4]
Quantitative Finance , volume =
Aurélien Alfonsi and Antje Fruth and Alexander Schied , title =. Quantitative Finance , volume =
- [5]
- [6]
-
[7]
Olivier Gueant , title =
-
[8]
Proceedings of the 31st International Conference on Machine Learning , pages =
Deterministic Policy Gradient Algorithms , author =. Proceedings of the 31st International Conference on Machine Learning , pages =. 2014 , editor =
work page 2014
-
[9]
Proceedings of the 4th International Conference on Learning Representations , year =
Continuous Control with Deep Reinforcement Learning , author =. Proceedings of the 4th International Conference on Learning Representations , year =
-
[10]
Bergault, Philippe and Drissi, Fay. Multi-asset Optimal Execution and Statistical Arbitrage Strategies under Ornstein--Uhlenbeck Dynamics , journal =
-
[11]
Journal of Financial Markets , volume =
Dimitris Bertsimas and Andrew Lo , title =. Journal of Financial Markets , volume =
- [12]
-
[13]
Available at SSRN: https://ssrn.com/abstract=4639959 , year =
Alvaro Cartea and Patrick Chang and Gabriel Garcıa-Arenas , title =. Available at SSRN: https://ssrn.com/abstract=4639959 , year =
-
[14]
Almgren, Robert and Thum, Chee and Hauptmann, Emmanuel and Li, Hong , title =. Risk , volume =
-
[15]
T.G. Andersen and T. Bollerslev and F.X. Diebold and P. Labys , title =. Econometrica , year =
-
[16]
Zhelezniak, Vitalii and Savkov, Aleksandar and Shen, April and Hammerla, Nils , title =. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , year =
work page 2019
-
[17]
Pang-Ning Tan and Michael Steinbach and Vipin Kumar , title =
-
[18]
Finance and Stochastics , year =
Alexander Schied and Torsten Schöneborn , title =. Finance and Stochastics , year =
-
[19]
Bouchaud, Jean-Philippe and Gefen, Yuval and Potters, Marc and Wyart, Matthieu , title =. Quantitative Finance , year =
-
[20]
Doyne and Gerig, Austin and Lillo, Fabrizio and Mike, Szabolcs , title =
Farmer, J. Doyne and Gerig, Austin and Lillo, Fabrizio and Mike, Szabolcs , title =. Quantitative Finance , year =
-
[21]
Journal of Financial Markets , year =
Obizhaeva, Anna and Wang, Jiang , title =. Journal of Financial Markets , year =
-
[22]
Gatheral, Jim and Schied, Alexander and Slynko, Alla , title =. Mathematical Finance , year =
-
[23]
Proceedings of the 23rd International Conference on Machine Learning (ICML) , year =
Nevmyvaka, Yuriy and Feng, Yi and Kearns, Michael , title =. Proceedings of the 23rd International Conference on Machine Learning (ICML) , year =
-
[24]
Hendricks, Dieter and Wilcox, Diane , title =. 2014 IEEE Conference on Computational Intelligence for Financial Engineering & Economics (CIFEr) , pages =. 2014 , publisher =
work page 2014
-
[25]
IEEE Transactions on Neural Networks , year =
Moody, John and Saffell, Matthew , title =. IEEE Transactions on Neural Networks , year =
-
[28]
Quantitative finance , volume=
Fluctuations and response in financial markets: the subtle nature of `random' price changes , author=. Quantitative finance , volume=
-
[29]
Finance and Stochastics , volume=
Reinforcement learning and stochastic optimisation , author=. Finance and Stochastics , volume=. 2022 , publisher=
work page 2022
-
[31]
Kraft, Dieter , title =
-
[32]
Measuring price impact and information content of trades in a time-varying setting
Measuring price impact and information content of trades in a time-varying setting , author=. arXiv preprint arXiv:2212.12687 , year=
work page internal anchor Pith review Pith/arXiv arXiv
-
[33]
Doyne Farmer and Fabrizio Lillo , title =
Elia Zarinelli and Michele Treccani and J. Doyne Farmer and Fabrizio Lillo , title =. Market Microstructure and Liquidity , volume =
-
[34]
Quantitative Finance , volume =
Natalia Bershova and Dmitry Rakhlin , title =. Quantitative Finance , volume =
-
[35]
Esteban Moro and Javier Vicente and Luis G. Moyano and Austin Gerig and J. Doyne Farmer and Salvatore Vaglica and Fabrizio Lillo and Rosario N. Mantegna , title =. Physical Review E , volume =
-
[36]
Physical Review Letters , volume =
Fabio Bucci and Michael Benzaquen and Fabrizio Lillo and Jean-Philippe Bouchaud , title =. Physical Review Letters , volume =
-
[37]
Alessio Azzutti and Wolf-Georg Ringe and H. Siegfried Stiehl , title =. University of Pennsylvania Journal of International Law , year =
-
[38]
IEEE Symposium Series on Computational Intelligence , pages=
Can an AI Perform Market Manipulation at Its Own Discretion? -- A Genetic Algorithm Learns in an Artificial Market Simulation , author=. IEEE Symposium Series on Computational Intelligence , pages=
-
[39]
IEEE Conference on Evolving and Adaptive Intelligent Systems , pages=
Learning Unfair Trading: A Market Manipulation Analysis from the Reinforcement Learning Perspective , author=. IEEE Conference on Evolving and Adaptive Intelligent Systems , pages=
-
[40]
Ethical Issues for Autonomous Trading Agents , author=. Minds and Machines , volume=
-
[41]
Journal of Financial Markets , volume=
Microstructure-Based Manipulation: Strategic Behavior and Performance of Spoofing Traders , author=. Journal of Financial Markets , volume=
-
[42]
Fabrizio Lillo and J.D. Farmer and Rosario N. Mantegna , title =. Nature , volume =
-
[43]
Lillo, Fabrizio and Farmer, J. Doyne and Mantegna, Rosario N. Price Impact Function of a Single Transaction. The Complex Dynamics of Economic Interaction. 2004
work page 2004
-
[44]
Quantitative Finance , volume =
Wei-Xing Zhou , title =. Quantitative Finance , volume =. 2012 , publisher =
work page 2012
-
[45]
M. Potters and J.P. Bouchaud , title =. Physica A: Statistical Mechanics and its Applications , volume =
-
[46]
Quantifying stock-price response to demand fluctuations , author =. Phys. Rev. E , volume =. 2002 , publisher =
work page 2002
-
[47]
R. Almgren and N. Chriss. Optimal execution of portfolio transactions. Journal of Risk, 3: 0 5--39, 2000
work page 2000
-
[48]
Direct estimation of equity market impact
Robert Almgren, Chee Thum, Emmanuel Hauptmann, and Hong Li. Direct estimation of equity market impact. Risk, 18: 0 58--62, 2005
work page 2005
-
[49]
T.G. Andersen, T. Bollerslev, F.X. Diebold, and P. Labys. Modelling and forecasting realized volatility. Econometrica, 71: 0 579--625, 2003
work page 2003
-
[50]
Alessio Azzutti, Wolf-Georg Ringe, and H. Siegfried Stiehl. Machine learning, market manipulation, and collusion on capital markets: Why the ``black box" matters. University of Pennsylvania Journal of International Law, 43: 0 79--136, 2021
work page 2021
-
[51]
Spoofing and manipulating order books with learning algorithms
Alvaro Cartea, Patrick Chang, and Gabriel Garcıa-Arenas. Spoofing and manipulating order books with learning algorithms. Available at SSRN: https://ssrn.com/abstract=4639959, 2023
work page 2023
-
[52]
Optimal execution with non-linear transient market impact
Gianbiagio Curato, Jim Gatheral, and Fabrizio Lillo. Optimal execution with non-linear transient market impact. Quantitative Finance, 17: 0 41--54, 2017
work page 2017
- [53]
-
[54]
The Financial Mathematics of Market Liquidity
Olivier Gueant. The Financial Mathematics of Market Liquidity. Chapman & Hall/CRC Financial Mathematics Series, 2016
work page 2016
-
[55]
A reinforcement learning extension to the almgren-chriss framework for optimal trade execution
Dieter Hendricks and Diane Wilcox. A reinforcement learning extension to the almgren-chriss framework for optimal trade execution. In 2014 IEEE Conference on Computational Intelligence for Financial Engineering & Economics (CIFEr), pages 457--464. IEEE, 2014
work page 2014
-
[56]
G. Huberman and W. Stanzl. Price manipulation and quasi-arbitrage. Econometrica, 72: 0 1247–1275, 2004
work page 2004
-
[57]
A Deep Reinforcement Learning Framework for the Financial Portfolio Management Problem
Zhengyao Jiang, Dixing Xu, and Jinjun Liang. A deep reinforcement learning framework for the financial portfolio management problem. arXiv preprint arXiv:1706.10059, 2017
work page internal anchor Pith review Pith/arXiv arXiv 2017
-
[58]
u r Dynamik der Flugsysteme, Deutsche Forschungs- und Versuchsanstalt f \
Dieter Kraft. A software package for sequential quadratic programming. Technical Report DFVLR-FB 88-28, Institut f \"u r Dynamik der Flugsysteme, Deutsche Forschungs- und Versuchsanstalt f \"u r Luft- und Raumfahrt, 1988
work page 1988
-
[59]
Microstructure-based manipulation: Strategic behavior and performance of spoofing traders
Eun Jung Lee, Kyoung-Soo Eom, and Kyung Suh Park. Microstructure-based manipulation: Strategic behavior and performance of spoofing traders. Journal of Financial Markets, 16: 0 227--252, 2013
work page 2013
-
[60]
T. P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. Continuous control with deep reinforcement learning. In Proceedings of the 4th International Conference on Learning Representations, Proceedings of Machine Learning Research. PMLR, 2016
work page 2016
-
[61]
Fabrizio Lillo, J.D. Farmer, and Rosario N. Mantegna. Master curve for price-impact function. Nature, 421: 0 129–130, 2003
work page 2003
-
[62]
Fabrizio Lillo, J. Doyne Farmer, and Rosario N. Mantegna. Price impact function of a single transaction. In The Complex Dynamics of Economic Interaction, pages 153--160, Berlin, Heidelberg, 2004. Springer Berlin Heidelberg
work page 2004
-
[63]
Reinforcement learning for optimal execution when liquidity is time-varying
Andrea Macrì and Fabrizio Lillo. Reinforcement learning for optimal execution when liquidity is time-varying. Applied Mathematical Finance, 31: 0 312–342, 2025
work page 2025
-
[64]
Deep reinforcement learning for optimal trading with partial information
Andrea Macrì, Sebastian Jaimungal, and Fabrizio Lillo. Deep reinforcement learning for optimal trading with partial information. arXiv:2511.00190, 2025
-
[65]
Learning unfair trading: A market manipulation analysis from the reinforcement learning perspective
Enrique Mart \' nez-Miranda, Frank McGroarty, and Edward Tsang. Learning unfair trading: A market manipulation analysis from the reinforcement learning perspective. In IEEE Conference on Evolving and Adaptive Intelligent Systems, pages 103--108, 2016
work page 2016
-
[66]
Deep Reinforcement Learning for Online Optimal Execution Strategies
Alessandro Micheli and M \'e lodie Monod. Deep reinforcement learning for online optimal execution strategies. arXiv preprint arXiv:2410.13493, 2024
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[67]
Takanobu Mizuta. Can an ai perform market manipulation at its own discretion? -- a genetic algorithm learns in an artificial market simulation. In IEEE Symposium Series on Computational Intelligence, pages 407--414, 2020
work page 2020
-
[68]
Learning to trade via direct reinforcement
John Moody and Matthew Saffell. Learning to trade via direct reinforcement. IEEE Transactions on Neural Networks, 12: 0 875--889, 2001
work page 2001
-
[69]
Reinforcement learning for optimized trade execution
Yuriy Nevmyvaka, Yi Feng, and Michael Kearns. Reinforcement learning for optimized trade execution. In Proceedings of the 23rd International Conference on Machine Learning (ICML), pages 673--680, 2006
work page 2006
-
[70]
Double deep q-learning for optimal execution
Brian Ning, Franco Ho Ting Lin, and Sebastian Jaimungal. Double deep q-learning for optimal execution. Applied Mathematical Finance, 28: 0 361--380, 2021
work page 2021
-
[71]
M. Potters and J.P. Bouchaud. More statistical properties of order books and price impact. Physica A: Statistical Mechanics and its Applications, 324: 0 133--140, 2003
work page 2003
-
[72]
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller. Deterministic policy gradient algorithms. In Eric P. Xing and Tony Jebara, editors, Proceedings of the 31st International Conference on Machine Learning, volume 32 of Proceedings of Machine Learning Research, pages 387--395, Bejing, China, 22-24 Jun 2014. PMLR
work page 2014
-
[73]
Pang-Ning Tan, Michael Steinbach, and Vipin Kumar. Introduction to Data Mining. Pearson Education Limited, 2013
work page 2013
-
[74]
Michael P. Wellman and Uday Rajan. Ethical issues for autonomous trading agents. Minds and Machines, 27: 0 609--624, 2017
work page 2017
-
[75]
Correlation coefficients and semantic textual similarity
Vitalii Zhelezniak, Aleksandar Savkov, April Shen, and Nils Hammerla. Correlation coefficients and semantic textual similarity. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 951--962. Association for Computational Li...
work page 2019
-
[76]
Universal price impact functions of individual trades in an order-driven market
Wei-Xing Zhou. Universal price impact functions of individual trades in an order-driven market. Quantitative Finance, 12 0 (8): 0 1253--1263, 2012
work page 2012
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.