REVIEW 4 major objections 5 minor 32 references
ARL-Based Multi-Action Market Making with Hawkes Processes and Variable Volatility
T0 review · 4 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read A market-making agent that chooses when to quote, rather than recomputing prices, keeps its edge when volatility jumps 100-fold.
desk verdict The 4-action adaptation claim is a real empirical finding, but the reported wealth variance at vol=200 looks too small for the modeled random walk, which makes me worry the volatility change isn't actually taking effect as described. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The core mechanism is a two-level decision structure. A high-level discrete policy chooses one of four actions: no quote, bilateral quotes, ask-only, or bid-only; whenever it chooses to quote, the actual bid and ask offsets come from a frozen Always Quoting agent trained at volatility 2. Execution intensity is $$\$lambda_n^{{\pm}}$ = \text{Hawkes\_intensity}\,$e^{{-k_n^{\pm}}$\$delta_n^{{\pm}}$},$$ where $\delta_n^{\pm}$ are the bid/ask offsets and $k_n^{\pm}$ is the volume-decay parameter; the intensity itself mean-reverts toward a baseline and jumps upward on each fill, which is the self-exciting feature that replaces Poisson arrivals. The adversary controls drift, baseline arrival rate, and volume de
What would settle it
Measure realized fill rates, time-to-fill, and fill-conditional P&L at $\sigma=200$ for the four-action policy trained at $\sigma=2$, with the same offset bounds $[0,3]$. If fills are rare or systematically one-sided, the stable terminal wealth is an artifact of non-execution; if fills occur at reasonable rates and capture spread, the adaptation claim is supported.
Extended reading notes
Core claim
The central claim is that a four-action market maker trained with adversarial reinforcement learning under low volatility generalizes to a high-volatility regime without retraining. The agent chooses among four discrete quoting modes — no quote, bilateral, ask-only, bid-only — and, when quoting, uses offsets from a frozen Always Quoting agent trained at volatility 2. Execution follows a self-exciting Hawkes intensity that mean-reverts and jumps on fills, replacing the Poisson arrivals of earlier work. Across seven risk-coefficient tables and five adversary types, the policy trained at $\sigma=2$ and tested at $\sigma=200$ keeps terminal wealth and Sharpe close to the low-volatility baseline
Load-bearing premise
The price-setting sub-policy trained at volatility 2 must still produce sensible bid/ask prices when the price process moves at volatility 200; if its offsets become stale or nonmarketable, the four-action agent's stable numbers could simply reflect that its quotes rarely trade.
Editorial extensions
If this is right
- A four-action quoting space generalizes across a 100-fold volatility change: the agent trained at $\sigma=2$ and tested at $\sigma=200$ keeps Sharpe ratio and terminal wealth close to baseline across all seven risk-coefficient tables.
- The same agent meets exchange-style quoting obligations under stress: bilateral quotes appear at least 92% of the time, above the 90% continuous-quoting requirement the paper cites for regulated market makers.
- The mismatch appears in price setting, not the action space: the agent trained and tested at $\sigma=200$ with a frozen $\sigma=2$ price policy quotes conservatively and has lower Sharpe.
- Hawkes arrivals change quoting incentives: because a fill raises the intensity of further fills, agents quote more aggressively after trades, which the paper argues supports liquidity in low-liquidity states.
Reading between the lines
- The paper does not report fill rates or spread capture at $\sigma=200$; computing those would separate genuine adaptation from quotes that are posted but rarely marketable.
- A testable extension: retrain only the price-setting policy at $\sigma=200$, keep the four-action supervisor frozen, and check whether bilateral quoting and Sharpe improve; this would isolate whether the discrete policy or the offset policy carries the robustness.
- Because quote offsets are bounded in $[0,3]$ while price increments scale with $\sigma$, scaling offset bounds with volatility could turn the reported stability into higher profitability; this is an editorial extrapolation, not a claim in the paper.
- The paper acknowledges that Hawkes arrivals lower mean wealth and Sharpe relative to Poisson; a Poisson-baseline comparison at $\sigma=200$ would show whether the volatility-transfer result is an effect of the arrival process or of the flexible action space alone.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper extends the authors' prior ARL market-making framework ([30]) by replacing Poisson execution arrivals with a Hawkes process and by testing a four-action market maker (no quote, bilateral, ask-only, bid-only) under volatility levels 2 and 200. The headline claim is that a 4-action MM trained at volatility 2 and tested at volatility 200 adapts effectively, preserving Sharpe-type performance and providing two-sided quotes at least 92% of the time. The evaluation is reported in seven tables across risk-coefficient configurations, with several adversarial training regimes.
Significance. If established, the result would be a useful robustness finding for RL-based market making: a discrete quoting policy trained in a calm market could remain profitable and liquid in a much more volatile market. The paper also responsibly reports that the Hawkes-process agents perform worse in absolute terms than the earlier Poisson-process results, which is a useful caveat. The tables provide a large amount of comparative information. However, several load-bearing issues prevent the stated contribution from being accepted as written: the 92% quoting claim is contradicted by multiple entries in the authors' own tables, there is an internal inconsistency about the test volatility of one agent, and the economic meaning of the quoting decision at volatility 200 is not established without fill and spread diagnostics.
major comments (4)
- [Abstract; Tables 1–7, column '4-Action MM (Train @ vol=2, Test @ vol=200)'] The abstract and Section 1 claim that the 4-action MM trained at vol=2 and tested at vol=200 provides two-sided quotes at least 92% of the time. This is contradicted by the tables. Bilateral quote ratios below 92% include: Table 1, row A: 77.27%; Table 2, Fix: 69.74%; Table 5, Fix: 46.96%; Table 6, Fix: 31.44%; Table 7, Fix: 48.77% and All: 74.64%. In these rows the agent instead chooses no-quote or unilateral quoting. The central claim must be restricted, or the heterogeneity must be explained and the abstract revised.
- [Section 5.1 vs Table headers and Section 5.2] Section 5.1 states that the 4-Action MM (Train & Test @ vol=200) is 'trained in an environment with Volatility=200 ... but ultimately tested in an environment with Volatility=2.' The table headers label this column '4-Action MM (Train @ vol=200, Test @ vol=200)', and Section 5.2 discusses it as a high-volatility-tested agent. This is a direct inconsistency. The test environment determines whether the reported conservatism (e.g., bid-only 21.80% in Table 1 All row, or unilateral quoting in Table 7) supports the paper's conclusion about high-volatility training. The text and labels must be reconciled.
- [Section 5.2, '2-Action MM ... & 4-Action MM ...'; Sections 3.1.1 and 3.2.2] The 4-action MM's 'quote' decision delegates price setting to the Always Quoting MM trained at vol=2. The sub-policy's offsets are bounded in [0,3] (Section 3.2.2), while at vol=200 with dt=0.005 the per-step price standard deviation is about 14.14 (Section 3.1.1). The paper does not report average offsets, fill counts, or spread widths at test time, so the reader cannot tell whether the observed bilateral quote ratios represent executable two-sided liquidity or nearly inert quotes. I do not press the strongest version of this concern: terminal wealth and inventory in Tables 1–7 are nonzero, so some trading does occur. But without these diagnostics, the 'effective adaptation' claim is not fully supported.
- [Section 5.1; Tables 1–7] The 'stable performance' claim rests on point estimates. Sharpe ratios are E/σ computed from the evaluation runs, but no confidence intervals, standard errors, or significance tests are reported. For example, Table 1 Fix gives Sharpe 0.6810 for 4-Action MM (Train & Test @ vol=2) and 0.6837 for 4-Action MM (Train @ vol=2, Test @ vol=200); many differences are of this size and will be within sampling noise. Please provide uncertainty quantification, at least for the key train@2/test@200 versus train&test@2 comparison.
minor comments (5)
- [Section 5.1] The phrase 'evaluated 100 times, each with 1000 episodes' is ambiguous: does this mean 100 independent evaluations of 1000 episodes each, or 100 episodes in 1000 runs? Please clarify the total number of evaluation trajectories and how the reported means and standard deviations are aggregated.
- [Section 3.1.2, Eq. (1)] The quantity 'last_match_result' is used in the Hawkes intensity update but is not formally defined. State that it is the indicator of a trade in the previous time step, and specify whether the market maker observes its own fills or all market trades.
- [Section 5.1] Terminal Wealth is described as 'mean ± variance' but the tables report standard deviations (e.g., 3.2613 when the mean is 2.1945). Please use 'mean ± standard deviation' for clarity.
- [Section 3.1.2 / Section 4] The Hawkes parameters (mean_reversion_speed=60, baseline_arrival_rate=10, jump_size=40, dt=0.005) are taken from the mbt_gym defaults without sensitivity analysis or calibration. At least one robustness check would strengthen the claim that the results are not tied to this particular parameter choice.
- [Figures 2–7] The description of the figures is terse. In particular, Figure 2 is said to contain 'six curves' and color-coded Sharpe ratios, but the caption and text do not identify which curve corresponds to which agent/adversary configuration. Please add a legend or a more complete caption.
Circularity Check
No construction-level circularity; the headline result is an empirical simulation observation, not a derivation from its own inputs.
full rationale
The paper's central claim—that a 4-action MM trained at volatility 2 adapts to volatility 200—is an empirical result obtained from the authors' own simulator, not a mathematical derivation. The price-setting sub-policy being the vol=2-trained Always Quoting MM is an architectural choice, not an equation-level identity: the 4-action agent still learns when to quote, which sides to quote, and whether to refrain, and is then evaluated at vol=200. No fitted parameter is renamed as a prediction: wealth, Sharpe ratio, inventory, and quoting ratios are measured post hoc. The only self-citation is the prior action-space/ARL framework [30], but it is non-load-bearing because all agents are retrained in this paper and the new volatility and Hawkes components are this paper's own contribution. The Hawkes module is cited from non-overlapping authors [20]. No uniqueness theorem, ansatz, or parametric equivalence is imported from same-author work. The serious issues in the paper are correctness/consistency rather than circularity: the abstract's 'at least 92%' two-sided quoting claim is contradicted by many table entries (e.g., Table 5, Fix: 45.45% no-quote and 46.96% bilateral for 4-Action MM (Train @ vol=2, Test @ vol=200)), and the bounded offset space [0,3] versus the much larger per-step price dispersion at vol=200 raises an artifact concern about whether the agent is genuinely providing executable liquidity. These concerns should be weighed as validity risks, but they do not make the derivation circular.
Assumptions & free parameters
free parameters (2)
- Hawkes intensity parameters: mean_reversion_speed, baseline_arrival_rate, jump_size, dt =
60.0, 10.0, 40.0, 0.005
- Quote offset bounds =
[0, 3]
assumptions (4)
- domain assumption Price dynamics follow Brownian motion with drift: Z_{n+1} = Z_n + b_n Δt + σ_n W_n.
- domain assumption Execution intensity follows a Hawkes process with exponential decay of intensity toward baseline and a jump after each trade (Equation 1), with parameters from mbt_gym.
- domain assumption Zero-sum game between market maker and adversary, with the adversary's reward equal to the negative of the market maker's reward.
- ad hoc to paper The Always Quoting MM sub-policy, trained at volatility 2, is a valid price setter when the 4-action MM operates at volatility 200.
Cite this review
Pith. "Pith review of ARL-Based Multi-Action Market Making with Hawkes Processes and Variable Volatility." pith.science (2026). https://pith.science/paper/H4IDN2UT
@misc{pith2026250816589,
author = {Pith},
title = {Pith review of: ARL-Based Multi-Action Market Making with Hawkes Processes and Variable Volatility},
year = {2026},
howpublished = {\url{https://pith.science/paper/H4IDN2UT}},
note = {Machine review of arXiv:2508.16589}
}
read the original abstract
We advance market-making strategies by integrating Adversarial Reinforcement Learning (ARL), Hawkes Processes, and variable volatility levels while also expanding the action space available to market makers (MMs). To enhance the adaptability and robustness of these strategies -- which can quote always, quote only on one side of the market or not quote at all -- we shift from the commonly used Poisson process to the Hawkes process, which better captures real market dynamics and self-exciting behaviors. We then train and evaluate strategies under volatility levels of 2 and 200. Our findings show that the 4-action MM trained in a low-volatility environment effectively adapts to high-volatility conditions, maintaining stable performance and providing two-sided quotes at least 92\% of the time. This indicates that incorporating flexible quoting mechanisms and realistic market simulations significantly enhances the effectiveness of market-making strategies.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[30]
Ziyi Wang, Carmine Ventre, and Maria Polukarov. 2023. Robust Market Mak- ing: To Quote, or not To Quote. InProceedings of the Fourth ACM International Conference on AI in Finance . 664–672
work page 2023
-
[1]
Frédéric Abergel and Aymen Jedidi. 2015. Long-time behavior of a hawkes process–based limit order book. SIAM Journal on Financial Mathematics 6, 1 (2015), 1026–1043
work page 2015
-
[2]
Jacob Abernethy and Satyen Kale. 2013. Adaptive market making via online learning. Advances in Neural Information Processing Systems 26 (2013)
2013
-
[3]
Marco Avellaneda and Sasha Stoikov. 2008. High-frequency trading in a limit order book. Quantitative Finance 8, 3 (2008), 217–224
work page 2008
-
[4]
Emmanuel Bacry, Sylvain Delattre, Marc Hoffmann, and Jean-François Muzy
-
[5]
Emmanuel Bacry, Thibault Jaisson, and Jean-François Muzy. 2016. Estimation of slowly decreasing hawkes kernels: application to high-frequency order book dynamics. Quantitative Finance 16, 8 (2016), 1179–1201
work page 2016
-
[6]
Emmanuel Bacry, Iacopo Mastromatteo, and Jean-François Muzy. 2015. Hawkes processes in finance. Market Microstructure and Liquidity 1, 01 (2015), 1550005
2015
-
[7]
Emmanuel Bacry and Jean-François Muzy. 2014. Hawkes model for price and trades high-frequency dynamics. Quantitative Finance 14, 7 (2014), 1147–1166
work page 2014
Show all 32 references
-
[8]
Álvaro Cartea, Sebastian Jaimungal, and José Penalva. 2015. Algorithmic and high-frequency trading. Cambridge University Press
2015
-
[9]
Nicholas Tung Chan and Christian Shelton. 2001. An electronic market-maker. (2001)
2001
-
[10]
London Stock Exchange. 2022. ETF & ETP Market Maker Obliga- tions. https://docs.londonstockexchange.com/sites/default/files/documents/ Market%20Maker%20Obligations%20Factsheet_18.05.2022.pdf
2022
-
[11]
Pietro Fodra and Mauricio Labadie. 2012. High-frequency market-making with inventory constraints and directional bets. arXiv preprint arXiv:1206.4810 (2012)
2012 arXiv
-
[12]
Olivier Guéant. 2017. Optimal market making. Applied Mathematical Finance 24, 2 (2017), 112–154
2017
-
[13]
Olivier Guéant, Charles-Albert Lehalle, and Joaquin Fernandez-Tapia. 2013. Deal- ing with the inventory risk: a solution to the market making problem. Mathe- matics and financial economics 7, 4 (2013), 477–507
2013
-
[14]
Qi Guo, Bruno Remillard, and Anatoliy Swishchuk. 2020. Multivariate general compound point processes in limit order books. Risks 8, 3 (2020), 98
2020
-
[15]
Qi Guo and Anatoliy Swishchuk. 2020. Multivariate general compound Hawkes processes and their applications in limit order books. Wilmott 2020, 107 (2020), 42–51
2020
-
[16]
Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen, George Tucker, Sehoon Ha, Jie Tan, Vikash Kumar, Henry Zhu, Abhishek Gupta, Pieter Abbeel, et al. 2018. Soft actor-critic algorithms and applications. arXiv preprint arXiv:1812.05905 (2018)
2018 arXiv
-
[17]
Alan G Hawkes. 1971. Spectra of some self-exciting and mutually exciting point processes. Biometrika 58, 1 (1971), 83–90
1971
-
[18]
Patrick Hewlett. 2006. Clustering of order arrivals, price impact and trade path optimisation. In Workshop on Financial Modeling with Jump processes, Ecole Poly- technique. Citeseer, 6–8
2006
-
[19]
Thomas Ho and Hans R Stoll. 1981. Optimal dealer pricing under transactions and return uncertainty. Journal of Financial economics 9, 1 (1981), 47–73
1981
-
[20]
Joseph Jerome, Leandro Sánchez-Betancourt, Rahul Savani, and Martin Herdegen
-
[21]
Jeremy Large. 2007. Measuring the resiliency of an electronic limit order book. Journal of Financial Markets 10, 1 (2007), 1–25
2007
-
[22]
Xiaofei Lu and Frédéric Abergel. 2017. Limit order book modelling with high dimensional Hawkes processes. (2017)
2017
-
[23]
Xiaofei Lu and Frédéric Abergel. 2018. High-dimensional Hawkes processes for limit order books: modelling, empirical analysis and numerical calibration. Quantitative Finance 18, 2 (2018), 249–264
2018
-
[24]
Deutsche Börse Cash Market. 2019. MiFID II: Regulated Market Maker . https: //www.xetra.com/resource/blob/2450614/def8bfa58c17bd971bb69ec14f2a024f/ data/Regulated-Market-Maker-Handbook.pdf
2019
-
[25]
Ana Roldan Contreras and Anatoliy Swishchuk. 2022. Optimal Liquidation, Acquisition and Market Making Problems in HFT under Hawkes Models for LOB. Risks 10, 8 (2022), 160
2022
-
[26]
Thomas Spooner, John Fearnley, Rahul Savani, and Andreas Koukorinis. 2018. Market making via reinforcement learning. arXiv preprint arXiv:1804.04216 (2018)
2018 arXiv
-
[27]
Thomas Spooner and Rahul Savani. 2020. Robust market making via adversarial reinforcement learning. arXiv preprint arXiv:2003.01820 (2020)
2020 arXiv
-
[28]
Anatoliy Swishchuk and Aiden Huffman. 2020. General compound Hawkes processes in limit order books. Risks 8, 1 (2020), 28
2020
-
[29]
Ioane Muni Toke and Fabrizio Pomponio. 2012. Modelling trades-through in a limit order book using Hawkes processes. Economics 6, 1 (2012), 20120022
2012
-
[2013]
Quantitative finance 13, 1 (2013), 65–77
Modelling microstructure noise with mutually exciting point processes. Quantitative finance 13, 1 (2013), 65–77
2013
-
[2023]
In Proceedings of the Fourth ACM International Conference on AI in Finance
Mbt-gym: Reinforcement learning for model-based limit order book trading. In Proceedings of the Fourth ACM International Conference on AI in Finance . 619– 627
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.