Pith. sign in

REVIEW 4 major objections 5 minor 32 references

ARL-Based Multi-Action Market Making with Hawkes Processes and Variable Volatility

T0 review · 4 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read A market-making agent that chooses when to quote, rather than recomputing prices, keeps its edge when volatility jumps 100-fold.

desk verdict The 4-action adaptation claim is a real empirical finding, but the reported wealth variance at vol=200 looks too small for the modeled random walk, which makes me worry the volatility change isn't actually taking effect as described. read the letter →

arxiv 2508.16589 v1 pith:H4IDN2UT submitted 2025-08-07 q-fin.TR cs.AIecon.GNq-fin.EC

classification q-fin.TRcs.AIecon.GNq-fin.EC
keywords marketmakingadversarialreinforcementlearningHawkesprocesseslimitorderbookhigh-frequencytradingvolatilityrobustnessflexiblequotingstochasticoptimalcontrol
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that a market-making agent with a discrete choice — quote both sides, quote one side, or stop quoting — can absorb a hundredfold increase in volatility without being retrained. Trained at volatility 2 under an adversary, with executions arriving through a self-exciting Hawkes process, the four-action agent tested at volatility 200 keeps its Sharpe ratio and terminal wealth close to the low-volatility baseline and provides two-sided quotes at least 92% of the time. The load-bearing design is separation: the high-level policy decides when to quote, while a frozen price-setting policy trained at volatility 2 chooses the actual bid/ask offsets. The authors interpret this as evidence that flexible quoting plus realistic order flow yields robustness to regime change. The paper also acknowledges that replacing Poisson arrivals with Hawkes arrivals lowers absolute performance, so the gain is in robustness rather than raw profitability.

What carries the argument

The core mechanism is a two-level decision structure. A high-level discrete policy chooses one of four actions: no quote, bilateral quotes, ask-only, or bid-only; whenever it chooses to quote, the actual bid and ask offsets come from a frozen Always Quoting agent trained at volatility 2. Execution intensity is $$\$lambda_n^{{\pm}}$ = \text{Hawkes\_intensity}\,$e^{{-k_n^{\pm}}$\$delta_n^{{\pm}}$},$$ where $\delta_n^{\pm}$ are the bid/ask offsets and $k_n^{\pm}$ is the volume-decay parameter; the intensity itself mean-reverts toward a baseline and jumps upward on each fill, which is the self-exciting feature that replaces Poisson arrivals. The adversary controls drift, baseline arrival rate, and volume de

What would settle it

Measure realized fill rates, time-to-fill, and fill-conditional P&L at $\sigma=200$ for the four-action policy trained at $\sigma=2$, with the same offset bounds $[0,3]$. If fills are rare or systematically one-sided, the stable terminal wealth is an artifact of non-execution; if fills occur at reasonable rates and capture spread, the adaptation claim is supported.

Watch

Extended reading notes

Core claim

The central claim is that a four-action market maker trained with adversarial reinforcement learning under low volatility generalizes to a high-volatility regime without retraining. The agent chooses among four discrete quoting modes — no quote, bilateral, ask-only, bid-only — and, when quoting, uses offsets from a frozen Always Quoting agent trained at volatility 2. Execution follows a self-exciting Hawkes intensity that mean-reverts and jumps on fills, replacing the Poisson arrivals of earlier work. Across seven risk-coefficient tables and five adversary types, the policy trained at $\sigma=2$ and tested at $\sigma=200$ keeps terminal wealth and Sharpe close to the low-volatility baseline

Load-bearing premise

The price-setting sub-policy trained at volatility 2 must still produce sensible bid/ask prices when the price process moves at volatility 200; if its offsets become stale or nonmarketable, the four-action agent's stable numbers could simply reflect that its quotes rarely trade.

Editorial extensions

If this is right

  • A four-action quoting space generalizes across a 100-fold volatility change: the agent trained at $\sigma=2$ and tested at $\sigma=200$ keeps Sharpe ratio and terminal wealth close to baseline across all seven risk-coefficient tables.
  • The same agent meets exchange-style quoting obligations under stress: bilateral quotes appear at least 92% of the time, above the 90% continuous-quoting requirement the paper cites for regulated market makers.
  • The mismatch appears in price setting, not the action space: the agent trained and tested at $\sigma=200$ with a frozen $\sigma=2$ price policy quotes conservatively and has lower Sharpe.
  • Hawkes arrivals change quoting incentives: because a fill raises the intensity of further fills, agents quote more aggressively after trades, which the paper argues supports liquidity in low-liquidity states.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper does not report fill rates or spread capture at $\sigma=200$; computing those would separate genuine adaptation from quotes that are posted but rarely marketable.
  • A testable extension: retrain only the price-setting policy at $\sigma=200$, keep the four-action supervisor frozen, and check whether bilateral quoting and Sharpe improve; this would isolate whether the discrete policy or the offset policy carries the robustness.
  • Because quote offsets are bounded in $[0,3]$ while price increments scale with $\sigma$, scaling offset bounds with volatility could turn the reported stability into higher profitability; this is an editorial extrapolation, not a claim in the paper.
  • The paper acknowledges that Hawkes arrivals lower mean wealth and Sharpe relative to Poisson; a Poisson-baseline comparison at $\sigma=200$ would show whether the volatility-transfer result is an effect of the arrival process or of the flexible action space alone.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper extends the authors' prior ARL market-making framework ([30]) by replacing Poisson execution arrivals with a Hawkes process and by testing a four-action market maker (no quote, bilateral, ask-only, bid-only) under volatility levels 2 and 200. The headline claim is that a 4-action MM trained at volatility 2 and tested at volatility 200 adapts effectively, preserving Sharpe-type performance and providing two-sided quotes at least 92% of the time. The evaluation is reported in seven tables across risk-coefficient configurations, with several adversarial training regimes.

Significance. If established, the result would be a useful robustness finding for RL-based market making: a discrete quoting policy trained in a calm market could remain profitable and liquid in a much more volatile market. The paper also responsibly reports that the Hawkes-process agents perform worse in absolute terms than the earlier Poisson-process results, which is a useful caveat. The tables provide a large amount of comparative information. However, several load-bearing issues prevent the stated contribution from being accepted as written: the 92% quoting claim is contradicted by multiple entries in the authors' own tables, there is an internal inconsistency about the test volatility of one agent, and the economic meaning of the quoting decision at volatility 200 is not established without fill and spread diagnostics.

major comments (4)
  1. [Abstract; Tables 1–7, column '4-Action MM (Train @ vol=2, Test @ vol=200)'] The abstract and Section 1 claim that the 4-action MM trained at vol=2 and tested at vol=200 provides two-sided quotes at least 92% of the time. This is contradicted by the tables. Bilateral quote ratios below 92% include: Table 1, row A: 77.27%; Table 2, Fix: 69.74%; Table 5, Fix: 46.96%; Table 6, Fix: 31.44%; Table 7, Fix: 48.77% and All: 74.64%. In these rows the agent instead chooses no-quote or unilateral quoting. The central claim must be restricted, or the heterogeneity must be explained and the abstract revised.
  2. [Section 5.1 vs Table headers and Section 5.2] Section 5.1 states that the 4-Action MM (Train & Test @ vol=200) is 'trained in an environment with Volatility=200 ... but ultimately tested in an environment with Volatility=2.' The table headers label this column '4-Action MM (Train @ vol=200, Test @ vol=200)', and Section 5.2 discusses it as a high-volatility-tested agent. This is a direct inconsistency. The test environment determines whether the reported conservatism (e.g., bid-only 21.80% in Table 1 All row, or unilateral quoting in Table 7) supports the paper's conclusion about high-volatility training. The text and labels must be reconciled.
  3. [Section 5.2, '2-Action MM ... & 4-Action MM ...'; Sections 3.1.1 and 3.2.2] The 4-action MM's 'quote' decision delegates price setting to the Always Quoting MM trained at vol=2. The sub-policy's offsets are bounded in [0,3] (Section 3.2.2), while at vol=200 with dt=0.005 the per-step price standard deviation is about 14.14 (Section 3.1.1). The paper does not report average offsets, fill counts, or spread widths at test time, so the reader cannot tell whether the observed bilateral quote ratios represent executable two-sided liquidity or nearly inert quotes. I do not press the strongest version of this concern: terminal wealth and inventory in Tables 1–7 are nonzero, so some trading does occur. But without these diagnostics, the 'effective adaptation' claim is not fully supported.
  4. [Section 5.1; Tables 1–7] The 'stable performance' claim rests on point estimates. Sharpe ratios are E/σ computed from the evaluation runs, but no confidence intervals, standard errors, or significance tests are reported. For example, Table 1 Fix gives Sharpe 0.6810 for 4-Action MM (Train & Test @ vol=2) and 0.6837 for 4-Action MM (Train @ vol=2, Test @ vol=200); many differences are of this size and will be within sampling noise. Please provide uncertainty quantification, at least for the key train@2/test@200 versus train&test@2 comparison.
minor comments (5)
  1. [Section 5.1] The phrase 'evaluated 100 times, each with 1000 episodes' is ambiguous: does this mean 100 independent evaluations of 1000 episodes each, or 100 episodes in 1000 runs? Please clarify the total number of evaluation trajectories and how the reported means and standard deviations are aggregated.
  2. [Section 3.1.2, Eq. (1)] The quantity 'last_match_result' is used in the Hawkes intensity update but is not formally defined. State that it is the indicator of a trade in the previous time step, and specify whether the market maker observes its own fills or all market trades.
  3. [Section 5.1] Terminal Wealth is described as 'mean ± variance' but the tables report standard deviations (e.g., 3.2613 when the mean is 2.1945). Please use 'mean ± standard deviation' for clarity.
  4. [Section 3.1.2 / Section 4] The Hawkes parameters (mean_reversion_speed=60, baseline_arrival_rate=10, jump_size=40, dt=0.005) are taken from the mbt_gym defaults without sensitivity analysis or calibration. At least one robustness check would strengthen the claim that the results are not tied to this particular parameter choice.
  5. [Figures 2–7] The description of the figures is terse. In particular, Figure 2 is said to contain 'six curves' and color-coded Sharpe ratios, but the caption and text do not identify which curve corresponds to which agent/adversary configuration. Please add a legend or a more complete caption.

Circularity Check

0 steps flagged · score 0.0 of 10

No construction-level circularity; the headline result is an empirical simulation observation, not a derivation from its own inputs.

full rationale

The paper's central claim—that a 4-action MM trained at volatility 2 adapts to volatility 200—is an empirical result obtained from the authors' own simulator, not a mathematical derivation. The price-setting sub-policy being the vol=2-trained Always Quoting MM is an architectural choice, not an equation-level identity: the 4-action agent still learns when to quote, which sides to quote, and whether to refrain, and is then evaluated at vol=200. No fitted parameter is renamed as a prediction: wealth, Sharpe ratio, inventory, and quoting ratios are measured post hoc. The only self-citation is the prior action-space/ARL framework [30], but it is non-load-bearing because all agents are retrained in this paper and the new volatility and Hawkes components are this paper's own contribution. The Hawkes module is cited from non-overlapping authors [20]. No uniqueness theorem, ansatz, or parametric equivalence is imported from same-author work. The serious issues in the paper are correctness/consistency rather than circularity: the abstract's 'at least 92%' two-sided quoting claim is contradicted by many table entries (e.g., Table 5, Fix: 45.45% no-quote and 46.96% bilateral for 4-Action MM (Train @ vol=2, Test @ vol=200)), and the bounded offset space [0,3] versus the much larger per-step price dispersion at vol=200 raises an artifact concern about whether the agent is genuinely providing executable liquidity. These concerns should be weighed as validity risks, but they do not make the derivation circular.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central claim rests on a simulation model whose parameters (Hawkes intensity, offset bounds) are chosen a priori from a library, not calibrated to data. The most fragile premise is that the price-setting sub-policy trained at low volatility remains effective at high volatility. The paper introduces no new entities.

free parameters (2)
  • Hawkes intensity parameters: mean_reversion_speed, baseline_arrival_rate, jump_size, dt = 60.0, 10.0, 40.0, 0.005
    Fixed values adopted from the mbt_gym HawkesArrivalModel; not fitted to real order flow. The self-exciting intensity directly controls execution probabilities and therefore the quoting behavior and wealth statistics that support the central claim.
  • Quote offset bounds = [0, 3]
    The bid/ask offsets are continuous actions in [0,3] for the Always Quoting sub-policy. The same bound is used at volatility 200, which may make quotes non-competitive when per-step price shocks are roughly 14, but this choice is fixed a priori and not tuned. It is a modeling constraint that affects the interpretation of adaptation.
assumptions (4)
  • domain assumption Price dynamics follow Brownian motion with drift: Z_{n+1} = Z_n + b_n Δt + σ_n W_n.
    Standard assumption from Avellaneda-Stoikov; not validated against any real price series in this paper. The volatility regime shift from 2 to 200 is applied within this model, so the whole stress test is model-internal.
  • domain assumption Execution intensity follows a Hawkes process with exponential decay of intensity toward baseline and a jump after each trade (Equation 1), with parameters from mbt_gym.
    The paper does not calibrate the Hawkes kernel to market data; it relies on the mbt_gym implementation as a realistic order-arrival model. The central claim's external validity depends on this being a faithful proxy.
  • domain assumption Zero-sum game between market maker and adversary, with the adversary's reward equal to the negative of the market maker's reward.
    Adversarial RL robustness results are only defined relative to this game structure; real markets are not zero-sum exactly with a market maker.
  • ad hoc to paper The Always Quoting MM sub-policy, trained at volatility 2, is a valid price setter when the 4-action MM operates at volatility 200.
    The 4-action agent chooses actions but not prices, so the entire conclusion depends on the fixed price-setting policy retaining competence after a 100x volatility jump. The paper provides no analysis of quote competitiveness at volatility 200.

how reviews work

0 comments
Cite this review

Pith. "Pith review of ARL-Based Multi-Action Market Making with Hawkes Processes and Variable Volatility." pith.science (2026). https://pith.science/paper/H4IDN2UT

@misc{pith2026250816589,
  author       = {Pith},
  title        = {Pith review of: ARL-Based Multi-Action Market Making with Hawkes Processes and Variable Volatility},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/H4IDN2UT}},
  note         = {Machine review of arXiv:2508.16589}
}
read the original abstract

We advance market-making strategies by integrating Adversarial Reinforcement Learning (ARL), Hawkes Processes, and variable volatility levels while also expanding the action space available to market makers (MMs). To enhance the adaptability and robustness of these strategies -- which can quote always, quote only on one side of the market or not quote at all -- we shift from the commonly used Poisson process to the Hawkes process, which better captures real market dynamics and self-exciting behaviors. We then train and evaluate strategies under volatility levels of 2 and 200. Our findings show that the 4-action MM trained in a low-volatility environment effectively adapts to high-volatility conditions, maintaining stable performance and providing two-sided quotes at least 92\% of the time. This indicates that incorporating flexible quoting mechanisms and realistic market simulations significantly enhances the effectiveness of market-making strategies.

Figures

Figures reproduced from arXiv: 2508.16589 by the authors.

Figure 1
Figure 1. Training and Testing Multi-Action Market Makers under Different Volatility Levels [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 3
Figure 3. Performance of Always Quoting MM (𝜂 = 0.0, 𝜁 = 0.0) [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Performance of 2-Action MM (Train & Test @ vol=2) (𝜂 = 0.0, 𝜁 = 0.0) [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figures from the paper (3 more)
Figure 2
Figure 2. Figure 2: Performance of Five Types of MM Agents (𝜂 = 0.0, 𝜁 = 0.0) Despite these observations, the market maker in this study per￾forms worse overall in terms of wealth mean 𝐸( Î 𝑁 ) and Sharpe ratio compared to results using the Poisson process. This could be due to several fa…
Figure 6
Figure 6. Figure 6: Performance of 4-Action MM (Train @ vol=2, Test @ vol=200) (𝜂 = 0.0, 𝜁 = 0.0) [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 7
Figure 7. Figure 7: Performance of 4-Action MM (Train & Test @ vol=200) (𝜂 = 0.0, 𝜁 = 0.0) [PITH_FULL_IMAGE:figures/full_fig_p007_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

32 extracted references · 27 canonical work pages

  1. [30]

    Ziyi Wang, Carmine Ventre, and Maria Polukarov. 2023. Robust Market Mak- ing: To Quote, or not To Quote. InProceedings of the Fourth ACM International Conference on AI in Finance . 664–672

  2. [1]

    Frédéric Abergel and Aymen Jedidi. 2015. Long-time behavior of a hawkes process–based limit order book. SIAM Journal on Financial Mathematics 6, 1 (2015), 1026–1043

  3. [2]

    Jacob Abernethy and Satyen Kale. 2013. Adaptive market making via online learning. Advances in Neural Information Processing Systems 26 (2013)

  4. [3]

    Marco Avellaneda and Sasha Stoikov. 2008. High-frequency trading in a limit order book. Quantitative Finance 8, 3 (2008), 217–224

  5. [4]

    Emmanuel Bacry, Sylvain Delattre, Marc Hoffmann, and Jean-François Muzy

  6. [5]

    Emmanuel Bacry, Thibault Jaisson, and Jean-François Muzy. 2016. Estimation of slowly decreasing hawkes kernels: application to high-frequency order book dynamics. Quantitative Finance 16, 8 (2016), 1179–1201

  7. [6]

    Emmanuel Bacry, Iacopo Mastromatteo, and Jean-François Muzy. 2015. Hawkes processes in finance. Market Microstructure and Liquidity 1, 01 (2015), 1550005

  8. [7]

    Emmanuel Bacry and Jean-François Muzy. 2014. Hawkes model for price and trades high-frequency dynamics. Quantitative Finance 14, 7 (2014), 1147–1166

Show all 32 references
  1. [8]

    Álvaro Cartea, Sebastian Jaimungal, and José Penalva. 2015. Algorithmic and high-frequency trading. Cambridge University Press

  2. [9]

    Nicholas Tung Chan and Christian Shelton. 2001. An electronic market-maker. (2001)

  3. [10]

    London Stock Exchange. 2022. ETF & ETP Market Maker Obliga- tions. https://docs.londonstockexchange.com/sites/default/files/documents/ Market%20Maker%20Obligations%20Factsheet_18.05.2022.pdf

  4. [11]

    Pietro Fodra and Mauricio Labadie. 2012. High-frequency market-making with inventory constraints and directional bets. arXiv preprint arXiv:1206.4810 (2012)

  5. [12]

    Olivier Guéant. 2017. Optimal market making. Applied Mathematical Finance 24, 2 (2017), 112–154

  6. [13]

    Olivier Guéant, Charles-Albert Lehalle, and Joaquin Fernandez-Tapia. 2013. Deal- ing with the inventory risk: a solution to the market making problem. Mathe- matics and financial economics 7, 4 (2013), 477–507

  7. [14]

    Qi Guo, Bruno Remillard, and Anatoliy Swishchuk. 2020. Multivariate general compound point processes in limit order books. Risks 8, 3 (2020), 98

  8. [15]

    Qi Guo and Anatoliy Swishchuk. 2020. Multivariate general compound Hawkes processes and their applications in limit order books. Wilmott 2020, 107 (2020), 42–51

  9. [16]

    Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen, George Tucker, Sehoon Ha, Jie Tan, Vikash Kumar, Henry Zhu, Abhishek Gupta, Pieter Abbeel, et al. 2018. Soft actor-critic algorithms and applications. arXiv preprint arXiv:1812.05905 (2018)

  10. [17]

    Alan G Hawkes. 1971. Spectra of some self-exciting and mutually exciting point processes. Biometrika 58, 1 (1971), 83–90

  11. [18]

    Patrick Hewlett. 2006. Clustering of order arrivals, price impact and trade path optimisation. In Workshop on Financial Modeling with Jump processes, Ecole Poly- technique. Citeseer, 6–8

  12. [19]

    Thomas Ho and Hans R Stoll. 1981. Optimal dealer pricing under transactions and return uncertainty. Journal of Financial economics 9, 1 (1981), 47–73

  13. [20]

    Joseph Jerome, Leandro Sánchez-Betancourt, Rahul Savani, and Martin Herdegen

  14. [21]

    Jeremy Large. 2007. Measuring the resiliency of an electronic limit order book. Journal of Financial Markets 10, 1 (2007), 1–25

  15. [22]

    Xiaofei Lu and Frédéric Abergel. 2017. Limit order book modelling with high dimensional Hawkes processes. (2017)

  16. [23]

    Xiaofei Lu and Frédéric Abergel. 2018. High-dimensional Hawkes processes for limit order books: modelling, empirical analysis and numerical calibration. Quantitative Finance 18, 2 (2018), 249–264

  17. [24]

    Deutsche Börse Cash Market. 2019. MiFID II: Regulated Market Maker . https: //www.xetra.com/resource/blob/2450614/def8bfa58c17bd971bb69ec14f2a024f/ data/Regulated-Market-Maker-Handbook.pdf

  18. [25]

    Ana Roldan Contreras and Anatoliy Swishchuk. 2022. Optimal Liquidation, Acquisition and Market Making Problems in HFT under Hawkes Models for LOB. Risks 10, 8 (2022), 160

  19. [26]

    Thomas Spooner, John Fearnley, Rahul Savani, and Andreas Koukorinis. 2018. Market making via reinforcement learning. arXiv preprint arXiv:1804.04216 (2018)

  20. [27]

    Thomas Spooner and Rahul Savani. 2020. Robust market making via adversarial reinforcement learning. arXiv preprint arXiv:2003.01820 (2020)

  21. [28]

    Anatoliy Swishchuk and Aiden Huffman. 2020. General compound Hawkes processes in limit order books. Risks 8, 1 (2020), 28

  22. [29]

    Ioane Muni Toke and Fabrizio Pomponio. 2012. Modelling trades-through in a limit order book using Hawkes processes. Economics 6, 1 (2012), 20120022

  23. [2013]

    Quantitative finance 13, 1 (2013), 65–77

    Modelling microstructure noise with mutually exciting point processes. Quantitative finance 13, 1 (2013), 65–77

  24. [2023]

    In Proceedings of the Fourth ACM International Conference on AI in Finance

    Mbt-gym: Reinforcement learning for model-based limit order book trading. In Proceedings of the Fourth ACM International Conference on AI in Finance . 619– 627

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.