Pith. sign in

REVIEW 3 major objections 5 minor 112 references

Robustness in Sequential Decision Making under Evolving Uncertainty: Evidence from High-Frequency Market Making

T0 review · 3 major / 5 minor · reviewed 2026-07-10 · grok-4.5

Pith's one-line read In high-frequency market making, how you respond to uncertainty reshapes quoting more than how much uncertainty you admit.

desk verdict Clean two-knob robustness story for HFT market making with solid directional evidence that δ dominates ε̄ and liquidity modulates value; the shared exponential fill model is a real but disclosed soft spot, not a collapse of the claim. read the letter →

arxiv 2607.08291 v1 pith:463REQGS submitted 2026-07-09 q-fin.TR q-fin.MFq-fin.RM

classification q-fin.TRq-fin.MFq-fin.RM
keywords RobustReinforcementLearningSequentialDecisionMakingModelUncertaintyHigh-FrequencyMarketDistributionallyOptimizationSinkhornAmbiguitySetsActionRobustnessTolerance
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper studies how a market maker should set quotes when the true market dynamics keep changing and any fixed model is therefore misspecified. It argues that robustness is not just a statistical safety net: it is a two-dimensional, state-dependent rule that changes how the agent quotes, manages inventory and controls risk. One dimension (uncertainty tolerance) sets how large a neighbourhood of alternative transition laws the agent is willing to entertain; the other (action robustness) sets how conservatively the agent redistributes probability inside that neighbourhood. Simulation under controlled stresses and out-of-sample tests on liquid and illiquid equities show that the second dimension moves spreads, quantities and risk-adjusted performance far more than the first. The same evidence shows that the value of robustness is liquidity-dependent: in deep markets it improves Sharpe ratios and drawdowns while still allowing execution; in thin markets excessive robustness can starve the agent of fills and cut profitability.

What carries the argument

Sinkhorn-based ambiguity sets around a reference transition law, dualized into a robust Bellman operator whose two parameters (shifted radius ¯ε and entropic regularisation δ) separately control uncertainty tolerance and action robustness; the resulting fitted actor–critic learns adaptive bid/ask spreads and quantities under that operator.

What would settle it

Re-estimate fill rates from the same order-book data under an adverse-selection or multi-agent model; if the liquidity ranking of robust versus non-robust Sharpe ratios reverses or disappears once fills deviate from the exponential form, the claim that action robustness is the dominant and liquidity-dependent lever is falsified.

Watch

Extended reading notes

Core claim

Robustness in sequential market making has two economically distinct dimensions—uncertainty tolerance (how much model deviation is admitted) and action robustness (how conservatively decisions respond inside the admitted set)—and action robustness exerts a substantially larger effect on quoting, inventory paths and risk-adjusted performance. Robustness therefore reshapes the state-to-action map itself rather than merely protecting terminal P&L, and its net value is positive mainly when execution opportunities remain plentiful.

Load-bearing premise

Fill probabilities are treated as known exponential functions of quoted distance that are identical in simulation and on real data and independent of adverse selection or competing market makers.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper develops a distributionally robust RL framework for high-frequency market making that places Sinkhorn ambiguity sets on the stochastic innovation law of a finite-horizon MDP. It decomposes robustness into two parameters: a shifted radius ¯ε (uncertainty tolerance, size of the ambiguity set) and an entropic regularization δ (action robustness, how the worst-case measure is tilted). A dual Sinkhorn Bellman operator yields a fitted actor–critic algorithm. Simulation under six modular stress regimes and a 2×2 empirical design (AAPL/TSLA/MKC/TWLO × 2019 vs COVID-2020) are used to argue that (i) robustness reshapes state-dependent quoting and inventory paths rather than only terminal metrics, (ii) δ has a substantially larger behavioral and performance impact than ¯ε, and (iii) robustness improves risk-adjusted outcomes mainly in liquid markets and can reduce profitability when execution opportunities are scarce.

Significance. If the two-dimensional robustness interpretation and the liquidity-dependent value of robustness hold under more realistic execution, the paper would give practitioners separate, economically interpretable levers for model risk versus decision conservatism, and would push robust RL beyond one-parameter worst-case guarantees. Strengths include a clean deterministic/stochastic state split, an explicit dual Bellman form with value-iteration contraction, modular stress scenarios that isolate distinct misspecifications, a held-out COVID distribution-shift design, and extensive policy diagnostics (intraday paths, state-conditional maps, Pareto frontiers). The contribution is therefore potentially useful for both robust sequential decision theory and market-microstructure practice, provided the main comparative claims are more tightly quantified and stress-tested against the shared fill map.

major comments (3)
  1. [Appendix A.1.4; §4.2; Table 3; §5] Appendix A.1.4 and the empirical protocol: executed quantities in both simulation and real-data evaluation are generated from the same exponential fill map λ = A e^{−κδ} (and the same square-root terminal liquidation). Section 5 correctly lists adverse selection, competing makers, and depth-dependent fills as limitations, but the central claims—that δ dominates ¯ε and that excessive robustness “limits execution opportunities” in illiquid names (abstract, §4.2, Table 3)—are conditioned on this fixed fill map. Because the Sinkhorn adversary only perturbs innovations around that map, comparative rankings of (¯ε, δ) and the liquidity interaction may be artifacts of fill misspecification. At minimum the paper should re-evaluate greedy vs robust policies under alternative fill specifications (e.g., depth-dependent or adverse-selection-adjusted fills) or show that the δ-vs-¯ε ranking and the HL
  2. [§4.1.1–4.1.2; Table 1; Figure 1; Appendix B heatmaps] The claim that “action robustness has a substantially larger impact than uncertainty tolerance” (abstract, finding (2), §4.1.2) rests mainly on visual comparison of four (¯ε, δ) corners in Figure 1 and qualitative reading of heatmaps (Figs. 5–6, 23–24). Table 1 further reports that the same two pairs—(0.0001, 0.1) for Val PnL and (4, 1) for Val Sharpe—are selected in all six simulation scenarios, which weakens the claim of state-dependent calibration and makes the δ-dominance story look like a corner-effect of the grid. A load-bearing revision should quantify the relative contribution of δ versus ¯ε (e.g., partial derivatives of policy/performance along each axis, ANOVA-style decomposition over the full grid, or hold-one-fixed sweeps with confidence bands) rather than relying on four labeled policies and identical selected pairs across environments.
  3. [§4.2.1; Table 3; Figures 4, 26–27] In the low-liquidity cells of Table 3 (MKC LL–LV, TWLO LL–HV), robust policies often improve P&L or MDD modestly but leave Sharpe near zero or negative (e.g., MKC 2020 Sharpe ≈ −0.7 to −0.8 across greedy and robust). The narrative that “excessive robustness may reduce profitability in illiquid markets by limiting execution opportunities” is therefore only partially supported: the paper shows weaker gains, not a clean demonstration that robustness itself rationed fills. Direct evidence—fill rates, participation rates, or opportunity counts by liquidity tier under high-δ policies—should be reported so that the mechanism is distinguished from simply “harder markets where no policy works well.”
minor comments (5)
  1. [§2.3.2; Table 1 notes] Notation for the shifted radius switches between ¯ε, ¯ε_{x,a}, and ϵ in tables/notes (e.g., Table 1 notes write (¯ϵ, δ)). Unify symbols and state once whether the reported grid is the shifted or original Sinkhorn radius.
  2. [Lemma 2.1; §2.2.2] Lemma 2.1 assumes martingale mid-price and fill–price independence; these are used to justify the additive reward but are not revisited when interpreting inventory risk under price stress. A short remark on when the decomposition fails would help.
  3. [Appendices B–C] Several appendix figures (e.g., Shapley panels, distribution plots) are dense; consider moving a subset to an online supplement and keeping only the diagnostics that directly support δ-vs-¯ε and liquidity claims in the main text.
  4. [References; Appendix C.4] Typos and wording: “stategies” (Cartea–Jaimungal citation), “countvvgbherparts” in C.4, and occasional missing spaces around math. A careful copy-edit pass is needed.
  5. [§2.1.3; §3.2] Clarify early that order quantities are participation rates in [0,1] of contemporaneous volume (footnote in §2.1.3 / §3.2); some readers will otherwise misread absolute share sizes.

Circularity Check

0 steps flagged · score 1.0 of 10

No load-bearing circularity: ¯ε and δ are free Sinkhorn design parameters whose relative impact is measured out-of-sample, not forced by definition or self-citation.

full rationale

The paper’s central claims—that robustness has two economically distinct dimensions (uncertainty tolerance ¯ε vs action robustness δ), that δ has a substantially larger behavioral impact, and that excessive robustness can hurt in illiquid markets—are not derived by construction from their inputs. ¯ε and δ enter as free parameters of the Sinkhorn dual Bellman operator (Eqs. 2.26–2.29, Corollary 2.3); the paper then varies them on a grid and reports empirical/simulation differences in quoting, inventory, Sharpe, and P&L against held-out stress scenarios and real 2019/2020 LOB data (Tables 1–3, Figs. 1–4). Validation selection and out-of-sample test windows (including the COVID shift) are disclosed; the ranking “δ matters more than ¯ε” is a comparative static, not an identity. The shared exponential fill map (Appendix A.1.4) is a modeling assumption that can misspecify real LOBs, but that is a correctness risk, not circularity: fills are not fitted to the same performance metrics later reported as predictions. Self-citations (e.g., Lu–Sester–Zhang DRO RL) supply algorithmic background and are not used as uniqueness theorems that force the two-dimensional claim. The derivation chain is therefore self-contained against external benchmarks (greedy, AS, random, fixed); score 1 only for ordinary non-load-bearing self-reference in the methods lineage.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

The central empirical claims rest on a standard inventory-based market-making MDP, a data-driven deep-ensemble reference transition, Sinkhorn duality taken from prior work, an assumed exponential fill model, and a large set of free robustness and training hyper-parameters selected on validation performance. No new physical entities are postulated; the free parameters and domain modeling choices are the main load-bearing inputs.

free parameters (6)
  • shifted Sinkhorn radius ¯ε (uncertainty tolerance)
    Grid-searched and selected by validation PnL or Sharpe; directly controls size of ambiguity set and therefore the reported robustness effects.
  • Sinkhorn regularization δ (action robustness)
    Grid-searched; claimed to dominate policy behavior; free design choice whose magnitude is not derived from data.
  • risk-aversion schedule γ_t and baseline γ
    Hand-chosen functional form that increases inventory penalty near close; shapes terminal inventory behavior.
  • fill intensity A and elasticity κ
    Calibrated from LOB data then held fixed for both simulation and empirical evaluation; determine execution probabilities.
  • deep-ensemble size K, network widths, learning rates, discount α, look-back m
    Training hyper-parameters that affect the reference transition and the learned policy; listed in Table 5 and Appendix.
  • liquidation impact factor η
    Square-root impact coefficient used for terminal inventory cost; free scaling constant.
assumptions (5)
  • domain assumption One-step mid-price increments are martingale and fills are conditionally independent of price innovations (Lemma 2.1).
    Required to rewrite terminal wealth as additive one-period rewards; standard but strong for high-frequency data.
  • standard math Strong duality for the Sinkhorn primal-dual pair holds under the stated cost and support conditions (Wang et al. 2025).
    Invoked to replace the infinite-dimensional worst-case expectation by a finite-dimensional dual (Eqs. 2.26–2.29).
  • domain assumption Reference transition is adequately captured by a deep ensemble of Gaussian/log-normal conditionals; ambiguity is only around the stochastic innovation.
    Defines the center of every Sinkhorn ball; misspecification of the ensemble is not itself robustified.
  • domain assumption Single representative market maker, no adverse selection, no competing liquidity providers.
    Explicitly adopted in §2.1 and listed as a limitation in §5; isolates robustness effects but omits a first-order market-making risk.
  • ad hoc to paper Exponential fill probability λ = A exp(−κδ) is the true execution mechanism for both simulation and real-data evaluation.
    Assumed known and identical across environments (Appendix A.1.4); never estimated jointly with the policy.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Robustness in Sequential Decision Making under Evolving Uncertainty: Evidence from High-Frequency Market Making." pith.science (2026). https://pith.science/paper/463REQGS

@misc{pith2026260708291,
  author       = {Pith},
  title        = {Pith review of: Robustness in Sequential Decision Making under Evolving Uncertainty: Evidence from High-Frequency Market Making},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/463REQGS}},
  note         = {Machine review of arXiv:2607.08291}
}
read the original abstract

We study sequential decision making under evolving uncertainty in high-frequency financial markets, where changing market dynamics continually challenge static decision policies. We show that robustness has two economically meaningful dimensions: uncertainty tolerance, which determines how much uncertainty the decision maker allows, and action robustness, which governs how conservatively decisions respond. Robustness is not merely protection against model misspecification, but a state-dependent mechanism that reshapes sequential decision behaviors. Simulation and empirical evidence show that action robustness has a substantially larger impact than uncertainty tolerance. Moreover, excessive robustness may reduce profitability in illiquid markets by limiting execution opportunities.

Figures

Figures reproduced from arXiv: 2607.08291 by the authors.

Figure 1
Figure 1. Intraday average quoting behavior under the Price Stress scenario for different robustness parameter combinations. The four panels report bid spread, ask spread, bid quantity, and ask quantity [PITH_FULL_IMAGE:figures/full_fig_p025_1.png] view at source ↗
Figure 2
Figure 2. Cross-scenario intraday quoting behavior for the best robust policy selected using validation Sharpe ratio [PITH_FULL_IMAGE:figures/full_fig_p026_2.png] view at source ↗
Figure 3
Figure 3. Cross-scenario state-conditional policy responses for the best robust policy selected using validation Sharpe ratio. the quoted spread. For both simulation and empirical evaluation, any inventory IT remaining at the terminal step is liquidated at a cost given by the square-root price-impact model described in Appendix A.1 as well. Our empirical analysis is organized around two key market characteristics: liquidity a… view at source ↗
Figures from the paper (38 more)
Figure 4
Figure 4. Figure 4: Out-of-sample risk–return frontiers during 2020 stress regime under robustness. Robustness substantially expands the attainable frontier for the highly liquid assets AAPL (HL–LV) and TSLA (HL–HV), while the gains are more modest for MKC (LL–LV) and TWLO (LL–HV). Highli…
Figure 5
Figure 5. Figure 5: Test-period mean PnL and Sharpe ratio over the (¯ε, δ) grid for the stable, price-stress, and liquidity-dry-out scenarios. The first row reports mean PnL and the second row reports Sharpe ratio, with columns corresponding to the three scenarios; the greedy benchmark va…
Figure 6
Figure 6. Figure 6: Test-period mean PnL and Sharpe ratio over the (¯ε, δ) grid for the buy-arrival-imbalance, sell-arrival-imbalance, and fill-stress scenarios. The first row reports mean PnL and the second row reports Sharpe ratio, with columns corresponding to the three scenarios; the …
Figure 7
Figure 7. Figure 7: Test-period Pareto frontiers in mean PnL-versus-Sharpe-ratio space for the six simulation stress scenarios. The panels show that the profitability–risk trade-off varies across scenarios, with some robustness choices delivering better risk-adjusted performance than othe…
Figure 8
Figure 8. Figure 8: Intraday average bid spread, ask spread, bid quantity, and ask quantity in the stable baseline scenario for the greedy benchmark and the four robust policies. The figure shows that increasing δ has a stronger effect than increasing ¯ε, mainly through competitive spread…
Figure 9
Figure 9. Figure 9: State-conditional mean bid quantity, ask quantity, bid spread, and ask spread of the robust and greedy agents in the stable baseline scenario, with ±1 s.d. bands. The robust policy preserves the same broad state dependence as the greedy benchmark but becomes more selec…
Figure 10
Figure 10. Figure 10: showcases the quoting behavior in the liquidity dry-out scenario [PITH_FULL_IMAGE:figures/full_fig_p044_10.png]
Figure 11
Figure 11. Figure 11: State-conditional mean bid quantity, ask quantity, bid spread, and ask spread of the robust and greedy agents in the liquidity dry-out scenario, with ±1 s.d. bands. The robust policy remains sensitive to the same broad state signals as the greedy benchmark but quotes …
Figure 12
Figure 12. Figure 12: Intraday average bid spread, ask spread, bid quantity, and ask quantity in the buy-arrival-imbalance scenario for the greedy benchmark and the four robust policies. The figure shows that stronger robustness mainly affects the policy through δ, with the best robust pol…
Figure 13
Figure 13. Figure 13: State-conditional mean bid quantity, ask quantity, bid spread, and ask spread of the robust and greedy agents in the buy-arrival-imbalance scenario, with ±1 s.d. bands. The robust policy preserves the directional dependence on the state variables but adjusts the ask s…
Figure 14
Figure 14. Figure 14: Intraday average bid spread, ask spread, bid quantity, and ask quantity in the sell-arrival-imbalance scenario for the greedy benchmark and the four robust policies. The figure shows that stronger robustness mainly affects the policy through δ, with the best robust po…
Figure 15
Figure 15. Figure 15: State-conditional mean bid quantity, ask quantity, bid spread, and ask spread of the robust and greedy agents in the sell-arrival-imbalance scenario, with ±1 s.d. bands. The robust policy preserves the directional dependence on the state variables but adjusts the bid …
Figure 16
Figure 16. Figure 16: Intraday average bid spread, ask spread, bid quantity, and ask quantity in the fill-stress scenario for the greedy benchmark and the four robust policies. The figure shows that stronger robustness mainly affects the policy through δ, with the best robust policy wideni…
Figure 17
Figure 17. Figure 17: State-conditional mean bid quantity, ask quantity, bid spread, and ask spread of the robust and greedy agents in the fill-stress scenario, with ±1 s.d. bands. The robust policy preserves the same broad state dependence as the greedy benchmark but adjusts both spreads …
Figure 18
Figure 18. Figure 18: Intraday mean inventory paths of the robust and greedy agents in the six simulation scenarios. The panels show that the robust policy generally stabilizes inventory more effectively, with the largest reductions in inventory exposure appearing in the stressed scenarios…
Figure 19
Figure 19. Figure 19: Train, validation, and test distributions of all state variables for AAPL in 2019 and 2020. The comparison makes the stronger cross-split distributional shift in 2020 visually apparent [PITH_FULL_IMAGE:figures/full_fig_p053_19.png]
Figure 20
Figure 20. Figure 20: Train, validation, and test distributions of all state variables for TSLA in 2019 and 2020. The comparison makes the stronger cross-split distributional shift in 2020 visually apparent [PITH_FULL_IMAGE:figures/full_fig_p054_20.png]
Figure 21
Figure 21. Figure 21: Train, validation, and test distributions of all state variables for MKC in 2019 and 2020. The comparison makes the stronger cross-split distributional shift in 2020 visually apparent [PITH_FULL_IMAGE:figures/full_fig_p055_21.png]
Figure 22
Figure 22. Figure 22: Train, validation, and test distributions of all state variables for TWLO in 2019 and 2020. The comparison makes the stronger cross-split distributional shift in 2020 visually apparent. C.2. Sensitivity to Uncertainty Tolerance (ε¯) and Action Robustness (δ). Figures …
Figure 23
Figure 23. Figure 23: Test-period Sharpe ratio Difference between Greedy Policy and Robust Policy across the (¯ε, δ) grid, with the 2019 stock panels in the first row and the 2020 panels in the second row. The greedy benchmark is reported below each panel. The panels show that the region o…
Figure 24
Figure 24. Figure 24: Test-period mean P&L Difference between Greedy Policy and Robust Policy across the (¯ε, δ) grid, with the 2019 stock panels in the first row and the 2020 panels in the second row. The panels show that robustness often trades off some average profitability against impr…
Figure 25
Figure 25. Figure 25: Out-of-sample test Pareto frontiers during the 2019 regime. Each point corresponds to a robust hyperparameter configuration, and highlighted points denote the policies selected from the validation frontier. The 2019 frontiers are generally tighter than their 2020 coun…
Figure 26
Figure 26. Figure 26: Mean quote (bid quantity, ask quantity, bid spread, ask spread) of robust strategies against baselines for AAPL and TSLA. Robust hyperparameters are those selected from the best validation Sharpe and best validation PnL. For the AS strategy, the maximum spread is 0.6.…
Figure 27
Figure 27. Figure 27: Mean quote (bid quantity, ask quantity, bid spread, ask spread) of robust strategies against baselines for MKC and TWLO. Robust hyperparameters are those selected from the best validation Sharpe and best validation PnL. For the AS strategy, the maximum spread is 0.6. …
Figure 28
Figure 28. Figure 28: Average test Sharpe ratio of the robust agent conditional on the original Sinkhorn radius ε (markers), with the greedy benchmark shown as a dashed line. Across panels, larger values of ε often deliver comparable or higher Sharpe ratios, especially in 2020, although th…
Figure 29
Figure 29. Figure 29: Distribution of absolute Shapley values for greedy and robust policies with ¯ε = 1 and δ = 0.1 on AAPL, 2019, where the robust policy is the best test Sharpe Pareto configuration. The figure shows that both policies are driven primarily by trade-flow and timing variab…
Figure 30
Figure 30. Figure 30: Distribution of absolute Shapley values for greedy and robust policies with ¯ε = 1 and δ = 1 on AAPL, 2020, where the robust policy is the best test Sharpe Pareto configuration. The figure shows that both policies are driven primarily by trade-flow and timing variable…
Figure 31
Figure 31. Figure 31: Distribution of absolute Shapley values for greedy and robust policies with ¯ε = 1 and δ = 0.1 on TSLA, 2019, where the robust policy is the best test Sharpe Pareto configuration. As in AAPL, the dominant drivers are trade-flow and timing variables, but the spread com…
Figure 32
Figure 32. Figure 32: Distribution of absolute Shapley values for greedy and robust policies with ¯ε = 1 and δ = 0.1 on TSLA, 2020, where the robust policy is the best test Sharpe Pareto configuration. The figure again points to trade-flow and timing variables as the main drivers, with rob…
Figure 33
Figure 33. Figure 33: Distribution of absolute Shapley values for greedy and robust policies with ε¯ = 0.0001 and δ = 0.1 on MKC, 2019, where the robust policy is the best test Sharpe Pareto configuration. The importance profiles are comparatively stable across the greedy and robust polici…
Figure 34
Figure 34. Figure 34: Distribution of absolute Shapley values for greedy and robust policies with ¯ε = 1 and δ = 0.01 on MKC, 2020, where the robust policy is the best test Sharpe Pareto configuration. The same broad set of trade-flow and timing variables remains dominant, and the robust s…
Figure 35
Figure 35. Figure 35: Distribution of absolute Shapley values for greedy and robust policies with ¯ε = 0.01 and δ = 0.01 on TWLO, 2019, where the robust policy is the best test Sharpe Pareto configuration. The main drivers are again trade-flow and timing variables, although the relative we…
Figure 36
Figure 36. Figure 36: Distribution of absolute Shapley values for greedy and robust policies with ε¯ = 0.0001 and δ = 1 on TWLO, 2020, where the robust policy is the best test Sharpe Pareto configuration. The decomposition remains concentrated on a small set of trade-flow and timing variab…
Figure 37
Figure 37. Figure 37: Mean quote (bid quantity, ask quantity, bid spread, ask spread) of the robust and greedy agents conditional on each state feature for AAPL, with ±1 s.d. bands. The robust policy corresponds to the best test Sharpe Pareto configuration and responds structurally to stat…
Figure 38
Figure 38. Figure 38: Mean quote (bid quantity, ask quantity, bid spread, ask spread) of the robust and greedy agents conditional on each state feature for TSLA, with ±1 s.d. bands. The robust policy corresponds to the best test Sharpe Pareto configuration and responds structurally to stat…
Figure 39
Figure 39. Figure 39: Mean quote (bid quantity, ask quantity, bid spread, ask spread) of the robust and greedy agents conditional on each state feature for MKC, with ±1 s.d. bands. The robust policy corresponds to the best test Sharpe Pareto configuration and responds structurally to state…
Figure 40
Figure 40. Figure 40: Mean quote (bid quantity, ask quantity, bid spread, ask spread) of the robust and greedy agents conditional on each state feature for TWLO, with ±1 s.d. bands. The robust policy corresponds to the best test Sharpe Pareto configuration and responds structurally to stat…
Figure 41
Figure 41. Figure 41: Validation Pareto frontier (left) and out-of-sample test performance (right) in Sharpe vs. mean-P&L space. Colour encodes log10(δ); dashed line is the greedy benchmark. Robust configurations dominate the greedy benchmark out-of-sample across all stocks and periods, wi…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

112 extracted references · 112 canonical work pages

  1. [1]

    Mathematics and financial economics , volume=

    Dealing with the inventory risk: a solution to the market making problem , author=. Mathematics and financial economics , volume=. 2013 , publisher=

  2. [2]

    2011 , publisher=

    Robustness , author=. 2011 , publisher=

  3. [3]

    Applied Mathematical Finance , volume=

    Deep reinforcement learning for market making in corporate bonds: beating the curse of dimensionality , author=. Applied Mathematical Finance , volume=. 2019 , publisher=

  4. [4]

    Mathematics of Operations Research , volume=

    Distributionally robust stochastic optimization with Wasserstein distance , author=. Mathematics of Operations Research , volume=. 2023 , publisher=

  5. [5]

    Mathematical Programming , volume=

    Data-driven distributionally robust optimization using the Wasserstein metric: Performance guarantees and tractable reformulations , author=. Mathematical Programming , volume=. 2018 , publisher=

  6. [7]

    Deep Reinforcement Learning for Market Making Under a Hawkes Process-Based Limit Order Book Model , year=

    Gašperov, Bruno and Kostanjčar, Zvonko , journal=. Deep Reinforcement Learning for Market Making Under a Hawkes Process-Based Limit Order Book Model , year=

  7. [8]

    The Review of Financial Studies , volume=

    On the Sensitivity of Mean-Variance-Efficient Portfolios to Changes in Asset Means: Some Analytical and Computational Results , author=. The Review of Financial Studies , volume=

  8. [9]

    The Journal of Portfolio Management , volume=

    The Effect of Errors in Means, Variances, and Covariances on Optimal Portfolio Choice , author=. The Journal of Portfolio Management , volume=

Show all 112 references
  1. [10]

    Journal of Financial Economics , volume=

    Stock Return Predictability and Model Uncertainty , author=. Journal of Financial Economics , volume=

  2. [11]

    Applied Mathematical Finance , volume=

    Pricing and Hedging Derivative Securities in Markets with Uncertain Volatilities , author=. Applied Mathematical Finance , volume=

  3. [12]

    Review of Economic Dynamics , volume=

    Model Uncertainty and Liquidity , author=. Review of Economic Dynamics , volume=

  4. [13]

    Management Science , volume=

    Robust Mean-Variance Portfolio Selection: A Distributionally Robust Optimization Approach , author=. Management Science , volume=

  5. [14]

    Acta Numerica , volume=

    Distributionally robust optimization , author=. Acta Numerica , volume=. 2025 , publisher=

  6. [15]

    Management Science , volume=

    A Generalized Approach to Portfolio Optimization: Improving Performance by Constraining Portfolio Norms , author=. Management Science , volume=

  7. [16]

    Robust Control and Model Uncertainty , author=

  8. [17]

    Intraday market making with overnight inventory costs , journal =

    Adrian,Tobias and Capponi, Agostino and Fleming, Michael and Vogt, Erik and Zhang, Hongzhong , keywords =. Intraday market making with overnight inventory costs , journal =. 2020 , issn =. doi:https://doi.org/10.1016/j.finmar.2020.100564 , url =

  9. [18]

    2013 , publisher=

    Convergence of probability measures , author=. 2013 , publisher=

  10. [19]

    2004 , publisher=

    Probability essentials , author=. 2004 , publisher=

  11. [20]

    Journal of the Operational Research Society , volume=

    Dynamic programming and optimal control , author=. Journal of the Operational Research Society , volume=

  12. [22]

    2009 , publisher=

    Optimal transport: old and new , author=. 2009 , publisher=

  13. [23]

    Management Science , volume=

    Distributionally robust mean-variance portfolio selection with Wasserstein distances , author=. Management Science , volume=. 2022 , publisher=

  14. [25]

    Management Science , volume =

    Chen, Ningyuan and Kou, Steven and Wang, Chun , title =. Management Science , volume =. 2018 , doi =. https://doi.org/10.1287/mnsc.2016.2639 , abstract =

  15. [26]

    Quantitative Finance , volume =

    Avellaneda,Marco and Stoikov,Sasha , title =. Quantitative Finance , volume =. 2008 , publisher =

  16. [27]

    2013 , issn =

    Optimal trading strategy and supply/demand dynamics , journal =. 2013 , issn =. doi:https://doi.org/10.1016/j.finmar.2012.09.001 , author =

  17. [28]

    2013 , issn =

    Dealing with the inventory risk: a solution to the market making problem , journal =. 2013 , issn =. doi:https://doi.org/10.1007/s11579-012-0087-0 , author =

  18. [29]

    Optimal Make-Take Fees in a Multi Market-Maker Environment , journal =

    Baldacci, Bastien and Possama\". Optimal Make-Take Fees in a Multi Market-Maker Environment , journal =. 2021 , doi =

  19. [30]

    Applied Mathematical Finance , volume =

    Guéant,Olivier , title =. Applied Mathematical Finance , volume =. 2017 , publisher =

  20. [32]

    Buy Low, Sell High: A High Frequency Trading Perspective , journal =

    Cartea, \'. Buy Low, Sell High: A High Frequency Trading Perspective , journal =. 2014 , doi =

  21. [33]

    Applied Mathematical Finance , volume =

    Campi,Luciano and Zabaljauregui,Diego , title =. Applied Mathematical Finance , volume =. 2020 , publisher =

  22. [34]

    Mathematical Finance , volume =

    Guilbaud, Fabien and Pham, Huyên , title =. Mathematical Finance , volume =. doi:https://doi.org/10.1111/mafi.12042 , abstract =

  23. [35]

    Quantitative Finance , volume =

    Guilbaud,Fabien and Pham,Huyên , title =. Quantitative Finance , volume =. 2013 , publisher =

  24. [36]

    Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems , pages =

    Spooner, Thomas and Fearnley, John and Savani, Rahul and Koukorinis, Andreas , title =. Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems , pages =. 2018 , publisher =

  25. [37]

    Applied Mathematical Finance , volume =

    Guéant,Olivier and Manziuk,Iuliia , title =. Applied Mathematical Finance , volume =. 2019 , publisher =

  26. [38]

    Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence , articleno =

    Spooner, Thomas and Savani, Rahul , title =. Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence , articleno =. 2021 , isbn =

  27. [39]

    2019 , eprint=

    Reinforcement Learning for Market Making in a Multi-agent Dealer Market , author=. 2019 , eprint=

  28. [40]

    Online Robust Reinforcement Learning with Model Uncertainty , volume =

    Wang, Yue and Zou, Shaofeng , booktitle =. Online Robust Reinforcement Learning with Model Uncertainty , volume =

  29. [41]

    Kernel-Based Reinforcement Learning in Robust

    Lim, Shiau Hong and Autef, Arnaud , booktitle =. Kernel-Based Reinforcement Learning in Robust. 2019 , editor =

  30. [42]

    Reinforcement Learning under Model Mismatch , url =

    Roy, Aurko and Xu, Huan and Pokutta, Sebastian , booktitle =. Reinforcement Learning under Model Mismatch , url =

  31. [43]

    Proceedings of the 31st International Conference on Machine Learning , pages =

    Scaling Up Robust MDPs using Function Approximation , author =. Proceedings of the 31st International Conference on Machine Learning , pages =. 2014 , editor =

  32. [44]

    Robust Markov Decision Processes , journal =

    Wiesemann, Wolfram and Kuhn, Daniel and Rustem, Ber. Robust Markov Decision Processes , journal =

  33. [45]

    Robust Markov Decision Processes: Beyond Rectangularity , journal =

    Goyal, Vineet and Grand-Cl\'. Robust Markov Decision Processes: Beyond Rectangularity , journal =

  34. [46]

    and Singh, Sundeep and Moore, David , title =

    Goh, Joel and Bayati, Mohsen and Zenios, Stefanos A. and Singh, Sundeep and Moore, David , title =. Operations Research , volume =

  35. [47]

    Mathematics of Operations Research , volume =

    Xu, Huan and Mannor, Shie , title =. Mathematics of Operations Research , volume =

  36. [48]

    Strens, Malcolm J. A. , title =. Proceedings of the Seventeenth International Conference on Machine Learning , pages =. 2000 , isbn =

  37. [49]

    , title =

    Iyengar, Garud N. , title =. Mathematics of Operations Research , volume =

  38. [50]

    Operations Research , volume=

    Robust control of Markov decision processes with uncertain transition matrices , author=. Operations Research , volume=. 2005 , publisher=

  39. [51]

    Neural computation , volume=

    Robust reinforcement learning , author=. Neural computation , volume=. 2005 , publisher=

  40. [52]

    and Border, Kim C

    Aliprantis, Charalambos D. and Border, Kim C. , TITLE =. 2006 , PAGES =

  41. [53]

    High Frequency Traders and Asset Prices , journal =

    Cvitanic, Jaksa and Kirilenko, Andrei , year =. High Frequency Traders and Asset Prices , journal =

  42. [54]

    Quantitative Finance , volume =

    Marcello Rambaldi and Emmanuel Bacry and Fabrizio Lillo , title =. Quantitative Finance , volume =. 2017 , publisher =

  43. [55]

    Market Making With Signals Through Deep Reinforcement Learning , year=

    Gašperov, Bruno and Kostanjčar, Zvonko , journal=. Market Making With Signals Through Deep Reinforcement Learning , year=

  44. [56]

    Proceedings of the AAAI conference on artificial intelligence , volume=

    Deep reinforcement learning with double q-learning , author=. Proceedings of the AAAI conference on artificial intelligence , volume=

  45. [57]

    1921 , publisher=

    Risk, uncertainty and profit , author=. 1921 , publisher=

  46. [58]

    Mathematics , VOLUME =

    Gašperov, Bruno and Begušić, Stjepan and Posedel Šimović, Petra and Kostanjčar, Zvonko , TITLE =. Mathematics , VOLUME =. 2021 , NUMBER =

  47. [59]

    Proceedings of the Second ACM International Conference on AI in Finance , articleno =

    Zhao, Muchen and Linetsky, Vadim , title =. Proceedings of the Second ACM International Conference on AI in Finance , articleno =. 2022 , isbn =

  48. [60]

    Advances in Neural Information Processing Systems , volume=

    Adaptive market making via online learning , author=. Advances in Neural Information Processing Systems , volume=

  49. [61]

    2013 , publisher=

    Probability theory: a comprehensive course , author=. 2013 , publisher=

  50. [62]

    2007 , publisher=

    Measure theory , author=. 2007 , publisher=

  51. [63]

    1981 , author =

    Optimal dealer pricing under transactions and return uncertainty , journal =. 1981 , author =

  52. [64]

    Mathematical Finance , volume=

    Markov decision processes under model uncertainty , author=. Mathematical Finance , volume=

  53. [65]

    The Review of Financial Studies , volume =

    Korajczyk, Robert A and Murphy, Dermot , title = ". The Review of Financial Studies , volume =. 2018 , month =

  54. [68]

    Advances in Neural Information Processing Systems , volume=

    Simple and scalable predictive uncertainty estimation using deep ensembles , author=. Advances in Neural Information Processing Systems , volume=

  55. [69]

    Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models , booktitle =

    Kurtland Chua and Roberto Calandra and Rowan McAllister and Sergey Levine , editor =. Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models , booktitle =. 2018 , url =

  56. [70]

    Journal of Financial Econometrics , year =

    Cont, Rama and Kukanov, Arseniy and Stoikov, Sasha , title =. Journal of Financial Econometrics , year =

  57. [71]

    Quantitative Finance , volume =

    Sasha Stoikov , title =. Quantitative Finance , volume =. 2018 , publisher =. doi:10.1080/14697688.2018.1489139 , URL =

  58. [72]

    and Bollerslev, Tim , title =

    Andersen, Torben G. and Bollerslev, Tim , title =. Journal of Finance , year =

  59. [73]

    2015 , publisher=

    Algorithmic and High-Frequency Trading , author=. 2015 , publisher=

  60. [75]

    International Conference on Learning Representations , year =

    Loshchilov, Ilya and Hutter, Frank , title =. International Conference on Learning Representations , year =

  61. [76]

    Risk , volume =

    Almgren, Robert and Thum, Chee and Hauptmann, Emmanuel and Li, Hong , title =. Risk , volume =

  62. [77]

    Anomalous Price Impact and the Critical Nature of Liquidity in Financial Markets , journal =

    T. Anomalous Price Impact and the Critical Nature of Liquidity in Financial Markets , journal =

  63. [78]

    Direct estimation of equity market impact

    Robert Almgren, Chee Thum, Emmanuel Hauptmann, and Hong Li. Direct estimation of equity market impact. Risk, 18 0 (7): 0 58--62, 2005

  64. [79]

    Andersen and Tim Bollerslev

    Torben G. Andersen and Tim Bollerslev. Deutsche mark--dollar volatility: Intraday activity patterns, macroeconomic announcements, and longer run dependencies. Journal of Finance, 53 0 (1): 0 219--265, 1998

  65. [80]

    High-frequency trading in a limit order book

    Marco Avellaneda and Sasha Stoikov. High-frequency trading in a limit order book. Quantitative Finance, 8 0 (3): 0 217--224, 2008. doi:https://doi.org/10.1080/14697680701381228

  66. [81]

    Pricing and hedging derivative securities in markets with uncertain volatilities

    Marco Avellaneda, Arnon Levy, and Antonio Par \'a s. Pricing and hedging derivative securities in markets with uncertain volatilities. Applied Mathematical Finance, 2 0 (2): 0 73--88, 1995

  67. [82]

    Stock return predictability and model uncertainty

    Doron Avramov. Stock return predictability and model uncertainty. Journal of Financial Economics, 64 0 (3): 0 423--458, 2002

  68. [83]

    Optimal make-take fees in a multi market-maker environment

    Bastien Baldacci, Dylan Possama\" , and Mathieu Rosenbaum. Optimal make-take fees in a multi market-maker environment. SIAM Journal on Financial Mathematics, 12 0 (1): 0 446--486, 2021. doi:10.1137/19M1277412

  69. [84]

    Best and Robert R

    Michael J. Best and Robert R. Grauer. On the sensitivity of mean-variance-efficient portfolios to changes in asset means: Some analytical and computational results. The Review of Financial Studies, 4 0 (2): 0 315--342, 1991

  70. [85]

    Convergence of probability measures

    Patrick Billingsley. Convergence of probability measures. John Wiley & Sons, 2013

  71. [86]

    Algorithmic and High-Frequency Trading

    \'A lvaro Cartea, Sebastian Jaimungal, and Jos \'e Penalva. Algorithmic and High-Frequency Trading. Cambridge University Press, 2015

  72. [87]

    Risk metrics and fine tuning of high-frequency trading stategies

    Álvaro Cartea and Sebastian Jaimungal. Risk metrics and fine tuning of high-frequency trading stategies. Mathematical Finance, 25 0 (3): 0 576--611, 2015. doi:https://doi.org/10.1111/mafi.12023

  73. [88]

    Chopra and William T

    Vijay K. Chopra and William T. Ziemba. The effect of errors in means, variances, and covariances on optimal portfolio choice. The Journal of Portfolio Management, 19 0 (2): 0 6--11, 1993

  74. [89]

    Deep reinforcement learning in a handful of trials using probabilistic dynamics models

    Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine. Deep reinforcement learning in a handful of trials using probabilistic dynamics models. In Samy Bengio, Hanna M. Wallach, Hugo Larochelle, Kristen Grauman, Nicol \` o Cesa - Bianchi, and Roman Garnett, edito...

  75. [90]

    The price impact of order book events

    Rama Cont, Arseniy Kukanov, and Sasha Stoikov. The price impact of order book events. Journal of Financial Econometrics, 12 0 (1): 0 47--88, 2014

  76. [91]

    Distributionally robust optimization under moment uncertainty with application to data-driven problems

    Erick Delage and Yinyu Ye. Distributionally robust optimization under moment uncertainty with application to data-driven problems. Operations Research, 58 0 (3): 0 595--612, 2010. doi:10.1287/opre.1090.0741. URL https://doi.org/10.1287/opre.1090.0741

  77. [92]

    Distributionally robust stochastic optimization with wasserstein distance

    Rui Gao and Anton Kleywegt. Distributionally robust stochastic optimization with wasserstein distance. Mathematics of Operations Research, 48 0 (2): 0 603--655, 2023

  78. [93]

    Limit theorems for markovian hawkes processes with a large initial intensity

    Xuefeng Gao and Lingjiong Zhu. Limit theorems for markovian hawkes processes with a large initial intensity. Stochastic Processes and their Applications, 128 0 (11): 0 3807--3839, 2018. ISSN 0304-4149. doi:https://doi.org/10.1016/j.spa.2017.12.001. URL https://www.sciencedirec...

  79. [94]

    Zenios, Sundeep Singh, and David Moore

    Joel Goh, Mohsen Bayati, Stefanos A. Zenios, Sundeep Singh, and David Moore. Data uncertainty in markov chains: Application to cost-effectiveness analyses of medical innovations. Operations Research, 66 0 (3): 0 697--715, 2018

  80. [95]

    Robust markov decision processes: Beyond rectangularity

    Vineet Goyal and Julien Grand-Cl\' e ment. Robust markov decision processes: Beyond rectangularity. Mathematics of Operations Research, 48 0 (1): 0 203--226, 2023

  81. [96]

    Deep reinforcement learning for market making in corporate bonds: beating the curse of dimensionality

    Olivier Gu \'e ant and Iuliia Manziuk. Deep reinforcement learning for market making in corporate bonds: beating the curse of dimensionality. Applied Mathematical Finance, 26 0 (5): 0 387--452, 2019

  82. [97]

    Dealing with the inventory risk: a solution to the market making problem

    Olivier Gu \'e ant, Charles-Albert Lehalle, and Joaquin Fernandez-Tapia. Dealing with the inventory risk: a solution to the market making problem. Mathematics and financial economics, 7 0 (4): 0 477--507, 2013

  83. [98]

    Optimal high-frequency trading with limit and market orders

    Fabien Guilbaud and Huyên Pham. Optimal high-frequency trading with limit and market orders. Quantitative Finance, 13 0 (1): 0 79--94, 2013. doi:10.1080/14697688.2012.708779

  84. [99]

    Robustness

    Lars Peter Hansen and Thomas J Sargent. Robustness. Princeton university press, 2011

  85. [100]

    Garud N. Iyengar. Robust dynamic programming. Mathematics of Operations Research, 30 0 (2): 0 257--280, 2005

  86. [101]

    Distributionally robust optimization

    Daniel Kuhn, Soroosh Shafiee, and Wolfram Wiesemann. Distributionally robust optimization. Acta Numerica, 34: 0 579--804, 2025

  87. [102]

    Simple and scalable predictive uncertainty estimation using deep ensembles

    Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. Simple and scalable predictive uncertainty estimation using deep ensembles. In Advances in Neural Information Processing Systems, volume 30, pages 6405--6416, 2017

  88. [103]

    Kernel-based reinforcement learning in robust M arkov decision processes

    Shiau Hong Lim and Arnaud Autef. Kernel-based reinforcement learning in robust M arkov decision processes. In Kamalika Chaudhuri and Ruslan Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learnin...

  89. [104]

    Decoupled weight decay regularization

    Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. International Conference on Learning Representations, 2019

  90. [105]

    Distributionally robust deep q-learning

    Chung I Lu, Julian Sester, and Aijia Zhang. Distributionally robust deep q-learning. arXiv preprint arXiv:2505.19058, 2025

  91. [106]

    Price fluctuations from the order book perspective—empirical facts and a simple model

    Sergei Maslov and Mark Mills. Price fluctuations from the order book perspective—empirical facts and a simple model. Physica A: Statistical Mechanics and its Applications, 299 0 (1): 0 234--246, 2001. ISSN 0378-4371. doi:https://doi.org/10.1016/S0378-4371(01)00301-6. URL https...

  92. [107]

    Data-driven distributionally robust optimization using the wasserstein metric: Performance guarantees and tractable reformulations

    Peyman Mohajerin Esfahani and Daniel Kuhn. Data-driven distributionally robust optimization using the wasserstein metric: Performance guarantees and tractable reformulations. Mathematical Programming, 171 0 (1): 0 115--166, 2018

  93. [108]

    Markov decision processes under model uncertainty

    Ariel Neufeld, Julian Sester, and Mario S iki \'c . Markov decision processes under model uncertainty. Mathematical Finance, 33 0 (3): 0 618--665, 2023

  94. [109]

    Robust control of markov decision processes with uncertain transition matrices

    Arnab Nilim and Laurent El Ghaoui. Robust control of markov decision processes with uncertain transition matrices. Operations Research, 53 0 (5): 0 780--798, 2005

  95. [110]

    Prajit Ramachandran, Barret Zoph, and Quoc V. Le. Searching for activation functions. arXiv preprint arXiv:1710.05941, 2018

  96. [111]

    Market making via reinforcement learning

    Thomas Spooner, John Fearnley, Rahul Savani, and Andreas Koukorinis. Market making via reinforcement learning. In Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems, AAMAS '18, page 434–442, Richland, SC, 2018. International Foundation...

  97. [112]

    Malcolm J. A. Strens. A bayesian framework for reinforcement learning. In Proceedings of the Seventeenth International Conference on Machine Learning, ICML '00, page 943–950, San Francisco, CA, USA, 2000. Morgan Kaufmann Publishers Inc. ISBN 1558607072

  98. [113]

    Scaling up robust mdps using function approximation

    Aviv Tamar, Shie Mannor, and Huan Xu. Scaling up robust mdps using function approximation. In Eric P. Xing and Tony Jebara, editors, Proceedings of the 31st International Conference on Machine Learning, volume 32 of Proceedings of Machine Learning Research, pages 181--189, Bej...

  99. [114]

    Anomalous price impact and the critical nature of liquidity in financial markets

    Bence T \'o th, Yves Lemp \'e ri \`e re, Cyril Deremble, Joachim De Lataillade, Julien Kockelkoren, and Jean-Philippe Bouchaud. Anomalous price impact and the critical nature of liquidity in financial markets. Physical Review X, 1: 0 021006, 2011

  100. [115]

    Optimal transport: old and new, volume 338

    C \'e dric Villani et al. Optimal transport: old and new, volume 338. Springer, 2009

  101. [116]

    Sinkhorn distributionally robust optimization

    Jie Wang, Rui Gao, and Yao Xie. Sinkhorn distributionally robust optimization. Operations Research, 2025. doi:10.1287/opre.2023.0294. URL https://doi.org/10.1287/opre.2023.0294

  102. [117]

    Online robust reinforcement learning with model uncertainty

    Yue Wang and Shaofeng Zou. Online robust reinforcement learning with model uncertainty. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems, volume 34, pages 7193--7206. Curran Associates, Inc., 2021

  103. [118]

    Robust markov decision processes

    Wolfram Wiesemann, Daniel Kuhn, and Ber c Rustem. Robust markov decision processes. Mathematics of Operations Research, 38 0 (1): 0 153--183, 2013

  104. [119]

    Distributionally robust markov decision processes

    Huan Xu and Shie Mannor. Distributionally robust markov decision processes. Mathematics of Operations Research, 37 0 (2): 0 288--300, 2012

Pith tools

Reviewed July 10, 2026 · model on record in the stance chip above.