Pith. sign in

REVIEW 3 major objections 5 minor 19 references

Strategic OTC market making with reputation feedback

T0 review · 3 major / 5 minor · reviewed 2026-08-02 · deepseek-v4-flash

Pith's one-line read A dealer's performance-based access to flow can lock in either a marginal or a leading market-making role.

desk verdict A genuinely new mechanism—future flow access as an endogenous reputation state—wrapped in an unvalidated reduction; the idea deserves serious referee time, but the bistability claim is not yet pinned down. read the letter →

arxiv 2607.11328 v2 pith:S2KFNSXC submitted 2026-07-13 q-fin.MF q-fin.RM

classification q-fin.MFq-fin.RM MSC 91G8093E20
keywords marketmakingreputationfeedbackstochasticcontrolwinratiofilllastlookadiabaticapproximationbistability
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that a dealer's future order-flow access, when tied to recent win and fill ratios through performance gates, turns market making into a franchise problem rather than a series of independent quote optimizations. In the model, optimal quotes alternate between reputation-building campaigns (tight spreads near promotion thresholds) and franchise-monetization phases (wider spreads once access is secure). The central result is that the induced reputation dynamics can exhibit two stable equilibria: a low-score marginal dealer and a high-score leading liquidity provider, with an unstable boundary between them. If correct, this shows that reputation feedback alone—without competition or learning—is sufficient to create persistent bistable client-flow regimes in a single-dealer setting, which matters for how dealers manage scores and for how platforms allocate flow.

What carries the argument

The central mechanism is the slow-fast (adiabatic) reduction: the fast inventory-control problem is solved with the reputation scores frozen, and the resulting policy is averaged over the stationary inventory law to produce a deterministic slow drift on the two-dimensional score space. This reduction turns the HJB equation into a stationary Riccati relation for the inventory slope, and the phase portrait of the score drift carries the bistability. Cross-tier spillovers enter through the common inventory and continuation value.

What would settle it

Run a full two-scale simulation (or solve the HJB without the frozen-score averaging) for the baseline parameters and check whether the two attracting fixed points survive; increasing the score-update size alpha from 0.001 to, say, 0.01 and observing the collapse of the two stable branches would confirm the separation assumption is load-bearing.

Watch

Extended reading notes

Core claim

Within a two-tier stochastic-control model (RFQ and streaming), the paper defines exponentially weighted moving-average scores for win ratio and fill ratio, multiplies baseline request intensities by logistic gates of these scores, and solves the HJB problem under a slow-fast approximation. The optimal controls show a campaign phase near the gate threshold where the dealer tightens quotes to improve the score, and a monetization phase where spreads widen. Averaging over the stationary inventory distribution yields a deterministic slow drift on the score space, and its phase portrait exhibits two attracting equilibria separated by an unstable fixed point; the RFQ nullcline can fold, so the or

Load-bearing premise

The entire phase-portrait analysis rests on the adiabatic separation of time scales: the inventory process must reach stationarity much faster than the scores move, and first-order terms in the score-update size must be sufficient; if that separation fails, the multiple equilibria may be artifacts of the approximation.

Editorial extensions

If this is right

  • Quotes around a promotion threshold tighten below the level that a myopic spread-capture policy would choose, because winning the current RFQ moves the score closer to the gate.
  • A dealer with a strong streaming franchise can afford tighter RFQ quotes thanks to better inventory mixing, so the two channels must be managed jointly.
  • In the bistable regime, a dealer can remain stuck in a low-flow, low-score equilibrium unless a transient shock or a deliberate campaign pushes the score past the unstable separatrix.
  • The model implies that score-management decisions should be evaluated at the portfolio or client-channel level rather than per quote.
  • The folded nullcline geometry means the sequence of repair matters: sometimes streaming must recover first, sometimes RFQ.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial extension: the same feedback mechanism should generalize to multi-dealer competition, where the gate is a relative ranking rather than an absolute score; this could turn the bistability into a market-share bifurcation, suggesting that a small initial edge can compound into durable leadership.
  • Editorial extension: the model makes a testable prediction: dealers just below a routing gate's midpoint should quote systematically tighter than dealers just above it, which could be checked with RFQ-level data.
  • Editorial extension: a natural stress test is to increase the memory coefficient alpha; if the bistability disappears for alpha near 0.01, the phenomenon is tied to slow scores, while persistence at larger alpha would broaden its applicability.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes a stochastic-control model of a single OTC market maker serving two electronic tiers: RFQ flow (tier A) and streaming flow (tier B). In each tier, realised execution success updates a slow reputation score (win ratio for RFQ, fill ratio for streaming), and future request intensity is multiplied by a logistic gate of that score. The dealer optimises quote offsets and, for streaming, a last-look rejection threshold, subject to inventory risk and a terminal penalty. The author derives the HJB equation, introduces an adiabatic slow–fast approximation in which the fast inventory problem is solved at frozen scores, expands to first order in the score-update sizes α, and then studies a deterministic two-dimensional drift for the reputation scores. Numerically, the continuation value is obtained from a discounted stationary score problem by policy iteration. The reported results show reputation-dependent quoting regimes, campaigns near the gate midpoint, monetisation above it, cross-tier spillovers, and, for the baseline parameter set, a phase portrait with two attracting reputation equilibria separated by a saddle-like branch.

Significance. If the central claim is correct, the paper makes a useful conceptual contribution: it shows that performance-based flow feedback alone, even without competition or learning, can produce endogenous reputation-building and monetisation phases and multiple stable client-flow regimes in a parsimonious single-dealer model. The model is technically clean in parts: the Hamiltonians for RFQ and streaming protocols are derived explicitly, the streaming protocol is given an option-like interpretation, and closed-form or Lambert-function expressions are provided for the RFQ and streaming controls. However, the main bistability result is obtained through a double numerical/analytical reduction that is not validated against the original finite-horizon stochastic problem, and the phenomenon is demonstrated for one baseline parameter set only. The paper does not provide code, but the derivations are sufficiently detailed to be checked; the missing numerical validation and sensitivity analysis are the main obstacles to accepting the central claim as established.

major comments (3)
  1. [§4 'Adiabatic approximation' and 'Reduced stationary score problem'] The phase portrait in Fig. 4 is computed from the averaged drift (32)–(34), which is obtained by expanding the HJB in α and keeping only O(α) terms (Eqs. (16)–(21)) and by replacing the finite-horizon slow problem with a discounted stationary problem using an arbitrary discount ρ=0.05. No error bound or time-scale separation check is given for α=0.001 with λ_A=100/day and λ_B=500/day. Since U from the discounted stationary problem enters the controls through p_sτ in Eq. (30), it directly determines μ^R and the nullclines. The original finite-horizon problem (2) has a unique value, so the bistability is currently a property of the approximate reduction, not of the original model. Please validate by simulating the original score-update process (1) with the computed controls, or by solving the finite-horizon HJB, and report sensitivity to α and ρ.
  2. [§5 'Numerical results', Fig. 4 and Table 1] The bistability claim is supported only by a single phase portrait for the baseline parameters in Table 1. The text asserts near Fig. 2 that 'as υ decreases, a stable + unstable pair annihilates leaving one stable branch,' but no bifurcation diagram, parameter sweep, or sensitivity analysis is shown. Since the central claim is that performance-based flow gates 'can create' bistable regimes, the reader needs to know how robust the two-basin portrait is to the gate steepness υ, midpoint R0, minimum Gmin, and to the other free parameters. Please add a systematic parameter study, e.g. bifurcation diagrams in υ and Gmin, and if the dependence is drawn from the author's prior work [17], reproduce the key result here.
  3. [§4.1 'Reduced stationary score problem' / Eqs. (32)–(34)] The drift integrals average over the invariant law μ^R of the fast inventory process on a truncated grid of ±50M. The paper does not report convergence of μ^R or of the resulting nullclines with respect to grid size/truncation, nor does it discuss uniqueness of the invariant law under the computed policy. Because the fixed points in Fig. 4 are read off these integrals, a truncation or non-uniqueness artifact could directly affect the bistability conclusion. Please provide a grid-convergence check and, ideally, a check that the policy iteration converges to a unique stationary policy on the score grid used.
minor comments (5)
  1. [Section heading] The heading 'F rozen-score ergodic fast problem' contains an unintended space; please correct to 'Frozen-score ergodic fast problem'.
  2. [§5 'Numerical results'] The numerical section states that the reduced value is computed by damped policy iteration with discount ρ=0.05, but the damping factor, iteration tolerance, and stopping criterion are not specified. Please report these details for reproducibility.
  3. [§5 'Numerical results'] The tier-B spread is constrained from below to 0.3 bp. This constraint is not part of the HJB formulation in §3; please clarify whether it is imposed on the controls in the numerical optimisation and how it affects the reported optimality conditions.
  4. [Fig. 4 caption] The caption says 'colored curves show representative trajectories' but the figure does not appear to distinguish them in the printed version; please add a legend or clearer labels.
  5. [§5, near Eq. (36)] The Lambert-W solution (36) is a nice closed-form element. Please state the branch used and the domain of validity, since W is multi-valued and the argument can be negative depending on a_A, b_A, and c_A.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the bistability claim is computed from the paper's own model, and self-citations are ancillary.

full rationale

The derivation chain is self-contained for the paper's stated purpose. The model explicitly defines the score updates (1), the gate function (39), the HJB equation (7), the adiabatic expansions (16)-(21), the frozen-score Riccati relation (29), and the slow drift (32)-(34). The phase portrait and fixed-point analysis are computed from these equations, not imported from a fitted dataset. The fixed-point condition R_A = f-bar_A(R_A) follows from the definition of the EWMA score (R -> (1-alpha)R + alpha I), so stationary points are by construction self-consistent win ratios; however, the economic content that the optimal policy makes f-bar_A intersect the identity three times is an output of the control problem through the continuation value U, not an input. The self-citations [17] and [18] point to the author's earlier models for the win-score concept, the fair-protocol last-look, and a qualitative bifurcation remark; they do not carry the central numerical demonstration, which is reproduced in this paper's Figures 2 and 4. The main weaknesses are approximation and calibration choices: the adiabatic reduction is not accompanied by error bounds, the discounted stationary replacement uses an arbitrary rho=0.05, and Table 1 parameters are hand-picked. These are correctness and robustness concerns, not circular reductions, because the results are not fitted to, nor defined by, the target claim. Under the standard that circularity requires a specific reduction of the claimed output to an input, no such step is found.

Assumptions & free parameters 5 free parameters · 3 assumptions · 2 invented entities

The central results are driven by the choice of logistic gates with steepness above a certain threshold. The model rests on the adiabatic approximation, which is asserted but not validated. The parameters are plausible but not fitted to data. The two-level score state is a structural assumption. The paper does not claim empirical fit, so these are acknowledged assumptions, but they reduce the external validity of the bistability prediction.

free parameters (5)
  • Gate steepness υτ = 60 (tier A), 10 (tier B)
    The steepness of the logistic gate controls whether bistability appears. The paper explicitly notes that decreasing υ annihilates the stable/unstable pair, so the central phenomenon is parameter-driven.
  • Gate midpoint Rτ0 = 0.35 (tier A), 0.8 (tier B)
    The location of the promotion threshold determines where the policy changes from campaign to monetization and influences the fixed point structure.
  • Gate minimum Gminτ = 0.1 for both tiers
    Sets the floor on how much flow a dealer retains even with score zero; a higher floor would reduce bistability.
  • Latency mark mean mτ = -0.1 bp (A), -0.2 bp (B)
    Adverse selection level; affects the value of rejecting and the streaming threshold. Not fitted to data, but chosen.
  • Score memory coefficients ατ = 0.001 for both tiers
    Sets the time scale of reputation; must be small for the adiabatic approximation to hold. Not validated.
assumptions (3)
  • domain assumption The fast inventory process reaches a stationary distribution μR for each frozen score R, and the adiabatic approximation to first order in α is accurate.
    Used in equations (16)-(21) and the slow drift (32)-(34). No error bound is provided. This is the most fragile assumption.
  • ad hoc to paper The score update is a simple exponential moving average and the gate function is logistic as in Eq. (39).
    The specific functional forms are chosen for tractability, not derived from data or theory. The bistability result is sensitive to the gate shape.
  • domain assumption The fair-protocol streaming kernel with Gaussian latency mark Y ~ N(m,ν²) is representative of real last-look protocols.
    The Gaussian assumption and the fair protocol are borrowed from [6], but the validity for FX or bond markets is not established here.
invented entities (2)
  • Score-dependent flow gates Gτ(Rτ) multiplying baseline intensities
    purpose: Represents platforms allocating more or less flow based on a dealer's score. This is the key new mechanism.
    The gates are a modeling device. No empirical evidence is presented that real platforms use exactly such logistic gates, or that the parameters (υ, R0) match reality. However, the qualitative idea of performance-based routing is plausible and consistent with the cited literature.
  • Two-dimensional score state (R_A, R_B)
    purpose: Tracks win ratio and fill ratio as slow state variables. This is a modeling construct, not a new physical entity.
    Scores are defined as exponential moving averages of success indicators. Real platforms do have scores, but the exact update rule (1) is a stylized assumption.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Strategic OTC market making with reputation feedback." pith.science (2026). https://pith.science/paper/S2KFNSXC

@misc{pith2026260711328,
  author       = {Pith},
  title        = {Pith review of: Strategic OTC market making with reputation feedback},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/S2KFNSXC}},
  note         = {Machine review of arXiv:2607.11328}
}
read the original abstract

Electronic over-the-counter (OTC) liquidity provision is increasingly shaped not only by the price of the next quote, but also by a dealer's accumulated standing with clients and platforms. We develop a stochastic-control model in which request-for-quote (RFQ) win ratios and streaming fill ratios feed back into future flow through performance-based flow gates, creating an explicit trade-off between immediate spread capture and long-term franchise value. The resulting policy naturally alternates between reputation-building campaigns and franchise monetization phases, and can generate multiple stable client-flow regimes even in a parsimonious single dealer control problem.

Figures

Figures reproduced from arXiv: 2607.11328 by the authors.

Figure 1
Figure 1. Promotion values ∆Uτ (R) = U(Rτ +) − U(Rτ −) over the two-dimensional score space [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. RFQ pricing and score fixed points. Left: bid offset at zero inventory as a function of the RFQ score [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Streaming tier pricing and acceptance policy. Left: total streamed spread ∆ [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Two-dimensional slow reputation dynamics. Arrows show the averaged score drift, blue and green lines [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

19 extracted references · 4 linked inside Pith

  1. [17]

    Barzykin, Win-score promotion gates in aggregator-routed RFQ markets: A two-tier stochastic control model, arXiv:2603.10569, 2026

    A. Barzykin, Win-score promotion gates in aggregator-routed RFQ markets: A two-tier stochastic control model, arXiv:2603.10569, 2026

  2. [1]

    Avellaneda and S

    M. Avellaneda and S. Stoikov. High-frequency trading in a limit order book.Quant. Finance, 8:217–224, 2008

  3. [2]

    Gu´ eant, C.-A

    O. Gu´ eant, C.-A. Lehalle, and J. Fern´ andez-Tapia. Dealing with the inventory risk: a solution to the market making problem.Math. Financ. Econ.,7:477–507, 2013

  4. [3]

    Cartea, S

    ´A. Cartea, S. Jaimungal, and J. Ricci, Buy low, sell high: A high frequency trading perspective. SIAM J. Financ. Math.,5:415–444, 2014

  5. [4]

    Bergault and O

    P. Bergault and O. Gu´ eant, Size matters for OTC market makers: general results and dimensionality reduction techniques.Math. Finance,31:279–322, 2021

  6. [5]

    Oomen, Execution in an aggregator.Quant

    R. Oomen, Execution in an aggregator.Quant. Finance,17:383–404, 2017

  7. [6]

    Oomen, Last look.Quant

    R. Oomen, Last look.Quant. Finance,17:1057–1070, 2017

  8. [7]

    Cartea, S

    ´A. Cartea, S. Jaimungal, and J. Walton, Foreign exchange markets with last look.Math. Fi- nanc. Econ.,13:, 1–30, 2019

Show all 19 references
  1. [8]

    Barzykin, P

    A. Barzykin, P. Bergault, O. Gu´ eant, and M. Lemmel, Optimal quoting under adverse selection and price reading. arXiv:2508.20225v3, 2025

  2. [9]

    Cartea and L

    ´A. Cartea and L. S´ anchez-Betancourt, A simple strategy to deal with toxic flow. arXiv:2503.18005, 2025

  3. [10]

    P. Bank, I. Ekren, and J. Muhle-Karbe. Liquidity in competitive dealer markets.Math. Finance, 31:827–856, 2021

  4. [11]

    Wang, The limits of multi-dealer platforms,J

    C. Wang, The limits of multi-dealer platforms,J. Financ. Econ.,149:434–450, 2023

  5. [12]

    Boyce, M

    R. Boyce, M. Herdegen, and L. S´ anchez-Betancourt. Market making with exogenous competition. arXiv:2407.17393, 2024

  6. [13]

    Eisler and J

    Z. Eisler and J. Muhle-Karbe. Optimizing broker performance evaluation through intraday modeling of execution cost. arXiv:2405.18936, 2024

  7. [14]

    Cont and W

    R. Cont and W. Xiong. A study of algorithmic collusion in multi-dealer-to-client platforms. SSRN:5463594, 2024

  8. [15]

    Mar ´ ın, S

    P. Mar ´ ın, S. Ardanza-Trevijano, and J. Sabio. Causal Interventions in Bond Multi-Dealer-to-Client Platforms. arXiv:2506.18147v2, 2025

  9. [16]

    Boyce and E

    R. Boyce and E. Neuman. Competition in Dealer Markets with Internalisation and Externalisation. SSRN:6844540, 2026

  10. [18]

    Barzykin, Dynamic slippage control and rejection feedback in spot FX market making, arXiv:2603.07752, 2026

    A. Barzykin, Dynamic slippage control and rejection feedback in spot FX market making, arXiv:2603.07752, 2026

  11. [19]

    Bergault, D

    P. Bergault, D. Evangelista, O. Gu´ eant and D. Vieira, Closed-form approximations in multi-asset market making.Appl. Math. Finance, 2021,28, 101–142. 11

Pith tools

Reviewed August 2, 2026 · model on record in the stance chip above.