{"id":"4185b6ec-3c87-4d9b-844d-a85337138c76","arxiv_id":"2607.11328","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Performance-based flow gates make a single dealer's optimal quoting policy alternate between reputation-building and monetization, and can produce two stable reputation regimes.","lead":"This paper builds a math model where a dealer's past success in winning requests or filling trades changes how much future business they get, and studies the optimal pricing strategy that balances short-term profit against building a good reputation. A generalist reader might care because it shows how simple reputation feedback can create two very different possible futures for a dealer: staying marginal or becoming a leader.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Bistability claim rests on an unvalidated double reduction: the original finite-horizon control problem is replaced by an adiabatic averaged drift plus an ad hoc discounted-stationary value, so the two-basin phase portrait may not describe the actual stochastic process.","rationale":"The reader's weakest assumption is the adiabatic approximation and the ad hoc replacement of the finite-horizon problem by a discounted stationary one; my concern is the same, focused on the consequence for the central bistability claim. The numerical results are illustrative, and the paper's derivations are largely internally consistent, but the strongest claim is not robustly established. A conditional accept, requiring validation of the reduction (e.g. by direct simulation or full-HJB comparison), remains the right verdict; the paper should not be rejected, since the modelling idea is plausible and testable. The one concrete test I propose would resolve whether the two-mode phase portrait is a genuine property of the original process or an artifact.","tokens_in":9663,"tokens_out":7337,"duration_ms":73534,"concrete_test":"Simulate the original discrete event system at baseline parameters with the policy obtained from the reduced problem: for initial scores (0.1,0.1) and (0.9,0.9), run long paths (e.g. 250 trading days × 10^4 replications) and record the empirical joint distribution of (R_A,R_B). If the distributions from the two initial conditions are unimodal and overlap (no bimodality or no separation), the two-basin phase portrait is an artifact of the adiabatic/discounted-stationary reduction; if they form two distinct modes with rare transitions, the bistability claim survives in the original stochastic process.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that performance-based flow gates can create bistable client-flow regimes. This is evidenced by the phase portrait in Fig. 4, computed from the deterministic averaged drift (32)–(34). That drift is obtained via a double reduction that is not justified against the original model. First, the HJB equation (7) is replaced by a frozen-score ergodic fast problem using expansions (16)–(21) that keep only O(α) terms; no error bound or time-scale separation check is given for α=0.001. Second, and more specifically, the paper states 'we replace the finite-horizon slow equation by a discounted stationary problem' (Reduced stationary score problem) with an arbitrary discount ρ=0.05. The continuation value U from that discounted stationary problem enters the controls through p_sτ (30), hence determines the policy, the invariant law μ^R, and the nullclines in Fig. 4. The original finite-horizon problem (2) has a unique value and is not a bistable dynamical system; 'bistability' is a property of the approximate deterministic reduction. Without a check that the actual stochastic score process with α=0.001 and finite horizon has two metastable modes (or that the phase portrait persists for the true finite-horizon HJB), the strongest claim is unsupported. This is a real soft spot: the central conclusion could be an artifact of the discounted-stationary/adiabatic reduction.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a stochastic-control model of a single OTC market maker serving two electronic tiers: RFQ flow (tier A) and streaming flow (tier B). In each tier, realised execution success updates a slow reputation score (win ratio for RFQ, fill ratio for streaming), and future request intensity is multiplied by a logistic gate of that score. The dealer optimises quote offsets and, for streaming, a last-look rejection threshold, subject to inventory risk and a terminal penalty. The author derives the HJB equation, introduces an adiabatic slow–fast approximation in which the fast inventory problem is solved at frozen scores, expands to first order in the score-update sizes α, and then studies a deterministic two-dimensional drift for the reputation scores. Numerically, the continuation value is obtained from a discounted stationary score problem by policy iteration. The reported results show reputation-dependent quoting regimes, campaigns near the gate midpoint, monetisation above it, cross-tier spillovers, and, for the baseline parameter set, a phase portrait with two attracting reputation equilibria separated by a saddle-like branch.","tokens_in":10043,"tokens_out":5607,"duration_ms":55133,"significance":"If the central claim is correct, the paper makes a useful conceptual contribution: it shows that performance-based flow feedback alone, even without competition or learning, can produce endogenous reputation-building and monetisation phases and multiple stable client-flow regimes in a parsimonious single-dealer model. The model is technically clean in parts: the Hamiltonians for RFQ and streaming protocols are derived explicitly, the streaming protocol is given an option-like interpretation, and closed-form or Lambert-function expressions are provided for the RFQ and streaming controls. However, the main bistability result is obtained through a double numerical/analytical reduction that is not validated against the original finite-horizon stochastic problem, and the phenomenon is demonstrated for one baseline parameter set only. The paper does not provide code, but the derivations are sufficiently detailed to be checked; the missing numerical validation and sensitivity analysis are the main obstacles to accepting the central claim as established.","major_comments":[{"comment":"The phase portrait in Fig. 4 is computed from the averaged drift (32)–(34), which is obtained by expanding the HJB in α and keeping only O(α) terms (Eqs. (16)–(21)) and by replacing the finite-horizon slow problem with a discounted stationary problem using an arbitrary discount ρ=0.05. No error bound or time-scale separation check is given for α=0.001 with λ_A=100/day and λ_B=500/day. Since U from the discounted stationary problem enters the controls through p_sτ in Eq. (30), it directly determines μ^R and the nullclines. The original finite-horizon problem (2) has a unique value, so the bistability is currently a property of the approximate reduction, not of the original model. Please validate by simulating the original score-update process (1) with the computed controls, or by solving the finite-horizon HJB, and report sensitivity to α and ρ.","section":"§4 'Adiabatic approximation' and 'Reduced stationary score problem'"},{"comment":"The bistability claim is supported only by a single phase portrait for the baseline parameters in Table 1. The text asserts near Fig. 2 that 'as υ decreases, a stable + unstable pair annihilates leaving one stable branch,' but no bifurcation diagram, parameter sweep, or sensitivity analysis is shown. Since the central claim is that performance-based flow gates 'can create' bistable regimes, the reader needs to know how robust the two-basin portrait is to the gate steepness υ, midpoint R0, minimum Gmin, and to the other free parameters. Please add a systematic parameter study, e.g. bifurcation diagrams in υ and Gmin, and if the dependence is drawn from the author's prior work [17], reproduce the key result here.","section":"§5 'Numerical results', Fig. 4 and Table 1"},{"comment":"The drift integrals average over the invariant law μ^R of the fast inventory process on a truncated grid of ±50M. The paper does not report convergence of μ^R or of the resulting nullclines with respect to grid size/truncation, nor does it discuss uniqueness of the invariant law under the computed policy. Because the fixed points in Fig. 4 are read off these integrals, a truncation or non-uniqueness artifact could directly affect the bistability conclusion. Please provide a grid-convergence check and, ideally, a check that the policy iteration converges to a unique stationary policy on the score grid used.","section":"§4.1 'Reduced stationary score problem' / Eqs. (32)–(34)"}],"minor_comments":[{"comment":"The heading 'F rozen-score ergodic fast problem' contains an unintended space; please correct to 'Frozen-score ergodic fast problem'.","section":"Section heading"},{"comment":"The numerical section states that the reduced value is computed by damped policy iteration with discount ρ=0.05, but the damping factor, iteration tolerance, and stopping criterion are not specified. Please report these details for reproducibility.","section":"§5 'Numerical results'"},{"comment":"The tier-B spread is constrained from below to 0.3 bp. This constraint is not part of the HJB formulation in §3; please clarify whether it is imposed on the controls in the numerical optimisation and how it affects the reported optimality conditions.","section":"§5 'Numerical results'"},{"comment":"The caption says 'colored curves show representative trajectories' but the figure does not appear to distinguish them in the printed version; please add a legend or clearer labels.","section":"Fig. 4 caption"},{"comment":"The Lambert-W solution (36) is a nice closed-form element. Please state the branch used and the domain of validity, since W is multi-valued and the argument can be negative depending on a_A, b_A, and c_A.","section":"§5, near Eq. (36)"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern is valid and is the main reason for major revision. The bistability result is interesting but currently rests on an unvalidated double reduction; the absence of any sensitivity analysis around the single baseline parameter set makes it impossible to judge how structural the two-basin phase portrait is. I would encourage the editor to ask for either a Monte Carlo validation of the original process or a careful comparison with the finite-horizon HJB, plus a bifurcation study in the gate parameters. The paper also leans on self-citations [17,18] for the dependence of fixed points on gate properties; those claims should be summarised in the present manuscript."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth a careful read. It adds something the classical market-making literature mostly lacks: future flow is not exogenous, it is a controlled state variable driven by a score of recent execution quality. The two-tier setup (RFQ win ratio and streaming fill ratio) is clean, the HJB derivation is careful, and the interpretation of optimal policies as campaigning, defending, and monetizing is economically sensible. The phase portrait in Figure 4 is an elegant way to summarize the coupled score dynamics.\n\nWhat is actually new is the feedback mechanism itself: a score-dependent gate multiplies baseline flow, so the dealer trades off immediate spread capture against future access. That is a real idea, and the paper shows in a tractable single-dealer model how it can generate multiple stable regimes without invoking competition or learning. Credit where due: the algebra looks internally consistent, the optimal-control derivations are standard but correctly executed, and the numerical illustrations match the stated mechanism.\n\nThe soft spots are real but fixable. The central bistability claim is demonstrated only through a double reduction: the fast inventory problem is frozen at fixed scores, and the finite-horizon slow problem is replaced by a discounted stationary problem with an arbitrary discount rho=0.05. No error bound is given for the adiabatic expansion, no convergence check is reported, and there is no Monte Carlo or full HJB comparison. So the phase portrait in Figure 4 describes the approximate deterministic drift, not necessarily the actual stochastic score process. The paper also shows bistability for a single baseline parameter set; it mentions that lowering gate steepness annihilates a stable/unstable pair, but does not show a bifurcation diagram or sensitivity analysis. The absence of code or data makes it hard to check. These are not fatal flaws—the mechanism is plausible and the approximations are standard in spirit—but they mean the strongest conclusion is not yet robustly established.\n\nThe reader's conditional-accept verdict and the stress-test concern both land. The stress-test is right that the finite-horizon problem has a unique value and that 'bistability' is a property of the reduction; the paper would be stronger if it showed that the true stochastic process has two metastable modes, or at least that the phase portrait persists under the approximation. That is the one load-bearing concern, and it is proportionate.\n\nWho gets value from this? Anyone working on OTC market making, dealer platforms, or execution quality and routing. It is a good modeling template and a useful source of hypotheses, even if the quantitative findings are not yet conclusive. I would send it to peer review and ask for validation of the adiabatic reduction, sensitivity analysis over the gate parameters, and a reproducible numerical recipe or code release. I would also bring it to a reading group; the mechanism is thought-provoking and the weaknesses are instructive.","headline":"A genuinely new mechanism—future flow access as an endogenous reputation state—wrapped in an unvalidated reduction; the idea deserves serious referee time, but the bistability claim is not yet pinned down.","tokens_in":10502,"tokens_out":2365,"would_cite":true,"duration_ms":25490,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91G80","93E20"],"pacs":[],"model":"deepseek-v4-flash","headline":"A dealer's performance-based access to flow can lock in either a marginal or a leading market-making role.","keywords":["market making","reputation feedback","stochastic control","win ratio","fill ratio","last look","adiabatic approximation","bistability"],"falsifier":"Run a full two-scale simulation (or solve the HJB without the frozen-score averaging) for the baseline parameters and check whether the two attracting fixed points survive; increasing the score-update size alpha from 0.001 to, say, 0.01 and observing the collapse of the two stable branches would confirm the separation assumption is load-bearing.","tokens_in":9513,"feed_emoji":"📊","tokens_out":5323,"duration_ms":44718,"temperature":0.7,"pith_summary":"The paper tries to establish that a dealer's future order-flow access, when tied to recent win and fill ratios through performance gates, turns market making into a franchise problem rather than a series of independent quote optimizations. In the model, optimal quotes alternate between reputation-building campaigns (tight spreads near promotion thresholds) and franchise-monetization phases (wider spreads once access is secure). The central result is that the induced reputation dynamics can exhibit two stable equilibria: a low-score marginal dealer and a high-score leading liquidity provider, with an unstable boundary between them. If correct, this shows that reputation feedback alone—without competition or learning—is sufficient to create persistent bistable client-flow regimes in a single-dealer setting, which matters for how dealers manage scores and for how platforms allocate flow.","feed_headline":"Reputation gates can split dealers into leaders and laggards","feed_subtitle":"Performance-based flow access alone can trap a dealer in a low-flow state or make it a leader.","key_machinery":"The central mechanism is the slow-fast (adiabatic) reduction: the fast inventory-control problem is solved with the reputation scores frozen, and the resulting policy is averaged over the stationary inventory law to produce a deterministic slow drift on the two-dimensional score space. This reduction turns the HJB equation into a stationary Riccati relation for the inventory slope, and the phase portrait of the score drift carries the bistability. Cross-tier spillovers enter through the common inventory and continuation value.","core_discovery":"Within a two-tier stochastic-control model (RFQ and streaming), the paper defines exponentially weighted moving-average scores for win ratio and fill ratio, multiplies baseline request intensities by logistic gates of these scores, and solves the HJB problem under a slow-fast approximation. The optimal controls show a campaign phase near the gate threshold where the dealer tightens quotes to improve the score, and a monetization phase where spreads widen. Averaging over the stationary inventory distribution yields a deterministic slow drift on the score space, and its phase portrait exhibits two attracting equilibria separated by an unstable fixed point; the RFQ nullcline can fold, so the or","pith_inferences":["Editorial extension: the same feedback mechanism should generalize to multi-dealer competition, where the gate is a relative ranking rather than an absolute score; this could turn the bistability into a market-share bifurcation, suggesting that a small initial edge can compound into durable leadership.","Editorial extension: the model makes a testable prediction: dealers just below a routing gate's midpoint should quote systematically tighter than dealers just above it, which could be checked with RFQ-level data.","Editorial extension: a natural stress test is to increase the memory coefficient alpha; if the bistability disappears for alpha near 0.01, the phenomenon is tied to slow scores, while persistence at larger alpha would broaden its applicability."],"forward_implications":["Quotes around a promotion threshold tighten below the level that a myopic spread-capture policy would choose, because winning the current RFQ moves the score closer to the gate.","A dealer with a strong streaming franchise can afford tighter RFQ quotes thanks to better inventory mixing, so the two channels must be managed jointly.","In the bistable regime, a dealer can remain stuck in a low-flow, low-score equilibrium unless a transient shock or a deliberate campaign pushes the score past the unstable separatrix.","The model implies that score-management decisions should be evaluated at the portfolio or client-channel level rather than per quote.","The folded nullcline geometry means the sequence of repair matters: sometimes streaming must recover first, sometimes RFQ."],"fun_headline_variants":["Reputation feedback splits OTC dealers into leaders and laggards","Performance-based flow gates create two dealer states","Reputation campaigns and monetization phases in OTC market making","Two stable flow regimes emerge from dealer reputation feedback","How reputation feedback creates OTC dealer leader-laggard dynamics"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The entire phase-portrait analysis rests on the adiabatic separation of time scales: the inventory process must reach stationarity much faster than the scores move, and first-order terms in the score-update size must be sufficient; if that separation fails, the multiple equilibria may be artifacts of the approximation.","fun_headline_variants_meta":{"raw":{"variants":["Reputation feedback splits OTC dealers into leaders and laggards","Performance-based flow gates create two dealer states","Reputation campaigns and monetization phases in OTC market making","Two stable flow regimes emerge from dealer reputation feedback","How reputation feedback creates OTC dealer leader-laggard dynamics"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000662,"raw_usage":{"total_tokens":2789,"prompt_tokens":597,"completion_tokens":2192,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":341,"completion_tokens_details":{"reasoning_tokens":2112}},"tokens_in":341,"tokens_out":2192,"duration_ms":15110,"temperature":1.0,"reasoning_tokens":2112,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T06:54:54.828628+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a full two-scale simulation (or solve the HJB without the frozen-score averaging) for the baseline parameters and check whether the two attracting fixed points survive; increasing the score-update size alpha from 0.001 to, say, 0.01 and observing the collapse of the two stable branches would confirm the separation assumption is load-bearing.","supporting_citations":[],"review_version":2}