{"id":"d27edd8d-c527-44e3-82c7-8434f1dae5ce","arxiv_id":"2608.02002","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A Volterra-Riccati approximation lets OTC market makers incorporate Hawkes-type persistence in RFQ flow into quote decisions, tracking the exact solution in exponential benchmarks and producing endogenous long-memory quote impact.","lead":"A market maker whose client request-for-quotes arrive in persistent bursts can improve quotes by conditioning on the recent request history, using a new low-dimensional approximation of the optimal control problem. The paper derives and validates this Volterra-Riccati approximation, showing it tracks the exact solution in benchmark cases and converts long-memory request flow into persistent quote skew.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The power-law results depend on an approximation whose error is uncontrolled in the intended long-memory regime; the exponential benchmark does not validate it, and Section 7 concedes no convergence theory.","rationale":"The reader's weakest assumption and my stress-test identify the same load-bearing concern: the Section 6 long-memory claims are computed under an approximation with no established convergence theory, and the exponential benchmark does not close that gap. The paper is honest about this limitation, and the exponential validation is a genuine positive result, so the appropriate outcome is not rejection. However, because the central advertised application to power-law-like memory is exactly where the approximation is unvalidated, the verdict should remain conditional rather than unconditional acceptance.","tokens_in":19231,"tokens_out":8155,"duration_ms":113083,"concrete_test":"Re-run the Section 6 power-law-like experiment with a small exponential-mixture kernel (e.g., N=3 or N=4 factors) that has the same side asymmetry and branching ratio, and solve the exact lifted HJB on a gridded (inventory, Hawkes-memory) state space. Compare exact optimal quotes, Monte Carlo objectives, and inventory paths against the state-feedback Volterra-Riccati policy under a 50-RFQ directional burst. If the relative paired regret stays below roughly 1% and the quote-skew direction matches, the approximation transfers to longer-memory settings; if regret grows substantially or the skew differs, Section 6's claims are artifacts of the surrogate. Varying N (2, 3, 4) would also indicate whether errors grow with memory dimensionality.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline long-memory application in Section 6 is evaluated with the Volterra-Riccati surrogate, but the paper provides no error bound or convergence result for this surrogate for non-exponential kernels. Section 7 explicitly states that 'a general error or convergence theory remains open' and that a frozen conditional forecast 'does not satisfy an intertemporal dynamic-programming principle.' The exponential validation in Section 5 covers only a scalar Hawkes memory, so it cannot certify the N=8-mixture, power-law-like setting of Section 6. All improvements there are relative to a misspecified Poisson dealer; an approximate policy can beat that baseline while still being far from the true path-dependent optimum. As a result, the persistent quote skew and inventory/P&L risk reductions may be properties of the approximate rule rather than of the exact optimal policy. The empirical branching ratios are also from proprietary HSBC data without error bars, but the missing long-memory benchmark is the decisive scientific issue.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a Volterra–Riccati approximation hierarchy for OTC market making when RFQ arrivals are modelled by general Hawkes kernels. It sets up an exact path-dependent HJB, then replaces the history state with conditional future-intensity forecast curves and uses a quadratic/Riccati expansion, adding a covariance correction and a state-feedback update after observed RFQs. In an exponential Hawkes benchmark, the state-feedback policy is shown numerically to closely track the exact lifted HJB, while a memory-free Poisson policy suffers large regret. The same rule is then applied to a power-law-like N=8 exponential mixture; a directional RFQ burst produces a persistent quote skew and reduces inventory/P&L risk relative to a Poisson benchmark. The paper explicitly states that no general error or convergence theory is provided and that the frozen conditional forecast does not satisfy an intertemporal dynamic-programming principle.","tokens_in":19612,"tokens_out":6839,"duration_ms":82743,"significance":"The exponential validation is a genuine strength: the exact Markovian lift is solved and compared with common random numbers and paired regrets, and the state-feedback mechanism is shown to matter, especially in directional regimes. If the approximation can be trusted for long-memory kernels, the paper offers a tractable bridge between persistent RFQ flow and market-making controls. However, the headline long-memory application is not yet supported to the same standard. The quote-impact tail in Eq. (53) is inherited by construction from the forecast response in the approximate rule, and Section 6.2 lacks any non-Poisson reference solution; these gaps must be addressed before the long-memory claims can be accepted.","major_comments":[{"comment":"The power-law experiment cannot support the central long-memory claim because the only comparator is the Poisson policy. Section 6.2 states that no lifted HJB is solved for the N=8 exponential mixture, and Section 7 concedes that 'a general error or convergence theory remains open' and that the frozen conditional forecast does not satisfy an intertemporal dynamic-programming principle. Consequently Figs. 5–7 establish properties of the approximate state-feedback rule, not of the exact optimal policy; an approximate policy can beat a misspecified Poisson policy while being far from the true path-dependent optimum. Please add a reference benchmark or error calibration in this regime, e.g., an exact lifted HJB for a smaller exponential mixture (N=2 or N=3) under the same economic parameters, a convergence study in N_f and N, or an a posteriori bound using the covariance/curvature terms in S","section":"Section 6.2, Section 7"},{"comment":"The long-memory impact tail is hardwired into the approximate rule. Eq. (50) defines the incremental shadow-price impact as a finite-lag convolution of the resolvent response rho_{i0} with the sensitivity difference D(t,q;u)-D(t,q+epsilon_i z_i;u). If that sensitivity kernel is integrable, the convolution automatically preserves the tail order of rho, so Eq. (53) is a mathematical consequence of the definition of the state-feedback rule, not a testable prediction of the exact optimal policy. The abstract's wording, 'endogenous OTC quote impact inherits the long-memory decay of the RFQ forecast response', is therefore accurate for the approximate policy, but it should be labelled as such; an independent optimality argument would be needed to claim that the true optimal policy has this property.","section":"Section 6.1, Eqs. (50)–(53)"},{"comment":"The noise-aware correction and the linearized impact formulas require twice functional differentiability of the map m -> V0(t,q;m), which is assumed but not proved or even stated with conditions. Since V0 is obtained from a Riccati system with coefficients linear in m, the derivatives are likely available under mild integrability conditions, but the paper does not derive the resulting sensitivity ODEs or state the required function space for m. This matters for Eq. (50), where integrability in u of the sensitivity difference is precisely what allows the tail to be pulled through the convolution. Please state the assumptions and verify the derivative equations used in the implementation, at least in the exponential benchmark where the exact lifted HJB is available.","section":"Section 4.3, Eqs. (36)–(38), (42), (50)"},{"comment":"The validation relies on truncated inventory and memory grids (qmax=50, Xmax=800, 1000 memory points) but no grid-convergence or truncation-sensitivity study is reported. Since the exact benchmark is itself discretized, the claims about 0.08–0.17% relative regret for the state-feedback policy should be accompanied by a check that the discretization is fine enough and that the truncation does not bias the comparison. This is not a fatal issue, but it is needed to make the numerical validation fully convincing.","section":"Section 5.1, numerical truncation"}],"minor_comments":[{"comment":"The fitted branching ratios and mixture weights are reported without standard errors, confidence intervals, or goodness-of-fit diagnostics, and the data are proprietary. Since Table 1 is the empirical motivation for the entire model, even a brief indication of estimation uncertainty would help the reader assess the strength of the persistence evidence.","section":"Section 2, Table 1"},{"comment":"There are small notational inconsistencies in the sigmoid win-probability formulas: the parameterization in Section 5.1 (f(δ) = (1+exp(δ-1.2))^{-1}) and in Section 6.2 (f(δ) = (1+exp(3(δ-1)))^{-1}) differ, and the displayed formula in Section 5.1 has a parenthesis imbalance. These should be cleaned up.","section":"Section 5.1 and Section 6.2"},{"comment":"The paper is commendably explicit about its limitations, but the concluding section could more directly state that the power-law results in Section 6 are properties of the approximate policy, not of the exact solution of the path-dependent HJB, unless the proposed convergence/benchmark work is added.","section":"Section 7"}],"recommendation":"major_revision","confidential_remarks":"The exponential benchmark is solid and the paper is honest about its limitations, but the abstract and Section 6 overweight the long-memory results. The missing non-Poisson benchmark or convergence estimate is the decisive issue for the paper's central claim. A major revision with an additional reference computation for the power-law-like setting, or a substantial reframing of Section 6 as a property of the approximate rule, would be needed. I do not think rejection is warranted, because the validated exponential hierarchy is a useful contribution in itself."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this paper earns a referee. The genuinely new piece is the state-feedback Volterra–Riccati quote rule for the OTC request/fill separation—the dealer updates the continuation-value shadow price with the post-request forecast curve, not just current intensity. I don't know that combination in the prior literature, and the paper situates it correctly against Jusselin and Bergault et al.\n\nThe exponential benchmark is the strongest part. Solving the exact lifted HJB on a grid and measuring paired regret against it is the right test. The Poisson comparator is itself an exact HJB solution within the wrong model, so the regret numbers cleanly separate misspecification error from approximation error. The near-critical and directional regimes show the mean forecast does a lot, and state feedback closes most of the gap. That is a real, reproducible numerical result.\n\nThe soft spot is Section 6. There is no exact benchmark for the power-law-like mixture, no convergence bound, and the paper admits in Section 7 that the conditional forecast surrogate does not satisfy a dynamic-programming principle and that a general error theory is open. On top of that, the persistent quote impact is largely constructed: the state-feedback shadow price builds the Volterra response rho_i into the quote, so the long-memory tail of the quote skew is inherited from the kernel by design. The inventory and P&L improvements over a Poisson dealer are consistent with conditioning on information, but they are not evidence that the approximate policy is near the true optimum in that regime. The exponential validation is one-dimensional and mean-reverting; it should not be oversold as certification for an 8-factor slowly decaying mixture. This is an addressable limitation—more tests, even a small exact lift or a different approximation, would help—but Section 6 should not be framed as validation.\n\nThe empirical motivation is honest: proprietary HSBC data, no error bars, two-way activity rather than signed flow, and the paper says so. That is acceptable for motivation, but keep it in that role.\n\nOverall, the hierarchy is clearly formulated and the benchmark is well executed. The citation pattern is appropriate. I'd send it to a serious referee, with the instruction to focus on whether Section 6's claims are appropriately qualified. If the authors add a long-memory sanity check or soften the framing, this would be a publishable contribution.","headline":"A solid approximation paper for Hawkes-driven OTC market making, with a strong exponential benchmark and an honest but unvalidated long-memory section—worth refereeing.","tokens_in":19958,"tokens_out":4037,"would_cite":true,"duration_ms":50024,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91G80","60G55","93E20"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper shows that a state-feedback Volterra–Riccati quote rule, driven by the conditional forecast of future RFQ flow, closely approximates the exact optimal market-making policy for Hawkes-distributed requests and generates an endogenou","keywords":["OTC market making","Hawkes processes","Volterra–Riccati approximation","RFQ flow","market impact","stochastic control","long memory","optimal quoting"],"falsifier":"Solve the power-law-like Hawkes model with a high-fidelity numerical method (e.g., a very large exponential mixture lift, or a path-dependent PDE solver) and compare the resulting optimal quotes and value against the state-feedback Volterra–Riccati policy; if the relative regret grows with the kernel tail exponent, the approximation's long-memory claims fail. Alternatively, from real RFQ data, measure the decay of quote skew after a directional burst and check whether it matches the fitted kernel's tail exponent.","tokens_in":19145,"feed_emoji":"📈","tokens_out":5383,"duration_ms":55953,"temperature":0.7,"pith_summary":"OTC market makers receive request-for-quote (RFQ) arrivals whose clustering and persistence contain information about future client demand. This paper argues that a low-dimensional quote rule built from the conditional Volterra forecast of future RFQ intensity can capture most of the value of that information. The author develops a hierarchy of Volterra–Riccati approximations, culminating in a state-feedback rule that updates the forecast after each observed request, and validates it against an exact lifted HJB benchmark in an exponential Hawkes model. The state-feedback rule tracks the exact optimal quotes with small relative regret, especially when flow is directional, whereas a memory-free Poisson policy loses up to about 27% of the objective. In a long-memory power-law model, a directional burst becomes a persistent quote skew via the continuation-value shadow price, an endogenous OTC impact that reduces inventory and P&L risk.","feed_headline":"State-feedback quotes nearly match exact OTC market making","feed_subtitle":"Volterra–Riccati rule turns persistent RFQ memory into near-optimal quoting and lower inventory risk.","key_machinery":"The machinery is the Volterra forecast curve m_t(u) = E[λ_{t+u} | F_t], a deterministic curve that summarizes the Hawkes memory. A hierarchy of approximations is built on it: the mean Volterra–Riccati surrogate solves a backward Riccati system with m_t as the driving intensity; the noise-aware version adds a covariance correction; the state-feedback version replaces m_t with the post-request forecast m_t + ρ_i, using the functional sensitivity of the value map to form the post-request shadow price. Quotes are always produced by the exact Hamiltonian optimizer applied to the approximated shadow price. In the exponential benchmark the forecast curve is represented by the finite-dimensional Haw","core_discovery":"The central discovery is that the path-dependent Hawkes market-making problem, which needs the full RFQ history as state, can be reduced to a tractable quote rule operating on the conditional forecast curve of future request intensities. The paper proves numerically, in an exponential Hawkes benchmark where the exact lifted HJB can be solved, that the state-feedback Volterra–Riccati policy — which updates the forecast curve after each request and recomputes the quadratic continuation value — closely tracks the exact optimal quotes, with relative regret around 0.08%–0.17% across the tested regimes, while a Poisson policy that ignores memory loses up to 27%. The same rule, when applied to a po","pith_inferences":["One could use the linearized impact formula (50) to invert observed quote-skew decay into an estimate of the Hawkes memory kernel, giving a new empirical test on RFQ data.","The approximation hierarchy suggests a modular production design: a Hawkes/Volterra forecast engine, a Riccati solver, and a tabulated Hamiltonian optimizer can replace a full dynamic-programming solver at low latency.","If the forecast response decays as a power law, the model predicts OTC quote impact should also decay as a power law with the same exponent; this is testable with dealer-level or platform-level RFQ data.","The method could be extended to treat win probabilities as state-dependent (e.g., competition or toxicity), but then the state-feedback update would need to propagate through the response function as well; the paper's validation does not cover that case."],"forward_implications":["Ignoring request-flow history when flow is directional leads to systematically misplaced quotes and larger inventory risk, even if the dealer solves the misspecified Poisson problem optimally.","The state-feedback Volterra–Riccati rule provides a closed, low-dimensional quoting algorithm that remains applicable when exact Markovian lifting is impractical, such as power-law or very high-dimensional mixture kernels.","An observed directional RFQ burst should alter quotes persistently, not just momentarily, because the conditional forecast of future flow is itself persistent; the quote skew then decays with the forecast response tail.","Conditioning on the RFQ memory primarily improves risk control (reduces inventory and P&L variance), more than it changes mean P&L, in the long-memory experiment."],"fun_headline_variants":["0.2% regret vs 27%: memory-aware OTC quotes","Volterra-Riccati: near-optimal OTC market making","Hawkes OTC quoting: forecast curve beats Poisson by 27%","Path-dependent OTC quotes tamed by Volterra-Riccati","Near-optimal OTC quotes from Hawkes memory alone"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The receding-horizon conditional surrogate is assumed to remain an accurate approximation of the true path-dependent value function for general long-memory kernels, but exact validation is only supplied for exponential kernels, and Section 7 explicitly states that a general error or convergence theory remains open.","fun_headline_variants_meta":{"raw":{"variants":["0.2% regret vs 27%: memory-aware OTC quotes","Volterra-Riccati: near-optimal OTC market making","Hawkes OTC quoting: forecast curve beats Poisson by 27%","Path-dependent OTC quotes tamed by Volterra-Riccati","Near-optimal OTC quotes from Hawkes memory alone"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001392,"raw_usage":{"total_tokens":5521,"prompt_tokens":849,"completion_tokens":4672,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":593,"completion_tokens_details":{"reasoning_tokens":4594}},"tokens_in":593,"tokens_out":4672,"duration_ms":40620,"temperature":1.0,"reasoning_tokens":4594,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T16:53:34.839716+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Solve the power-law-like Hawkes model with a high-fidelity numerical method (e.g., a very large exponential mixture lift, or a path-dependent PDE solver) and compare the resulting optimal quotes and value against the state-feedback Volterra–Riccati policy; if the relative regret grows with the kernel tail exponent, the approximation's long-memory claims fail. Alternatively, from real RFQ data, measure the decay of quote skew after a directional burst and check whether it matches the fitted kernel's tail exponent.","supporting_citations":[],"review_version":1}