{"id":"d128ebdf-4c5e-436f-90aa-c19f9b58299d","arxiv_id":"1908.02847","paper_version":2,"verdict":"REJECT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"A proposed order-book-based instantaneous volatility estimator, derived from an empirical invariant that averages near one but fails a formal equality test.","lead":"The authors report a new market \"invariant\" linking price swings, trading volume, price spread, and pending orders, and use it to estimate volatility from a few minutes of data. The idea could be useful for algorithmic trading, but the paper's own statistical tests reject the exact invariant and the estimate is biased.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper rejects its own invariant: Section 3 says the strong null γ=1 is statistically rejected, and the replacement normality test only checks distribution shape, not equality to one; 'no significant violation' is therefore unsupported.","rationale":"The reader's verdict is correct and I agree with the rejection, but I locate the load-bearing weakness differently from the reader's weakest_assumption. The P(n) function in Eq. (5) is indeed calibrated only on LSE limit-order simulations and is a real external-validity risk. However, the paper's own Section 3 contains a more direct internal problem: the strong null γ=1 is rejected, and the authors then change the hypothesis to normality. That move cannot establish the invariant. A normal distribution with mean different from 1 would still be compatible with the reported normality tests, so the tests do not test Eq. (8). The GARCH comparison in Section 5 is also not apples-to-apples, as the reader notes, and the one-step-ahead test (Table 3, σξ=1.454) shows a 45% underprediction; these further weaken the applied claims but are secondary to the invariant itself. I credit the paper for reporting the rejection and for the cross-exchange fungibility check in Fig. 3, which is a reasonable falsifiable design. But the central claim, as stated in the abstract and Section 3, is not supported by the evidence actually presented. A one-sample test on the underlying daily gamma values would settle whether the invariant can be salvaged as an approximate regularity.","tokens_in":9935,"tokens_out":8316,"duration_ms":90376,"concrete_test":"Use the daily γ series behind Table 1 (one γ per day per contract) to run a one-sample t-test of H0: E[γ]=1 for each contract, and a combined test across contracts, for example an inverse-variance weighted mean of the nine contract means. Report the p-values and 95% confidence intervals. If the combined interval excludes 1, or if most contracts individually reject at the 5% level, the invariant as stated in Eq. (8) fails; the conclusion should be downgraded to 'γ is approximately, but not exactly, 1,' and the abstract's no-violation claim should be withdrawn.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing problem is internal to the paper's own statistical test. In Section 3, after defining the invariant as γ=1 in Eq. (8), the authors state that the strong null hypothesis '⟨γ⟩=1' does not pass the statistical significance test and should be rejected. They do not report the test statistic or p-value, but they concede rejection. The response is to substitute a weaker null hypothesis that the daily γ values are normally distributed and to rely on the fact that the means are within about 2σ of one. Normality is irrelevant to Eq. (8): a Gaussian centered at 0.85 or 1.26 can pass the S-W/K-S test without supporting γ=1. Moreover, Table 1 shows a systematic downward pattern: eight of nine derivative means are below unity, and the single above-unity value (Buxl, 1.125) is the only one above. The exchange-level averages in Table 2 are not a joint test of equality; averaging can mask deviations such as OMX (1.169) and S&P/TSX (1.264). The abstract's statement 'we did not find significant violation of the invariant' is therefore contradicted by the paper's own significance test. Because Eq. (12) is derived from this invariant, the volatility estimator inherits the unsupported equality. This concern is independent of the P(n) calibration question in Section 2.2: even a perfectly calibrated correction term would not rescue a central claim whose direct statistical test is rejected.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a new market invariant, Eq. (8), asserting that gamma, a dimensionless combination of price volatility, spread, traded volume, and order-book depth, equals one. It then derives from this invariant an 'instantaneous volatility' estimator, Eq. (12), and compares this estimator with realised volatility and a one-day-ahead GARCH(1,1) forecast. The empirical tests cover derivatives (E-mini, crude oil, Treasury and German bond contracts) and equities from several European, Japanese, and Canadian exchanges in 2016 and early 2017. The manuscript reports that the strong null hypothesis gamma = 1 is rejected by its own significance test, and instead appeals to normality tests and to closeness of means to one. The paper concludes that the invariant holds for liquid markets and that the resulting volatility estimator is practically useful.","tokens_in":10256,"tokens_out":4180,"duration_ms":46867,"significance":"If the invariant were established, the paper would offer a practically valuable, microstructure-based volatility estimator requiring only short-window spreads, volumes, and order-book depth, and the cross-exchange and cross-asset empirical scope would be a useful contribution. The paper also has strengths: it directly formulates a falsifiable null hypothesis, uses high-frequency data from multiple venues, and includes a fungible-instrument check across trading venues. However, the central claim is not supported by the evidence actually presented. The paper's own statistical test rejects gamma = 1, and the fallback normality tests do not test equality to one; the correction coefficient P(n) is fitted in-sample to LSE simulations and then used for LSE stocks; the equity test uses only one quarter of data and a liquidity filter; and the GARCH comparison is not a same-horizon forecast comparison. Because Eq. (12) is derived from the rejected equality, the paper's main practical contribution inherits the unsupported claim.","major_comments":[{"comment":"The paper's own significance test rejects the central hypothesis. The text states that the strong null hypothesis 'gamma = 1' does not pass the statistical significance test and should be rejected, but then substitutes a weaker null hypothesis that the daily gamma values are normally distributed. Normality is irrelevant to Eq. (8): a Gaussian centred at 0.832 or 1.125 can pass the Shapiro-Wilk and Kolmogorov-Smirnov tests without supporting equality to one. Table 1 shows eight of nine derivative means below unity, with only Buxl above unity (1.125), so the abstract's claim that no significant violation of the invariant was found is contradicted by the manuscript's own test. Since Eq. (12) is obtained by setting the left side of Eq. (8) to one, the volatility estimator inherits this unsupported equality.","section":"Section 3, Eq. (8) and Table 1"},{"comment":"The correction coefficient P(n) is fitted to LSE limit-order simulations and then used when testing the invariant on LSE stocks, making part of the apparent fit self-referential. The exponential form and the decay exponent -0.5 are empirical, with boundary conditions P(1)=1 and P(infinity)=1/2, but no out-of-sample validation is provided for other venues or asset classes. Because P(n) enters Eq. (6), and hence T_Volume, Eq. (8), and Eq. (12), a miscalibrated P(n) would invalidate the invariant and every volatility estimate built on it.","section":"Section 2.2, Eq. (5) and Fig. 1"},{"comment":"The equity test is based on one quarter of 2016 and on a liquidity filter T_Price < 15 min, and the exchange-level averages conceal instrument-level deviations such as OMX 30 at 1.169 and S&P/TSX 60 at 1.264. Averaging over exchanges is not a joint test of gamma = 1; no test statistic for equality is reported. The statement that the invariant holds 'for statistical averages in a wide range of markets' is therefore not supported by the reported evidence.","section":"Section 3, Table 2"},{"comment":"The GARCH comparison is not a same-horizon benchmark: the instantaneous volatility is computed from same-day data, whereas the one-day-ahead GARCH(1,1) forecast is made before the day begins, and the Brexit outlier is removed before computing the MSE. The reported MSE ratio is therefore not a meaningful comparison of forecasting accuracy. In addition, Fig. 5 reports sigma_xi = 1.454 for 5-minute predictions, implying an average 45% underestimation of realised volatility, which is difficult to reconcile with the claim that Eq. (8) holds with gamma = 1.","section":"Section 5, Eq. (13), Figs. 4 and 5"}],"minor_comments":[{"comment":"There are several typos and typesetting errors: 'Ki ngdom' in the author affiliation, 'T V olumre' in the Table 1 and Table 2 headers, 'Instrumens' in Table 2, and 'Andrsen' in the references.","section":"Throughout"},{"comment":"The radical notation in Eqs. (5) and (8) is garbled, with literal 'radicaltp' and 'radicalvertex' tokens that need proper mathematical typesetting.","section":"Eqs. (5), (8), and surrounding text"},{"comment":"The text states P is proportional to exp(-sqrt(n)) and then writes P(n) = 0.5(1 + exp(-(n-1)/sqrt(n))); the relationship between the proportional form and the final normalised form should be made explicit.","section":"Section 2.2"},{"comment":"The phrase 'one tick quote and trade data' is ambiguous; the paper should clarify whether the data are top-of-book or include multiple depth levels.","section":"Section 1"},{"comment":"Table 3 is based on a single stock (Barclays) and one month of observations; the claim that 5-10 minutes of history is sufficient should be caveated as a single-instrument result.","section":"Table 3"},{"comment":"The 'weaker null hypothesis' is misnamed: the Shapiro-Wilk and Kolmogorov-Smirnov tests check distribution shape, not whether the mean equals one. The paper should either test mean equality directly or rename the hypothesis.","section":"Section 3"}],"recommendation":"reject","confidential_remarks":"The manuscript reads like an industry application note, and the central invariant claim is not supported by its own statistical tests. The rejection of gamma = 1 in Section 3 is a load-bearing problem that affects Eq. (12) and the entire volatility estimation framework, so I do not see a path to acceptance without fundamentally reframing the paper as an approximate empirical relation and adding out-of-sample validation with a properly specified test of the equality."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nI read the Danyliv-Bland paper with interest. The headline finding—a dimensionless invariant linking volatility, spread, traded volume, and book depth—is genuinely new and has a practical feel. Equation (8) is not something I have seen in the cited literature, and the authors test it across several venues and asset classes, which is more than most microstructure papers do. The cross-exchange comparison for Barclays (Figure 3) is a nice natural experiment: the same stock on fungible venues with very different volumes gives similar instantaneous volatility estimates. That part of the paper is worth taking seriously.\n\nThe problem is the core claim. In Section 3 the authors outright reject the strong null that <γ>=1; they report no test statistic, but the rejection is unambiguous. Then they switch to a normality test, which is beside the point. A Gaussian centered at 0.85 or 1.26 can pass Shapiro-Wilk or Kolmogorov-Smirnov without supporting γ=1. And the numbers themselves tell the story: eight of nine derivative means sit below unity, and exchange averages run from 0.938 (FTSE All Share) to 1.264 (S&P/TSX). The abstract's assertion that the invariant shows no significant violation is contradicted by the paper's own significance test.\n\nSecondary issues compound this. The correction coefficient P(n) in Eq. (5) is fitted to LSE order-flow simulations and then used to test the invariant on LSE stocks—part of the fit is self-referential. The GARCH comparison in Section 5 is not a fair one: same-day instantaneous volatility is being compared with a one-day-ahead forecast, and the one-step-ahead test in Figure 5 shows a 45% underprediction (σξ=1.454). That is a big miss for a forecasting claim.\n\nI want to be fair: the paper is transparent about the rejection and the underprediction. It reads like an honest empirical study, not a fabrication. The underlying relation may hold approximately, and with a fitted constant it could be a useful intraday volatility proxy for execution algorithms. But the invariant claim, as stated, is not supported by the evidence.\n\nWho should read this? Practitioners and researchers interested in microstructure-based volatility estimation will want to know the idea, even if the strict invariance fails. It deserves a serious referee because the new ratio and the cross-exchange validation are testable and worth replicating. I would send it to review, but expect the authors to either soften the claim or produce a properly calibrated version with out-of-sample tests.\n\nMy take: do not cite it as evidence for the invariant, but do look at the estimator idea.","headline":"The paper's central invariant fails its own significance test, but the new ratio and the cross-exchange validation make it worth a serious referee.","tokens_in":10753,"tokens_out":3137,"would_cite":false,"duration_ms":32649,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":null,"created_at":"2026-08-14T14:32:02.842489+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":null,"supporting_citations":[],"review_version":1}