{"id":"54b7a579-cea6-4bff-bb40-6abf3abdb012","arxiv_id":"2602.07018","paper_version":3,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"Extreme sentiment regimes show higher estimated spreads and uncertainty than neutral ones in Bitcoin data, but the effect is sensitive to controls and overlaps mechanically with volatility.","lead":"This paper reports that extreme 'fear' or 'greed' readings on the Crypto Fear & Greed Index go with wider estimated Bitcoin spreads, even after some volatility controls -- an 'extremity premium.' The finding is honestly hedged, but the pre-specified statistical test fails and part of the effect may be mechanical because the index and the spread estimator share volatility inputs.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'beyond realized volatility' claim is not cleanly identified: CS spreads share a high-low input with the F&G index's volatility/momentum components, the pre-specified within-quintile endpoint fails correction, and the only surviving controlled test is a post-hoc pooled comparison.","rationale":"The reader identifies the same load-bearing weakness I find: the paper never cleanly separates F&G's non-volatility sentiment content from its embedded volatility/momentum inputs, and the CS spread estimator shares a high-low input with both the volatility proxy and the index. This is not a manufactured objection—the paper's own limitations sections concede the extended Granger result is 'partly mechanical' and that the premium is not 'conclusively separable' from F&G's embedded volatility. The central quantitative version of the claim rests on (i) a pre-specified within-quintile test that fails multiple-testing correction in the primary sample and (ii) a post-hoc pooled version in the extended sample, while flexible regression controls absorb the regime effects. Given that the unique content of the claim is 'beyond realized volatility,' the absence of a clean measurement or a sentiment-only index construction is decisive. I therefore do not change the reader's REJECT verdict. The paper deserves credit for unusually candid self-disclosure and for reporting null results, but candor does not establish the central claim. A weaker claim—that extreme F&G regimes co-move with estimated spreads and volatility-related uncertainty—is supported, but the adverse-selection 'extremity premium beyond volatility' is not.","tokens_in":34818,"tokens_out":7715,"duration_ms":91139,"concrete_test":"Compute the Table 24 extreme-vs-neutral within-volatility-quintile gap using actual tick-level quoted/effective spreads from Binance and Bybit over at least one full calendar year overlapping all F&G regimes, instead of Corwin-Schultz estimates. If the within-quintile gap is absent or substantially smaller than the CS-based 4–9 bps (d≈0.8), the premium is a CS high-low artifact and the central claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Central claim (Abstract, §5.4) requires that F&G sentiment extremity, not the index's embedded 25% volatility and 25% momentum/volume components, drives higher spreads/uncertainty after realized-volatility control. That identification is not achieved. In the primary sample, the pre-specified within-volatility-quintile endpoint does not survive Holm-Bonferroni: Table 24 shows only Q3 remains significant (padj=0.024); Q1/Q2/Q5 are raw p≈0.013–0.029 but adjust to [0.051,0.058]. The extended sample's headline p=2.7e-14 (Table 30) is a raw pooled extreme-vs-neutral gap; the sole controlled extended result is the post-hoc pooled within-quintile test (t=3.36, p=0.0008, d=0.21, §7), while the comprehensive Table 10 regressions—Model 5 with RV, RV², |returns|, log volume, and FE—leave all regime coefficients insignificant (p>0.25). The mechanical channel is concrete: the Corwin-Schultz estimator (Eq. 21) is a function of two-day high-low ratios; the Parkinson volatility control uses daily high-low; and F&G's volatility component uses 30/90-day ranges while the momentum component is volume/price-trend based. Extreme F&G days are therefore high-range/high-momentum days by construction. Stratifying on daily Parkinson quintiles does not remove all shared range information—overnight gaps and two-day range persistence can still inflate CS within a quintile. The paper concedes this twice: §5.10.11 calls the extended Granger F=211 'partly mechanical, sharing a high-low input with the spread measure,' and §7 states the premium is not 'conclusively separable from the F&G Index's embedded volatility component.' The DVOL non-replication (§5.10.12) does not rescue the claim because DVOL is a different volatility object, and no sentiment-only F&G variant is constructed. The LOB validation (§5.9.7, Table 22) is only 61–90 days and reports correlations, not the extreme-vs-neutral regime gap. So the distinctive content of the claim—beyond realized volatility—rests on a proxy that shares inpu","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper uses the Crypto Fear & Greed Index and daily Binance OHLCV data for Bitcoin (and ETH) to argue that sentiment extremity — extreme fear or extreme greed — predicts wider Corwin-Schultz spreads and higher uncertainty after controlling for realized volatility, an effect it calls the “extremity premium” and interprets as adverse selection. The manuscript includes an uncertainty-decomposition framework, an agent-based model, SMM calibration, and extensive robustness tests: within-volatility-quintile stratification, Granger causality, placebo tests, Monte Carlo weight robustness, LOB validation, cross-asset replication, and an extended 2018–2026 sample. The paper is unusually transparent: it explicitly labels several key tests as exploratory, concedes that the ABM spread-uncertainty link is coded rather than emergent, and acknowledges that the premium is “not conclusively separable from the F&G Index’s embedded volatility component” (Section 7) and that the extended-sample Granger test is “partly mechanical, sharing a high-low input with the spread measure” (Section 5.10.11).","tokens_in":35248,"tokens_out":4138,"duration_ms":49611,"significance":"If the extremity premium were cleanly identified, the paper would make a meaningful contribution to cryptocurrency market microstructure: it would show that sentiment intensity, not direction, predicts liquidity withdrawal beyond standard volatility controls, consistent with adverse-selection models. The paper has genuine strengths: public data and code, multiple spread estimators, direct order-book validation over a limited window, cross-asset replication, Monte Carlo robustness of the heuristic uncertainty weights, and a commendable level of self-criticism. However, the central quantitative claim is not supported by the pre-specified endpoint, and the main alternative explanation — that the result is driven by shared high-low/range inputs and the F&G index’s 25% volatility and 25% momentum/volume components — is explicitly conceded in the manuscript. As a result, the current version cannot support the strong claims made in the abstract and conclusions.","major_comments":[{"comment":"The pre-specified primary endpoint — the within-volatility-quintile extremity premium with multiple-testing correction — fails. After Holm-Bonferroni correction, only Quintile 3 survives (p_adj = 0.024); Q1, Q2, and Q5 have raw p ≈ 0.013–0.029 but adjust to [0.051, 0.058], and Q4 fails outright. The abstract’s headline “p < 0.001” refers not to this endpoint but to a pooled extreme-versus-neutral comparison that the paper itself labels post-hoc and exploratory (Section 7: t = 3.36, p = 0.0008, d = 0.21). The central claim is therefore not established by the analysis that was designed to test it.","section":"§5.10.3, Table 24; Abstract; §7"},{"comment":"The “beyond realized volatility” interpretation is contaminated by a mechanical channel that the paper partially concedes. The Corwin-Schultz spread estimator is a function of two-day high-low ratios; the Parkinson volatility control uses daily high-low ranges; and the F&G index embeds 25% volatility (30/90-day ranges) and 25% momentum/volume. An extreme F&G day is therefore by construction a high-range/high-momentum day. Stratifying on daily Parkinson quintiles does not remove all shared range information, because two-day range persistence and overnight gaps can still inflate CS spreads within a quintile. The paper states that the extended-sample Granger F = 211 is “partly mechanical, sharing a high-low input with the spread measure” and that the premium cannot be “conclusively separable” from the F&G index’s embedded volatility component. Without component-level F&G data or a full-samp","section":"Eq. (21); §3.1.1; §4.3; §5.10.11; §7"},{"comment":"The comprehensive regression analysis contradicts the robust-premium narrative. In Table 10, Model 5 — which includes realized volatility, volatility squared, |returns|, log volume, and day/month/year fixed effects — leaves all regime coefficients insignificant (all p > 0.25). The paper’s response is to prefer nonparametric stratification because regression controls may impose the wrong functional form. But the preferred within-quintile result does not survive multiple-testing correction in the primary sample, and the extended-sample stratified test is explicitly post-hoc. A result that is significant only under post-hoc selected stratifications and insignificant under comprehensive parametric controls is not robust enough to support the abstract’s causal and directional claims.","section":"Table 10; §5.10.4; §7"},{"comment":"The extended-sample evidence does not rescue the central claim. The headline p = 2.7e-14 is a raw pooled extreme-vs-neutral gap in spreads (62 bps), not a volatility-controlled effect. The only controlled extended result is the post-hoc pooled within-quintile test (t = 3.36, p = 0.0008, d = 0.21), which is labeled exploratory. In addition, the extended-sample Granger causality is acknowledged to be partly mechanical due to the shared high-low input. Thus, the “extended validation” does not provide an independently pre-specified confirmation of the extremity premium; at most it provides a suggestive but confounded replication.","section":"Table 30; §5.10.11; §7"}],"minor_comments":[{"comment":"The abstract simultaneously says the finding is “confirmed” and that the within-volatility-quintile endpoint is exploratory and does not survive multiple-testing correction. Please revise the abstract so that the headline claim matches the inferential status of the evidence.","section":"Abstract"},{"comment":"Table 8 reports Extreme Fear as +0.039 without significance stars in Model 2, while Table 6 reports the same coefficient with ** (p < 0.01). This inconsistency should be fixed.","section":"Table 8 vs Table 6"},{"comment":"The Limitation section refers to “Section 5.7” for the expanding-window normalization, but the actual discussion is in Section 5.10.10. Similarly, “Section 5.8” for the ABM ablation should be Section 5.11.1. Please correct cross-references.","section":"§5.10.10; §7"},{"comment":"The sample sizes differ across tables (N = 715, N = 739, N = 739). The text notes that Table 5 excludes 24 missing-lag observations, but the table footnotes should state this explicitly to avoid apparent inconsistencies.","section":"Table 5, Table 12, Table 13"}],"recommendation":"reject","confidential_remarks":"The paper is unusually honest about its limitations, and the empirical machinery is transparent and reproducible. However, the central claim is not supported by the pre-specified endpoint, the mechanical high-low/range confound is conceded in the manuscript itself, and the comprehensive regressions eliminate the regime effects. These are load-bearing identification problems, not presentation issues. A future revision that reframes the results as exploratory or that obtains component-level sentiment data and direct intraday spreads over the full sample could address them, but the current version does not meet the bar for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I'll get straight to it: this paper has a real phenomenon buried under a headline it can't support, and the authors know it. The 'extremity premium'—higher spreads/uncertainty in both extreme fear and extreme greed than in neutral—is a new and interesting pattern, and the paper documents it with more honesty than most. But the claim that it exists 'beyond realized volatility' is not established, and the paper's own disclosures mostly admit this.\n\nWhat's genuinely good: the raw pattern is consistent across samples (62 bps gap in extended data, d=0.40 pooled), replicates on ETH, and survives several placebo checks. The authors are unusually candid: they concede the ABM's spread-uncertainty link is coded, the SMM only validates a simplified version, the uncertainty weights are heuristic, and the extended-sample Granger F-stat is 'partly mechanical.' That level of self-disclosure is rare and earns credit.\n\nThe soft spots are load-bearing. The pre-specified endpoint—within-volatility-quintile t-tests—does not survive Holm-Bonferroni: only one of five quintiles remains significant after correction. The headline extended-sample p=2.7e-14 is a raw pooled comparison, not a volatility-controlled one. The only controlled extended test is a post-hoc pooled within-quintile test (d=0.21, p=0.0008), but the comprehensive regression with RV, RV², |returns|, volume, and fixed effects kills all regime coefficients (p>0.25). Residual-on-residual correlation is r=0.04. And the mechanical channel is concrete: Corwin-Schultz spreads and Parkinson volatility both use high-low ranges, and F&G's built-in volatility and momentum components ensure extreme F&G days are high-range days. Stratifying on daily Parkinson quintiles doesn't eliminate two-day range persistence or overnight gaps. The paper concedes twice that the premium isn't conclusively separable from the index's embedded volatility.\n\nThe DVOL non-replication doesn't rescue the claim, because DVOL is a different object, and no sentiment-only F&G variant is built.\n\nWho's this for: anyone working on crypto sentiment and liquidity, and anyone teaching what can go wrong when predictors and outcomes share mechanical inputs. The raw pattern is probably worth knowing, but the 'beyond volatility' claim needs a cleaner identification strategy—a sentiment index without the embedded volatility/momentum components, or intraday data around regime transitions. I'd send this to peer review with a clear expectation of major revision: restate the claim as an association, present the shared-input problem prominently, and stop selling the pooled test as the core evidence. The authors have the skills to fix it; the paper as it stands overreaches.","headline":"Honest, thorough paper that documents a plausible raw pattern but fails to identify the 'beyond volatility' claim; send to peer review with major-revision expectations.","tokens_in":35869,"tokens_out":3054,"would_cite":true,"duration_ms":33098,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Sentiment extremity, not direction, predicts wider Bitcoin spreads after volatility is controlled.","keywords":["extremity premium","sentiment regimes","adverse selection","bid-ask spread","Crypto Fear & Greed Index","market microstructure","Bitcoin","uncertainty decomposition"],"falsifier":"Construct a version of the Crypto Fear & Greed Index with its volatility and momentum components removed, then recompute the volatility-quintile-stratified extreme-versus-neutral spread gap; if the premium disappears, the claim that sentiment extremity itself—rather than embedded volatility—drives liquidity withdrawal is falsified.","tokens_in":34653,"feed_emoji":"🪙","tokens_out":4006,"duration_ms":46100,"temperature":0.7,"pith_summary":"This paper tries to establish that in cryptocurrency markets, the intensity of sentiment—not whether it is bullish or bearish—drives spread-widening by market makers. Using the Crypto Fear & Greed Index and daily Bitcoin data, the author finds that both extreme fear and extreme greed regimes show higher uncertainty and wider bid-ask spreads than neutral periods, even after controlling for realized volatility. The proposed mechanism is adverse selection: when the crowd commits strongly to a directional view, market makers face greater risk of informed trading and withdraw liquidity. The paper reports the premium across an extended 2018–2026 sample as a 62-basis-point raw gap (p = 2.7e-14) and on Ethereum, while also conceding that the effect is sensitive to functional form and is not conclusively separable from the volatility component embedded in the sentiment index.","feed_headline":"Sentiment intensity, not direction, drives Bitcoin spreads","feed_subtitle":"Extreme fear and extreme greed show a 62-basis-point spread premium over neutral days, even after volatility controls.","key_machinery":"The central object is the extremity premium itself: the gap in mean spread or uncertainty between extreme sentiment regimes (Crypto Fear & Greed Index below 25 or above 75) and neutral regimes (45–55). The argument is carried by classifying days into sentiment regimes and then comparing extreme versus neutral days within realized-volatility quintiles, which holds volatility constant nonparametrically. A continuous distance-from-neutral version (|F&G − 50|/50) performs comparably, reinforcing that extremity, not a specific threshold, is the operative variable. The interpretation is pinned to adverse-selection logic from market microstructure: spread widening is a defensive response to the ris","core_discovery":"This paper claims to document a previously unnamed phenomenon, the 'extremity premium': extreme values of the Crypto Fear & Greed Index—whether extreme fear or extreme greed—are associated with higher market-maker spreads than neutral readings, over and above what realized volatility predicts. The author interprets the premium as adverse selection: when the crowd commits strongly to a directional view, the risk of trading against an informed counterpart rises, so liquidity providers withdraw. The load-bearing evidence is a volatility-controlled regime effect (extreme greed +5.5, extreme fear +3.9 percentage points of uncertainty) and, in the extended sample, a 62-basis-point raw spread gap b","pith_inferences":["A testable extension not developed in the paper: if intensity, not direction, strips liquidity, then intraday order-book spreads should widen systematically on days when the sentiment index crosses extreme thresholds, even when realized volatility is matched tick-by-tick.","The same concept could be tested in other retail-driven markets (e.g., equity sentiment indices) where a composite score reaches extreme values while separable volatility components can be removed.","The paper's residual-on-residual result implies that many baseline spread-uncertainty correlations may be inflated by shared volatility; an implication left implicit is that sentiment-based market-making models should use regime indicators rather than continuous scores to avoid absorbing the effect in linear controls.","If the premium survives a volatility-free sentiment index, it would support a behavioral mechanism—conviction itself being risky—rather than a pure volatility-embedding artifact."],"forward_implications":["Market makers should widen quotes when sentiment is extreme in either direction, not merely when it is bullish or bearish.","Directional sentiment alone is a weak spread predictor (correlation near 0.085), so intensity should be the primary input to liquidity provision models.","Simple, coarse sentiment regimes may outperform more elaborate uncertainty-decomposition models because aleatoric noise dominates in crypto markets.","The premium replicates on Ethereum and in 6 of 7 market cycles, suggesting it is a structural feature of cryptocurrency markets rather than a Bitcoin-specific artifact.","Granger tests indicate that uncertainty predicts spreads, though the extended-sample result is acknowledged to be partly mechanical due to shared high-low inputs, with a weaker reverse channel during crises."],"fun_headline_variants":["Intensity over direction: Extreme sentiment widens Bitcoin spreads","Extreme fear and greed both stretch Bitcoin spreads—intensity matters","Bitcoin's 'extremity premium': Sentiment intensity widens spreads","Sentiment extremes, not their sign, push Bitcoin spreads wider","Market-makers widen Bitcoin spreads in extreme sentiment regimes"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the non-volatility components of the Crypto Fear & Greed Index drive the regime effect, and that the high-low construction shared between the index and the spread estimator does not mechanically create the association.","fun_headline_variants_meta":{"raw":{"variants":["Intensity over direction: Extreme sentiment widens Bitcoin spreads","Extreme fear and greed both stretch Bitcoin spreads—intensity matters","Bitcoin's 'extremity premium': Sentiment intensity widens spreads","Sentiment extremes, not their sign, push Bitcoin spreads wider","Market-makers widen Bitcoin spreads in extreme sentiment regimes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001087,"raw_usage":{"total_tokens":4444,"prompt_tokens":875,"completion_tokens":3569,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":619,"completion_tokens_details":{"reasoning_tokens":3484}},"tokens_in":619,"tokens_out":3569,"duration_ms":30952,"temperature":1.0,"reasoning_tokens":3484,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T05:46:33.994465+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct a version of the Crypto Fear & Greed Index with its volatility and momentum components removed, then recompute the volatility-quintile-stratified extreme-versus-neutral spread gap; if the premium disappears, the claim that sentiment extremity itself—rather than embedded volatility—drives liquidity withdrawal is falsified.","supporting_citations":[],"review_version":1}