REVIEW 7 major objections 6 minor 1 cited by
ASRI: An Aggregated Systemic Risk Index for Cryptocurrency Markets
T0 review · 7 major / 6 minor · reviewed 2026-08-03 · deepseek-v4-flash
Pith's one-line read ASRI, a weighted composite of four crypto-risk channels, is argued to be worth building for channel-by-channel attribution, not for outpredicting simple benchmarks — its own tests find it matches a standalone VIX series.
desk verdict The paper's candid abstract and its body tell opposite stories about whether ASRI actually detects crises; the framework is worth a look, the validation is not. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the linear aggregation identity ASRI_t = 0.30·SCR_t + 0.25·DLR_t + 0.25·CR_t + 0.20·OR_t, with every sub-index bounded to [0,100] by construction. Linearity does two jobs: it guarantees decomposability — every reading has a unique attribution c_i = w_i·S_i, so 'why is the index up?' always has a component-level answer — and it fuses indicators of very different frequency and provenance (daily TVL, monthly attestations, quarterly filings proxied at daily frequency) onto one comparable scale. On the validation side, the machinery is a constant-mean event-study protocol — 60-day estimation window, 41-day event window, cumulative abnormal signal whose standard error is
What would settle it
Recompute the four event-study tests with autocorrelation-robust standard errors (Newey–West, or a block bootstrap with block length of 20 or more days matched to the measured AR(1) ≈ 0.8–0.9) on the released ASRI series. If Terra/Luna's t-statistic falls below 2, or if placebo dates clear the significance threshold about as often as crisis dates — as the paper's own abstract states — then the event-study claims fail and 'detection' is a threshold-design artifact. Second, fit the same binary crisis classification to a standalone VIX series and to the strongest sub-index on the next out-of-samp
Extended reading notes
Core claim
On the authors' own terms, the finding is stated directly: 'We read aggregation's value as interpretive — channel attribution, lead time, and regime structure in one auditable composite — not as discriminative gain.' The index is the linear identity ASRI_t = 0.30·SCR_t + 0.25·DLR_t + 0.25·CR_t + 0.20·OR_t, where each sub-index is itself a weighted sum of observable indicators (stablecoin TVL drawdown, Treasury yield level, issuer concentration, peg volatility; protocol concentration, TVL volatility, audit coverage, flash-loan and leverage proxies; real-world-asset share, a Treasury–VIX bank-stress composite, yield-curve spread, BTC–equity correlation, bridge count; unregulated volume, issuer
Load-bearing premise
The load-bearing premise is that the daily abnormal ASRI readings entering the event study are independent, so the cumulative-signal standard error is the daily volatility times √41; the paper's own abstract says the series is heavily autocorrelated (AR(1) ≈ 0.8–0.9), and if that is true the reported t-statistics shrink by a factor of roughly four, collapsing the 'all four crises significant' finding into something near noise.
Editorial extensions
If this is right
- Practitioners should treat ASRI as a diagnostic overlay, not an alarm: when the index rises, the decomposition says whether to inspect stablecoin reserves, DeFi liquidity, TradFi contagion channels, or opacity — and which components lead rather than confirm.
- The walk-forward 4/4 out-of-sample detection and the near-zero degradation under simulated publication lags support the claim that the detection record is not an artifact of look-ahead bias — while the high false-positive cost makes the honest reading 'no leakage,' not 'clean prediction.'
- The benchmarking implies the connectedness index and ASRI are complementary layers of one monitoring stack: the former as a model-free first-stage filter for any spillover intensification, the latter as the channel-specific diagnostic; neither, at the operational threshold, catches Terra/Luna-style algorithmic-stablecoin reflexivity.
- Because ASRI's day-level discrimination is statistically indistinguishable from a standalone VIX series and from its own strongest sub-index, the paper's position implies a crisis/non-crisis line alone does not justify the composite — it earns its cost only through attribution, lead time, and audibility.
- The Bybit out-of-sample result implies the framework measures transmission channels rather than raw magnitude: a $1.5 billion theft with no contagion path should not move the index, and the authors present that specificity as a design feature.
Reading between the lines
- If the paper's own AR(1) ≈ 0.8–0.9 figure is right, the event-study standard error is understated by a factor of about √((1+ρ)/(1−ρ)) ≈ 4.3, putting Terra/Luna's headline t-statistic near 1.3 — my read is that the 'all four crises significant' claim is the casualty of the internal contradiction, not a robust result.
- The VIX-equivalence result points to a cheap testable extension: a two-input monitor (VIX plus a stablecoin peg-deviation series) may carry nearly all of ASRI's day-level discriminative information, so the composite's defensible value must show up in lead time and attribution rather than classification.
- The paper leaves its most distinctive claim — channel attribution — largely untested as a prediction. A pre-registered test that each crisis type is dominated by its expected sub-index (SCR for stablecoin stress, CR for counterparty contagion, DLR for liquidity-driven stress) would turn the decomposition from a narrative device into a checkable claim.
- The candid framing suggests a deflationary inference with a constructive edge: if a transparent composite cannot beat a single volatility index with four labeled crises, the binding constraint is event scarcity, not aggregation — pooling crisis episodes across markets or applying the same channels at protocol level would give the interpretive claim real statistical power.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces the Aggregated Systemic Risk Index (ASRI), a daily composite of four sub-indices — Stablecoin Concentration Risk (30%), DeFi Liquidity Risk (25%), Contagion Risk (25%), and Regulatory Opacity Risk (20%) — intended to monitor DeFi-TradFi interconnection risk. The full-text abstract claims: event studies detect statistically significant abnormal signals for all four crises (t = 5.47–32.64, all p < 0.01), threshold-based detection identifies three of four events, walk-forward validation detects 4/4 out-of-sample, and ASRI outperforms a Diebold-Yilmaz benchmark. The paper also includes a much more cautious arXiv abstract stating the event-study signal is inconclusive, heavily serially correlated, no better than a standalone VIX, and that fixed-threshold detection works at high false-positive cost. The body itself contains multiple candid caveats: §5.14.3 reports walk-forward R² ≈ −13,800 and validation_passed=False; §6.3 admits the theoretical weights were specified using the full 2021–2024 sample including the validation crises; Table 37 documents fixed placeholder proxies for important components. These two narratives cannot both describe the same validated instrument, and the manuscript does not reconcile them.
Significance. If the strong version of the claims were correct, ASRI would be a useful, interpretable early-warning monitor for crypto-native systemic risk. The paper has real strengths: a transparent four-channel architecture, reproducible code and data links, an explicit proxy-validity table, and a limitations section that candidly acknowledges several of the problems raised here. However, the central empirical claim — that ASRI detects crises with statistical significance and out-of-sample validity — is not reliably established. The event-study significance rests on an independence assumption contradicted by the paper's own autocorrelation discussion; the out-of-sample test uses weights informed by the full sample; and the paper contains two abstracts that assert opposite conclusions. As submitted, the manuscript cannot support its headline claims, and it is not merely a presentation issue: the validation evidence is internally contradictory.
major comments (7)
- [§5.4.1, Eq. (17); §5.4.3; §F.3.2] The event-study significance claims rest on independence of abnormal signals. Eq. (17) sets SE(CAS) = σ̂_AS√41, justified by 'Ljung-Box p > 0.10 for all events' (§F.3.2). But §5.4.3 calibrates the bootstrap block size to ASRI's autocorrelation structure and says residual autocorrelation is insignificant only beyond lag 15–20, implying significant autocorrelation at shorter lags; the two statements cannot both be true. The arXiv abstract reports AR(1) ≈ 0.8–0.9. Under ρ ≈ 0.85, the variance inflation factor is roughly (1+ρ)/(1−ρ) ≈ 12, so the effective SE is about 3.5× larger and the Terra/Luna t-statistic falls from 5.47 to about 1.5; none of the 'highly significant' results in Table 5 survive. Because every event-study detection claim uses Eq. (17), the central empirical claim is unsupported until autocorrelation-robust inference is supplied.
- [§5.14.1; §6.3; Table 35] The walk-forward 'out-of-sample' detection uses the theoretical weights, but §6.3 states those weights 'were specified using domain knowledge accumulated from observing the full 2021–2024 sample, including the crises used for validation.' Fixing standardization parameters on pre-crisis data does not remove the information leakage from weight choice. Thus Table 35's 4/4 OOS detection and the full-text abstract's statement that walk-forward 'confirms detection performance is not an artifact of look-ahead bias' overstate what was actually tested. The manuscript itself concedes this in §5.14.3 ('A more rigorous test would re-derive data-driven weights using only pre-crisis data'). This is a load-bearing gap for the no-look-ahead claim.
- [Full-text Abstract; arXiv Abstract; §5.14.3] The paper presents two different summaries of the same results. The full-text abstract asserts statistically significant event-study detection (t = 5.47–32.64, all p < 0.01) and a walk-forward result that 'confirms detection performance is not an artifact of look-ahead bias.' The arXiv abstract says the event-study signal is 'inconclusive,' that ASRI's discrimination is statistically indistinguishable from a standalone VIX (0.875, p = 0.58), and that walk-forward thresholds flag 4/4 'at high false-positive cost.' The body's §5.14.3 reports validation_passed = False, walk-forward R² ≈ −13,800, and OOS R² ≈ −21,812. These are not alternative interpretations of the same evidence; they are mutually incompatible claims. The authors must decide which version is the paper and provide a consistent account.
- [Table 5; §6.3; Table 3; Table 27; Table 28; Table 32; §5.4.5] Numerous reported quantities do not match. Terra/Luna CAS is 100.3 in Table 5 but 394.3 in §6.3, and its t-statistic is 5.47 in Table 5 but 7.18 in §6.3. ASRI sample min/max is 25.8/84.7 in Table 3, but 14.2/74.6 in Table 27 and max 81.1 in Table 28. AUROC is 0.918 in Table 32, 0.890 in §5.4.5, and 0.866 in the arXiv abstract. §5.16 says all event-study t-statistics exceed 6.6, contradicting Table 5's 5.47. These discrepancies make it impossible to know which specification underlies the headline claims and prevent independent replication.
- [§5.12.1; Table 31] The Diebold-Yilmaz comparator is constructed from the four ASRI sub-indices: 'we compute the Diebold and Yılmaz (2012) connectedness index using the four ASRI sub-indices as inputs.' Comparing ASRI to a D-Y index built from ASRI's own components cannot validate ASRI's incremental information; it is a circular benchmark. The conclusion in Table 31 that ASRI achieves comparable coverage with higher precision is therefore not evidence of superiority over an independent systemic-risk measure. A non-circular benchmark, for example D-Y estimated on independent crypto and TradFi asset prices, is needed before claiming ASRI outperforms connectedness approaches.
- [§5.4.4; §F.6.2; §3.8] The false-positive evidence undermines the discrimination narrative even setting aside the autocorrelation problem. Table 8 shows that at threshold 50 there are 497 false-positive alert days and only 69 true-positive days (precision 12.2%). The placebo analysis in §F.6.2 reports 1 of 10 placebo dates significant at the 5% level, consistent with nominal rates, but the fixed-threshold operational rule generates a very large number of false alerts. The full-text abstract's phrase 'statistically significant abnormal signals for all four crises' coexists with an abstract that says 'placebo dates clearing the nominal threshold as often as crises.' The operational claims need to be restated in terms of the precision-recall trade-off, not as clean 4/4 detection.
- [Appendix A, Table 37] Important components are placeholders: Unreg_t is fixed at 35.0, Sent_t is fixed at 50.0, and several other components are proxies with 'Low' or 'Medium' validity. §5.16 acknowledges 30–40% of components use proxies or fixed defaults. This is candid, but it means the precise t-statistics and AUROC values in the main text convey a spurious degree of precision. The headline numbers should be clearly labeled as conditional on these placeholder assumptions, and the claimed 4/4 event-study significance should be repeated under extreme alternatives for the placeholders, as is done only for Sent_t in Appendix A.4.4.
minor comments (6)
- [§3.8, Eqs. (11)–(12)] The text introduces full-sample min-max normalization in Eq. (12) but the footnote says 'empirical analyses use raw weighted aggregates directly.' Clarify which definition is used in every table.
- [§5.5.1, Tables 10 and 11] Elastic Net weights differ between Table 10 (SCR 0.145, DLR 0.842, CR 0.000, OR 0.013) and Table 11 (SCR 0.34, DLR 0.21, CR 0.45, OR 0.00). The reader cannot tell which specification generated the reported text.
- [Figure 5] The figure labels 'Post: 73.0' for Terra/Luna, Celsius/3AC and FTX alike, while Table 5 reports peaks 48.7, 71.4 and 84.7. Define whether 'Post' is a peak-over-window, a regime mean, or something else.
- [§5.9.2 and §3.8] Table 24 calls threshold 60 'optimal' by F1 while §3.8 and Table 8 adopt threshold 50 as the operational alert level. Explain how practitioners should reconcile these choices.
- [§F.8] The robustness table says an AR(1) normal model gives 'equivalent significance conclusions (all p < 0.01)' but no statistics or SE correction are shown. Given the autocorrelation concern, this statement needs to be quantified.
- [§5.12.2, Table 31] Lead times in Table 31 (31d, 50d, 60d) differ from those in Table 5 (30, 30, 30, 29) and Table 35 (3, 36, 4, 28). Each table should state its exact detection definition and search window.
Circularity Check
Out-of-sample crisis detection is circular: theoretical weights were set using the validation crises, and the D-Y benchmark is built from ASRI's own sub-indices.
-
fitted input called prediction
[Section 5.14.1 and Section 6.3 Limitations]
"For each window, we retain the theoretical weights (which are based on ex-ante domain knowledge rather than statistical optimization) ... The theoretical weights are based on domain knowledge accumulated from observing the full 2021–2024 sample, which includes the crisis events used for validation. Walk-forward validation (Section 5.14) addresses this concern: using only pre-crisis data to calibrate standardization parameters, ASRI achieves 100% out-of-sample detection."
The paper first asserts the theoretical weights are ex-ante domain knowledge, then admits in §6.3 that those weights were specified using the full sample containing the four validation crises. The walk-forward test freezes these same weights and recalibrates only standardization parameters, so the 'out-of-sample' 4/4 detection is not out-of-sample with respect to the most load-bearing input, the weighting scheme. The detected crises are exactly the crises that informed the weights; the prediction is in-sample by construction.
-
self definitional
[Section 5.12.1–5.12.2; arXiv abstract]
"we compute the Diebold and Yılmaz (2012) connectedness index using the four ASRI sub-indices as inputs. ... The full-sample D-Y total connectedness is only 0.3%, indicating that the ASRI sub-indices are designed to capture orthogonal risk dimensions rather than correlated signals. ... ASRI's day-level discrimination (AUROC 0.866) beats only the circular D-Y comparator (0.670)."
The D-Y benchmark is defined on exactly the four ASRI sub-indices, then used both as the comparator ASRI 'beats' and as evidence that the sub-indices are orthogonal. Comparing ASRI with a function of ASRI's own components is not an independent benchmark, and the low connectedness is a property of the same inputs rather than external validation. The paper's own arXiv abstract explicitly labels it the 'circular D-Y comparator'.
full rationale
The central claimed achievements — 4/4 walk-forward out-of-sample detection and superiority over D-Y connectedness — reduce, by the paper's own statements, to inputs known to the model. The weights are admitted in §6.3 to have been set using full-sample domain knowledge including the validation crises, so fixing them in walk-forward does not eliminate look-ahead for the most important parameters; the 4/4 result is a fitted input labeled an out-of-sample prediction. The D-Y benchmark is constructed from the same four sub-indices, which the paper itself calls circular; any detection-rate comparison or orthogonality confirmation derived from it is self-referential. I do not count the many other inconsistencies (AR(1)≈0.8–0.9 in the arXiv abstract versus Ljung-Box p>0.10 in §F.3.2; AUROC 0.866 vs 0.918; Terra/Luna CAS 100.3 vs 394.3) as circularity per se — they are internal contradictions that affect correctness — but they compound the problem by undermining the statistical bridge used to inflate the detection claims. Because the main validation claim and the benchmark are circular, the score is 7 rather than lower.
Assumptions & free parameters
free parameters (5)
- Sub-index weights (SCR/DLR/CR/OR) =
0.30 / 0.25 / 0.25 / 0.20
- Internal component weights within sub-indices =
SCR: 0.4/0.3/0.2/0.1; DLR: 0.35/0.25/0.20/0.10/0.10; CR: 0.30/0.25/0.20/0.15/0.10; OR: 0.25/0.25/0.20/0.15/0.15
- Alert thresholds =
50 (Elevated), 70 (High); F1-optimal threshold reported as 60 in Table 24
- Placeholder / proxy fixed values =
Unreg=35.0, Sent=50.0, Corr default=0.5
- Normalization bounds =
Treasury 2–6%, VIX 12–40, yield-spread 0–2%, RWA 0–10%, TVLVol 0–20%, 1-day TVL change 0–20%
assumptions (7)
- standard math Constant-mean event-study CLT: abnormal signals are independent and SE(CAS)=σ_AS×√T_event
- standard math ADF stationarity justifies level-based inference without differencing
- domain assumption DeFi Llama, FRED, and CoinGecko data are accurate, complete, and daily
- domain assumption Treasury yield, VIX, and yield-curve spread proxy bank stress and TradFi-DeFi linkage (Bankt, Linkt)
- domain assumption Audit coverage and lending TVL share proxy smart-contract risk and leverage
- ad hoc to paper Operational crisis definition thresholds (15% drawdown, correlation surge 0.20, 5-day duration) select the true systemic events
- ad hoc to paper The Bybit hack is non-systemic based on asserted absence of contagion channels
invented entities (3)
-
Bank exposure proxy (Treasury-VIX composite)
-
TradFi linkage proxy (yield curve spread)
-
Regulatory sentiment (Sentt) and unregulated exposure (Unregt) placeholders
Cite this review
Pith. "Pith review of ASRI: An Aggregated Systemic Risk Index for Cryptocurrency Markets." pith.science (2026). https://pith.science/paper/AHO4GWAU
@misc{pith2026260203874,
author = {Pith},
title = {Pith review of: ASRI: An Aggregated Systemic Risk Index for Cryptocurrency Markets},
year = {2026},
howpublished = {\url{https://pith.science/paper/AHO4GWAU}},
note = {Machine review of arXiv:2602.03874}
}
abstract
Cryptocurrency markets exceed USD 3 trillion in capitalisation, yet practitioners lack an interpretable, channel-decomposed composite for characterising crypto-native systemic stress. We introduce the Aggregated Systemic Risk Index (ASRI), built from four weighted sub-indices -- Stablecoin Concentration Risk (30%), DeFi Liquidity Risk (25%), Contagion Risk (25%, implemented as a TradFi-stress proxy), and Regulatory Opacity Risk (20%) -- with a Diebold--Yilmaz connectedness series computed on the sub-indices as network benchmark. We evaluate ASRI retrospectively against four crises (Terra/Luna, Celsius/3AC, FTX, SVB) and give a methodological account of how autocorrelation- and block-structure-robust inference reshapes apparent crisis-detection strength. The event-study signal is inconclusive: heavily serially correlated (AR(1) $\approx 0.8$--$0.9$), with placebo dates clearing the nominal threshold as often as crises. Fixed-threshold detection flags three of four events with $\approx$19-day average lead ($\approx$5 days under a responsive specification); walk-forward thresholds flag 4/4 but at high false-positive cost -- evidence against look-ahead bias, not a clean prediction record. ASRI's day-level discrimination (AUROC 0.866) beats only the circular D--Y comparator (0.670); it is statistically indistinguishable from its strongest sub-index (0.851), PC1 (0.858), and a standalone VIX series (0.875, $p=0.58$). We read aggregation's value as interpretive -- channel attribution, lead time, and regime structure in one auditable composite -- not as discriminative gain. With four crisis events the binding power limit, ASRI is a transparent, reproducible, retrospective monitoring framework targeting crypto-native vulnerabilities that SRISK and CoVaR are not built to capture, not a validated early-warning system. Out of sample it classifies the 2025 Bybit hack as non-systemic.
Figures
Figures from the paper (4 more)
Forward citations
Cited by 1 Pith paper
-
The Extremity Premium: Sentiment Regimes and Adverse Selection in Cryptocurrency Markets
Extreme sentiment regimes show higher estimated spreads and uncertainty than neutral ones in Bitcoin data, but the effect is sensitive to controls and overlaps mechanically with volatility.
Reference graph
Works this paper leans on
-
[1]
Integration of exploit database (DeFi Rekt, Immunefi) for dynamic SCt and Bridget com- ponents
-
[2]
Protocol deployment timestamp extraction from blockchain explorers for age-based risk weighting
-
[3]
GDELT/SEC filing NLP pipeline for automated regulatory sentiment scoring
-
[4]
Chain-level regulatory classification for dynamic Unregt calculation
-
[5]
B API Documentation Summary Table 38 provides endpoint documentation for primary data sources
Enterprise analytics partnership (Chainalysis/TRM) for stablecoin flow analysis. B API Documentation Summary Table 38 provides endpoint documentation for primary data sources. Table 38: Primary API Endpoints Source Endpoint Rate Limit Authentication DeFi Llamaapi.llama.fi/v2/tvl300/5min None DeFi Llamastablecoins.llama.fi/stablecoins300/5min None FREDapi....
2022
-
[9]
Stablecoin concentration in USDT/USDC (HHI = 0.52)
-
[10]
Treasury exposure through stablecoin reserves ($80B+ in T-bills)
-
[11]
Emerging RWA tokenization growth (+45% YoY) Regime Classification The HMM classifies current market conditions asRegime 2 (Moderate)with 78% probability. Transition probabilities indicate: •3.9% probability of moving to Elevated regime •2.3% probability of moving to Low Risk regime •93.8% probability of remaining in Moderate regime Alert Status No immedia...
Show all 20 references
-
[12]
Stablecoin reserve composition changes
-
[13]
Cross-market correlation shifts (BTC-equity)
-
[14]
Bridge vulnerability and exploit frequency 79 F Event Study Protocol Specification This appendix provides the complete methodological specification for the event study analy- sis presented in Section 5.4, addressing reviewer concerns regarding pre-registration, window selectio...
2022
-
[15]
Estimation windows remain independent
-
[16]
CeFi lending)
The events represent distinct crisis mechanisms (algorithmic stablecoin vs. CeFi lending)
-
[17]
Separate CAS calculations use event-specific baselines F.6 Placebo Testing To assess false positive rates under the null hypothesis of no crisis, we conduct placebo analysis on 10 randomly selected non-crisis dates. F.6.1 Placebo Date Selection Dates were drawn uniformly from ...
2021
-
[18]
False positive rate is consistent with nominalαlevels
-
[19]
Crisis events produce dramatically largert-statistics than placebo dates
-
[20]
Captures earliest structural warning signal
The 19×difference in mean|t|between crisis and placebo dates demonstrates genuine discriminative ability F.7 Lead Time Measurement Two complementary lead time definitions are employed: Definition 1 (First-crossing): Days between first observation exceeding 1.5σabove base- line...
-
[2022]
Examines stablecoin behavior and spillovers during turbulent market conditions
doi: 10.1186/s40854-023-00492-4. Examines stablecoin behavior and spillovers during turbulent market conditions. DeFi Llama. Defi tvl and protocol data, 2025. URLhttps://defillama.com/. Accessed: December 2025. Francis X. Diebold and Kamil Yılmaz. Better to give than to receiv...
2025
-
[2025]
stablecoin flows to TradFi-connected entities
doi: 10.1016/j.frl.2025.108927. David Vidal-Tomás, Antonio Briola, and Tomaso Aste. FTX’s downfall and Binance’s consol- idation: The fragility of centralised digital finance.Physica A: Statistical Mechanics and its Applications, 2023. doi: 10.2139/ssrn.4368806. Post-FTX marke...
2025
-
[2026]
Lewis Gudgeon, Daniel Perez, Dominik Harz, Benjamin Livshits, and Arthur Gervais
Models fire sale dynamics when stablecoins become systemically important. Lewis Gudgeon, Daniel Perez, Dominik Harz, Benjamin Livshits, and Arthur Gervais. The decentralized financial crisis. InCrypto Valley Conference on Blockchain Technology, 2020. doi: 10.1109/CVCBT50464.20...
2020
Reviewed August 3, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.