{"id":"b35ce790-b12b-4324-8678-ed14c260fe24","arxiv_id":"2608.09378","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"USDT and USDC transfer-value distributions are heavy-tailed with fitted power-law exponents around 1.4 to 1.8, and smart-contract-to-smart-contract transfers show consistently larger exponents than transfers involving regular accounts.","lead":"This paper fits power-law curves to the sizes of hundreds of millions of USDT and USDC transfers on Ethereum. It finds that transfers between smart contracts have a different tail exponent than transfers involving ordinary accounts, and this difference persists over time.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Missing goodness-of-fit validation: the central 'power-law scaling' claim rests on visual CCDF inspection and MLE fits alone, so the abstract overstates what Section 4.1 establishes.","rationale":"The reader's verdict identified exactly this gap, and I agree it is the load-bearing assumption. Without the goodness-of-fit test, the distinction between a genuine power-law tail and a generically heavy-tailed distribution is unsupported; the exponent estimates remain descriptive, but the categorical claim in the abstract ('Power-law tail behavior is observed throughout stablecoin transaction activity') is stronger than the evidence. The paper deserves credit for the sensitivity design and for the explicit caveat in the conclusion, but those do not substitute for model validation. A secondary issue is the counterfactual: Section 4.3 correctly labels the comparison as a range comparison, while the abstract's 'accounts for only 10%-35%' overstates it, and a pooled tail exponent is not generally a weighted average of component exponents; this affects the counterfactual claim but not the two-regime finding. If the goodness-of-fit test and Vuong comparison largely support the power-law fits, the conditional verdict could be upgraded to acceptance; if not, the claim should be softened. For now, the reader's conditional verdict is appropriate, so no adjustment is needed.","tokens_in":16066,"tokens_out":6287,"duration_ms":66517,"concrete_test":"Implement, for every category-period-stablecoin fit (6 periods × 2 stablecoins × 5 categories), the Clauset et al. (2009) goodness-of-fit test on the paper's estimated xmin values: generate 1000 synthetic datasets from the fitted power law above xmin_hat, recompute the KS statistic, and record the p-value. Then fit a lognormal to the same tail region and run Vuong's test against the power-law fit, using the same candidate threshold grid as Section 2.1. Report the fraction of fits with p<0.05 and the fraction where lognormal is preferred; if most fits reject the power-law null, the headline should be softened from 'power-law scaling' to 'heavy-tailed scaling'.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.1 asserts that the empirical distributions 'are well described by the power-law form above xmin,' but the only evidence offered is the visual straightness of the CCDFs in Fig. 1 and the MLE exponent values. Section 2.1 explicitly adopts the Clauset et al. framework, whose standard protocol includes a semi-parametric bootstrap goodness-of-fit test of the fitted power law against synthetic data and a comparison against alternative heavy-tailed families; neither step is performed, and no p-values are reported. Many heavy-tailed families (lognormal, Weibull, truncated power laws) can appear linear on a log-log plot over the fitted range, so the fitted alpha values do not by themselves establish that the tail is power-law rather than merely heavy-tailed. The paper's own conclusion concedes that confirmation would require formal tail-model validation, yet the abstract and strongest claim state the power-law conclusion categorically. The sensitivity analysis is well designed, but it tests stability of the fitted exponent under sampling, not whether the power-law family is the right model.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper analyzes approximately 370 million USDT and USDC transfers on the Ethereum blockchain across six periods from June 2024 to February 2026. Transactions are classified into four interaction categories (EOA–EOA, EOA–SC, SC–EOA, SC–SC), and the paper fits power-law tails to transaction-value distributions using maximum-likelihood estimation with a KS-distance-based x_min selection. The authors report that all categories and periods exhibit heavy-tailed power-law scaling, with EOA-involved categories clustering around alpha ≈ 1.45–1.60 and SC–SC transactions around alpha ≈ 1.72–1.73. Robustness is assessed by varying fitting sample sizes and random seeds, and a composition-only counterfactual exponent is constructed to test whether changing category weights can explain the temporal variation of the overall exponent. The paper concludes that interaction-specific scaling regimes exist and that composition changes alone cannot reproduce the observed overall exponent path.","tokens_in":16266,"tokens_out":5511,"duration_ms":51085,"significance":"If the power-law claim survives formal validation, this would be the first large-scale empirical characterization of stablecoin transaction-value tails, with a valuable decomposition by account interaction type. The dataset is large, the estimation is standard, and the sensitivity analysis is thorough: the authors test five fitting sample sizes and five random seeds per condition and report stable mean exponents. The counterfactual is explicitly labeled as descriptive, and the conclusion appropriately notes that universality would require formal tail-model validation. However, the manuscript's central claim that the tails are 'well described by the power-law form' currently rests on visual inspection and MLE fits alone, without the goodness-of-fit tests and alternative-family comparisons that the cited Clauset et al. framework prescribes. The two-regime separation is also weaker in later periods than the averaged numbers suggest. These gaps are fixable and do not undermine the value of the data construction, but they temper the strength of the current conclusions.","major_comments":[{"comment":"The central claim that the distributions 'are well described by the power-law form above xmin' is not supported by any goodness-of-fit test. Section 2.1 explicitly adopts the Clauset et al. framework, whose standard protocol includes a semi-parametric bootstrap p-value for the power-law null and a comparison against alternative heavy-tailed families; neither is performed, and no p-values are reported. Many heavy-tailed families (e.g., lognormal, truncated power law) can appear approximately linear on the fitted range, so the MLE exponents alone do not establish that the tail is power-law rather than merely heavy-tailed. The conclusion itself concedes that 'confirmation of universality would require formal tail-model validation,' which is in tension with the categorical wording of the abstract and Section 4.1. I recommend either adding the bootstrap p-values and alternative-model comparisons, or softening the claims throughout to 'heavy-tailed' and 'power-law-like.'","section":"Section 4.1, Fig. 1"},{"comment":"The two-regime separation is weaker than the averaged numbers suggest. For USDT in Period 5 the SC-SC exponent is 1.593 and in Period 6 it is 1.621; for USDC the corresponding values are 1.560 and 1.637. These values overlap the upper end of the EOA-involved range (~1.43–1.60), so the claim that SC-SC transactions exhibit exponents of approximately 1.72–1.73 relies on earlier periods and on averaging across periods with large standard deviations (0.120 for USDT, 0.148 for USDC in the N=200,000 row of Table 4). The authors should report per-period confidence intervals and a formal test of whether the regime separation is statistically significant, rather than comparing period-averaged point estimates.","section":"Section 4.1, Fig. 1(e)-(f), Table 4"},{"comment":"The counterfactual conclusion is stated more strongly than the analysis supports. The quantity C_range = 100 * Delta_alpha_comp / Delta_alpha_actual compares only the temporal ranges of two different objects: the pooled-distribution exponent and a weighted average of category-specific exponents. As the authors note, a pooled exponent is not equal to a weighted average of component exponents, so the '10%–35%' figure is a descriptive range ratio, not a measure of explained variation. The abstract and conclusion nevertheless say composition changes 'cannot explain' the observed variation. I suggest either presenting a direct comparison of the two paths (e.g., correlation or mean absolute difference) and/or consistently using language such as 'is not reproduced by a composition-only counterfactual in this descriptive sense.'","section":"Sections 2.2 and 4.3, Table 5"}],"minor_comments":[{"comment":"The abstract states 'February 202' but should read 'February 2026' to match Table 1 and the body text.","section":"Abstract"},{"comment":"The estimated x_min values are never reported in the text or tables; reporting them would allow readers to assess the tail fraction and reproduce the fits.","section":"Section 4.1"},{"comment":"The USDC total of 158,945,393 does not match the sum of the four category totals (158,944,393), an apparent arithmetic error of 1,000.","section":"Table 3"},{"comment":"The inequality symbol renders as '6' in 'P(X 6 x)'; the typesetting of the less-than-or-equal sign should be corrected.","section":"Figure 1 caption"},{"comment":"The phrase 'the three categories involving EOAs interaction cluster' is awkward; suggest 'the three EOA-involved categories'.","section":"Section 4.1"},{"comment":"References [30] and [36] are the same Clauset et al. paper; please consolidate or clearly distinguish the SIAM Review version from the arXiv preprint.","section":"References"},{"comment":"The statement that preprocessed data are 'available from the author' should be 'the authors' or list a corresponding author.","section":"Data availability"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within scope for q-fin.ST and the empirical work is substantial, but the categorical power-law claim is not yet backed by the formal validation the paper's own methodology section references. This is fixable by adding goodness-of-fit tests, alternative-model comparisons, and a statistically grounded test of regime separation. I would not recommend rejection, as the underlying data construction and sensitivity analysis are useful and the limitations are explicitly acknowledged in the conclusion."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing you should know: the paper is better than the abstract. It analyzes roughly 370 million USDT/USDC transfers, fits MLE power-law tails per period and interaction category, and finds a consistent split: EOA-involved categories around alpha 1.45-1.60 versus SC-SC around 1.72-1.73. With six periods, two stablecoins, five fitting sample sizes, and five random seeds, the separation survives. That is a genuinely new large-scale empirical result. The sensitivity analysis is well designed and honestly reported, including the caveat that the confidence intervals are conditional on the estimated xmin. The counterfactual is clearly labeled as a descriptive range comparison, not an explained-variance measure, and the text explicitly notes that a pooled exponent need not equal a weighted average of category exponents.\n\nThe soft spot is real: the title and abstract say power-law, but the paper never runs the Clauset goodness-of-fit test or compares against lognormal, Weibull, truncated power-law, or other heavy-tailed alternatives. Visual straightness of the CCDF plus MLE exponent values does not establish that the power-law family is the right model. The stress-test note has this right. It is not fatal, though: the weaker claim of heavy-tailed scaling survives the gap, and the authors' own conclusion concedes that confirmation would require formal tail-model validation. The fix is either to add p-values and Vuong tests (which could go either way on this data) or to soften the abstract and title to \"heavy-tailed scaling\" and remove the categorical power-law wording. The \"accounts for 10-35%\" phrase in the abstract is also a bit loose; it is a range ratio, not explained variance, though the body is careful about this.\n\nNothing circular here. The exponents are fitted from the same data used to report the pattern, which is normal for empirical characterization. The citations include the authors' own prior work on ERC20 token networks, but that work is directly relevant, and self-citation is not a flaw in this context. The one practical weakness is that no code or processed data is shipped; \"available from the author upon request\" is weaker than it should be, though the raw blockchain data is public.\n\nVerdict: a conditional accept for peer review. I would send it out. The empirical contribution is substantial and the missing goodness-of-fit tests are a fixable methodological requirement, not a desk-reject reason. A good referee should ask for formal tail-model validation or softened claims, plus code and data release. I would cite this for the stablecoin-specific exponents and the interaction-category split even in its current form.","headline":"Solid large-scale empirical mapping of stablecoin tail exponents with a real EOA/SC split; needs formal tail-model validation before the abstract can say 'power-law'.","tokens_in":16787,"tokens_out":2761,"would_cite":true,"duration_ms":27682,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"USDT and USDC transaction values on Ethereum follow heavy-tailed power laws, with exponents that split by whether a smart contract is involved.","keywords":["Stablecoin","Ethereum","USDT","USDC","Power-law distribution","Heavy tails","Scaling exponent","Blockchain transactions"],"falsifier":"Run a bootstrap goodness-of-fit test for power-law fits on the USDT and USDC tails; if most of the fitted category-period combinations fail the test, the claim of power-law scaling would be weakened, and a direct comparison of log-likelihoods against a lognormal fit would decide whether the tail is actually a power law or merely heavy.","tokens_in":15877,"feed_emoji":"📊","tokens_out":6963,"duration_ms":63196,"temperature":0.7,"pith_summary":"This paper tries to establish that the dollar-denominated values of USDT and USDC transfers on Ethereum follow heavy-tailed power-law distributions in every period and every interaction category, with tail exponents in the range 1.4 to 1.8. More specifically, it claims the exponent separates into two regimes: transactions involving at least one externally owned account (a user-controlled wallet) cluster near 1.45–1.60, while smart-contract-to-smart-contract transfers sit higher, near 1.72–1.73. The reason this would matter is that it turns an unexplored corner of blockchain finance into a statistical object comparable to stock returns, trade sizes, and other quantities in the financial-scaling literature, while showing that the interaction mechanism, not just the asset, shapes the tail. The paper also argues on the basis of a counterfactual exercise that shifts in the mix of transaction categories explain only 10–35% of the temporal variation of the overall exponent, so the observed drift in the pooled exponent reflects changes inside categories rather than mere composition changes.","feed_headline":"Stablecoin transfers split into two power-law regimes","feed_subtitle":"370 million USDT/USDC transactions on Ethereum show heavy tails whose slope depends on whether a smart contract is involved.","key_machinery":"The load-bearing object is the tail exponent $\\alpha$ of the power-law density $p(x) = C x^{-\\alpha}$ for $x \\geq x_{\\min}$, estimated by maximum likelihood after selecting $x_{\\min}$ to minimize the Kolmogorov–Smirnov distance between the empirical and fitted cumulative distributions. The second essential piece is the interaction classification of Ethereum accounts into externally owned accounts (user-controlled wallets) and smart contracts (program-controlled accounts), which produces the four categories whose exponents are compared. The third piece is the composition-only counterfactual exponent $\\alpha_{\\mathrm{comp},t} = \\sum_i w_i(t) \\alpha_{i,\\mathrm{ref}}$, where the category weights $w_i(t)$ change over time but the category exponents are held at their six-period averages; this counterfactual is the test that separates composition effects from within-category tail changes.","core_discovery":"Working with roughly 370 million USDT and USDC transactions on Ethereum across six periods from June 2024 to February 2026, the paper reports that the complementary cumulative distribution of transaction values is “well described by the power-law form above $x_{\\min}$” for both stablecoins, all four interaction categories (EOA–EOA, EOA–SC, SC–EOA, SC–SC), and all six periods. The maximum-likelihood estimates of the tail exponent $\\alpha$ fall in $1.4 \\lesssim \\alpha \\lesssim 1.8$, consistent with earlier estimates for trade sizes and share volumes. Averaged over periods, the three EOA-involved categories give $\\alpha \\approx 1.45$\\u2013$1.60$, whereas the SC–SC category gives $\\alpha \\approx 1.72$\\u2013$1.73$, and this separation persists across fitting sample sizes from 50,000 to 800,000 and across five random seeds. The composition-only counterfactual, which holds category exponents fixed and lets only category weights vary, reproduces only 10–35% of the temporal range of the directly fitted overall exponent. The paper's own conclusion is careful: the results show recurring power-law-like tails with robust interaction-specific differences, not a single universal tail exponent for stablecoins.","pith_inferences":["Beyond the paper, the same classification and counterfactual machinery could be applied to other blockchains and stablecoins (the paper lists DAI, BUSD, USDP, BSC, Solana, and Polygon as future work), which would test whether the EOA-versus-SC split is specific to Ethereum or a general property of smart-contract settlement.","One mechanism the paper leaves implicit is that SC–SC transfers are dominated by DeFi protocol operations—swaps, liquidity provisioning, rebalancing—which may process amounts on a different economic scale than human peer-to-peer payments; that difference could explain the higher exponent.","If the tails were instead better described by a lognormal or truncated Pareto distribution, the fitted exponents would remain useful descriptive numbers but would not by themselves support a scaling-law interpretation; this is a testable alternative, not a failure of the empirical regularities reported."],"forward_implications":["The overall pooled exponent is a mixture of category-specific exponents, so studies of stablecoin scaling should stratify by account type rather than fit a single tail.","Because the exponent separation is stable across four fitting sample sizes and five random seeds, the two-regime structure is unlikely to be an artifact of subsampling.","The counterfactual range result (10–35%) implies that temporal changes in the overall exponent signal changes in the shape of category-level tails, not just changes in which category dominates the transaction count.","With exponents between 1.4 and 1.8, the tail distributions have infinite variance, so standard deviation-based risk measures for stablecoin transfer sizes are not well defined and tail-focused metrics are needed."],"supporting_citations":[{"why":"Supplies the maximum-likelihood estimator and the Kolmogorov–Smirnov minimization for choosing $x_{\\min}$, the procedure behind every exponent in the paper.","marker":"[36]"},{"why":"Supplies the XBlock-ETH dataset from which the roughly 370 million USDT and USDC transactions are taken.","marker":"[42]"},{"why":"Provides the official USDT and USDC contract addresses used to identify the two stablecoins in the Ethereum data.","marker":"[44]"},{"why":"Provides the benchmark power-law exponents for trade sizes and share volumes (about 1.5 and 1.7) that the paper compares its 1.4–1.8 range against.","marker":"[13]"},{"why":"Gives the inverse-cubic stock-return scaling result that frames the paper's expectation of heavy tails in financial quantities.","marker":"[18]"},{"why":"Supplies prior Bitcoin return power-law exponents in the range 2–2.5, the nearest cryptocurrency comparison for stablecoin tail behavior.","marker":"[21]"},{"why":"Reports power-law scaling in ERC20 token transfers on Ethereum, the prior result this paper extends to stablecoin transfers stratified by account type.","marker":"[8]"}],"fun_headline_variants":["Stablecoin tails split by smart contract involvement","Two power laws govern stablecoin transfers","Smart contracts shift stablecoin tail exponents","USDT and USDC show dual scaling regimes","Stablecoin values follow distinct power-law tails"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that a maximum-likelihood power-law fit with a Kolmogorov–Smirnov-chosen threshold and a visual check of the complementary cumulative distribution is sufficient to establish that the tails are power-law, and it does not test the power-law null against other heavy-tailed distribution families.","fun_headline_variants_meta":{"raw":{"variants":["Stablecoin tails split by smart contract involvement","Two power laws govern stablecoin transfers","Smart contracts shift stablecoin tail exponents","USDT and USDC show dual scaling regimes","Stablecoin values follow distinct power-law tails"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000411,"raw_usage":{"total_tokens":2233,"prompt_tokens":1157,"completion_tokens":1076,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":773,"completion_tokens_details":{"reasoning_tokens":1009}},"tokens_in":773,"tokens_out":1076,"duration_ms":9281,"temperature":1.0,"reasoning_tokens":1009,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T18:24:23.776742+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a bootstrap goodness-of-fit test for power-law fits on the USDT and USDC tails; if most of the fitted category-period combinations fail the test, the claim of power-law scaling would be weakened, and a direct comparison of log-likelihoods against a lognormal fit would decide whether the tail is actually a power law or merely heavy.","supporting_citations":[{"cited_title":"Zheng, Z","cited_arxiv_id":null,"evidence_quote":"Supplies the XBlock-ETH dataset from which the roughly 370 million USDT and USDC transactions are taken."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the official USDT and USDC contract addresses used to identify the two stablecoins in the Ethereum data."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the benchmark power-law exponents for trade sizes and share volumes (about 1.5 and 1.7) that the paper compares its 1.4–1.8 range against."},{"cited_title":"Plerou, H","cited_arxiv_id":null,"evidence_quote":"Gives the inverse-cubic stock-return scaling result that frames the paper's expectation of heavy tails in financial quantities."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies prior Bitcoin return power-law exponents in the range 2–2.5, the nearest cryptocurrency comparison for stablecoin tail behavior."},{"cited_title":"Universal Patterns in the Blockchain: Analysis of EOAs and Smart Contracts in ERC20 Token Networks","cited_arxiv_id":"2508.04671","evidence_quote":"Reports power-law scaling in ERC20 token transfers on Ethereum, the prior result this paper extends to stablecoin transfers stratified by account type."}],"review_version":1}