{"id":"34999c93-85f5-4d2d-b8d3-9813e7db79d8","arxiv_id":"2411.10249","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Natural fork rates in Proof-of-Work blockchains are governed, to a good approximation, by the product of the propagation-to-mining time ratio and one minus the Herfindahl-Hirschman concentration index.","lead":"Bitcoin miners occasionally solve the puzzle for the next block almost simultaneously, causing stale blocks and wasted energy. This paper derives a simple formula for how often that happens, based on the number of miners, the concentration of computing power, and the time it takes blocks to spread across the network.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed fork-rate formula Eq. (8) depends on synchronized miner starts; realistic propagation desynchronizes starts and may break the HHI factorization, so the empirical comparability claim is not yet established.","rationale":"The paper's mathematical core is clear and internally consistent under its stated assumptions: the derivation of Eqs. (6)–(8) is correct for synchronized exponential miners, and the simulations of Fig. 10 validate exactly that model. The empirical comparison is honest in showing all three propagation thresholds, but the 50% threshold is the only one that agrees, and the fork-rate series has been rescaled to match an external value, so the 'comparable with empirical' claim is not as strong as it first appears. The most load-bearing problem, however, is the synchronized-start assumption, flagged by the authors themselves and by the reader. In real Bitcoin, miners cannot mine the next block until they receive the previous one, so the race starts at staggered times; the order-statistics derivation and the HHI factorization do not obviously survive this change. This concern directly threatens the central quantitative claim, not merely the interpretation. The truncated power-law 'rich-get-richer' justification is an overclaim but is secondary; it does not support the fork-rate formula. Because the concern is real but the model may still be a useful baseline, keeping the CONDITIONAL verdict is appropriate without moving to accept or reject.","tokens_in":19094,"tokens_out":6100,"duration_ms":60792,"concrete_test":"Re-run the simulation in Appendix D with staggered starts: for each round, draw each miner's start time S_i from an empirical block-propagation distribution (e.g., KASTEL 50% arrival times, or an exponential with mean Δ0), then draw mining times E_i ~ Exp(λ_i) and count a fork if two effective mining times T_i = S_i + E_i differ by less than Δ0. Compare the simulated fork rate with Eq. (8) across the 24 periods of Table 1. If the relative error exceeds 30% for realistic S_i distributions, the synchronized-start simplification is quantitatively load-bearing and the paper's headline comparability claim requires qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central formula C(Δ0) ≈ Δ0Λ(1−HHI) (Eq. 8) is derived from the joint distribution of the first two order statistics of N exponential clocks that all start at t=0 (Eqs. 1–6). The paper's own Discussion admits this is 'the relatively strong assumption that miners start mining at the same time.' In Bitcoin, miners begin work on the next block only when they receive the previous block; propagation delays therefore desynchronize starts by amounts comparable to Δ0. When starts are staggered, each miner's effective time to solution is T_i = S_i + E_i, where S_i is the arrival time of the previous block at miner i and E_i ~ Exp(λ_i). The joint distribution of the two smallest T_i is no longer the simple exponential order-statistics structure of Eqs. (1)–(6), and the HHI factorization in Eq. (7) relies on the memoryless common-start property. Thus the predicted fork rate and its concentration dependence may not hold in the very regime the empirical test is meant to validate. The empirical comparison is additionally softened because the observed fork-rate series is rescaled by a factor 1.476 to match a single external value, and Δ0 = Δ0^(50) is selected after seeing the agreement, so Fig. 4b is not an independent confirmation of Eq. (8).","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an analytical model for the rate of soft forks in Proof-of-Work blockchains with heterogeneous miners. The central result is Eq. (8), C(Δ0) ≈ Δ0 Λ (1 − HHI), which expresses the fork probability as the product of the propagation-time threshold, the total hash rate, and a concentration term. The model is tested against Bitcoin data on mining-pool shares, block propagation times, and stale blocks, and is used to argue that both greater hash-rate concentration and lower propagation time reduce fork rates. The paper also fits exponential, log-normal, and truncated power-law distributions to empirical hash rates and uses the model to extrapolate fork rates under different miner counts and hash-rate dispersions.","tokens_in":19333,"tokens_out":10091,"duration_ms":95687,"significance":"The analytical result is clean and potentially useful: if valid, Eq. (8) gives a simple, falsifiable relationship between an observable fork rate and measurable aggregates (total hash rate, propagation time, and HHI). The derivation from exponential order statistics to Eqs. (6)–(8) is mathematically sound under the stated assumptions, and the simulations in Fig. 10 confirm the analytical formulas for several hash-rate distributions. The empirical contribution is more mixed: the hash-rate distribution analysis and the historical concentration trend are informative, but the central validation claim that model-estimated fork rates are 'comparable with the empirical stale blocks rate' is weakened by the rescaling of the fork-rate data and by the posterior choice of the 50% propagation-time percentile. The stress-test concern about synchronized miner starts lands: it identifies a potentially load-bearing assumption that is acknowledged but not quantified.","major_comments":[{"comment":"The central approximation C(Δ0) ≈ Δ0 Λ (1 − HHI) is derived from N exponential clocks that all start at t = 0. In a real network, miner i can begin working on a new block only after receiving it, so its completion time is better described as T_i = S_i + E_i, where S_i is the arrival time of the previous block at miner i and E_i ~ Exp(λ_i). Once the S_i are non-degenerate, the joint law of the two smallest T_i is not the common-start order-statistic law in Eqs. (1)–(6), and the HHI factorization in Eq. (7) relies on the memoryless property of synchronized exponentials. The paper acknowledges this in the Discussion ('the relatively strong assumption that miners start mining at the same time'), but it does not quantify the error. Please add a sensitivity analysis, for example simulating T_i = S_i + E_i with S_i drawn from an empirically plausible propagation-time distribution or a two-miner calculation with fixed start-time offsets, and report how C(Δ0) and its HHI dependence change. Without such a check, the claim that Eq. (8) is validated in the regime of interest is not established.","section":"§2, Eqs. (1)–(8); §3 Discussion"},{"comment":"The empirical validation is not independent as presented. The fork-rate series is rescaled by a constant factor 1.476 so that its value in February 2016 matches the 0.41% value from Gervais et al. [18], and the comparison in Fig. 4b uses Δ0 = Δ0^(50), a choice made after inspecting the agreement with the model. Therefore the match in Fig. 4b is not a free prediction: one point is imposed by construction, and the time-series agreement is not out-of-sample. Please show the raw, unrescaled fork-rate series, and either preselect Δ0 on a training period and evaluate on the remainder or otherwise provide an out-of-sample test. The current Fig. 4b supports only the weaker statement that the model is compatible with the rescaled data for one particular propagation threshold.","section":"§4.1.2, Fig. 4"},{"comment":"The choice of Δ0 as the time to reach 50% of nodes is not derived from the model; in the model, Δ0 should be the time within which a second miner must solve the puzzle for a fork to occur, i.e., the time by which the first block reaches the second successful miner, not a network percentile. The paper shows in Fig. 4a that using 50%, 90%, or 99% propagation times changes the predicted fork rate by up to an order of magnitude, and Fig. 6b shows that the implied HHI is sometimes negative, which the authors correctly attribute to Δ0 being too large. The authors should clarify the interpretation of Δ0 in Eq. (8) in terms of the peer-to-peer propagation process, and report how the conclusions change if a different propagation-time percentile or a fitted Δ0 is used. This is necessary to substantiate the 'comparable with empirical stale blocks rate' claim.","section":"Fig. 4a, §2, Fig. 6b"}],"minor_comments":[{"comment":"There is a sign error or typo in the exponent of the joint density before Eq. (5): as printed, e^{λ_i(t−t′)+Σ_k λ_k t′} does not integrate to Eqs. (5)–(6). The correct exponent should be e^{λ_i(t′−t)−Σ_k λ_k t′} (or an equivalent decaying form). Please fix this to make the derivation self-contained for readers.","section":"Eq. (5) and the preceding joint density"},{"comment":"The method-of-moments estimators for the truncated power-law parameters are stated as α = 1 − (m/s)^2 and β = m/s^2; the expression for α should be 1 − (m/s)^2 only if the variance is such that s > m, which holds here, but the notation is likely a typo for 1 − (m/s)^2. Please clarify the formula and its domain of validity.","section":"Eq. (12), §4.1.5"},{"comment":"The fork rate is defined as the number of block heights with at least one stale or orphan block divided by the total number of blocks, but the term 'fork rate' is later used interchangeably with the probability that a block becomes a stale block. These are not the same quantity, and the distinction should be stated explicitly.","section":"§4.1.2"},{"comment":"The 90% confidence band around the historical fork-rate series is mentioned but its construction is not described. Please specify the source of uncertainty (e.g., Poisson sampling, resampling, or measurement noise) used to produce the shaded area.","section":"Fig. 4b, Fig. 5"},{"comment":"The 'truncated power law' in Eq. (24) is a Gamma-like density with an exponential cutoff, not a truncated power law in the sense of Burroughs and Tebbens [22]; please justify the naming or use a standard term such as 'power law with exponential cutoff.' Also, the text in §C.3.4 refers to Fig. 9 when describing the scenarios of zero miners, but the fork-rate results are in Fig. 8; please correct the cross-references.","section":"§C.3.3 and Fig. 8/9"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the journal's scope and the analytical contribution is solid under its stated assumptions. The empirical validation is the main weakness: the rescaling of fork-rate data and the posterior selection of the 50% propagation percentile make the headline agreement less compelling than the abstract suggests. I would ask for the robustness checks described in the major comments before further consideration; the central formula and the simulation results justify a revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth a serious look for one result: a closed-form approximation for natural fork rates in PoW ledgers with heterogeneous miners, C ≈ Δ0Λ(1−HHI). That is a real extension beyond the homogeneous-miner models of Decker-Wattenhofer and Fadda et al., and it gives a simple, interpretable diagnostic for anomaly detection. The derivation from exponential mining times to Eqs. (6)–(8) is clean, and the Monte Carlo in Fig. 10 matches the analytics across the tested distributions. The empirical work also shows genuine care: long time series, multiple distribution fits, and a sensible discussion of implied versus measured propagation times.\n\nThe soft spots are in the empirical validation, not the math. The fork-rate series is rescaled by a factor of 1.476 to pin one point to Gervais et al., and the headline comparison uses the 50% propagation threshold, selected after seeing the agreement. So Fig. 4b is suggestive, not independent confirmation. The stress-test concern about synchronized miner starts is legitimate: the HHI factorization in Eq. (8) inherits the common-start, memoryless exponential setup, and realistic propagation delays desynchronize starts. The authors acknowledge this in the Discussion, but they do not show how much the formula degrades under staggered starts. That is a gap, not a fatal flaw—the model is explicitly a baseline and the paper is honest about it.\n\nOne more issue: the abstract says the truncated power law is “justified by a rich-get-richer effect,” but the paper never derives that. It is a verbal justification, consistent with prior literature, not a result of this model. It should be softened or backed up.\n\nOverall, the central contribution—the HHI dependence of the natural fork rate—is solid and useful. The paper deserves a serious referee. I would send it to peer review with the expectation that the empirical claims get robustness checks (alternative thresholds, no rescaling) and the TPL claim gets toned down or derived. If those are fixed, it is a solid contribution for a blockchain or distributed-systems venue.","headline":"A clean analytical fork-rate formula with a genuinely new HHI factorization, softened by an acknowledged synchronized-start assumption and a somewhat post-hoc empirical validation.","tokens_in":19925,"tokens_out":2184,"would_cite":false,"duration_ms":22640,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60G70","62G30"],"pacs":[],"model":"deepseek-v4-flash","headline":"Bitcoin's natural fork rate reduces, to first order, to one formula in propagation time, total hash power, and mining concentration.","keywords":["Bitcoin","proof-of-work","soft forks","block propagation time","hash rate concentration","Herfindahl-Hirschman index","order statistics","mining pools"],"falsifier":"Measure, over many blocks on the live network, the actual times at which each mining pool begins hashing the new tip (for example from stratum traffic or pool-published timestamps), and compare the empirical distribution of the gap between the two fastest solutions with the prediction of Eq. (5) built from the same $\\{\\lambda_i\\}$; if the measured gap distribution is systematically wider when propagation is slow, the synchronized-start premise fails and the first-order fork formula with it. A cheaper, already available check: in periods where the paper's Fig. 6b yields negative implied HHI, the inversion of Eq. (8) is meaningless, and those episodes are the places a staggered-start correction would have to bite.","tokens_in":18838,"feed_emoji":"⛓️","tokens_out":12766,"duration_ms":119173,"temperature":0.7,"pith_summary":"This paper sets out to explain how often a Proof-of-Work blockchain like Bitcoin naturally produces competing blocks — soft forks — and to do it analytically rather than by simulation. Its central claim is that the fork probability is, to first order, $C(\\Delta_0) \\approx \\Delta_0 \\Lambda (1 - \\mathrm{HHI})$: the block propagation time $\\Delta_0$ times the total hash rate $\\Lambda$, scaled down by a factor that grows with how concentrated mining power is, measured by the Herfindahl-Hirschman index. Plugging in the time a block needs to reach half of the network, the model reproduces the empirically observed stale-block rate over a decade of Bitcoin data. If the paper is right, fork frequency — and the wasted energy it represents — can be forecast and partly controlled through propagation speed and the distribution of hash rates, and the efficiency-versus-concentration trade-off becomes explicit: fewer natural forks come at the cost of greater mining centralization.","feed_headline":"One formula predicts Bitcoin fork rates from mining concentration","feed_subtitle":"Model ties soft forks to block propagation speed and hash-rate spread, matching real data from 2015 to 2024.","key_machinery":"The load-bearing object is the joint distribution of the first two order statistics of $N$ independent exponential mining clocks. The paper computes the density of the gap $\\Delta$ between the two fastest miners (Eq. 5); evaluating that density at $\\Delta = 0$ gives $\\Lambda(1-\\mathrm{HHI})$, and the first-order Taylor expansion of the gap's cumulative distribution then yields $C(\\Delta_0) \\approx \\Delta_0 \\Lambda (1 - \\mathrm{HHI})$ (Eq. 8). The Herfindahl-Hirschman index $\\mathrm{HHI} = \\sum_i (\\lambda_i/\\Lambda)^2$ is the single summary of hash-rate heterogeneity that enters the formula, and the empirical pipeline that feeds it — estimating $\\lambda_i$ from mined-block counts, fitting exponential, log-normal, and truncated power-law null distributions by the method of moments, plus a semi-empirical Bayesian posterior for the hash rates — is what allows the comparison with observed fork rates.","core_discovery":"The paper derives the probability that two miners solve a block within a propagation-time window. Writing each miner's mining time as an independent exponential clock with rate $\\lambda_i$, the gap between the first two successful miners has a distribution whose density at zero equals $\\Lambda(1-\\mathrm{HHI})$, so for small propagation delays the fork rate is $C(\\Delta_0) \\approx \\Delta_0 \\Lambda (1 - \\mathrm{HHI})$. The authors validate this against Bitcoin data from 2015 to 2024: with $\\Delta_0$ taken as the time to reach 50% of the network, predicted fork rates are comparable with the measured stale-block rate. They further show that empirical hash rates over the decade follow a truncated power law, consistent with a rich-get-richer dynamic capped by total energy supply, that the number of active mining pools fell from about 100 to about 35 while concentration rose, and that these two trends partly offset each other, leaving falling propagation time as the dominant driver of the decline in fork rates. On the security side, the model's inversion — implied $\\Delta_0$ and implied HHI — gives a diagnostic: periods where the implied values contradict the measured ones indicate miners with better-than-median connectivity, intra-pool coordination, or strategic behaviour such as selfish mining.","pith_inferences":["The identity $p(0) = \\Lambda(1-\\mathrm{HHI})$ is a statement about any race between exponential clocks, not only blockchains: wherever the gap between the first two finishers determines wasted work or a tie-break, such as sharded block production, DAG ordering, or replicated commit protocols, the same concentration discount should appear.","Because Eq. (8) needs only the fork rate, $\\Delta_0$, and $\\Lambda$, it could serve as a passive monitor of effective mining concentration in ledgers that do not publish pool attribution — a use the paper gestures at but does not develop.","A direct way to probe the model's limits is to relax the synchronized-start assumption in simulation by drawing each miner's start time from the observed propagation kernel; the deviation from $\\Delta_0\\Lambda(1-\\mathrm{HHI})$ should grow with the spread of start times, and the episodes in the paper's Fig. 6b where implied HHI turns negative are a natural place to look for exactly that failure."],"forward_implications":["Fork rates in any Proof-of-Work ledger can be estimated directly from three measurable quantities — propagation time, total hash rate, and the HHI of the mining distribution — without simulating the consensus protocol.","The observed drop in Bitcoin's fork rate from roughly 1% in 2015 to about 0.1% today is attributable mainly to faster block propagation, because the opposing effects of fewer miners and greater concentration have largely cancelled out.","Higher hash-rate concentration suppresses natural forks, so a ledger designer faces an explicit trade-off: accept more waste and orphaned blocks for a more decentralized mining population, or accept centralization risk for smoother consensus.","Because the formula can be inverted, deviations between measured and implied fork rates become detection signals for non-competitive behaviour — selfish mining, block withholding, or a network core in which miners learn of blocks faster than the median node.","For faster blockchains where the ratio of propagation time to expected mining time exceeds about 0.4, the linear approximation degrades and the full order-statistics expression — including the hash-rate distribution's shape — must be used."],"supporting_citations":[{"why":"Defines the Bitcoin Proof-of-Work protocol and its target block time, the system the fork-rate model is built for.","marker":"[1]"},{"why":"Provides the original network-propagation explanation of forks and the empirical fork-rate benchmark that the model is measured against.","marker":"[5]"},{"why":"Introduces the homogeneous-miner order-statistics fork model that this paper generalizes to heterogeneous hash rates; it is the direct basis for Eqs. (1)-(6).","marker":"[7]"},{"why":"Supplies monitoring data on fork timing and miner attribution used to build the empirical fork-rate series and validate the model.","marker":"[8]"},{"why":"Provides the February 2016 fork-rate value (0.41%) used to rescale the crowd-sourced stale-block data into the empirical benchmark.","marker":"[18]"}],"fun_headline_variants":["Bitcoin fork rates predicted by mining concentration and block speed","Fork rate formula ties hash concentration to block propagation","Fewer miners, faster blocks = fewer Bitcoin forks","Concentration and propagation speed set Bitcoin fork rates","New model links pool size and latency to fork frequency"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The model hinges on every miner beginning to mine the next block at the same instant, so the first two solutions behave like the first two order statistics of independent exponential clocks; staggered starts in the real network would change the gap distribution and could break the clean $\\Delta_0 \\Lambda (1-\\mathrm{HHI})$ dependence — an assumption the authors explicitly flag as 'relatively strong' in the Discussion.","fun_headline_variants_meta":{"raw":{"variants":["Bitcoin fork rates predicted by mining concentration and block speed","Fork rate formula ties hash concentration to block propagation","Fewer miners, faster blocks = fewer Bitcoin forks","Concentration and propagation speed set Bitcoin fork rates","New model links pool size and latency to fork frequency"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00028,"raw_usage":{"total_tokens":1736,"prompt_tokens":1095,"completion_tokens":641,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":711,"completion_tokens_details":{"reasoning_tokens":565}},"tokens_in":711,"tokens_out":641,"duration_ms":7118,"temperature":1.0,"reasoning_tokens":565,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T19:48:48.518930+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure, over many blocks on the live network, the actual times at which each mining pool begins hashing the new tip (for example from stratum traffic or pool-published timestamps), and compare the empirical distribution of the gap between the two fastest solutions with the prediction of Eq. (5) built from the same $\\{\\lambda_i\\}$; if the measured gap distribution is systematically wider when propagation is slow, the synchronized-start premise fails and the first-order fork formula with it. A cheaper, already available check: in periods where the paper's Fig. 6b yields negative implied HHI, the inversion of Eq. (8) is meaningless, and those episodes are the places a staggered-start correction would have to bite.","supporting_citations":[{"cited_title":"Bitcoin: A Peer-to-Peer Electronic Cash System","cited_arxiv_id":null,"evidence_quote":"Defines the Bitcoin Proof-of-Work protocol and its target block time, the system the fork-rate model is built for."},{"cited_title":"Consensus formation on heterogeneous networks","cited_arxiv_id":null,"evidence_quote":"Introduces the homogeneous-miner order-statistics fork model that this paper generalizes to heterogeneous hash rates; it is the direct basis for Eqs. (1)-(6)."}],"review_version":1}