REVIEW 3 major objections 3 minor 60 references
Uncertainty in the Hot Hand Fallacy: Detecting Streaky Alternatives to Random Bernoulli Sequences
T0 review · 3 major / 3 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Four controlled basketball shooting experiments, including the classic study behind the hot hand fallacy, are too underpowered to detect realistic levels of streakiness, and only one shooter in the classic data is robustly non-random.
desk verdict A genuinely new asymptotic theory for run-based permutation tests of randomness, with a careful but calibration-dependent claim that existing hot hand experiments are underpowered. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is a parsimonious class of Markov-chain streaky alternatives: each shooter is either random or streaky, with a proportion $\zeta$ of streaky shooters, and a streaky shooter increases the chance of a make (and of a miss) by $\epsilon$ after a run of $m$ consecutive makes (or misses). Against this class the paper studies permutation tests based on the plug-in statistics $\hat{P}_{n,k}(X_i) - \hat{p}_{n,i}$ (success rate after $k$ consecutive makes minus overall success rate) and $\hat{D}_{n,k}(X_i)$ (success rate after $k$ consecutive makes minus success rate after $k$ consecutive misses), averaged over shooters for joint tests. An exact finite-sample theorem shows that permutation tests are the only tests with exact type I error control, avoiding the small-sample bias of the asymptotic approximations; asymptotic results then give the limiting permutation distribution and a closed-form local power approximation: for $m=k=1$, the required total sample size satisfies $ns \approx \big((z_{1-\alpha} - z_{1-\beta})/(2\zeta\epsilon)\big)^2$, with the general power given by $1 - \Phi(z_{1-\alpha} - \phi_T(k,m,h)\zeta)$. This machinery turns power analysis from a heavy simulation into an analytic calculation and is what allows the paper to assess all four experiments on one scale.
What would settle it
Simulate the classic experiment's design, about 26 shooters taking roughly 100 shots each, under the paper's own Markov-chain alternative with $\epsilon=0.038$, $\zeta=0.5$, and $m=3$, and run the same stratified permutation test; if more than 80% of simulated datasets reject randomness, the claim that these experiments cannot detect realistic streakiness is wrong.
Extended reading notes
Core claim
The paper's central claim is that a definitive empirical statement about the hot hand fallacy cannot be made from the four controlled shooting experiments currently available. The theoretical part establishes exact finite-sample permutation tests for the null hypothesis that shot outcomes are i.i.d. Bernoulli, shows that these are the only tests with exact type I error control, and characterizes their asymptotic power against a class of Markov-chain alternatives whose two parameters, $\epsilon$ (size of the streak effect) and $\zeta$ (share of streaky shooters), are calibrated from the distribution of NBA field-goal percentages. Applied to the data, the tests reject randomness for exactly one shooter in the classic experiment after controlling for multiple testing, and that shooter's sequence is genuinely extreme; the evidence against randomness is otherwise absent. Because all four experiments would detect the benchmark alternatives only with low probability, the paper concludes that the existing data cannot resolve whether shooting is streaky or whether people overestimate streakiness, and that substantially larger experiments are required.
Load-bearing premise
The underpowered conclusion rests on the assumption that realistic streakiness is no stronger than the amount implied by the spread of NBA shooting percentages; if the real hot-hand effect is much larger, some of the experiments would have enough power.
Editorial extensions
If this is right
- If the power conclusion is right, a failure to reject randomness in the existing experiments is not evidence that basketball shooting is random; it is evidence only that the studies were too small.
- The sample-size formula gives future experimenters a direct target: for the paper's benchmark alternative, the number of shooters times shots per shooter must be roughly $(z_{1-\alpha} - z_{1-\beta})^2 / (4\zeta^2\epsilon^2)$, which is far larger than any of the four experiments.
- The robust rejection for one shooter means that at least one player in the classic data shot in a way that is very unlikely under randomness, so the claim that nobody has a hot hand is not supported even by the experiment that founded the fallacy.
- A direct test of the fallacy requires measuring observers' probabilistic expectations of a make after a streak on the same scale as $\bar\theta^P_k$ or $\bar\theta^D_k$, not the hypothetical survey questions used so far.
Reading between the lines
- If real streakiness is as large as the single robust shooter's estimated $\theta_D \approx 0.38$, some of the experiments would have had adequate power, so the underpowered conclusion should be read as conditional on the calibrated $\epsilon$ range rather than as a universal statement.
- The same analytic power framework could be applied to other streak literatures, such as mutual-fund performance persistence or weak-form market efficiency, where small samples and null results may be underpowered in the same way.
- A natural next step is to estimate $\epsilon$ and $\zeta$ directly from large shot-level NBA tracking data instead of calibrating them from cross-player variation, which would replace the paper's modeling judgment with measured parameters.
- The proposed belief-elicitation design, asking observers to state the probability of the next make before each shot with proper scoring, could be piloted side-by-side with the old hypothetical surveys to see whether framing alone explains the gap between stated and revealed beliefs.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops asymptotic theory for a class of permutation tests of randomness of Bernoulli sequences, using test statistics that compare the proportion of successes after k consecutive successes with either the overall success proportion or the proportion after k consecutive failures. The main theoretical results characterize the asymptotic and permutation distributions under the null, under stationary alternatives, and under a class of Markov-chain "streaky" alternatives, and yield local asymptotic power approximations. The paper then applies these tools to four controlled basketball shooting experiments. It finds that the GVT data contain one shooter (Shooter 109) whose sequence is significantly non-random even after multiple-testing correction, but that the aggregate evidence against randomness is concentrated in that shooter. It argues that all four experiments are underpowered against conservative Markov-chain alternatives calibrated to the cross-player dispersion of NBA field-goal percentages, and concludes that substantially larger datasets are needed to measure streakiness in basketball shooting.
Significance. If the results hold, the paper makes a valuable contribution to the long-running hot-hand debate and to the statistics of testing Bernoulli sequences. The distinction between individual, joint, and simultaneous tests is important, and the analytic power approximations are a practical advance that substantially reduces the computational cost of power calculations. The replication package and the placement of proofs in Online Appendix K are strengths, as is the simulation evidence at n=100 in Figures 2 and 3, which supports the accuracy of the power approximation in the low-power region. The paper's central empirical conclusion, that the existing controlled shooting experiments cannot resolve the hot-hand question, is consequential for behavioral economics. However, as detailed below, the uniqueness theorem for permutation tests is false as stated, and the "realistic" effect-size calibration is a modeling judgment that is load-bearing for the underpowered claim.
major comments (3)
- [Section 3.2, Theorem 3.2] Theorem 3.2 is false as stated. A simple counterexample for n=2 and alpha=0.05 is phi(0,0)=phi(1,1)=0.05, phi(0,1)=0.10, phi(1,0)=0. Under Bernoulli(p), E[phi] = (1-p)^2*0.05 + p(1-p)*(0.10+0) + p^2*0.05 = 0.05 for every p in (0,1), so phi has exact level alpha, but phi is not invariant under permutations because phi(0,1) differs from phi(1,0). The completeness of the binomial sufficient statistic yields only E[phi | sum X_j] = alpha, not permutation invariance. The claim that permutation tests are the only tests with exact type 1 error control therefore needs correction. The later confidence-bound statement in Section 5.4 only requires exactness of the permutation test itself, but the theorem and its surrounding discussion should be revised, for instance by proving uniqueness within a restricted class of tests or by replacing the "only" claim with a conditional-exactness result.
- [Section 5.3 and Equation (4.2)] The central underpowered conclusion is carried by the calibration of epsilon and zeta from the between-player distribution of NBA field-goal percentages. The parameter theta_D = 2*epsilon is a within-player swing in make probability, and cross-player dispersion in average field-goal percentage does not by itself bound that within-player swing. Shooter 109's estimated theta_D of 0.379 in Table 4 shows that larger within-player swings are observable in these data. The paper's own formula (4.2) makes the conditional nature explicit: for zeta=0.5 and the NBA Three-Point contest (ns approximately 5,600), the test based on D_1 reaches 80% power at epsilon approximately 0.033, which is below the paper's stated upper value of 0.038. The paper should either provide a within-player calibration from repeated-session data, or explicitly report the boundary of parameter values for which each experiment has adequate power, and temper the language that the chosen parameterization is a "conservative upper bound."
- [Section 4.2, Theorem 4.1] Theorem 4.1 contains the typo "as n to 0" where the intended statement is clearly "as n to infinity". More substantively, the theorem's variance expression for the P-statistic and the subsequent Remark 4.1 are used to justify the local power formula, and the simulation evidence in Figures 2 and 3 supports the approximation in the low-power region. However, the text in Section 4.3 also notes that the approximation overestimates power when the true power is close to 0.9. Because Figure 7 uses the same asymptotic approximation to conclude that some experiments have "reasonable power" for m=1 and m=2, the paper should state this high-power caveat prominently in Section 5.3 and indicate the direction of the potential bias.
minor comments (3)
- [Section 4.2, Theorem 4.1] The phrase "as n to 0" in statement (ii) should be corrected to "as n to infinity".
- [Section 4.3 and Figure 3] The overestimation of power near 0.9 is acknowledged in the text but should be stated as a limitation of the analytic approximation in the main results, not only in the simulation section.
- [Data Availability Statement] "Zenondo" should be "Zenodo" in the data availability statement.
Circularity Check
No significant circularity: the power analysis is conditional on externally calibrated effect sizes, not derived from the experimental outcomes it assesses.
full rationale
The paper's theoretical results (asymptotic distributions under the null, permutation distribution limits, and local power against Markov-chain alternatives) are derived from explicit i.i.d. and Markov assumptions; they are not obtained by assuming the conclusions of the empirical application. The streaky-alternative parameters ε ∈ {0.024, 0.038} and ζ ∈ {0.25, 0.5} are calibrated from the cross-player distribution of NBA field-goal percentages (Section 5.3, Figure 6), which is external to the GVT, Miller-Sanjurjo, Jagacinski, and Three-Point Contest datasets. The underpowered conclusion in Section 5.4 is therefore a conditional statement: given these modeling choices about realistic effect sizes, the four experiments have low power. That is a judgment about effect-size calibration, not a circular derivation, and the paper explicitly flags the conditionality of the claim. The GVT empirical findings, including the Shooter 109 result, come from direct permutation tests of the observed sequences and are not fitted inputs. The few self-citations (Lehmann and Romano 2005; Romano et al. 2011; Politis and Romano 1994) are standard mathematical results or textbook methods, and none is load-bearing in a way that forces the paper's conclusions. No equation or fitted parameter is renamed as a prediction, and no uniqueness theorem is imported to rule out alternatives. The main fragility—whether NBA cross-player dispersion in average shooting percentage bounds within-player streak swings—is a substantive modeling assumption and a potential correctness concern, but not a circularity.
Assumptions & free parameters
free parameters (3)
- epsilon (streakiness magnitude) =
0.024 and 0.038 used in power analysis
- zeta (prevalence of streaky shooters) =
0.25 and 0.5 used in power analysis
- m (streak length in alternative) =
3 benchmark, with values 1 to 4 examined
assumptions (5)
- domain assumption Under H0, each sequence Xi is i.i.d. Bernoulli and the joint distribution is invariant under permutations (randomization hypothesis).
- domain assumption Shot outcomes across shooters are independent stationary Bernoulli processes, and under alternatives follow the class of Markov chain streaky processes in Section 4.1.
- domain assumption For the power analysis, p_i = 0.5 for all individuals.
- standard math The asymptotic results rely on standard CLTs for dependent sequences, including Rinott's Stein-method CLT and Ibragimov's results for alpha-mixing processes.
- domain assumption The NBA field goal percentage distribution is an appropriate external benchmark for calibrating realistic deviations from randomness.
Cite this review
Pith. "Pith review of Uncertainty in the Hot Hand Fallacy: Detecting Streaky Alternatives to Random Bernoulli Sequences." pith.science (2026). https://pith.science/paper/F3WCIVGX
@misc{pith2026190801406,
author = {Pith},
title = {Pith review of: Uncertainty in the Hot Hand Fallacy: Detecting Streaky Alternatives to Random Bernoulli Sequences},
year = {2026},
howpublished = {\url{https://pith.science/paper/F3WCIVGX}},
note = {Machine review of arXiv:1908.01406}
}
read the original abstract
We study a class of permutation tests of the randomness of a collection of Bernoulli sequences and their application to analyses of the human tendency to perceive streaks of consecutive successes as overly representative of positive dependence - the hot hand fallacy. In particular, we study permutation tests of the null hypothesis of randomness (i.e., that trials are i.i.d.) based on test statistics that compare the proportion of successes that directly follow k consecutive successes with either the overall proportion of successes or the proportion of successes that directly follow k consecutive failures. We characterize the asymptotic distributions of these test statistics and their permutation distributions under randomness, under a set of general stationary processes, and under a class of Markov chain alternatives, which allow us to derive their local asymptotic power. The results are applied to evaluate the empirical support for the hot hand fallacy provided by four controlled basketball shooting experiments. We establish that substantially larger data sets are required to derive an informative measurement of the deviation from randomness in basketball shooting. In one experiment, for which we were able to obtain data, multiple testing procedures reveal that one shooter exhibits a shooting pattern significantly inconsistent with randomness - supplying strong evidence that basketball shooting is not random for all shooters all of the time. However, we find that the evidence against randomness in this experiment is limited to this shooter. Our results provide a mathematical and statistical foundation for the design and validation of experiments that directly compare deviations from randomness with human beliefs about deviations from randomness, and thereby constitute a direct test of the hot hand fallacy.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Albright, S. C. (1993). A statistical analysis of hitting streaks in baseball. Journal of the American Statistical Association , 88(424):1175--1183
work page 1993
-
[2]
Appelbaum, B. (2015). Streaks like daniel murphy's aren't necessarily random. The New York Times
work page 2015
-
[3]
Bar-Hillel, M. and Wagenaar, W. A. (1991). The perception of randomness. Advances in Applied Mathematics , 12(4):428--454
work page 1991
-
[4]
Barberis, N. (2018). Psychology-based models of asset prices and trading volume. In Handbook of Behavioral Economics: Applications and Foundations 1 , volume 1, pages 79--175. Elsevier
work page 2018
-
[5]
Barberis, N., Greenwood, R., Jin, L., and Shleifer, A. (2015). X-capm: An extrapolative capital asset pricing model. Journal of Financial Economics , 115(1):1--24
work page 2015
-
[6]
Barberis, N. and Thaler, R. (2003). A survey of behavioral finance. Handbook of the Economics of Finance , 1:1053--1128
work page 2003
-
[7]
2018-19 NBA Player Stats: Totals
Basketball \ Reference (2019). 2018-19 NBA Player Stats: Totals. https://www.basketball-reference.com/leagues/NBA_2019_totals.html [Accessed: July 16, 2019]
work page 2019
-
[8]
Benjamin, D. J. (2019). Errors in probabilistic reasoning and judgment biases. In Handbook of Behavioral Economics: Applications and Foundations 1 , volume 2, pages 69--186. Elsevier
work page 2019
Show all 60 references
-
[9]
Bocskocsky, A., Ezekowitz, J., and Stein, C. (2014). The hot hand: A new approach to an old `fallacy'. In 8th Annual MIT Sloan Sports Analytics Conference . Citeseer
2014
-
[10]
Bradley, R. C. (1986). Basic properties of strong mixing conditions. In Dependence in Probability and Statistics , pages 165--192. Springer
1986
-
[11]
Carhart, M. M. (1997). On persistence in mutual fund performance. The Journal of Finance , 52(1):57--82
1997
-
[12]
Carlson, K. A. and Shu, S. B. (2007). The rule of three: How the third event signals the emergence of a streak. Organizational Behavior and Human Decision Processes , 104(1):113--121
2007
-
[13]
Y., Hoynes, H
Chay, K. Y., Hoynes, H. W., and Hyslop, D. R. (1999). A non-experimental analysis of true state dependence in monthly welfare participation sequences. Proceedings of the American Statistical Association , pages 9--17
1999
-
[14]
Cohen, B. (2015). The `hot hand' debate gets flipped on its head. The Wall Street Journal
2015
-
[15]
Fama, E. F. (1965). The behavior of stock-market prices. The Journal of Business , 38(1):34--105
1965
-
[16]
Fama, E. F. (1970). Efficient capital markets: A review of theory and empirical work. The Journal of Finance , 25(2):383--417
1970
-
[17]
Gilovich, T., Vallone, R., and Tversky, A. (1985). The hot hand in basketball: On the misperception of random sequences. Cognitive Psychology , 17(3):295--314
1985
-
[18]
and Shleifer, A
Greenwood, R. and Shleifer, A. (2014). Expectations of returns and expected returns. The Review of Financial Studies , 27(3):714--746
2014
-
[19]
Haberstroh, T. (2017). He's heating up, he's on fire! klay thompson and the truth about the hot hand. ESPN
2017
-
[20]
Harrison, G. W. and Rutstr \"o m, E. E. (2008). Experimental evidence on the existence of hypothetical bias in value elicitation methods. Handbook of Experimental Economics , 1:752--767
2008
-
[21]
Heckman, J. J. (1981). Heterogeneity and state dependence. Studies in Labor Markets , pages 91--140
1981
-
[22]
Hendricks, D., Patel, J., and Zeckhauser, R. (1993). Hot hands in mutual funds: Short-run persistence of relative performance, 1974--1988. The Journal of Finance , 48(1):93--130
1993
-
[23]
Ibragimov, I. A. (1962). Some limit theorems for stationary processes. Theory of Probability & Its Applications , 7(4):349--382
1962
-
[24]
J., Newel, K
Jagacinski, R. J., Newel, K. M., and Isaac, P. D. (1979). Predicting the success of a basketball shot at various stages of execution. Journal of Sport Psychology , 1(4):301 -- 310
1979
-
[25]
Jensen, M. C. (1968). The performance of mutual funds in the period 1945-1964. The Journal of Finance , 23(2):389--416
1968
-
[26]
Johnson, G. (2015). Gamblers, scientists and the mysterious hot hand. The New York Times
2015
-
[27]
Kahneman, D. (2011). Thinking, Fast and Slow . Macmillan
2011
-
[28]
Keane, M. P. (1997). Modeling heterogeneity and state dependence in consumer choice behavior. Journal of Business & Economic Statistics , 15(3):310--327
1997
-
[29]
Korb, K. B. and Stillwell, M. (2003). The story of the hot hand: Powerful myth or powerless critique. In International Conference on Cognitive Science
2003
-
[30]
K \"u nsch, H. R. (1989). The jackknife and the bootstrap for general stationary observations. Annals of Statistics , 17(3):1217--1241
1989
-
[31]
Lahiri, S. N. (2013). Resampling Methods for Dependent Data . Springer, NY
2013
-
[32]
Lantis, R. M. and Nesson, E. T. (2019). Hot shots: An analysis of the `hot hand' in nba field goal and free throw shooting. Technical report, National Bureau of Economic Research
2019
-
[33]
Lehmann, E. L. and Romano, J. P. (2005). Testing Statistical Hypotheses . Springer, NY, 3 ^ rd edition
2005
-
[34]
Liu, R. Y. and Singh, K. (1992). Moving blocks jackknife and bootstrap capture weak dependence. Exploring the Limits of Bootstrap , 225:248
1992
-
[35]
Malkiel, B. G. (2003). The efficient market hypothesis and its critics. Journal of Economic Perspectives , 17(1):59--82
2003
-
[36]
Manski, C. F. (2004). Measuring expectations. Econometrica , 72(5):1329--1376
2004
-
[37]
Miller, J. B. and Sanjurjo, A. (2017). A visible (hot) hand? expert players bet on the hot hand and win. University of Alicante mimeo
2017
-
[38]
Miller, J. B. and Sanjurjo, A. (2018a). A cold shower for the hot hand fallacy: Robust evidence that belief in the hot hand is justified. University of Alicante mimeo
2018
-
[39]
Miller, J. B. and Sanjurjo, A. (2018b). Momentum isn't magic--vindicating the hot hand with the mathematics of streaks. Scientific American
2018
-
[40]
Miller, J. B. and Sanjurjo, A. (2018c). Supplement to ``surprised by the hot hand fallacy? a truth in the law of small numbers.''. Econometrica , 86(6):2019--2047
2018
-
[41]
Miller, J. B. and Sanjurjo, A. (2018d). Surprised by the hot hand fallacy? a truth in the law of small numbers. Econometrica , 86(6):2019--2047
2018
-
[42]
Miller, J. B. and Sanjurjo, A. (2019). Is it a fallacy to believe in the hot hand in the nba three-point contest? University of Alicante mimeo
2019
-
[43]
Miyoshi, H. (2000). Is the ``hot-hands'' phenomenon a misperception of random events? Japanese Psychological Research , 42(2):128--133
2000
-
[44]
Mood, A. M. (1940). The distribution theory of runs. The Annals of Mathematical Statistics , 11(4):367--392
1940
-
[45]
Politis, D. N. and Romano, J. P. (1994). The stationary bootstrap. Journal of the American Statistical Association , 89(428):1303--1313
1994
-
[46]
N., Romano, J
Politis, D. N., Romano, J. P., and Wolf, M. (1999). Subsampling . Springer, NY
1999
-
[47]
Rabin, M. (2002). Inference by believers in the law of small numbers. The Quarterly Journal of Economics , 117(3):775--816
2002
-
[48]
and Vayanos, D
Rabin, M. and Vayanos, D. (2010). The gamblers and hot-hand fallacies: theory and applications. Review of Economic Studies , 77(2):730--778
2010
-
[49]
Rao, J. M. (2009). Experts' perceptions of autocorrelation: The hot hand fallacy among professional basketball players
2009
-
[50]
Remnick, D. (2017). Bob dylan and the `hot hand.'. The New Yorker
2017
-
[51]
Rinott, Y. (1994). On normal approximation rates for certain sums of dependent random variables. Computational and Applied Mathematics , 55(2):134--143
1994
-
[52]
P., Shaikh, A., and Wolf, M
Romano, J. P., Shaikh, A., and Wolf, M. (2011). Consonance and the closure method in multiple testing. The International Journal of Biostatistics , 7(1):1--25
2011
-
[53]
Romano, J. P. and Wolf, M. (2005). Exact and approximate stepdown methods for multiple hypothesis testing. Journal of the American Statistical Association , 100(469):94--108
2005
-
[54]
Stern, H. S. and Morris, C. N. (1993). A statistical analysis of hitting streaks in baseball: Comment. Journal of the American Statistical Association , 88(242):1189--1194
1993
-
[55]
Stone, D. F. (2012). Measurement error and the hot hand. Scientific American , 66(1):61--66
2012
-
[56]
Thaler, R. H. and Sunstein, C. R. (2009). Nudge: Improving Decisions about Health, Wealth, and Happiness . Penguin
2009
-
[57]
Torgovitsky, A. (2019). Nonparametric inference on state dependence in unemployment. Econometrica , 87(5):1475--1505
2019
-
[58]
and Kahneman, D
Tversky, A. and Kahneman, D. (1971). Belief in the law of small numbers. Psychological Bulletin , 76(2):105
1971
-
[59]
and Kahneman, D
Tversky, A. and Kahneman, D. (1981). The framing of decisions and the psychology of choice. Science , 211(4481):453--458
1981
-
[60]
Wardrop, R. L. (1999). Statistical tests for the hot-hand in basketball in a controlled setting
1999
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.