{"id":"9937cc47-4c45-439d-a313-76f1dbdff3b0","arxiv_id":"1908.01406","paper_version":6,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"Permutation tests of Bernoulli randomness are characterized asymptotically, and reanalysis of controlled basketball shooting experiments shows they are underpowered for realistic streaky alternatives, with evidence confined to one shooter.","lead":"This paper develops valid permutation tests and power calculations for detecting 'streaky' alternatives to random Bernoulli sequences, then re-examines four basketball shooting experiments. It finds that existing experiments are too small to measure streakiness reliably, with only one shooter showing significant non-randomness.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Underpowered conclusion rests on NBA-calibrated effect sizes; a direct within-player estimate of streakiness is needed to rule out adequate power.","rationale":"The paper's formal contribution—asymptotic theory for permutation tests and local power against Markov alternatives—appears sound, and the finite-sample simulations support the accuracy of the approximation in the low-power region. The empirical conclusion is conditional on the calibrated values of ε and ζ in Section 5.3. The reader's weakest_assumption identifies exactly this calibration. I find no internal inconsistency or technical error that would threaten the proof machinery. The typo in Theorem 4.1 ('as n → 0') is cosmetic. The main risk is external validity: the translation from NBA cross-player FG% dispersion to within-player streak dependence is not directly estimated, and the power results are sensitive to ε. However, the paper's language is appropriately hedged ('we argue would be consistent'), and the paper transparently states the modeling choice, making it easy to challenge. A direct estimate of ε from shot-level data would settle the matter. Because the concern is a modeling judgment rather than a demonstrated error, the verdict should remain unchanged.","tokens_in":30255,"tokens_out":6611,"duration_ms":67203,"concrete_test":"Fit a mixed-effects logistic regression to NBA play-by-play shot data (e.g., 2013-2019 SportVU), modeling next-shot make probability as a function of previous-shot outcome and streak length with player and game-situation fixed effects, and obtain an estimate of the Markov-chain ε parameter (θ_D/2) for m=1 and m=3, together with a lower 95% confidence bound. Then compute the power of the four controlled experiments at this estimated ε using the paper's formula (4.2) or its replication code. If the lower bound exceeds 0.038, the Section 5.4 conclusion that all existing experiments are insufficiently powered would be overturned for MS and 3PT; if the bound lies below 0.024, the claim survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the GVT, MS, Jagacinski, and 3-point experiments lack adequate power against 'realistic' streaky alternatives is carried by the calibration in Section 5.3: ε ∈ {0.024, 0.038} and ζ ∈ {0.25, 0.5} are anchored to the between-player dispersion of NBA field-goal percentages, not to any estimated within-player state dependence. The power-relevant parameter is θ_D = 2ε for a streaky shooter, and cross-player heterogeneity in average FG% is not a bound on the swing in make probability after a streak. If ε is moderately larger than 0.038, the conclusion reverses for at least two of the four experiments: with ζ=0.5, the Miller-Sanjurjo experiment (ns≈3000) and the 3-point contest (ns≈5644) reach 80% power at ε≈0.045 and ε≈0.033, respectively, and even GVT (ns≈2600) crosses 80% at ε≈0.049. Shooter 109's estimated θ_D of 0.38 in Table 4 shows that larger within-player swings are observable. The paper's power approximation is accurate in the low-power region, so the concern is not technical; it is that the 'realistic' effect size is a modeling judgment, and the underpowered claim is conditional on it.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops asymptotic theory for a class of permutation tests of randomness of Bernoulli sequences, using test statistics that compare the proportion of successes after k consecutive successes with either the overall success proportion or the proportion after k consecutive failures. The main theoretical results characterize the asymptotic and permutation distributions under the null, under stationary alternatives, and under a class of Markov-chain \"streaky\" alternatives, and yield local asymptotic power approximations. The paper then applies these tools to four controlled basketball shooting experiments. It finds that the GVT data contain one shooter (Shooter 109) whose sequence is significantly non-random even after multiple-testing correction, but that the aggregate evidence against randomness is concentrated in that shooter. It argues that all four experiments are underpowered against conservative Markov-chain alternatives calibrated to the cross-player dispersion of NBA field-goal percentages, and concludes that substantially larger datasets are needed to measure streakiness in basketball shooting.","tokens_in":30461,"tokens_out":10756,"duration_ms":109287,"significance":"If the results hold, the paper makes a valuable contribution to the long-running hot-hand debate and to the statistics of testing Bernoulli sequences. The distinction between individual, joint, and simultaneous tests is important, and the analytic power approximations are a practical advance that substantially reduces the computational cost of power calculations. The replication package and the placement of proofs in Online Appendix K are strengths, as is the simulation evidence at n=100 in Figures 2 and 3, which supports the accuracy of the power approximation in the low-power region. The paper's central empirical conclusion, that the existing controlled shooting experiments cannot resolve the hot-hand question, is consequential for behavioral economics. However, as detailed below, the uniqueness theorem for permutation tests is false as stated, and the \"realistic\" effect-size calibration is a modeling judgment that is load-bearing for the underpowered claim.","major_comments":[{"comment":"Theorem 3.2 is false as stated. A simple counterexample for n=2 and alpha=0.05 is phi(0,0)=phi(1,1)=0.05, phi(0,1)=0.10, phi(1,0)=0. Under Bernoulli(p), E[phi] = (1-p)^2*0.05 + p(1-p)*(0.10+0) + p^2*0.05 = 0.05 for every p in (0,1), so phi has exact level alpha, but phi is not invariant under permutations because phi(0,1) differs from phi(1,0). The completeness of the binomial sufficient statistic yields only E[phi | sum X_j] = alpha, not permutation invariance. The claim that permutation tests are the only tests with exact type 1 error control therefore needs correction. The later confidence-bound statement in Section 5.4 only requires exactness of the permutation test itself, but the theorem and its surrounding discussion should be revised, for instance by proving uniqueness within a restricted class of tests or by replacing the \"only\" claim with a conditional-exactness result.","section":"Section 3.2, Theorem 3.2"},{"comment":"The central underpowered conclusion is carried by the calibration of epsilon and zeta from the between-player distribution of NBA field-goal percentages. The parameter theta_D = 2*epsilon is a within-player swing in make probability, and cross-player dispersion in average field-goal percentage does not by itself bound that within-player swing. Shooter 109's estimated theta_D of 0.379 in Table 4 shows that larger within-player swings are observable in these data. The paper's own formula (4.2) makes the conditional nature explicit: for zeta=0.5 and the NBA Three-Point contest (ns approximately 5,600), the test based on D_1 reaches 80% power at epsilon approximately 0.033, which is below the paper's stated upper value of 0.038. The paper should either provide a within-player calibration from repeated-session data, or explicitly report the boundary of parameter values for which each experiment has adequate power, and temper the language that the chosen parameterization is a \"conservative upper bound.\"","section":"Section 5.3 and Equation (4.2)"},{"comment":"Theorem 4.1 contains the typo \"as n to 0\" where the intended statement is clearly \"as n to infinity\". More substantively, the theorem's variance expression for the P-statistic and the subsequent Remark 4.1 are used to justify the local power formula, and the simulation evidence in Figures 2 and 3 supports the approximation in the low-power region. However, the text in Section 4.3 also notes that the approximation overestimates power when the true power is close to 0.9. Because Figure 7 uses the same asymptotic approximation to conclude that some experiments have \"reasonable power\" for m=1 and m=2, the paper should state this high-power caveat prominently in Section 5.3 and indicate the direction of the potential bias.","section":"Section 4.2, Theorem 4.1"}],"minor_comments":[{"comment":"The phrase \"as n to 0\" in statement (ii) should be corrected to \"as n to infinity\".","section":"Section 4.2, Theorem 4.1"},{"comment":"The overestimation of power near 0.9 is acknowledged in the text but should be stated as a limitation of the analytic approximation in the main results, not only in the simulation section.","section":"Section 4.3 and Figure 3"},{"comment":"\"Zenondo\" should be \"Zenodo\" in the data availability statement.","section":"Data Availability Statement"}],"recommendation":"major_revision","confidential_remarks":"The false uniqueness theorem is a clear correctness issue, but it is local and correctable. The calibration concern is also fixable by adding robustness analysis or by softening the language about conservative upper bounds. The paper's core contribution, including the asymptotic power approximations and the reanalysis of the shooting experiments, remains valuable if these points are addressed. The fit to econ.EM is reasonable given the behavioral economics application."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The real contribution here is theoretical: explicit limiting distributions for the plug-in run statistics under randomness, limits for their permutation distributions under null and stationary alternatives, and analytic local power against a class of Markov streaky alternatives. That material is new — GVT, MS, Wardrop, and Miyoshi don't have it — and it makes the paper worth reading even if you never touch basketball. The simulation evidence in Figures 2 and 3 supports the approximations at the sample sizes that matter, and I appreciate that the authors actually ship a replication package.\n\nThe empirical reanalysis is careful. The distinction between individual, joint, and simultaneous tests is long overdue in this literature, and the finding that Shooter 109 survives standard multiple testing corrections is a real data point. But the headline claim — that the four controlled experiments lack adequate power against \"realistic\" streaky alternatives — rests on calibration choices in Section 5.3: ε ∈ {0.024, 0.038} and ζ ∈ {0.25, 0.5} are anchored to between-player dispersion in NBA field goal percentages. That is a modeling judgment, not an estimated within-player effect. If ε is moderately larger, say 0.05 with ζ = 0.5, the MS and three-point-contest experiments cross 80% power. So the underpowered conclusion is conditional on those conservative parameter values. The authors are transparent about this, and they do consider several parameterizations, but it remains the softest load-bearing spot. The stress-test note on this point is right.\n\nMinor issues: Theorem 4.1 says \"as n to 0\" instead of \"as n to infinity\" (a typo, not a substantive error), and the power approximation overestimates power in the high-power region near 0.9. The proofs live in Online Appendix K; they looked complete to me, but they are deferred.\n\nWho is this for? Econometricians and statisticians working on runs tests or permutation methods, and behavioral economists designing new shooting experiments. It deserves a serious referee — I would send it out, not desk-reject. The theoretical results are solid, the empirical claim is honestly conditional, and the limitations are stated rather than hidden. Recommend accept with minor revision and a request to sharpen the calibration discussion, perhaps by reporting power curves over a wider ε grid so readers can see how sensitive the \"inadequate power\" conclusion is.\n\nIn short: a genuinely useful paper with a clear-eyed view of its own assumptions. Cite it if you work on streak detection or permutation tests.","headline":"A genuinely new asymptotic theory for run-based permutation tests of randomness, with a careful but calibration-dependent claim that existing hot hand experiments are underpowered.","tokens_in":31025,"tokens_out":1214,"would_cite":true,"duration_ms":15428,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G10","62G09","62M02","62E20","62P20"],"pacs":[],"model":"deepseek-v4-flash","headline":"Four controlled basketball shooting experiments, including the classic study behind the hot hand fallacy, are too underpowered to detect realistic levels of streakiness, and only one shooter in the classic data is robustly non-random.","keywords":["hot hand fallacy","Bernoulli sequences","permutation tests","streak shooting","Markov chain alternatives","local asymptotic power","multiple testing","small-sample bias"],"falsifier":"Simulate the classic experiment's design, about 26 shooters taking roughly 100 shots each, under the paper's own Markov-chain alternative with $\\epsilon=0.038$, $\\zeta=0.5$, and $m=3$, and run the same stratified permutation test; if more than 80% of simulated datasets reject randomness, the claim that these experiments cannot detect realistic streakiness is wrong.","tokens_in":2016,"feed_emoji":"🏀","tokens_out":3047,"duration_ms":121550,"temperature":0.7,"pith_summary":"This paper develops a formal statistical framework for testing whether a collection of Bernoulli sequences is truly random or streaky, then applies it to the four controlled basketball shooting experiments that have shaped the hot hand debate. It shows that permutation tests comparing success rates after streaks of makes with overall success rates, or with success rates after streaks of misses, are the only tests with exact finite-sample type I error control, and it derives their asymptotic distributions and local power against a class of Markov-chain alternatives. The central empirical conclusion is that all four experiments lack adequate power to detect realistic deviations from randomness calibrated from NBA shooting variation. One shooter in the classic 1985 experiment is robustly streaky after multiple-testing corrections, but the evidence is confined to that shooter. The paper therefore argues that the existing experiments cannot settle whether the hot hand is real or whether people systematically overestimate it, and that a direct test requires larger samples and belief measurements on the same scale as the streaky parameters.","feed_headline":"Basketball shooting studies can't detect realistic hot-hand streaks","feed_subtitle":"Only one shooter in the classic experiment is robustly streaky; all four datasets are underpowered.","key_machinery":"The load-bearing object is a parsimonious class of Markov-chain streaky alternatives: each shooter is either random or streaky, with a proportion $\\zeta$ of streaky shooters, and a streaky shooter increases the chance of a make (and of a miss) by $\\epsilon$ after a run of $m$ consecutive makes (or misses). Against this class the paper studies permutation tests based on the plug-in statistics $\\hat{P}_{n,k}(X_i) - \\hat{p}_{n,i}$ (success rate after $k$ consecutive makes minus overall success rate) and $\\hat{D}_{n,k}(X_i)$ (success rate after $k$ consecutive makes minus success rate after $k$ consecutive misses), averaged over shooters for joint tests. An exact finite-sample theorem shows that permutation tests are the only tests with exact type I error control, avoiding the small-sample bias of the asymptotic approximations; asymptotic results then give the limiting permutation distribution and a closed-form local power approximation: for $m=k=1$, the required total sample size satisfies $ns \\approx \\big((z_{1-\\alpha} - z_{1-\\beta})/(2\\zeta\\epsilon)\\big)^2$, with the general power given by $1 - \\Phi(z_{1-\\alpha} - \\phi_T(k,m,h)\\zeta)$. This machinery turns power analysis from a heavy simulation into an analytic calculation and is what allows the paper to assess all four experiments on one scale.","core_discovery":"The paper's central claim is that a definitive empirical statement about the hot hand fallacy cannot be made from the four controlled shooting experiments currently available. The theoretical part establishes exact finite-sample permutation tests for the null hypothesis that shot outcomes are i.i.d. Bernoulli, shows that these are the only tests with exact type I error control, and characterizes their asymptotic power against a class of Markov-chain alternatives whose two parameters, $\\epsilon$ (size of the streak effect) and $\\zeta$ (share of streaky shooters), are calibrated from the distribution of NBA field-goal percentages. Applied to the data, the tests reject randomness for exactly one shooter in the classic experiment after controlling for multiple testing, and that shooter's sequence is genuinely extreme; the evidence against randomness is otherwise absent. Because all four experiments would detect the benchmark alternatives only with low probability, the paper concludes that the existing data cannot resolve whether shooting is streaky or whether people overestimate streakiness, and that substantially larger experiments are required.","pith_inferences":["If real streakiness is as large as the single robust shooter's estimated $\\theta_D \\approx 0.38$, some of the experiments would have had adequate power, so the underpowered conclusion should be read as conditional on the calibrated $\\epsilon$ range rather than as a universal statement.","The same analytic power framework could be applied to other streak literatures, such as mutual-fund performance persistence or weak-form market efficiency, where small samples and null results may be underpowered in the same way.","A natural next step is to estimate $\\epsilon$ and $\\zeta$ directly from large shot-level NBA tracking data instead of calibrating them from cross-player variation, which would replace the paper's modeling judgment with measured parameters.","The proposed belief-elicitation design, asking observers to state the probability of the next make before each shot with proper scoring, could be piloted side-by-side with the old hypothetical surveys to see whether framing alone explains the gap between stated and revealed beliefs."],"forward_implications":["If the power conclusion is right, a failure to reject randomness in the existing experiments is not evidence that basketball shooting is random; it is evidence only that the studies were too small.","The sample-size formula gives future experimenters a direct target: for the paper's benchmark alternative, the number of shooters times shots per shooter must be roughly $(z_{1-\\alpha} - z_{1-\\beta})^2 / (4\\zeta^2\\epsilon^2)$, which is far larger than any of the four experiments.","The robust rejection for one shooter means that at least one player in the classic data shot in a way that is very unlikely under randomness, so the claim that nobody has a hot hand is not supported even by the experiment that founded the fallacy.","A direct test of the fallacy requires measuring observers' probabilistic expectations of a make after a streak on the same scale as $\\bar\\theta^P_k$ or $\\bar\\theta^D_k$, not the hypothetical survey questions used so far."],"supporting_citations":[{"why":"Supplies the original controlled basketball shooting experiment, its data, and the null result that established the hot hand fallacy as consensus.","marker":"Gilovich et al. (1985)"},{"why":"Identifies the finite-sample bias in plug-in streak statistics and reopens the empirical question by rejecting randomness after bias correction.","marker":"Miller and Sanjurjo (2018d)"},{"why":"The source from which the paper obtained the GVT shooting data and replication programs.","marker":"Miller and Sanjurjo (2018c)"},{"why":"Provides the Spanish semi-professional shooting experiment, one of the four datasets assessed for power.","marker":"Miller and Sanjurjo (2018a)"},{"why":"Provides the NBA three-point contest dataset used as a fourth experiment in the power analysis.","marker":"Miller and Sanjurjo (2019)"},{"why":"Provides the six-player, nine-session controlled shooting dataset included among the four experiments.","marker":"Jagacinski et al. (1979)"},{"why":"Supplies the randomization-hypothesis theory and multiple-testing framework that justify exact permutation tests and FWER control.","marker":"Lehmann and Romano (2005)"},{"why":"Gives the Stein's-method central limit theorem used to derive the limiting permutation distributions.","marker":"Rinott (1994)"},{"why":"NBA field-goal percentage distribution used to calibrate the realistic values of epsilon and zeta in the Markov-chain alternatives.","marker":"Basketball Reference (2019)"}],"fun_headline_variants":["Hot hand data too weak to reveal shooting streaks","Classic hot hand experiments can't detect streaky shooters","Only one shooter shows streakiness in hot hand data","Basketball streak evidence: one shooter breaks randomness","Hot hand fallacy: current data cannot measure streakiness"],"cache_read_input_tokens":33152,"weakest_assumption_plain":"The underpowered conclusion rests on the assumption that realistic streakiness is no stronger than the amount implied by the spread of NBA shooting percentages; if the real hot-hand effect is much larger, some of the experiments would have enough power.","fun_headline_variants_meta":{"raw":{"variants":["Hot hand data too weak to reveal shooting streaks","Classic hot hand experiments can't detect streaky shooters","Only one shooter shows streakiness in hot hand data","Basketball streak evidence: one shooter breaks randomness","Hot hand fallacy: current data cannot measure streakiness"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000239,"raw_usage":{"total_tokens":1549,"prompt_tokens":1011,"completion_tokens":538,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":627,"completion_tokens_details":{"reasoning_tokens":461}},"tokens_in":627,"tokens_out":538,"duration_ms":6590,"temperature":1.0,"reasoning_tokens":461,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:14:00.752677+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate the classic experiment's design, about 26 shooters taking roughly 100 shots each, under the paper's own Markov-chain alternative with $\\epsilon=0.038$, $\\zeta=0.5$, and $m=3$, and run the same stratified permutation test; if more than 80% of simulated datasets reject randomness, the claim that these experiments cannot detect realistic streakiness is wrong.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the original controlled basketball shooting experiment, its data, and the null result that established the hot hand fallacy as consensus."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the NBA three-point contest dataset used as a fourth experiment in the power analysis."},{"cited_title":"J., Newel, K","cited_arxiv_id":null,"evidence_quote":"Provides the six-player, nine-session controlled shooting dataset included among the four experiments."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the randomization-hypothesis theory and multiple-testing framework that justify exact permutation tests and FWER control."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the Stein's-method central limit theorem used to derive the limiting permutation distributions."},{"cited_title":"2018-19 NBA Player Stats: Totals","cited_arxiv_id":null,"evidence_quote":"NBA field-goal percentage distribution used to calibrate the realistic values of epsilon and zeta in the Markov-chain alternatives."}],"review_version":1}