{"id":"60c94135-8152-4fc6-94e6-71b98b7c3c0f","arxiv_id":"1908.05745","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A Bayesian joint model of shot locations and shot outcomes shows a positive association between shot frequency and shooting accuracy for 80% of the NBA's 50 most frequent shooters in the 2017-2018 season.","lead":"This paper builds a Bayesian statistical model that describes where NBA players take shots and whether those shots go in. It finds that for 40 of the 50 most frequent shooters in the 2017-2018 season, players made shots at a higher rate in the locations where they shot more often.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 40-of-50 majority claim is not auditable: per-player DIC/LPML differences and ξ significance are not reported, and the paper's own thresholds suggest many 'favored' calls may be weak.","rationale":"The paper's most important empirical result is the aggregate statement in Section 6.4 that 40 of the top 50 shooters favor ξ≠0, which drives the abstract's 'majority' conclusion. To support that statement, one needs to know not merely that DIC and LPML pointed in the same direction, but that the preferences were strong enough to exceed the paper's own thresholds. Section 6.2 establishes those thresholds explicitly, yet Section 6.4 reports no magnitudes. The four-player table illustrates the risk: Durant is counted as a 'clear advantage' for ξ≠0 even though ΔDIC=8.6 and ΔLPML=4.2, below the stated 'substantial' and 'very strong' benchmarks. If many of the 40 players resemble Durant, the count could be dominated by weak evidence. Additionally, the text reports only that the 40 estimates were positive, not that their credible intervals excluded zero; positive posterior means can occur under models with weak support. This is not a claim that the model is wrong, but rather that the headline count is currently unverifiable from the manuscript. The shot clock confounder identified by the reader is real and is acknowledged in Section 7, but it is a second-order concern compared with the missing evidence for the count itself. A simple re-analysis reporting per-player comparison magnitudes and HPD intervals would settle whether the majority claim holds at the strength the paper implies. The simulation studies and four-player analyses are useful, but they do not substitute for the missing top-50 details.","tokens_in":15269,"tokens_out":11734,"duration_ms":123041,"concrete_test":"Rerun the Section 6.4 top-50 analysis (or obtain the authors' per-player output) and tabulate, for each player, ΔDIC = DIC(ξ=0) − DIC(ξ≠0), ΔLPML = LPML(ξ≠0) − LPML(ξ=0), and the 95% HPD interval for ξ. Then count how many players satisfy the paper's own thresholds (ΔDIC > 10 and/or ΔLPML > 4.5) and have HPD intervals excluding 0. If the count is materially below 40 (e.g., 30 or fewer), the 'majority positive association' claim should be weakened or qualified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim is the Section 6.4 statement that 40 of the top 50 most frequent shooters favored the intensity-dependent model (ξ≠0) in terms of DIC and LPML, with all estimated intensity coefficients positive. This count is the sole basis for the abstract's 'majority' conclusion, but the paper never reports the magnitudes of the DIC or LPML differences for those 40 players. Section 6.2 sets explicit thresholds: a DIC difference larger than 10 is substantial, a difference of 2–3 does not give evidence, and an LPML difference larger than 4.5 is 'very strong.' The four-player table shows the risk: Durant is called a 'clear advantage' for ξ≠0 even though ΔDIC=8.6 and ΔLPML=4.2, both below the paper's stated benchmarks. If a similar share of the 40 players has differences in the 1–10 DIC or 0.5–4.5 LPML range, the count could reflect weak or negligible model preferences rather than a robust majority. The text also reports only that the 40 estimated coefficients were positive, not that their 95% HPD intervals excluded zero; positive posterior means can occur even when evidence for ξ≠0 is weak. Without per-player comparison magnitudes and significance status, the headline count is not verifiable from the manuscript.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a Bayesian marked spatial point process model for NBA shot charts, in which shot locations follow a non-homogeneous Poisson process with intensity covariates derived from nonnegative matrix factorization bases, and the binary shot-success mark is modeled by a logistic regression that includes the shot intensity as a covariate. Inference is performed with an MCMC algorithm implemented in NIMBLE, and model comparison is based on DIC and LPML. The estimation procedure is evaluated in simulation studies that report small bias and generally reasonable coverage. The method is applied to four named players and then to the top 50 most frequent shooters of the 2017–2018 NBA regular season, with the central empirical claim being that 40 of the 50 players favored the intensity-dependent mark model and that all 40 estimated intensity coefficients were positive. The fitted coefficients are also used as inputs for hierarchical clustering of the players.","tokens_in":15587,"tokens_out":6140,"duration_ms":55871,"significance":"If the central empirical finding holds, the paper offers a quantitative, data-driven measure of shot-selection efficiency and a practical joint modeling framework for marked spatial point processes with applications beyond sports. The work is methodologically constructive: it combines NMF basis construction, a Bayesian joint model, and model comparison criteria in a reproducible pipeline, and the simulation studies are reported in detail with tables of bias, standard deviation, and coverage. The paper also produces falsifiable empirical claims about top NBA shooters, which are of interest to the sports analytics community. However, the significance of the headline 40-of-50 result is currently limited by the lack of per-player evidence in the manuscript, the inconsistent application of the paper's own model-comparison thresholds, and the acknowledged omission of shot clock information that could confound the intensity-accuracy association.","major_comments":[{"comment":"In addition, reporting only that the 40 estimated coefficients were positive is insufficient, because a positive posterior mean can occur even when the posterior mass overlaps zero substantially. The authors should report how many of the 40 players have HPD intervals for ξ that exclude zero, and ideally the posterior probabilities that ξ > 0.","section":"Section 6.4 and Section 6.2"},{"comment":"The paper should also clarify whether a player is counted as favoring the intensity-dependent model when both DIC and LPML favor it, or when either criterion does, since the counting rule directly affects the reported 40-of-50 figure.","section":"Section 6.2 and Table 1"},{"comment":"Additionally, the text in Section 6.2 says the reported results are from the second run, but Table 1 is presented before this caveat is explained in detail, which could mislead a reader who assumes the comparisons are from a single pre-specified model.","section":"Section 6.2 and Section 6.3"},{"comment":"The manuscript should also avoid language in the abstract and introduction that overstates the finding as a general positive association without explicitly acknowledging this confounding risk at the point of the claim.","section":"Section 7 and Section 6.1"}],"minor_comments":[{"comment":"The coverage summaries would be more interpretable if the Monte Carlo standard errors of the coverage rates were reported.","section":"Tables 3 and 4"},{"comment":"The sentence 'The might be due to his injury in that season and reduced time on court' in Section 6.3 contains a grammatical error ('The might be due') and appears to be an incomplete thought.","section":"Section 6.3"},{"comment":"The paper does not state what software was used for the hierarchical clustering beyond naming Ward's method and the R implementation; the specific function or package should be mentioned for reproducibility.","section":"Figure 4"},{"comment":"The introduction and abstract state that the intensity-dependent model is preferred for 'a majority' of the top 50 shooters, while the introduction later says 'about 80%'; these statements should be consistent throughout the manuscript.","section":"General"},{"comment":"The text in Section 6.2 says that convergence was confirmed by trace plots, but does not report any quantitative convergence diagnostics, such as Gelman-Rubin statistics or effective sample sizes; adding these would strengthen the reproducibility of the MCMC results.","section":"Section 6.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the journal's scope and the modeling framework is a useful contribution. The main concern is not the method itself but the verifiability of the headline empirical claim: the 40-of-50 count is not supported by the reported per-player statistics, and the paper's own thresholds are applied inconsistently in the four-player analysis. The issues are fixable with additional reporting and a more careful wording of the conclusions, so I recommend major revision rather than rejection. Please also note that the paper does not include a data or code availability statement; given the reliance on a public data source and the NIMBLE implementation, a reproducibility statement would be helpful."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Good paper, with a real gap between the headline and the evidence. The methodological core—a Bayesian joint model of shot intensity and shot outcome for basketball—is new for this literature, and it's competently executed. The point process intensity is built from NMF basis shot types, the mark model includes the intensity as a covariate, and the MCMC estimation is validated with simulations showing small bias and coverage near 0.95. The four-player analyses are readable and the expected-score maps are a nice touch.\n\nThe main problem is the 40-of-50 majority claim. The paper says 40 of the top 50 shooters favored the intensity-dependent model by DIC and LPML, and all 40 estimated intensity coefficients were positive. But it never reports the magnitudes of the DIC/LPML differences or the credible intervals for those coefficients. The paper itself sets thresholds: DIC differences above 10 are substantial, 2–3 are weak; LPML differences above 4.5 are 'very strong'. The four-player table shows Durant called a 'clear advantage' with ΔDIC=8.6 and ΔLPML=4.2, both below those benchmarks. If a similar share of the 40 have weak differences, the 80% headline is overstated. Positive posterior means alone don't establish association; you need to know how many of the 40 have HPD intervals that exclude zero.\n\nMinor issues: the reported results come from a second run after dropping insignificant covariates, which is post-hoc and makes the intervals a bit optimistic. The shot clock omission is acknowledged and is a genuine limitation, but not a fatal one.\n\nThe paper deserves a serious referee. The modeling and simulation are solid; the empirical claim needs more transparent reporting. I'd recommend sending it to review and asking for per-player summaries of model comparison and coefficient intervals.","headline":"Solid Bayesian point process application, but the 40-of-50 claim needs per-player DIC/LPML magnitudes and coefficient intervals before I'd trust it.","tokens_in":16104,"tokens_out":2755,"would_cite":true,"duration_ms":25856,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62M30","62F15"],"pacs":[],"model":"deepseek-v4-flash","headline":"For most top NBA shooters, shot accuracy rises where shot intensity is higher.","keywords":["Bayesian marked point process","basketball shot chart","intensity-dependent mark model","non-homogeneous Poisson process","field goal percentage","NBA shot data","shot selection","Bayesian model comparison"],"falsifier":"Re-fit the mark model with shot-clock time remaining included as a covariate for the same top-50 players; if the posterior mass of $\\xi$ moves toward zero for a substantial share of the 40 players, or if DIC and LPML no longer favor $\\xi\\neq 0$ in most cases, the claimed positive intensity-accuracy association is partly an artifact of shot-clock confounding.","tokens_in":15085,"feed_emoji":"🏀","tokens_out":8052,"duration_ms":74142,"temperature":0.7,"pith_summary":"This paper sets out to show that for most high-volume NBA players, a player's shot accuracy is higher in the court locations where he attempts shots more often, and that this association can be estimated directly from shot-chart data. The authors propose a Bayesian marked spatial point process in which shot locations follow a non-homogeneous Poisson process and the binary make/miss mark is a logistic function of the fitted shot intensity plus covariates. In the 2017-2018 regular season's top 50 most frequent shooters, 40 players' data favored the intensity-dependent model by both DIC and LPML, and every one of those 40 estimated coefficients was positive. If the claim holds, the fitted coefficient gives a per-player measure of shot-selection efficiency that coaches and analysts could use to decide where a player should shoot more.","feed_headline":"40 of 50 top NBA shooters hit more where they shoot more","feed_subtitle":"A joint Bayesian model of shot locations and outcomes links higher shot frequency to higher accuracy for 80% of top shooters.","key_machinery":"The central object is an intensity-dependent marked spatial point process. Shot locations are modeled by a non-homogeneous Poisson process with intensity $\\lambda(s)$, and each binary make/miss mark is modeled by a logistic regression that includes $\\lambda(s)$ as a covariate, $\\logit(\\theta(s)) = \\xi\\lambda(s)+Z(s)^\\top\\alpha$. Ten intensity bases built by nonnegative matrix factorization from prior-season shot data act as spatial covariates for $\\lambda$, giving each location an interpretation as a shot type such as corner three or restricted-area two. MCMC draws from the joint posterior, with the Poisson integral approximated on a grid, and DIC and LPML compare the model with $\\xi$ free against the model with $\\xi=0$.","core_discovery":"The central claim is that shot intensity and shot accuracy are positively associated for a majority of elite NBA shooters, and that the association can be identified in a joint model rather than in two separate analyses. In the mark model $\\logit(\\theta(s_i)) = \\xi\\lambda(s_i) + Z(s_i)^\\top\\alpha$, the parameter $\\xi$ is the quantity of interest: it says whether a player converts more often at locations where the fitted attempt intensity $\\lambda$ is higher. For the four featured players, Durant, Harden, and James favored $\\xi\\neq 0$ by DIC and LPML, while Curry favored the intensity-independent model. Across the top 50 most frequent shooters, 40 favored $\\xi\\neq 0$ and all of those 40 had positive estimated $\\xi$; the interaction between intensity and two-versus-three-point shot type was not significant in any of the 50, and for the 10 players preferring $\\xi=0$, shot distance was significantly negative in every case. These results are the paper's central empirical discovery, supported by simulation studies showing that the MCMC estimator has near-nominal coverage.","pith_inferences":["The authors do not test whether shot-clock pressure drives the association; adding it as a covariate is a natural check that could shrink $\\xi$ for late-clock shooters.","A hierarchical version pooling all players with random effects would let $\\xi$ vary by position and usage, testing whether the positive coupling is concentrated in high-usage stars rather than all shooters.","The same joint model could be run on play-by-play data with defender distance and shot-clock time, turning a descriptive association into a causal shot-selection diagnostic."],"forward_implications":["For each of the 40 players with positive $\\xi$, the coefficient provides a ranking of how efficiently that player converts his existing shot distribution into points at the spots he frequents.","Players whose estimated $\\xi$ falls below the paper's elite average of about 1.02 are candidates for shot-selection improvement: shifting attempts toward their own high-intensity, high-accuracy zones should raise expected scoring.","The fitted intensity and mark surfaces combine into an expected-score map over the whole court, which remains meaningful even in locations where a player took few or no shots.","The absence of a significant intensity-by-shot-type interaction indicates the positive frequency-accuracy link does not differ between two-point and three-point attempts for these players."],"supporting_citations":[{"why":"Supplies the nonnegative-matrix-factorization shot-type bases used as spatial covariates for shot intensity.","marker":"Miller et al. (2014)"},{"why":"Defines intensity-dependent mark models, the modeling approach the paper adapts to basketball shot charts.","marker":"Ho and Stoyan (2008)"},{"why":"Provides the LPML approximation for point-process intensity used to compare models with and without $\\xi$.","marker":"Hu et al. (2019)"},{"why":"Supplies DIC and the rule that a difference above 10 is substantial, used in model comparison.","marker":"Spiegelhalter et al. (2002)"},{"why":"Gives the LPML difference thresholds used to declare strong evidence for $\\xi\\neq 0$.","marker":"Kass and Raftery (1995)"},{"why":"Provides matching-law evidence that basketball shot selection tracks reinforcement rates, motivating the positive-association hypothesis.","marker":"Alferink et al. (2009)"},{"why":"Earlier hierarchical spatial shot-chart analysis that this joint marked-process model extends.","marker":"Reich et al. (2006)"},{"why":"Gives the Poisson point process framework and likelihood used for shot locations.","marker":"Diggle (2013)"}],"fun_headline_variants":["40 of 50 top NBA shooters hit more where they shoot more","For 80% of elite NBA shooters, shot volume and accuracy align","NBA study: 40 of top 50 scorers shoot better at high-volume spots","Where top NBA shooters shoot most, they score most for 40 of 50"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that no unmeasured factor, especially shot-clock time remaining, inflates the estimated link between shot intensity and accuracy; the data do not include shot clock time, which the paper names as an important covariate missing from the mark model.","fun_headline_variants_meta":{"raw":{"variants":["40 of 50 top NBA shooters hit more where they shoot more","For 80% of elite NBA shooters, shot volume and accuracy align","NBA study: 40 of top 50 scorers shoot better at high-volume spots","Where top NBA shooters shoot most, they score most for 40 of 50"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000588,"raw_usage":{"total_tokens":2779,"prompt_tokens":982,"completion_tokens":1797,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":598,"completion_tokens_details":{"reasoning_tokens":1711}},"tokens_in":598,"tokens_out":1797,"duration_ms":14159,"temperature":1.0,"reasoning_tokens":1711,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:05:18.538158+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-fit the mark model with shot-clock time remaining included as a covariate for the same top-50 players; if the posterior mass of $\\xi$ moves toward zero for a substantial share of the 40 players, or if DIC and LPML no longer favor $\\xi\\neq 0$ in most cases, the claimed positive intensity-accuracy association is partly an artifact of shot-clock confounding.","supporting_citations":[{"cited_title":"Bornn, R","cited_arxiv_id":null,"evidence_quote":"Supplies the nonnegative-matrix-factorization shot-type bases used as spatial covariates for shot intensity."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies DIC and the rule that a difference above 10 is substantial, used in model comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the LPML difference thresholds used to declare strong evidence for $\\xi\\neq 0$."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides matching-law evidence that basketball shot selection tracks reinforcement rates, motivating the positive-association hypothesis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Earlier hierarchical spatial shot-chart analysis that this joint marked-process model extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the Poisson point process framework and likelihood used for shot locations."}],"review_version":1}