REVIEW 3 major objections 5 minor 6 references
Efficient Frontier Optimization of NBA Shot Selection Using Markov Reward Processes
T0 review · 3 major / 5 minor · reviewed 2026-07-12 · grok-4.5
Pith's one-line read NBA shot diets already sit near the efficient frontier; the three-point boom has edged slightly past its optimum once variance and possession continuation are counted.
desk verdict Solid fusion of possession MRP and portfolio theory that lands near the league frontier, but the “slightly past optimal threes” claim rests on fixed make rates and free knee/lambda choices. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
An absorbing Markov reward process on possession states whose fundamental matrix yields transient possession value for a candidate shot diet π; that value enters a normalized objective that subtracts portfolio variance and a ridge-estimated timeout-response cost, with the frontier selected by a dx/dy-tail knee rule and fastbreaks reattached as a fixed observed mix.
What would settle it
Rebuild the frontier after letting category make rates change with shot share—using tracking-based look quality or an empirical efficiency-vs-volume curve—and check whether the risk-adjusted optimum still recommends less three-point volume than the league actually takes; if the optimum moves back to or past current three-point share, the overshoot claim fails.
Extended reading notes
Core claim
Under a possession-level Markov reward process combined with a mean-variance shot-allocation objective and a timeout-response cost, the league-wide shot diet already operates near the efficient frontier. Marginal improvement comes from increasing rim attempts while slightly decreasing midrange and non-corner three volume; the three-point revolution has therefore pushed perimeter volume slightly past its optimal boundary, and expected-value-only optimization overstates the case for threes.
Load-bearing premise
Make rates for each shot type are held fixed when the model reallocates attempt shares, so taking more rim shots does not lower rim efficiency and taking fewer threes does not raise three-point efficiency.
Editorial extensions
If this is right
- League and most team optima shift toward more rim attempts and away from midrange and non-corner threes relative to observed diets.
- Expected-value-only shot charts systematically overstate the benefit of high-variance perimeter volume once continuation and covariance enter the objective.
- Corner threes stay roughly stable while non-corner threes are cut, so three-point quality and location still matter under risk penalties.
- No single universal shot chart is optimal; team-specific transition matrices still point in the same directional direction but not to identical mixes.
- Half-court optima plus observed fastbreak rim bias still produce a full-game recommendation with higher rim share than today’s league diet.
Reading between the lines
- If defensive congestion lowers rim make rates as rim volume rises, the model’s recommended rim increase would shrink or reverse—the directional claim is most fragile exactly where selection effects are strongest.
- Optimizing for win probability rather than points per possession could justify keeping higher three volume when trailing, which is the future mapping the authors themselves flag.
- That the response-cost minimum nearly matches the empirical league diet suggests coaches already behave as if near a stability equilibrium, helping explain why observed diets hug the frontier.
- The anomaly teams under shrinkage (extreme rim-only or efficiency outliers) show that personnel-driven transition structure can dominate any league-wide prescription.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper models NBA shot selection as a dynamic portfolio problem. It represents a possession as an absorbing Markov reward process over rim, midrange, corner-three, and non-corner-three outcomes, with free throws and shot-type-specific offensive rebounds. Candidate diets π reallocate shot-outcome mass while holding empirical make rates fixed; value is the transient possession value VT = e_PS^T N(π)r, with portfolio variance from game-level category returns and a timeout-response cost ||Bs(π)||_2 from ridge-estimated pressure dynamics. Fastbreaks are excluded from the main optimizer and reintroduced as a fixed observed mix. Under a normalized mean–variance–response objective and a dx/dy-tail knee rule, the league is near the efficient frontier, with recommended shifts toward more rim and fewer midrange/non-corner threes, implying that EV-only optimization overstates perimeter volume.
Significance. If the ranking is robust, the paper supplies a concrete bridge between possession-level Markov reward models and macro shot-diet portfolio analysis, with a falsifiable directional claim: modern three-point volume has moved slightly past a risk- and continuation-adjusted optimum. Strengths include a clearly specified MRP (Q, N, VT), empirically grounded miss-to-OREB ordering, separation of corner vs non-corner threes, hybrid treatment of unobserved fastbreak geometry, and team-level frontiers plus shrinkage sensitivity. The contribution is incremental rather than foundational, but it is a useful applied synthesis for sports analytics if the fixed-efficiency assumption is stress-tested.
major comments (3)
- [§4.1, §7; Tables 2–5] §4.1 reallocation rule and §7: the central claim that the league has pushed three-point volume slightly past the optimal boundary (Abstract; Tables 2–5; §5–6) is produced under fixed empirical make rates p_z when π changes. VT then inherits rim’s higher OR continuation (Fig. 5; Pr(OR_R|R_X)≈0.38–0.39 vs ~0.26 for threes) and the variance term penalizes threes, so the dx/dy-tail optimum tilts rim-ward. If raising rim share lowers p_R or cutting NC3 raises p_NC3, the ranking can reverse. The authors flag this but provide no sensitivity (e.g., elasticities, congestion bounds, or tracking-based efficiency response). Without that, the headline “slightly past optimal” conclusion is not secured by the estimated MRP alone.
- [§4.5; Table 2] §4.5: frontier selection depends on free parameters λ1, λ2, min–max normalization of each objective term, and the dx/dy-tail knee rule. Table 2 reports selected λ values near 5.7/0.7, but there is no systematic robustness of the recommended diet (or of the “marginal rim increase / perimeter decrease” direction) across a neighborhood of λ, alternative knee rules, or unnormalized objectives. Because the paper’s claim is that the league is near the frontier with only small directional gains, the selected point must be shown not to be an artifact of the particular knee rule and scaling.
- [§4.3–4.4; Table 1] §4.3–4.4: Σ and B are estimated from the same league environment being judged, and Table 1 shows the response-cost minimizer nearly coincides with the actual full-sample diet. That does not force the optimizer to recover the actual diet (the optimum does move), but it does mean the “near frontier” finding partly reflects self-referential penalties fit on observed behavior. At minimum, report out-of-sample or leave-one-season validation for B and Σ, and show that the directional recommendation survives when these objects are estimated on held-out seasons.
minor comments (5)
- [§4.5] Notation for the objective is inconsistent: §4.5 writes max eV(π)−λ1 fσ^{2}(π)−λ2 eE(π) with unclear e/f prefixes; define the normalized terms once and use them uniformly.
- [§5.2–5.3] Figure captions and in-text references occasionally disagree (e.g., team-diet comparisons in Figs. 7–12 vs discussion of Fig. 11 in §5.2). Number and cross-reference carefully.
- [§3] Report sample sizes more completely: seasons 2021–22 through 2025–26, number of possessions/events after filters, and whether playoffs are excluded (text says regular season).
- [§4.4; Appendix B] Clarify ridge penalty α for B and the validation grid that selected K=1000 in Appendix B; currently only the chosen values appear.
- [Introduction; Abstract] Minor prose issues: “mid range” vs “midrange,” and a few long sentences in the Introduction that could be tightened without changing content.
Circularity Check
Mild self-reference via the timeout-response term (B fit on league play; argmin E ≈ actual diet), which soft-pulls the “near frontier” claim; the rim-ward directional optimum is not forced by construction.
-
fitted input called prediction
[§4.4 Timeout Response Cost; §5.1 Response Cost; Table 1; objective in §4.5]
"In Table 1, it can be seen that the π that minimizes ∥Bs(π)∥₂ is approximately the actual league shooting distribution. This means that the response cost yields approximately the same minimum as a standard L₂ penalty, an expected result given its derivation from aggregated empirical data. ... The final objective is max_π ẽV(π)−λ₁fσ²(π)−λ₂ẽE(π) ... Our results indicate that the league aggregate operates near the efficient frontier"
B is estimated by ridge regression on the league’s own timeout-separated pressure intervals; s(π) is then built from candidate diets and E(π)=∥Bs(π)∥₂ enters the objective with a minus sign. Because the fitted B has mean-reverting diagonals on the same environment that generated the actual diet, argmin E is empirically ~actual (Table 1). Including −λ₂ẽE therefore soft-pulls every selected frontier portfolio toward observed shares, so part of the headline finding that the league “operates near the efficient frontier” is statistically encouraged by a penalty fit to that same league rather than independently predicted. The reduction is partial (V and σ² still move the optimum off actual), not total.
full rationale
The paper is an empirical portfolio optimization over a possession-level MRP, not a first-principles derivation that reduces to its inputs. Transition structure Q, make rates p_z, rewards r, game-level covariance Σ, and timeout matrix B are all estimated from play-by-play; candidate diets π reallocate shot-outcome mass while holding those empirical rates fixed (§4.1), then VT(π), σ²(π), and E(π) are scored and optimized. That is standard plug-in estimation, not self-definitional circularity: the no-fastbreak transient optimum (0.415/0.233/0.095/0.257) differs materially from the actual diet (0.313/0.277/0.105/0.305), so the optimizer is not forced to recover observed shares. There is no load-bearing self-citation chain (references are to Fichman, Skinner, Sandholtz, Cervone, Pelechrinis, etc.), no uniqueness theorem imported from the authors, and no ansatz smuggled in via prior own work. The one genuine soft circularity is the response-cost channel: B is ridge-fit on the same league’s timeout pressure intervals, and the paper itself reports that argmin_π ∥Bs(π)∥₂ is approximately the full-sample actual diet (Table 1), so the −λ₂ẽE term in the objective systematically rewards looking like current play and therefore contributes by construction to the claim that the league sits near the efficient frontier. Even that term is only partial: the directional move toward more rim and fewer non-corner threes is driven by independent MRP continuation (higher OR after rim misses) and empirical three-point variance, not by E. Fixed make rates under reallocation are a modeling assumption (flagged in §7), not a circular reduction. Overall circularity is mild and localized; score 3.
Assumptions & free parameters
free parameters (6)
- λ1 (variance weight)
- λ2 (response-cost weight)
- ridge penalty α in timeout response regression
- dx/dy-tail knee selection rule
- shrinkage pseudo-count K=1000 (sensitivity)
- objective min-max normalization bounds
assumptions (7)
- domain assumption Category make rates pz stay constant when shot-share vector π is reallocated.
- domain assumption A possession is adequately modeled as an absorbing Markov reward process on the defined discrete state space.
- domain assumption Game-level category return covariance Σ is the right risk measure for in-game/playoff shot-diet risk.
- ad hoc to paper Timeout-interval pressure changes estimated by ridge matrix B proxy defensive/coaching response cost of a diet.
- ad hoc to paper Fastbreak geometry is unobserved, so optimal full-game diet is hybrid: optimize no-fastbreak π then mix with observed fastbreak ρFB at empirical weight α.
- domain assumption Mean-variance portfolio math on simplex π applies to mutually exclusive shot-location shares within possessions.
- standard math Standard linear algebra of fundamental matrix N=(I−Q)−1 gives expected cumulative reward before absorption.
invented entities (3)
-
Timeout-response stability cost E(π)=||B s(π)||2
-
dx/dy-tail knee rule for frontier portfolio selection
-
Hybrid no-fastbreak optimum + fixed observed fastbreak shot distribution
Cite this review
Pith. "Pith review of Efficient Frontier Optimization of NBA Shot Selection Using Markov Reward Processes." pith.science (2026). https://pith.science/paper/ON2B33Z5
@misc{pith2026260702933,
author = {Pith},
title = {Pith review of: Efficient Frontier Optimization of NBA Shot Selection Using Markov Reward Processes},
year = {2026},
howpublished = {\url{https://pith.science/paper/ON2B33Z5}},
note = {Machine review of arXiv:2607.02933}
}
read the original abstract
This paper asks whether modern NBA shot selection can be evaluated as a dynamic portfolio problem rather than as a collection of isolated shot attempts. We combine a possession-level Markov reward process with both a mean-variance shot allocation objective and a defense response cost. The model separates rim attempts, midrange attempts, corner threes, and noncorner threes, then values each candidate shot diet by immediate scoring, continuation after offensive rebounds, empirical game-level variance, and a timeout-response stability cost. Because fastbreak value depends on geometry that is not fully observed in standard play-by-play, we optimize the no-fastbreak sample first and then add back observed fastbreak shot selection as a fixed transition component. Our results indicate that the league aggregate operates near the efficient frontier, with marginal improvements to be made by slightly decreasing perimeter volume and mid range attempts while increasing rim attempts. Team-level analyses reflect this general trend. Ultimately, these findings suggest the three-point revolution has pushed volume slightly past its optimal boundary. By incorporating outcome variance and continuation effects, we show that optimizing solely for expected value mathematically overstates the benefit of perimeter shots.
Figures
Figures from the paper (17 more)
Reference graph
Works this paper leans on
-
[1]
Fichman, M. and O’Brien, J. R. (2019). Optimal shot selection strategies for the NBA.Journal of Quantitative Analysis in Sports, 15(3), 203–211. doi:10.1515/jqas-2017-0113
-
[2]
Fichman, M. and O’Brien, J. (2018). Three point shooting and efficient mixed strategies: A 19 portfolio management approach.Journal of Sports Analytics, 4, 107–120. doi:10.3233/JSA- 160154
-
[3]
Skinner, B. (2011). The problem of shot selection in basketball.Journal of Quantitative Analysis in Sports
2011
-
[4]
Cervone, D., D’Amour, A., Bornn, L., and Goldsberry, K. (2016). A multiresolution stochastic process model for predicting basketball possession outcomes.Journal of the American Statistical Association
2016
-
[5]
and Bornn, L
Sandholtz, N. and Bornn, L. (2020). Markov decision processes with dynamic transition prob- abilities: an analysis of shooting strategies in basketball.Annals of Applied Statistics
2020
-
[6]
Pelechrinis, K. and Goldsberry, K. (2021). The anatomy of corner 3s in the NBA: What makes them efficient, how are they generated and how can defenses respond?arXiv:2105.12785. 20
arXiv 2021
Reviewed July 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.