Pith. sign in

REVIEW 3 major objections 5 minor 6 references

Efficient Frontier Optimization of NBA Shot Selection Using Markov Reward Processes

T0 review · 3 major / 5 minor · reviewed 2026-07-12 · grok-4.5

Pith's one-line read NBA shot diets already sit near the efficient frontier; the three-point boom has edged slightly past its optimum once variance and possession continuation are counted.

desk verdict Solid fusion of possession MRP and portfolio theory that lands near the league frontier, but the “slightly past optimal threes” claim rests on fixed make rates and free knee/lambda choices. read the letter →

arxiv 2607.02933 v1 pith:ON2B33Z5 submitted 2026-07-03 stat.AP

classification stat.AP
keywords NBAshotselectionefficientfrontierMarkovrewardprocessmean-varianceportfoliothree-pointshootingpossessionvaluetimeoutresponseoffensiverebounds
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether modern NBA shot selection should be judged as a dynamic portfolio problem, not as a string of isolated expected-value shots. It models each possession as a Markov reward process over rim, midrange, corner-three, and non-corner-three outcomes, then scores candidate shot diets by immediate points, value after offensive rebounds, game-level return variance, and how defenses shift after timeouts. Fastbreaks are optimized only as a fixed observed mix because play-by-play data miss the geometry that makes them valuable. The league aggregate lands near the efficient frontier, with only marginal gains from more rim attempts and slightly fewer midrange and non-corner threes. The practical stake is clear: pure expected-value arguments that powered the three-point revolution mathematically overstate the benefit of perimeter volume once volatility and continuation are included.

What carries the argument

An absorbing Markov reward process on possession states whose fundamental matrix yields transient possession value for a candidate shot diet π; that value enters a normalized objective that subtracts portfolio variance and a ridge-estimated timeout-response cost, with the frontier selected by a dx/dy-tail knee rule and fastbreaks reattached as a fixed observed mix.

What would settle it

Rebuild the frontier after letting category make rates change with shot share—using tracking-based look quality or an empirical efficiency-vs-volume curve—and check whether the risk-adjusted optimum still recommends less three-point volume than the league actually takes; if the optimum moves back to or past current three-point share, the overshoot claim fails.

Watch

Extended reading notes

Core claim

Under a possession-level Markov reward process combined with a mean-variance shot-allocation objective and a timeout-response cost, the league-wide shot diet already operates near the efficient frontier. Marginal improvement comes from increasing rim attempts while slightly decreasing midrange and non-corner three volume; the three-point revolution has therefore pushed perimeter volume slightly past its optimal boundary, and expected-value-only optimization overstates the case for threes.

Load-bearing premise

Make rates for each shot type are held fixed when the model reallocates attempt shares, so taking more rim shots does not lower rim efficiency and taking fewer threes does not raise three-point efficiency.

Editorial extensions

If this is right

  • League and most team optima shift toward more rim attempts and away from midrange and non-corner threes relative to observed diets.
  • Expected-value-only shot charts systematically overstate the benefit of high-variance perimeter volume once continuation and covariance enter the objective.
  • Corner threes stay roughly stable while non-corner threes are cut, so three-point quality and location still matter under risk penalties.
  • No single universal shot chart is optimal; team-specific transition matrices still point in the same directional direction but not to identical mixes.
  • Half-court optima plus observed fastbreak rim bias still produce a full-game recommendation with higher rim share than today’s league diet.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If defensive congestion lowers rim make rates as rim volume rises, the model’s recommended rim increase would shrink or reverse—the directional claim is most fragile exactly where selection effects are strongest.
  • Optimizing for win probability rather than points per possession could justify keeping higher three volume when trailing, which is the future mapping the authors themselves flag.
  • That the response-cost minimum nearly matches the empirical league diet suggests coaches already behave as if near a stability equilibrium, helping explain why observed diets hug the frontier.
  • The anomaly teams under shrinkage (extreme rim-only or efficiency outliers) show that personnel-driven transition structure can dominate any league-wide prescription.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper models NBA shot selection as a dynamic portfolio problem. It represents a possession as an absorbing Markov reward process over rim, midrange, corner-three, and non-corner-three outcomes, with free throws and shot-type-specific offensive rebounds. Candidate diets π reallocate shot-outcome mass while holding empirical make rates fixed; value is the transient possession value VT = e_PS^T N(π)r, with portfolio variance from game-level category returns and a timeout-response cost ||Bs(π)||_2 from ridge-estimated pressure dynamics. Fastbreaks are excluded from the main optimizer and reintroduced as a fixed observed mix. Under a normalized mean–variance–response objective and a dx/dy-tail knee rule, the league is near the efficient frontier, with recommended shifts toward more rim and fewer midrange/non-corner threes, implying that EV-only optimization overstates perimeter volume.

Significance. If the ranking is robust, the paper supplies a concrete bridge between possession-level Markov reward models and macro shot-diet portfolio analysis, with a falsifiable directional claim: modern three-point volume has moved slightly past a risk- and continuation-adjusted optimum. Strengths include a clearly specified MRP (Q, N, VT), empirically grounded miss-to-OREB ordering, separation of corner vs non-corner threes, hybrid treatment of unobserved fastbreak geometry, and team-level frontiers plus shrinkage sensitivity. The contribution is incremental rather than foundational, but it is a useful applied synthesis for sports analytics if the fixed-efficiency assumption is stress-tested.

major comments (3)
  1. [§4.1, §7; Tables 2–5] §4.1 reallocation rule and §7: the central claim that the league has pushed three-point volume slightly past the optimal boundary (Abstract; Tables 2–5; §5–6) is produced under fixed empirical make rates p_z when π changes. VT then inherits rim’s higher OR continuation (Fig. 5; Pr(OR_R|R_X)≈0.38–0.39 vs ~0.26 for threes) and the variance term penalizes threes, so the dx/dy-tail optimum tilts rim-ward. If raising rim share lowers p_R or cutting NC3 raises p_NC3, the ranking can reverse. The authors flag this but provide no sensitivity (e.g., elasticities, congestion bounds, or tracking-based efficiency response). Without that, the headline “slightly past optimal” conclusion is not secured by the estimated MRP alone.
  2. [§4.5; Table 2] §4.5: frontier selection depends on free parameters λ1, λ2, min–max normalization of each objective term, and the dx/dy-tail knee rule. Table 2 reports selected λ values near 5.7/0.7, but there is no systematic robustness of the recommended diet (or of the “marginal rim increase / perimeter decrease” direction) across a neighborhood of λ, alternative knee rules, or unnormalized objectives. Because the paper’s claim is that the league is near the frontier with only small directional gains, the selected point must be shown not to be an artifact of the particular knee rule and scaling.
  3. [§4.3–4.4; Table 1] §4.3–4.4: Σ and B are estimated from the same league environment being judged, and Table 1 shows the response-cost minimizer nearly coincides with the actual full-sample diet. That does not force the optimizer to recover the actual diet (the optimum does move), but it does mean the “near frontier” finding partly reflects self-referential penalties fit on observed behavior. At minimum, report out-of-sample or leave-one-season validation for B and Σ, and show that the directional recommendation survives when these objects are estimated on held-out seasons.
minor comments (5)
  1. [§4.5] Notation for the objective is inconsistent: §4.5 writes max eV(π)−λ1 fσ^{2}(π)−λ2 eE(π) with unclear e/f prefixes; define the normalized terms once and use them uniformly.
  2. [§5.2–5.3] Figure captions and in-text references occasionally disagree (e.g., team-diet comparisons in Figs. 7–12 vs discussion of Fig. 11 in §5.2). Number and cross-reference carefully.
  3. [§3] Report sample sizes more completely: seasons 2021–22 through 2025–26, number of possessions/events after filters, and whether playoffs are excluded (text says regular season).
  4. [§4.4; Appendix B] Clarify ridge penalty α for B and the validation grid that selected K=1000 in Appendix B; currently only the chosen values appear.
  5. [Introduction; Abstract] Minor prose issues: “mid range” vs “midrange,” and a few long sentences in the Introduction that could be tightened without changing content.

Circularity Check

1 steps flagged · score 3.0 of 10

Mild self-reference via the timeout-response term (B fit on league play; argmin E ≈ actual diet), which soft-pulls the “near frontier” claim; the rim-ward directional optimum is not forced by construction.

  1. fitted input called prediction [§4.4 Timeout Response Cost; §5.1 Response Cost; Table 1; objective in §4.5]
    "In Table 1, it can be seen that the π that minimizes ∥Bs(π)∥₂ is approximately the actual league shooting distribution. This means that the response cost yields approximately the same minimum as a standard L₂ penalty, an expected result given its derivation from aggregated empirical data. ... The final objective is max_π ẽV(π)−λ₁fσ²(π)−λ₂ẽE(π) ... Our results indicate that the league aggregate operates near the efficient frontier"

    B is estimated by ridge regression on the league’s own timeout-separated pressure intervals; s(π) is then built from candidate diets and E(π)=∥Bs(π)∥₂ enters the objective with a minus sign. Because the fitted B has mean-reverting diagonals on the same environment that generated the actual diet, argmin E is empirically ~actual (Table 1). Including −λ₂ẽE therefore soft-pulls every selected frontier portfolio toward observed shares, so part of the headline finding that the league “operates near the efficient frontier” is statistically encouraged by a penalty fit to that same league rather than independently predicted. The reduction is partial (V and σ² still move the optimum off actual), not total.

full rationale

The paper is an empirical portfolio optimization over a possession-level MRP, not a first-principles derivation that reduces to its inputs. Transition structure Q, make rates p_z, rewards r, game-level covariance Σ, and timeout matrix B are all estimated from play-by-play; candidate diets π reallocate shot-outcome mass while holding those empirical rates fixed (§4.1), then VT(π), σ²(π), and E(π) are scored and optimized. That is standard plug-in estimation, not self-definitional circularity: the no-fastbreak transient optimum (0.415/0.233/0.095/0.257) differs materially from the actual diet (0.313/0.277/0.105/0.305), so the optimizer is not forced to recover observed shares. There is no load-bearing self-citation chain (references are to Fichman, Skinner, Sandholtz, Cervone, Pelechrinis, etc.), no uniqueness theorem imported from the authors, and no ansatz smuggled in via prior own work. The one genuine soft circularity is the response-cost channel: B is ridge-fit on the same league’s timeout pressure intervals, and the paper itself reports that argmin_π ∥Bs(π)∥₂ is approximately the full-sample actual diet (Table 1), so the −λ₂ẽE term in the objective systematically rewards looking like current play and therefore contributes by construction to the claim that the league sits near the efficient frontier. Even that term is only partial: the directional move toward more rim and fewer non-corner threes is driven by independent MRP continuation (higher OR after rim misses) and empirical three-point variance, not by E. Fixed make rates under reallocation are a modeling assumption (flagged in §7), not a circular reduction. Overall circularity is mild and localized; score 3.

Assumptions & free parameters 6 free parameters · 7 assumptions · 3 invented entities

The central claim rests on empirical transition/make structure plus several modeling choices that are not derived from first principles: fixed category efficiencies under reallocation, portfolio weights λ1/λ2 and a knee selection rule, a constructed timeout-response cost, and treating unobserved fastbreak geometry by freezing observed transition shot mix. Standard absorbing-Markov and mean-variance math is borrowed; the basketball-specific axioms and free parameters do the work that turns “EV of threes is high” into “slightly too many threes.”

free parameters (6)
  • λ1 (variance weight)
    Chosen via grid search over the normalized objective; reported near 5.7–5.8 for selected league optima (Table 2). Directly controls how far the solution moves from pure value toward lower variance.
  • λ2 (response-cost weight)
    Grid-searched; reported 0.7 at selected optima (Table 2). Scales the timeout-response penalty relative to value and variance.
  • ridge penalty α in timeout response regression
    B is fit by ridge regression Δs = Bs + ε; α is not reported as a data-independent constant and affects the response cost surface.
  • dx/dy-tail knee selection rule
    Authors choose the frontier point minimizing local |dx/dy| on the upper envelope rather than a utility-derived risk aversion; this discrete selection rule defines which diet is called “the” optimum.
  • shrinkage pseudo-count K=1000 (sensitivity)
    Used in appendix team-matrix smoothing; selected as largest candidate with lowest validation MSE. Changes some team optima substantially, showing free smoothing strength.
  • objective min-max normalization bounds
    Value, variance, and response cost are each rescaled by support min/max before weighting; those bounds depend on the candidate set and can alter trade-offs.
assumptions (7)
  • domain assumption Category make rates pz stay constant when shot-share vector π is reallocated.
    §4.1 transition construction and §7 Limitations; without this, VT(π) rankings can change under congestion/selection.
  • domain assumption A possession is adequately modeled as an absorbing Markov reward process on the defined discrete state space.
    §3–4; ignores continuous geometry, lineup identity, and non-Markov coaching state except via timeout intervals.
  • domain assumption Game-level category return covariance Σ is the right risk measure for in-game/playoff shot-diet risk.
    §4.3; risk is empirical across games, not derived from possession-level outcome distributions or win probability.
  • ad hoc to paper Timeout-interval pressure changes estimated by ridge matrix B proxy defensive/coaching response cost of a diet.
    §4.4 invents s(π) and E(π)=||Bs(π)||2 as stability penalty; not a standard basketball identity.
  • ad hoc to paper Fastbreak geometry is unobserved, so optimal full-game diet is hybrid: optimize no-fastbreak π then mix with observed fastbreak ρFB at empirical weight α.
    §4.7; freezes transition shot selection rather than optimizing it.
  • domain assumption Mean-variance portfolio math on simplex π applies to mutually exclusive shot-location shares within possessions.
    Objective in §4.5; inherits MPT framing from Fichman & O’Brien but with MRP returns.
  • standard math Standard linear algebra of fundamental matrix N=(I−Q)−1 gives expected cumulative reward before absorption.
    §4.1–4.2; classical absorbing Markov reward theory.
invented entities (3)
  • Timeout-response stability cost E(π)=||B s(π)||2
    purpose: Penalize shot diets that empirically induce large post-timeout pressure-vector moves, as a stand-in for defensive adaptation.
    Constructed from interval pressure features and ridge-fit B; independent evidence outside this paper is limited to the authors’ interpretation of mean reversion after timeouts.
  • dx/dy-tail knee rule for frontier portfolio selection
    purpose: Pick a single “selected” efficient diet from the λ-grid envelope as the risk-controlled optimum.
    A paper-specific selection heuristic; different knee rules would yield different recommended π.
  • Hybrid no-fastbreak optimum + fixed observed fastbreak shot distribution
    purpose: Produce a full-game recommendation without modeling unobserved transition geometry.
    Engineering composition of two samples; not an independently validated physical object.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Efficient Frontier Optimization of NBA Shot Selection Using Markov Reward Processes." pith.science (2026). https://pith.science/paper/ON2B33Z5

@misc{pith2026260702933,
  author       = {Pith},
  title        = {Pith review of: Efficient Frontier Optimization of NBA Shot Selection Using Markov Reward Processes},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ON2B33Z5}},
  note         = {Machine review of arXiv:2607.02933}
}
read the original abstract

This paper asks whether modern NBA shot selection can be evaluated as a dynamic portfolio problem rather than as a collection of isolated shot attempts. We combine a possession-level Markov reward process with both a mean-variance shot allocation objective and a defense response cost. The model separates rim attempts, midrange attempts, corner threes, and noncorner threes, then values each candidate shot diet by immediate scoring, continuation after offensive rebounds, empirical game-level variance, and a timeout-response stability cost. Because fastbreak value depends on geometry that is not fully observed in standard play-by-play, we optimize the no-fastbreak sample first and then add back observed fastbreak shot selection as a fixed transition component. Our results indicate that the league aggregate operates near the efficient frontier, with marginal improvements to be made by slightly decreasing perimeter volume and mid range attempts while increasing rim attempts. Team-level analyses reflect this general trend. Ultimately, these findings suggest the three-point revolution has pushed volume slightly past its optimal boundary. By incorporating outcome variance and continuation effects, we show that optimizing solely for expected value mathematically overstates the benefit of perimeter shots.

Figures

Figures reproduced from arXiv: 2607.02933 by the authors.

Figure 1
Figure 1. League shot-diet share by season. The three-point line shows the macro shift toward [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Shot probability and shot value move differently by category. Rim attempts have the [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. No-fastbreak empirical transition matrices. The left panel is the full transition matrix [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (17 more)
Figure 4
Figure 4. Figure 4: Full-sample empirical transition matrices. Adding fastbreaks increases the possession [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Offensive-rebound continuation probability conditional on a missed shot. Rim misses are [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: Estimated timeout response matrix B for the six pressure components. Negative diagonal entries indicate mean reversion within the same pressure channel, while off-diagonal entries show how one pressure source is associated with movement in another [PITH_FULL_IMAGE:fig…
Figure 7
Figure 7. Figure 7: No-fastbreak team diets compared with the no-fastbreak league actual diet. [PITH_FULL_IMAGE:figures/full_fig_p011_7.png]
Figure 8
Figure 8. Figure 8: No-fastbreak team diets compared with the no-fastbreak league transient frontier. [PITH_FULL_IMAGE:figures/full_fig_p012_8.png]
Figure 9
Figure 9. Figure 9: Full actual team diets compared with the hybrid league benchmark. The right-hand bar [PITH_FULL_IMAGE:figures/full_fig_p012_9.png]
Figure 10
Figure 10. Figure 10: Actual teams plotted on the full-sample league transient frontier. The dx/dy-tail point [PITH_FULL_IMAGE:figures/full_fig_p013_10.png]
Figure 11
Figure 11. Figure 11: No-fastbreak team shot diets compared with team-specific transient frontier portfolios. [PITH_FULL_IMAGE:figures/full_fig_p014_11.png]
Figure 12
Figure 12. Figure 12: Full actual team diets compared with team-specific hybrid benchmarks. The right-hand [PITH_FULL_IMAGE:figures/full_fig_p014_12.png]
Figure 13
Figure 13. Figure 13: Actual teams plotted on the full-sample league one-step frontier. The one-step model [PITH_FULL_IMAGE:figures/full_fig_p016_13.png]
Figure 14
Figure 14. Figure 14: Full-sample team shot diets compared with the league transient frontier portfolio selected [PITH_FULL_IMAGE:figures/full_fig_p016_14.png]
Figure 15
Figure 15. Figure 15: Full-sample team shot diets compared with team-specific transient frontier portfolios. [PITH_FULL_IMAGE:figures/full_fig_p017_15.png]
Figure 16
Figure 16. Figure 16: Full actual team diets compared with hybrid recompositions using the no-fastbreak [PITH_FULL_IMAGE:figures/full_fig_p018_16.png]
Figure 17
Figure 17. Figure 17: Hybrid team-frontier diets compared directly under [PITH_FULL_IMAGE:figures/full_fig_p018_17.png]
Figure 18
Figure 18. Figure 18: No-fastbreak team actual diets compared with empirical team-frontier diets and [PITH_FULL_IMAGE:figures/full_fig_p018_18.png]
Figure 19
Figure 19. Figure 19: No-fastbreak empirical team-frontier diets compared directly with [PITH_FULL_IMAGE:figures/full_fig_p019_19.png]
Figure 20
Figure 20. Figure 20: No-fastbreak frontier value gaps under empirical and [PITH_FULL_IMAGE:figures/full_fig_p019_20.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

6 extracted references · 2 canonical work pages

  1. [1]

    and O’Brien, J

    Fichman, M. and O’Brien, J. R. (2019). Optimal shot selection strategies for the NBA.Journal of Quantitative Analysis in Sports, 15(3), 203–211. doi:10.1515/jqas-2017-0113

  2. [2]

    and O’Brien, J

    Fichman, M. and O’Brien, J. (2018). Three point shooting and efficient mixed strategies: A 19 portfolio management approach.Journal of Sports Analytics, 4, 107–120. doi:10.3233/JSA- 160154

  3. [3]

    Skinner, B. (2011). The problem of shot selection in basketball.Journal of Quantitative Analysis in Sports

  4. [4]

    Cervone, D., D’Amour, A., Bornn, L., and Goldsberry, K. (2016). A multiresolution stochastic process model for predicting basketball possession outcomes.Journal of the American Statistical Association

  5. [5]

    and Bornn, L

    Sandholtz, N. and Bornn, L. (2020). Markov decision processes with dynamic transition prob- abilities: an analysis of shooting strategies in basketball.Annals of Applied Statistics

  6. [6]

    and Goldsberry, K

    Pelechrinis, K. and Goldsberry, K. (2021). The anatomy of corner 3s in the NBA: What makes them efficient, how are they generated and how can defenses respond?arXiv:2105.12785. 20

Pith tools

Reviewed July 12, 2026 · model on record in the stance chip above.