{"id":"89785c69-da30-4603-90b4-36089d87fc6d","arxiv_id":"2608.09824","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A Bayesian graphical framework with static, autoregressive, and hidden-Markov variants is applied to 76ers player-game data; the autoregressive variant wins on WAIC and supports player-level predictive inference.","lead":"Three Bayesian network models that describe how an NBA team's scoring depends on minutes, fouls, and shot choices are built and tested on the Philadelphia 76ers' 2005-06 season. The autoregressive variant fits best, and the framework yields per-player predictions of points and minutes for future games.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The dynamic LBN's WAIC advantage is not validated out of sample; WAIC is a full-data approximation and the 80-point margin is reported without uncertainty, so the central preference claim is not yet established.","rationale":"The strongest claim is specifically the Table 2 ranking. I looked for an internal inconsistency or a plausible alternative under the same framework that would overturn it. The most load-bearing point is not the fixed DAG, which the authors explicitly disclose, but that the evidence for the ranking is a single full-data WAIC number with no assessment of how stable the 80-point gap is. For hierarchical and latent-variable models, WAIC can depend on implementation choices, and the paper does not compare with a direct holdout. The proposed temporal split directly targets this concern: it compares the three models on data not used for fitting and on the actual outcomes emphasized in the paper. If the dynamic LBN retains the best out-of-sample log predictive density and RMSE, the concern is resolved; if not, the central claim should be downgraded to an in-sample finding. This does not question the model mechanics; it asks for the specific evidence that the headline comparison requires. The reader's weakest_assumption (fixed structure) is related but different: my concern would remain even if the DAG were learned, because the model-selection evidence is still not stress-tested out of sample.","tokens_in":16211,"tokens_out":9301,"duration_ms":94262,"concrete_test":"Split the 82-game season temporally: fit all three LBNs on games 1-60 with the same priors and MCMC settings, and score held-out games 61-82 by per-player-game log predictive density and RMSE for total points and minutes. Repeat with 5-fold blocks of games to avoid dependence on the particular split. If the dynamic LBN does not have the best average out-of-sample log predictive density and lower RMSE in most folds, the Table 2 ranking should be reported as an in-sample result rather than a validated predictive preference.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that the dynamic LBN is preferred because its WAIC (21194.13) is about 80 units lower than the static and hidden-Markov alternatives (Table 2, Section 5.1). WAIC is a full-data approximation to out-of-sample predictive performance, and for hierarchical or latent-variable models its value can depend on how random effects and latent states are handled in the pointwise predictive density. The paper reports no WAIC standard error, no repeated computation, and no holdout validation. The 80-unit gap over roughly 82 games times 13 players implies only a 7-8% average improvement in per-player-game geometric predictive density, so it is not self-evidently decisive. Moreover, after selecting the dynamic model, all subsequent predictive statements (Figures 4-6) are in-sample, based on the same data used to choose the model. If the WAIC gap is sensitive to estimation noise or to the treatment of random effects in the predictive density, the ranking could change.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a Bayesian graphical modeling framework for longitudinal analysis of NBA player-game data. Three models are specified: a static longitudinal Bayesian network (LBN), a dynamic LBN with a first-order autoregressive structure on minutes played, and a hidden Markov LBN with a team-level hot/cold latent state affecting shot success. The models are fitted to the Philadelphia 76ers' 2005-06 season with NIMBLE, and model selection via WAIC selects the dynamic LBN. The paper then reports posterior summaries for shooting success, shot attempts, minutes, and participation, and illustrates conditional posterior predictive distributions for points and minutes for two players. The GitHub repository provides code for reproduction.","tokens_in":16501,"tokens_out":4372,"duration_ms":44237,"significance":"If the model-selection and predictive claims hold, the paper offers a useful applied template for multivariate longitudinal sports data, with explicit likelihoods and priors that are easy to adapt. The main strengths are the full and internally consistent specification of the three models, the reproducible code, and the honest statement of limitations, including the fixed (not learned) network structure. The statistical novelty is modest: the components—Bayesian networks, Poisson regression, random effects, and hidden Markov states—are established, but the longitudinal graphical synthesis for basketball performance is a reasonable applied contribution. The results about shot-success probabilities and the difference between Iverson and Korver are plausible, but the paper's key model-selection conclusion is not yet rigorously supported.","major_comments":[{"comment":"The central claim that the dynamic LBN is preferred rests entirely on a single set of WAIC values with no measure of uncertainty. The dynamic model value is 21194.13 versus 21274.69 for the static model and 21274.60 for the hidden Markov model; the difference of about 80 units, when spread over roughly 82 games times 13 players, corresponds to only about 0.07-0.08 units in average per-observation log predictive density. WAIC is a full-data approximation to out-of-sample predictive performance, and for hierarchical or latent-variable models the computation of the pointwise predictive density can affect the value. The paper should report WAIC standard errors or comparisons under alternative pointwise predictive density treatments, and ideally support the ranking with k-fold or temporal cross-validation, before conditioning the entire remaining analysis on the dynamic model. The near-identical values for the static and hidden Markov models (difference 0.09) also suggest the latent state adds essentially nothing, which should be discussed.","section":"Section 5.1, Table 2"},{"comment":"The predictive analysis is entirely in-sample. Equation (16) defines the posterior predictive distribution for a new hypothetical match conditional on the observed data, and Figures 5 and 6 show conditional predictive distributions computed from the same 82-game season used to estimate and select the model. The Introduction states the objective of predicting points in the next match, but no holdout evaluation is provided. Please evaluate the predictive performance on held-out games (for example, a last-k-games temporal holdout or leave-one-game-out) and report calibration and accuracy measures such as interval coverage, RMSE, or ranked probability scores for points and minutes. Without this, the phrase 'predictive' is used in a purely conditional-reconstruction sense rather than a forecasting sense.","section":"Section 5.3, Eq. (16), Figures 5-6"},{"comment":"The model comparison confounds two separate modeling choices. The dynamic LBN adds autoregressive minutes but has no latent state, while the hidden Markov LBN adds a team-level hot/cold state but keeps the static minutes structure. In Section 4.3 the authors state that minutes, fouls, and attempts follow the same relationships as in the static network. Thus the better WAIC of the dynamic LBN may be due entirely to the autoregressive minutes component, and it does not establish that dynamic time dependence in general is preferred. A model that includes both the autoregressive minutes structure and the hidden state should be fitted and compared before drawing conclusions about the relative merits of the two dynamic mechanisms.","section":"Section 4.2 vs Section 4.3, Table 2"},{"comment":"No convergence diagnostics or effective sample sizes are reported. The paper states that three chains of 1,000,000 iterations were run with a 500,000 burn-in and thinning of 500, leaving 1,000 retained iterations per chain and 3,000 total. All posterior summaries, WAIC values, and predictive distributions depend on these samples. Please report R-hat statistics and effective sample sizes for the main parameters, or trace plots for representative parameters, to demonstrate that the posterior approximations are reliable despite the rather aggressive thinning.","section":"Section 5, MCMC details"}],"minor_comments":[{"comment":"The phrase 'whether or nor the player participates' should be 'whether or not the player participates'.","section":"Section 4.1"},{"comment":"The package is referred to as 'NIMBLER'; the correct name is NIMBLE (de Valpine et al., 2017). The spacing in 'W AIC' is also inconsistent with the standard notation 'WAIC'.","section":"Section 5"},{"comment":"The row label β_F^(Tk) in Table 5 is misleading for k=2,3, where the coefficient is associated with minutes played, not fouls. The text in the paragraph after Table 5 also refers to β_F^(T2) and β_F^(T3); these should be β_M^(T2) and β_M^(T3).","section":"Section 5.2, Table 5"},{"comment":"The transition probabilities q_CH and q_HC and the posterior distribution of the latent states are not reported. Since the hidden Markov model is one of the three main proposals, at least a summary of the inferred transition matrix or the posterior probability of being in the hot state across games would help assess the model.","section":"Section 4.3"},{"comment":"There is a typo: 'succesive' should be 'successive'. Also, the player name 'Dalembert' is spelled 'Dalmebert' in the discussion of Table 7.","section":"Section 6"},{"comment":"The data source is described as NBAstuffer accessed in 2022, but the data are not archived. For reproducibility, consider depositing the cleaned player-game dataset in a permanent repository alongside the code.","section":"Section 2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable applied contribution, but the central model-selection claim lacks uncertainty quantification or out-of-sample validation, and the predictive section is in-sample. The comparison between the dynamic and hidden Markov models is also confounded because they change different components of the model. These issues are fixable within the manuscript's scope, so I recommend major revision rather than rejection. The paper may be better suited to an applied statistics or sports analytics journal than to a methodology-focused journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a well-specified applied Bayesian paper that does what it says: it builds three longitudinal network models for basketball player-game data, fits them with NIMBLE, and reports posteriors and model selection. The genuinely new piece is the case-study comparison: on the 76ers 2005-06 data, the dynamic LBN with an AR(1) on minutes edges out the static and hidden-Markov variants (WAIC 21194 vs 21275 and 21275). The HMM variant is a null result relative to the static model, which is useful honest negative evidence about team-level hot-hand structure.\n\nThe model specification is complete and internally consistent: likelihoods, priors, and MCMC settings are all there, and the code is on GitHub. That is a real plus for reproducibility. The posterior summaries are plausible, and the player-level predictive distributions are a nice illustration of querying the graph in both directions.\n\nThe soft spots are real but not fatal. The model-preference claim rests on an 80-unit WAIC gap over roughly 82 games times 13 players, about 0.07 per player-game in log pointwise predictive density. No standard error or repeated computation is reported, so the gap is not obviously decisive. More importantly, all predictions in Section 5.3 are in-sample posterior predictives, not holdout forecasts, so the stated goal of predicting future matches is not actually tested. The DAG is fixed by the researcher, as the authors acknowledge, so the structure is assumed, not learned. The data are one team, one season, so the empirical reach is limited.\n\nNone of this sinks the paper. The framework is coherent, the exposition is clear, and the negative HMM result is worth having. But the central claim that the dynamic model 'is the preferred model' is too strong for the evidence presented. A holdout or cross-validation comparison, plus some quantification of WAIC uncertainty, would put the claim on solid ground.\n\nI'd send this to peer review rather than desk-reject. A good referee can push for the holdout analysis and for toning down the preference language. For sports analytics readers and methodologists interested in longitudinal BNs, it's a useful incremental contribution. I'd bring it to a reading group if the topic is on anyone's radar.","headline":"Clean applied Bayesian package of standard models; the dynamic-LBN preference rests on a narrow 80-unit WAIC margin with no uncertainty, and the predictions are in-sample.","tokens_in":17014,"tokens_out":2624,"would_cite":false,"duration_ms":24085,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A dynamic Bayesian network with an autoregressive minutes link is the preferred model for 76ers performance, beating static and hidden-Markov alternatives.","keywords":["Bayesian networks","longitudinal data","dynamic Bayesian networks","hidden Markov models","sports analytics","NBA","prediction","zero-inflated Poisson"],"falsifier":"A held-out predictive evaluation across multiple teams or seasons comparing the dynamic LBN with a version that uses a player-specific or nonlinear function of previous minutes; if the simple AR(1) term no longer improves out-of-sample scores (e.g., WAIC on a test season), the claim that the autoregressive minutes model is preferred would be refuted. A direct posterior predictive check on the 76ers, simulating game point totals and comparing them to actual outcomes, would also show whether the WAIC gain translates into calibrated predictions.","tokens_in":15993,"feed_emoji":"🏀","tokens_out":8340,"duration_ms":71213,"temperature":0.7,"pith_summary":"This paper proposes a Bayesian graphical framework for analysing basketball team performance across a season, treating player participation, minutes played, fouls drawn, and shots attempted and made as random variables in a longitudinal Bayesian network. Three baseline models are compared: a static network, a dynamic network with an autoregressive link between a player's minutes in consecutive games, and a dynamic network with a latent team-level hot/cold state. On the Philadelphia 76ers' 2005–06 season, the autoregressive dynamic model attains the lowest WAIC, and the selected model yields posterior predictive distributions for points and minutes at the player level. The contribution is a reusable modelling recipe for longitudinal team-sport data, not a universal claim about which temporal structure is best.","feed_headline":"Minutes carryover wins Bayesian model race for the 76ers","feed_subtitle":"A simple autoregressive minutes link outperforms static and hot-cold hidden-state models in predicting 76ers performance.","key_machinery":"The central object is the longitudinal Bayesian network: a directed acyclic graph whose nodes are random variables indexed by game, with a joint distribution factorised as local conditional models. The load-bearing mechanism is the dynamic LBN's minutes submodel, a zero-inflated Poisson for minutes played whose log mean is $\\log \\mu^{(M)}_{ij} = \\mu^{(M)}_0 + \\beta^{(+)}_M \\log(1+y^{(M)}_{i,j-1}) + b^{(M)}_i$, following the standard log-linear Poisson autoregressive specification for count time series. This single autoregressive edge between consecutive games, together with player random effects, is what distinguishes the preferred model; the hidden-Markov alternative instead shifts the logit of shooting success probabilities through a latent team state $Z_j \\in \\{C,H\\}$.","core_discovery":"The paper's central claim is that, for the 76ers data, a longitudinal Bayesian network in which minutes played in game $j$ depend on minutes played in game $j-1$ through $\\log(1+y^{(M)}_{i,j-1})$ provides better out-of-sample predictive performance than a static network and than a hidden-Markov network with a latent team hot/cold state. The dynamic LBN reports WAIC 21194.13 versus 21274.69 for the static model and 21274.60 for the hidden-Markov model. The authors take this as evidence that a simple autoregressive carryover captures meaningful temporal dependence in player workload, and they use the selected model to compute conditional predictions for individual players, such as the distribution of points scored when playing more than 30 minutes or the distribution of minutes played when scoring 10 or fewer points.","pith_inferences":["The tiny WAIC gap between the hidden-Markov and static models (21274.60 vs 21274.69) suggests the team-level hot/cold state adds little for this team and season; streakiness may operate at the player level or in different outcome variables, so this comparison should not be read as strong evidence against hot-hand effects in general.","The estimated autoregressive coefficient is small (mean 0.097 with 95% interval roughly 0.076–0.119), so the practical improvement in prediction may be modest; the WAIC advantage could be specific to the 76ers' workload patterns and should be re-tested on other teams or seasons.","Because the network structure is fixed rather than learned, the same framework could be used to compare alternative researcher-specified DAGs against the authors' three baselines, for example a model with player-specific carryover or with minutes feeding back into shot efficiency."],"forward_implications":["For longitudinal team-sport data where player workload carries over between games, the autoregressive minutes specification provides a better predictive baseline than a static network.","The fitted dynamic LBN can be queried for conditional posterior predictive distributions, enabling questions such as how many points a player scores if he plays more than 30 minutes, or how many minutes he plays if he scores fewer than 10 points.","The hidden-Markov LBN offers a Bayesian implementation of the team-level hot-hand effect, and while it did not win on this dataset, the model is available as a component for richer latent-state analyses.","The framework's modular structure supports adding covariates, new nodes, and extending to other leagues or seasons without changing the core inferential machinery."],"supporting_citations":[{"why":"Provides the MCMC computation software used for posterior inference in all three models.","marker":"de Valpine et al., 2017"},{"why":"Supplies the log-linear Poisson autoregression specification used for the dynamic minutes node.","marker":"Fokianos and Tjøstheim, 2011"},{"why":"Supplies the zero-inflated Poisson regression used to model game participation and minutes.","marker":"Lambert, 1992"},{"why":"Previous Bayesian hidden Markov implementation of the hot-hand effect that the hidden Markov LBN extends.","marker":"Calvo et al., 2025"},{"why":"Defines the WAIC criterion used to select the preferred model.","marker":"Watanabe, 2010"},{"why":"Source of the play-by-play data for the 76ers case study.","marker":"NBAstuffer, 2022"}],"fun_headline_variants":["Autoregressive minutes beat hot-cold states for 76ers","Simple minutes carryover wins Bayesian net race for 76ers","Minutes history outperforms static and Markov models in NBA net","For 76ers, minutes dependence beats hidden-state complexity","Bayesian model: minutes autocorrelation predicts better than hot-cold"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The DAG structure and the temporal forms (log(1+previous minutes) autoregression and team-level hot/cold state) are fixed by the researcher rather than learned from the data, so the model ranking and predictions could change if the true dependence structure differs.","fun_headline_variants_meta":{"raw":{"variants":["Autoregressive minutes beat hot-cold states for 76ers","Simple minutes carryover wins Bayesian net race for 76ers","Minutes history outperforms static and Markov models in NBA net","For 76ers, minutes dependence beats hidden-state complexity","Bayesian model: minutes autocorrelation predicts better than hot-cold"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000292,"raw_usage":{"total_tokens":1669,"prompt_tokens":878,"completion_tokens":791,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":494,"completion_tokens_details":{"reasoning_tokens":705}},"tokens_in":494,"tokens_out":791,"duration_ms":6754,"temperature":1.0,"reasoning_tokens":705,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T05:57:20.628275+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A held-out predictive evaluation across multiple teams or seasons comparing the dynamic LBN with a version that uses a player-specific or nonlinear function of previous minutes; if the simple AR(1) term no longer improves out-of-sample scores (e.g., WAIC on a test season), the claim that the autoregressive minutes model is preferred would be refuted. A direct posterior predictive check on the 76ers, simulating game point totals and comparing them to actual outcomes, would also show whether the WAIC gain translates into calibrated predictions.","supporting_citations":[{"cited_title":"Programming with models: writing statistical algorithms for general model structures with","cited_arxiv_id":null,"evidence_quote":"Provides the MCMC computation software used for posterior inference in all three models."},{"cited_title":"Log-linear","cited_arxiv_id":null,"evidence_quote":"Supplies the log-linear Poisson autoregression specification used for the dynamic minutes node."},{"cited_title":"Zero-inflated","cited_arxiv_id":null,"evidence_quote":"Supplies the zero-inflated Poisson regression used to model game participation and minutes."},{"cited_title":"Can the hot hand phenomenon be modelled?","cited_arxiv_id":null,"evidence_quote":"Previous Bayesian hidden Markov implementation of the hot-hand effect that the hidden Markov LBN extends."}],"review_version":1}