{"id":"22d4e848-9bdd-401f-b37d-432c696d99c3","arxiv_id":"1908.00939","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A team rating curve over game time is estimated by running the standard least-squares rating model separately at each second, yielding time-resolved strength-of-schedule-adjusted margins.","lead":"The paper fits a least-squares \"functional rating\" for each college basketball team at every second of the game, using all scoring data rather than only final scores. It argues the rating is a team's average point margin adjusted for schedule strength, and finds home-court advantage is real but roughly constant across teams.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim depends on play-by-play d(t) being exact at every second, but Section 3.1 documents systematic timing errors; without a sensitivity analysis, the functional ratings and Model 2 conclusion are not yet secure.","rationale":"The reader's weakest assumption identifies exactly the same load-bearing concern: the data accuracy of the interpolated score differential at every time. This is the right focus because the central assertion is a pointwise least-squares statement about d(t); any systematic error in d(t) propagates linearly into beta(t), the ranking, and the model-comparison tests. I considered alternative concerns, such as the multiple-comparison and temporal-correlation issues in the pointwise F-tests, but those affect the strength of the Model 2 conclusion rather than the core descriptive ratings. The data-accuracy issue is more fundamental: if d(t) is wrong, every downstream quantity inherits the bias. The paper is transparent about the three error types but provides no quantification or sensitivity analysis, so the central claim is not fully supported as stated. This does not make the paper unsound; the errors are documented and plausibly small for most games, and a sensitivity check could resolve the concern. Therefore the existing CONDITIONAL verdict is appropriate and no change is needed.","tokens_in":7700,"tokens_out":12296,"duration_ms":144008,"concrete_test":"Re-fit Model 2 on a corrected dataset: exclude the single game with no public scoring summary, remove the appended final scoring plays, and for each same-timestamp cluster either re-credit all points at the recorded time or distribute them uniformly across the preceding interval. Compare the resulting beta_i(t), the scalar rankings from Equation (5), and the Model 1-versus-Model 2 p-value functions. If any top-25 rank moves by more than a few positions, or the end-of-game home-court advantage alpha(T) changes by more than about half a point, the documented data errors are load-bearing.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing premise is that the pointwise least-squares fit in Section 2.1 and Equation (1) uses the true score differential d(t) at every time t. Section 3.1 shows this premise is violated on whole intervals: missing scoring plays are appended at the final timestamp, and simultaneous scoring plays are treated as if the pre-play lead persisted until the recorded time (Table 1 documents a 5-point error persisting over 213 seconds). Because the estimator is linear in d(t), these errors pass directly into the functional ratings and into the Section 4 ANOVA p-value functions used to select Model 2. The errors may be small on average, but the paper gives no count of affected games or seconds and no sensitivity analysis, so the strongest claim—that beta_i(t) - beta_j(t) is the expected point differential and that Model 2 with a single home-court advantage is appropriate—is conditioned on unverified data accuracy.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a functional rating model for sports teams. Given play-by-play score differences d(t) for m games, team ratings β_i(t) are estimated by solving Xβ(t)=d(t) in least squares at each second, with the identifiability constraint Σβ_i(t)=0. Three specifications are considered: no home-court advantage (Model 1), a common home-court advantage α(t) (Model 2), and team-specific advantages α_i(t) (Model 3). The model is applied to 5,603 NCAA Division 1 men's basketball games from the 2018–2019 season. Model selection is performed by computing ANOVA p-values pointwise in time, leading to the choice of Model 2. Ratings are smoothed with fourth-order B-splines, converted to scalar rankings by weighted integrals, and decomposed into average point differential and strength of schedule via the normal equations. The paper also illustrates in-sample 'predictions' of the expected point differential curve for a Duke–Virginia matchup.","tokens_in":7868,"tokens_out":4603,"duration_ms":46807,"significance":"The methodological idea is attractive, and the least-squares algebra is straightforward and largely self-contained. The strength-of-schedule decomposition in Section 5.2 is a nice consequence of the normal equations, and the explicit treatment of identifiability via the sum-to-zero constraint is sound. The paper also deserves credit for candidly documenting the three data-quality problems in Section 3.1. If the empirical claims were supported by out-of-sample validation and sensitivity analysis, the functional rating approach would be a useful addition to sports analytics, going beyond final-score ratings by making the entire time course of performance comparable. However, the current evidence for the central claims—that β_i(t)-β_j(t) is the expected point differential and that Model 2 is the appropriate specification—is in-sample and potentially sensitive to the documented data errors, so the significance is currently conditional.","major_comments":[{"comment":"The documented play-by-play timing errors directly affect the estimand. Because the estimator is linear in d(t), the convention of appending missing scoring plays at the final timestamp and carrying the pre-play lead across intervals with simultaneous scoring plays (Table 1) mechanically shifts β_i(t)-β_j(t) on the affected intervals. The paper states that these situations were not modified, but gives no count of games or seconds affected and no sensitivity analysis. Since the central claim in Eq. (1) is that the fitted difference equals the expected point differential at time t, the authors should report the proportion of affected data and re-estimate the ratings under conservative corrections (e.g., deleting or down-weighting affected intervals, or randomizing the positions of missing scoring plays) to show that rankings, Model 2 selection, and the home-court function are stable.","section":"§3.1, Eq. (1)"},{"comment":"The 'predicted game' in Figure 8 is not a prediction in the usual sense: it is the fitted difference β_Duke(t)-β_Virginia(t) evaluated on the same data used to estimate the parameters. No holdout data, cross-validation, or comparison with an external benchmark (e.g., Massey ratings, betting lines, or a simple final-score model) is provided. The abstract's claim that 'using two team's functional ratings we can predict the expected point differential at any time' is therefore unsupported. Please add an out-of-sample evaluation, for example by withholding a random subset of games, fitting ratings on the remainder, and comparing predicted end-of-game differentials or predicted curves to actual outcomes, with a baseline model.","section":"§5.3, Figure 8"},{"comment":"The ANOVA-based model choice is made from pointwise p-value functions evaluated at every second, but the analysis does not address multiple testing or temporal dependence. The statement 'consistently greater than .1 other than the first 20 seconds' in the Model 2 versus Model 3 comparison is a visual threshold, not a formal functional test, and the early-game exception coincides with the data-quality issues described in Section 3.1 (simultaneous scoring plays at the start of games). A joint test (e.g., an F-test on integrated squared errors or a permutation test) and a multiple-comparison adjustment are needed before concluding that Model 2 is appropriate over the whole game.","section":"§4, Figures 1–3"},{"comment":"The ranking comparisons in Table 3 are descriptive properties of the same fitted curves, not evidence of predictive or ranking accuracy. The paper does not quantify uncertainty in the ranks or scalar ratings (e.g., via bootstrap over games), and the claim that the top ten teams 'are the same, but in a different order' is a statement about one fitted model. A bootstrap or cross-validation would clarify whether rank movements such as Stanford's 40-position change are stable or noise.","section":"§5.1, Table 3"}],"minor_comments":[{"comment":"The word 'discus' should be 'discuss', and 'a teams average point differential' needs an apostrophe: 'a team's average point differential'.","section":"§2.1"},{"comment":"In the sentence after Eq. (6), 'slit' should be 'split'.","section":"§5.2"},{"comment":"The paper removes overtime data in Section 2.3, but Section 3.1 mentions an overtime score in the Jackson State–Alabama A&M game; please clarify whether the final recorded point for that game is regulation or overtime and how it is handled in the interpolation.","section":"§2.3 and §3.1"},{"comment":"The y-axis is labeled 'Score' but the plotted quantity is a predicted point differential; the caption 'which is filled to be their team color if winning' is unclear and should be rewritten.","section":"Figure 8"},{"comment":"The notation h_i and a_i is defined only after the equation; moving the definitions before the display would improve readability.","section":"Eq. (8)"},{"comment":"Reference [8] has a typo: 'Funtional Data Analysis' should be 'Functional Data Analysis'.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is honest about its limitations, which is to its credit. However, the central empirical claims are currently supported only by in-sample fits and by data that the authors themselves document as error-prone. I would not recommend publication before the sensitivity analysis and out-of-sample validation requested in the major comments are provided. The methodological core is sound and could become a useful contribution once those checks are in place."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing to know: this is a competent and clearly written extension of scalar least-squares sports ratings to a function-valued rating over game time. The pointwise least-squares algebra is correct, and the schedule-strength decomposition from the normal equations is a nice descriptive payoff. But the paper oversells the word 'prediction': the predictions in Section 5.3 are in-sample fitted differences, and the data errors documented in Section 3.1 are never tested for sensitivity.\n\nWhat's new: previous least-squares ratings (Stefani, Harville) used final scores; this runs the same regression at each second, producing rating curves. Because there is no temporal coupling, it is a pointwise collection of scalar ratings, not a new estimation principle. Still, the functional view gives a useful way to see when teams are strong and to break a rating into average margin and schedule strength. The authors are transparent about their data problems and about the fact that their 'predicted' game flow is not realistic.\n\nSoft spots, in proportion. The central descriptive model holds up. The bigger issues: first, the model-comparison ANOVA is run at every second without any adjustment for multiple comparisons or temporal correlation; the p-value functions in Figures 1 and 2 are highly correlated, so 'consistently less than .1' is not a formal inference. Second, the data errors are potentially load-bearing. The North Alabama/Samford example shows a 5-point lead assigned to a 213-second interval; if such intervals are common, the pointwise ratings and the home-court conclusion inherit those biases. The paper gives no count of affected games or seconds and no sensitivity analysis. This is fixable but not ignorable. Third, no code or data is provided, so the empirical claims cannot be checked.\n\nWho this is for: sports-analytics readers and functional-data-analysis folks looking for a simple, honest application. It deserves a serious referee, not a desk reject: the method is sound as a descriptive tool, and the deficiencies are addressable. I'd ask the authors for an out-of-sample or cross-validation check, a multiple-comparisons or time-series-aware treatment of the p-value functions, and sensitivity analysis around the data-cleaning choices. With those, the paper would be a reasonable contribution to an applied statistics or sports analytics venue.","headline":"Competent pointwise extension of least-squares ratings to functional data; the descriptive model works, but the predictive claims and inference need more support.","tokens_in":8383,"tokens_out":2172,"would_cite":true,"duration_ms":22099,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62J05","62P99"],"pacs":[],"model":"deepseek-v4-flash","headline":"Team strength as a curve, not a number: least squares fit to every second of game data yields ratings that predict point differential at any moment.","keywords":["functional data analysis","least squares","sports rankings","home-court advantage","point differential","basketball","strength of schedule","time-varying ratings"],"falsifier":"Re-run the analysis after removing the documented data errors—such as the game with only box-score points, missing final scoring plays appended at the final timestamp, and intervals where simultaneous scoring plays are treated as a persistent lead—and check whether the functional ratings and the model-selection P-values change substantially. If they do, the central claim that the ratings reflect true time-varying performance is undermined; if they barely change, the claim is supported.","tokens_in":7493,"feed_emoji":"🏀","tokens_out":3342,"duration_ms":34804,"temperature":0.7,"pith_summary":"This paper proposes a new way to rank sports teams: instead of using only final scores, it fits a least-squares model at every second of game time, producing a rating curve for each team. The central claim is that the difference of two teams' rating curves equals the expected point differential between them at that time, adjusted for strength of schedule. Using 2018-2019 NCAA men's basketball play-by-play data, the paper argues that a single common home-court advantage curve is the right specification: home advantage is statistically important but does not differ across teams. If correct, the method turns every scoring event into information about team strength and allows prediction of the expected score gap at any moment.","feed_headline":"Functional ratings predict score gap at any moment","feed_subtitle":"A least-squares model fit at every second shows home-court advantage is real but the same for every team.","key_machinery":"The load-bearing object is the design matrix $X$ together with the pointwise least-squares solution at each second. $X$ encodes each game as a row with $+1$ for the home team, $-1$ for the away team, and $0$ otherwise, and the constraint that the average rating is the zero function fixes the null space. The normal equation $X^\\top X\\beta(t) = X^\\top d(t)$ splits every rating into average point differential and strength of schedule. Model selection uses F-tests at every time point, producing a curve of $P$-values that shows when the constant home-court model is appropriate.","core_discovery":"The paper's central discovery is that the pointwise least-squares solution of $X\\beta(t) = d(t)$ with the constraint $\\sum_i \\beta_i(t) = 0$ yields a functional rating $\\beta_i(t)$ for each team such that $\\beta_i(t) - \\beta_j(t)$ is the expected point differential at time $t$ in a neutral-court game between teams $i$ and $j$. The rating naturally decomposes into the team's average point differential plus a time-varying strength-of-schedule component. Comparing three nested models with ANOVA at every second shows that the model with a constant home-court advantage $\\alpha(t)$ is needed, while a model with individual team home-court advantages is not, so the constant-home-advantage model is selected for the rest of the analysis.","pith_inferences":["The paper's P-value threshold is informal; a formal functional hypothesis test for the home-court advantage curves could sharpen the model choice and might alter the conclusion if early-game P-values were treated differently.","The decomposition $\\beta_i(t) = \\bar{d}_i(t) + \\mathrm{sos}_i(t)$ suggests a time-varying strength-of-schedule measure that could be compared with static strength-of-schedule rankings, potentially offering new insight into schedule difficulty.","Because the documented data errors (missing scoring plays, simultaneous plays treated as persistent leads) could bias specific intervals, a sensitivity analysis that drops or re-times those plays would test whether the ratings and model selection survive the noise."],"forward_implications":["Team rankings can be computed as weighted averages of the functional ratings, allowing users to favor early-game or late-game performance according to a chosen weight function.","The rating curve reveals a team's playing style, such as being a strong first-half or second-half team, as illustrated by Stanford improving late and Illinois leveling off.","Expected point differential can be predicted between any two teams at any time, even if they never played, and the home-court advantage adds about three points at the end of a game.","The framework extends naturally to other sports with scoring data, though further study is needed to simulate realistic game flows."],"supporting_citations":[{"why":"Supplies the functional data analysis framework and the pointwise least-squares solution method used to fit ratings at each second.","marker":"[8]"},{"why":"Provides the ANOVA model-comparison approach for testing home-court advantage specifications, which the paper adapts to a function of time.","marker":"[3]"},{"why":"Introduces least-squares rating models that this paper extends from scalar outcomes to time-dependent point differentials.","marker":"[9]"},{"why":"Offers improved least-squares prediction methodology that informs the model variations for home-court advantage.","marker":"[10]"},{"why":"Describes a related rating system that serves as a baseline for the least-squares approach and the strength-of-schedule interpretation.","marker":"[5]"}],"fun_headline_variants":["Least-squares ratings model point spread each second","Home-court edge uniform across NCAA teams","Score gap at any time from a single rating curve","Functional ratings: schedule-adjusted scoring curves","New ratings track point spread throughout the game"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"That the play-by-play scoring data, after interpolation to every second, accurately represents the true score differential at every time, so the pointwise least-squares ratings do not inherit bias from missing or simultaneous scoring plays.","fun_headline_variants_meta":{"raw":{"variants":["Least-squares ratings model point spread each second","Home-court edge uniform across NCAA teams","Score gap at any time from a single rating curve","Functional ratings: schedule-adjusted scoring curves","New ratings track point spread throughout the game"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000481,"raw_usage":{"total_tokens":2303,"prompt_tokens":797,"completion_tokens":1506,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":413,"completion_tokens_details":{"reasoning_tokens":1438}},"tokens_in":413,"tokens_out":1506,"duration_ms":10844,"temperature":1.0,"reasoning_tokens":1438,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:27:29.742327+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the analysis after removing the documented data errors—such as the game with only box-score points, missing final scoring plays appended at the final timestamp, and intervals where simultaneous scoring plays are treated as a persistent lead—and check whether the functional ratings and the model-selection P-values change substantially. If they do, the central claim that the ratings reflect true time-varying performance is undermined; if they barely change, the claim is supported.","supporting_citations":[{"cited_title":"Ramsay and B.W","cited_arxiv_id":null,"evidence_quote":"Supplies the functional data analysis framework and the pointwise least-squares solution method used to fit ratings at each second."},{"cited_title":"Harville and Michael H","cited_arxiv_id":null,"evidence_quote":"Provides the ANOVA model-comparison approach for testing home-court advantage specifications, which the paper adapts to a function of time."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces least-squares rating models that this paper extends from scalar outcomes to time-dependent point differentials."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Offers improved least-squares prediction methodology that informs the model variations for home-court advantage."},{"cited_title":"Statistical models applied to the rating of sports teams","cited_arxiv_id":null,"evidence_quote":"Describes a related rating system that serves as a baseline for the least-squares approach and the strength-of-schedule interpretation."}],"review_version":1}