{"id":"e7485cd2-f681-44bd-be6b-ffe26abf786a","arxiv_id":"2608.09887","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A junk-possession index and a video-based Space-Creation Index together classify sterile ball circulation versus space-creating possession, and the event flag adds outcome information beyond on-ball value models.","lead":"This paper builds a two-layer football analytics system to separate dead possession from possession that actually creates space. The first layer flags low-threat ball circulation from event data, and the second layer uses broadcast video to check whether the defending team's shape actually moved.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central non-reducibility claim in §4.4 rests on a self-implemented VAEP whose predictive validity is never reported, and the only VAEP-controlled regression is in-sample and score-state-gated; without an out-of-sample or independently validated VAEP comparator, 'beyond VAEP' is not established.","rationale":"The paper's own stated strongest contribution is the VAEP non-reducibility result, and that result depends on the comparator being credible. My concern is not that the authors are hiding something: Sections 4.4 and 7 explicitly flag that the same-match regression is descriptive, that conditional significance is not a coefficient-difference test, and that the xG model is fit in-sample. Those caveats are exactly what make the headline claim weaker than the abstract's 'not reducible' wording. The reader's weakest assumption was the spatial ghosting accuracy; that is a real limitation for the SCI proportions, but it is not load-bearing for the event-side quantitative claim, because the spatial layer is a separate case study whose 31-window purposive sample is already framed as descriptive. So I partially agree with the reader. I would keep the verdict CONDITIONAL (UNCHANGED), but the condition should include an external or out-of-fold VAEP validation before the non-reducibility claim is relied upon. The proposed test settles this: if junk-open remains significant with a held-out VAEP trained on disjoint matches, the concern is resolved; if not, the central claim needs to be downgraded to 'junk-open is associated with points in-sample conditional on our implementation of VAEP,' which is a much more modest claim.","tokens_in":9462,"tokens_out":6525,"duration_ms":61033,"concrete_test":"Re-run the §4.4 points regression in a leave-one-match-out protocol: train the VAEP model and estimate all corpus-wide constants (xG model, q percentile, box weight) on the other 102 matches, compute out-of-fold junk-open and VAEP for the held-out match, and pool the held-out predictions across matches. If junk-open no longer reaches p<1e-4 (or VAEP becomes significant) in the pooled match-clustered regression, the claim that the index is not reducible to on-ball value fails the paper's own out-of-sample standard.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.4's headline result — junk-open remains significant (p<1e-4) while VAEP does not (p=0.34) — is the paper's central quantitative claim, but the evidence for it is one same-match OLS using a VAEP implementation trained on the same corpus with approximated body-part/phase features and an in-sample xG model. The paper never reports this VAEP's own univariate association with points or xG, so a reader cannot tell whether 'not reducible to on-ball value' means the junk index captures something VAEP genuinely misses, or merely that this particular VAEP is too noisy. The cross-fitted splits in §4.5 are the genuinely predictive evidence, but they add no VAEP control; the score-state gating of junk-open means the same-match regression is descriptive, as §7 concedes. Conditional significance in a collinear model is also not an incremental-validity test: the paper itself notes it reports no formal test that the coefficients differ. If a stronger or out-of-fold VAEP were to absorb the junk-open association, the central claim would fail. This concern is about the event-side core of the paper, not the spatial ghosting limitation.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a two-layer possession-quality framework for broadcast football. The event-side junk-possession index segments each match into possession sequences, prices each by peak expected-threat gain plus a shot-xG term and a box-touch credit, normalizes by corpus-wide quantiles, and flags low-threat sequences in tied-or-losing game states. It reports that the resulting junk-open metric correlates negatively with points and xG difference, survives controls for field tilt and a self-implemented VAEP, and shows predictive associations in half-split and leave-one-match-out analyses. The spatial layer computes a Space-Creation Index (SCI) from broadcast video through a game-state-reconstruction pipeline with off-screen player imputation, and applies it to 31 flagged windows from nine World Cup matches, finding 74% non-space-creating, 19% weak progression, and 6% space-creating windows. The paper argues that the two layers together separate sterile from space-creating possession in a way that event-only on-ball value models cannot.","tokens_in":9680,"tokens_out":2984,"duration_ms":28134,"significance":"If the claims hold, the paper contributes a practical macro/micro hybrid for off-ball possession analysis from broadcast footage, an area where most event-based metrics are blind. The event-side index is simple, computable from public event feeds, and released with code; the paper is explicit about its limitations and reports honest nulls (e.g., the redeemed-junk refinement). The falsifiable outcome correlations and the transparent disclosure of corpus-wide fitted constants are strengths. However, the central quantitative claim of non-reducibility to on-ball value rests on an unvalidated VAEP implementation, and the spatial-layer results depend on an unevaluated off-screen imputation method; both are load-bearing for the paper's stated contributions.","major_comments":[{"comment":"The central claim that the junk flag 'is not reducible to on-ball possession value' is supported only by a same-match OLS regression using a self-implemented VAEP whose own predictive validity is never reported. The paper does not report VAEP's univariate association with points or xG, nor any cross-fitted or out-of-sample VAEP comparison. If the implemented VAEP is too noisy, the conditional non-significance of VAEP (p=0.34) is expected, and the conclusion that the junk index 'carries outcome-relevant variance that this on-ball action-value model does not capture' is not established. The manuscript should validate the VAEP implementation (e.g., report its standalone correlation with outcomes and compare it against a published VAEP baseline) and, ideally, repeat the VAEP-controlled analysis in an out-of-sample or cross-fitted setting. As written, Section 7's own acknowledgment that junk-open is constructed from live score state makes the same-match regression descriptive, so the predictive weight falls on the splits in §4.5, which do not include VAEP—leaving the non-reducibility claim without a predictive test.","section":"§4.4, Table 3"},{"comment":"The paper reports conditional significance of junk-open while VAEP is not significant, but it does not report a formal incremental-validity test, such as a nested-model comparison or a test that the junk-open coefficient differs from the VAEP coefficient. The conclusion that the index 'adds information beyond' VAEP requires demonstrating that adding junk-open to a model that already contains VAEP and tilt improves fit or that the coefficient difference is statistically significant; a non-significant VAEP coefficient in the joint model is not equivalent to such evidence. The manuscript should provide this test, preferably with match-clustered errors, and interpret it explicitly.","section":"§4.4 and §7"},{"comment":"The spatial-layer results—most importantly the 74% non-space-creating verdict across 31 flagged windows—depend on pitch-control estimates computed over off-screen player positions imputed by the 'ghosting layer' from companion work. The paper states that it did not filter windows on imputation sensitivity and does not evaluate the imputation's accuracy in this manuscript. If the ghosted positions are biased in even a subset of windows, the SCI classifications and the central off-ball message are unreliable. Since this is structurally different from the event-side validation, the manuscript should provide a sensitivity analysis (e.g., perturbing ghosted positions, comparing against a tracking-data subset, or reporting confidence intervals for SCI under imputation uncertainty) or at least quantify the potential impact on the 31-window classifications.","section":"§5 and §7"},{"comment":"The leave-one-match-out and half-split analyses are presented as the 'genuinely predictive evidence,' but the paper acknowledges that the q-normalizing 90th percentile, the q<0.15 threshold, and the box weight are corpus-wide constants, and the xG model is fit in-sample on the same corpus's shot–goal labels. This means the out-of-sample splits are cross-fitted only up to these global constants and that goal information enters the index weakly through the xG fit. The paper should state more explicitly how much of the reported out-of-sample correlation could be driven by these corpus-wide fitted components, and ideally report a version with the xG model trained out-of-fold for the split analyses.","section":"§4.5 and §7"}],"minor_comments":[{"comment":"Equation (1) contains a corrupted placeholder ('bracehtipupleft/bracehtipdownright/...') that obscures the formula; the weights 0.7 and 0.10 should be presented in a clean, rendered form.","section":"Eq. 1"},{"comment":"The 'junk dominators' versus 'efficient dominators' comparison (n=8 vs n=63) is correctly labeled descriptive, but the text should avoid phrasing like 'the eight averaged 28% of the efficient group's goals' without emphasizing the wide uncertainty; a confidence interval or a nonparametric test would help contextualize the tiny sample.","section":"§4.1"},{"comment":"The xG and xG-difference columns are acknowledged as partly mechanically coupled, but a table footnote should explicitly repeat this caveat so readers do not misread the r=-0.51 coefficient as independent evidence.","section":"Table 2"},{"comment":"The SCI bins (SCI>=+12, +4<=SCI<+12, SCI<+4) are described as 'prespecified operational bins from earlier internal work,' but no reference or rationale is provided; because these bins determine the 74%/19%/6% percentages, the paper should either cite the internal work or provide a justification for the cutoffs.","section":"§5"},{"comment":"Table 5 would benefit from listing the reasons for the four excluded windows (three kit contrast, one clip quality) directly in the table caption or a footnote, rather than only in the text.","section":"§6"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest about its limitations and provides a clear, reproducible event-side codebase, which is valuable. However, the central non-reducibility claim is the hinge of the paper's contribution, and it currently rests on an unvalidated VAEP comparator in a descriptive regression; the spatial layer's imputation sensitivity is also not addressed. These are fixable with additional validation and sensitivity analyses, so I recommend major revision rather than rejection. If the journal's scope is strictly computer vision, the spatial layer is modest in size, but the paper does present a viable broadcast-to-pitch pipeline; that fit concern is secondary to the statistical issues."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead the Choi paper on junk possession and the Space-Creation Index. Bottom line: this is a genuinely novel framing with a useful event-side metric, but the central \"not reducible to VAEP\" claim has a load-bearing weakness you should know about.\n\nWhat's new: the two-layer structure — an event-side junk-possession flag that scans whole matches cheaply, plus a video-based spatial index that adjudicates flagged windows — is a real contribution. The idea of separating \"space-creating but unconverted\" from \"sterile\" is exactly the distinction possession percentage and even VAEP miss, and the paper shows how to do it with broadcast footage, which is the realistic data source for most matches. The event-side index is transparent: sequence value is a simple combination of threat gain, shot xG, and box touch, and they ship code for the whole validation battery, including out-of-sample splits and match-clustered errors. That is reproducible work and it deserves credit.\n\nWhere it gets soft: the non-reducibility claim in §4.4 rests on a VAEP implementation trained on the same corpus, and the paper never reports whether that VAEP alone predicts points or xG. If it is too noisy, the conditional significance of the junk flag proves little. The authors are honest — they call the same-match regression descriptive and note they don't test whether the coefficients differ — but the reader is still asked to take \"beyond VAEP\" on faith. A stronger, or at least independently validated, on-ball baseline is needed before that claim lands. The spatial layer depends on an off-screen imputation method from companion work that is never validated here, and the 31-window sample is purposive. The authors flag both, but it means the spatial proportions should be read as a case study, not a prevalence estimate. Finally, the thresholds (0.7, 0.10, q<0.15, SCI bins) are operational values, partly data-informed but not calibrated; the out-of-sample splits are cross-fitted only up to these constants, as the paper concedes.\n\nNone of this kills the paper. The event-side analysis is a solid contribution to possession-quality metrics, and the spatial demonstration is a nice illustration. For a reader who works on football analytics or broadcast video, this is worth an afternoon.\n\nMy recommendation: send it to peer review rather than desk-reject. A referee should push on the VAEP validation and ask for a clearer separation of the descriptive same-match result from the out-of-sample evidence. With those holes patched, this could be a useful addition to the literature.","headline":"A genuinely novel two-layer possession-quality framework with an honest event-side core; the 'beyond VAEP' claim is real but under-supported by an unvalidated comparator.","tokens_in":10268,"tokens_out":2995,"would_cite":true,"duration_ms":26605,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a junk-possession index—low-threat circulation in tied or losing states—predicts points where on-ball value (VAEP) does not, and that a spatial layer from broadcast video can tell sterile possession from…","keywords":["possession quality","junk-possession index","expected threat","VAEP","pitch control","space creation","broadcast video analysis","football analytics"],"falsifier":"Run the same GSR pipeline on broadcast matches for which full commercial tracking data exist, compare the imputed off-screen player positions against true positions, and re-compute SCI verdicts on both; if the spatial classifications (non-creating vs weak vs space-creating) flip when imputed positions are perturbed or replaced with ground truth, the spatial layer's numbers are not trustworthy.","tokens_in":9147,"feed_emoji":"⚽","tokens_out":2972,"duration_ms":27060,"temperature":0.7,"pith_summary":"This paper argues that the most-cited football stat, possession percentage, hides a crucial off-ball distinction: whether holding the ball dragged the opponent's defensive block out of shape or simply circulated without threat. It builds a two-layer index that first flags \"junk\" possessions from event data—low-threat sequences in tied-or-losing game states—and then, for short flagged windows, uses broadcast video to measure whether the possession actually created space via a Space-Creation Index. The central quantitative claim is that the junk flag is not a repackaging of on-ball value: with team offensive VAEP and field tilt held fixed, the flag remains strongly negatively associated with points (p < $10^{-4}$) while VAEP is not significant (p = 0.34). The paper also reports that among 31 flagged windows from nine World Cup matches, 74% are spatially non-space-creating, 19% weak progression, and 6% space-creating—windows the event-only flag would score as failure. This matters because it offers a practical, broadcast-only route to measuring off-ball value at tournament scale.","feed_headline":"Off-ball possession index beats on-ball value at predicting points","feed_subtitle":"Two-layer system flags sterile circulation and reads space from broadcast video, separating dead possession from unlucky finishes.","key_machinery":"Two linked instruments carry the argument. First, the junk-possession index prices each possession sequence by value(s) = max(0, xT_max(s) − xT_start(s)) + 0.7·xG(s) + 0.1·1[box touch], normalized by a corpus-wide 90th percentile and clipped to [0,1]; sequences with q < 0.15 are flagged as junk, and the metric junk-open is the event share of such sequences in tied-or-losing states. Second, the Space-Creation Index (SCI) computes a net two-zone pitch-control change from broadcast video projected to pitch coordinates: SCI = Δ_own + Δ_opp, where Δ_own is the change in the possessing team's control of the attacking third and Δ_opp is the recession of the opponent's presence in the mirror zone. The event layer scans whole matches cheaply and plants flags; the spatial layer resolves why on short windows.","core_discovery":"The paper's core discovery is that the junk-possession flag—the on-ball-event share of low-threat possession sequences in tied-or-losing states—carries outcome-relevant variance that a standard on-ball action-value model (VAEP) does not capture on its own. In a same-match regression of 206 team-match rows from the 2026 FIFA World Cup, junk-open remains significantly negatively associated with points (p < $10^{-4}$, also under match-clustered errors) when team offensive VAEP and field tilt are held fixed, while VAEP is not significant (p = 0.34). The paper further shows that a spatial layer, computing a Space-Creation Index from pitch control on broadcast video, can adjudicate whether an event-flagged junk possession was spatially dead or space-creating, separating sterile domination (e.g., Germany's 73% possession in a penalty-shootout exit) from unlucky but genuinely space-creating play.","pith_inferences":["One implication the author leaves implicit: a tracking-fed VAEP or EPV that includes defensive shape might subsume some of the junk flag's predictive content, since the non-reducibility test uses an event-fed VAEP; a natural extension would test whether the junk flag survives against tracking-based on-ball value on the same matches.","The 6% space-creating windows suggest event-only metrics systematically misclassify a small but consequential set of possessions—those that improve pitch control without producing a shot—so expected-threat and VAEP-style models could be augmented with a pitch-control term.","The macro/micro hybrid is directly transferable to other broadcast-only competitions (club leagues, national teams outside top tracking leagues), and the SCI thresholds (+12/+4) are operational values that a larger study could calibrate against actual scoring rates.","A testable extension would be to run the full pipeline on matches for which commercial tracking data exist, checking whether the imputed off-screen positions and SCI verdicts match the tracking-derived ground truth."],"forward_implications":["If the index is right, teams can diagnose sterile possession from event data alone and target the off-ball problem without full tracking data, which most leagues lack.","The junk flag behaves like a team trait: its value in a team's other matches predicts points and xG difference in a held-out match, so it could be used for pre-match assessment.","The two-layer design shows that event-only possession-value models miss a real axis of quality—space creation—and that the distinction is measurable from broadcast video.","The case evidence indicates that dominant ball share can coexist with elimination (Germany 73% possession, two non-creating windows), so possession-based narratives need spatial correction."],"supporting_citations":[{"why":"Supplies the expected-threat grid used to price each possession sequence's peak threat gain, the core of the junk index.","marker":"[11]"},{"why":"Provides the VAEP on-ball action-value model that the paper's central non-reducibility claim is tested against.","marker":"[3]"},{"why":"Provides the physics-based pitch-control model from which the Space-Creation Index is computed.","marker":"[14]"},{"why":"Supplies the camera-calibration method (PnLCalib) used to project broadcast video to pitch coordinates in the GSR pipeline.","marker":"[7]"},{"why":"Supplies the BoT-SORT multi-object tracker used for player tracking in the broadcast GSR pipeline.","marker":"[1]"}],"fun_headline_variants":["New off-ball index beats VAEP at predicting football points","Junk possession flag predicts points better than on-ball value","Space-creation index separates dead possession from unlucky loss","Video-based index flags sterile possession, beats VAEP on points"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The spatial layer's Space-Creation Index depends on the off-screen player imputation 'ghosting layer' from companion work, and the paper does not evaluate the accuracy of that imputation or filter windows on imputation sensitivity; if the ghosted positions are biased, the pitch-control estimates and the SCI verdicts (74% non-creating, etc.) could be unreliable.","fun_headline_variants_meta":{"raw":{"variants":["New off-ball index beats VAEP at predicting football points","Junk possession flag predicts points better than on-ball value","Space-creation index separates dead possession from unlucky loss","Video-based index flags sterile possession, beats VAEP on points"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000705,"raw_usage":{"total_tokens":3272,"prompt_tokens":1129,"completion_tokens":2143,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":745,"completion_tokens_details":{"reasoning_tokens":2075}},"tokens_in":745,"tokens_out":2143,"duration_ms":13960,"temperature":1.0,"reasoning_tokens":2075,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T04:51:58.604033+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same GSR pipeline on broadcast matches for which full commercial tracking data exist, compare the imputed off-screen player positions against true positions, and re-compute SCI verdicts on both; if the spatial classifications (non-creating vs weak vs space-creating) flip when imputed positions are perturbed or replaced with ground truth, the spatial layer's numbers are not trustworthy.","supporting_citations":[{"cited_title":"Introducing expected threat (xt)","cited_arxiv_id":null,"evidence_quote":"Supplies the expected-threat grid used to price each possession sequence's peak threat gain, the core of the junk index."},{"cited_title":"Actions speak louder than goals: Valuing player actions in soccer","cited_arxiv_id":null,"evidence_quote":"Provides the VAEP on-ball action-value model that the paper's central non-reducibility claim is tested against."},{"cited_title":"Physics-based modeling of pass probabilities in soccer","cited_arxiv_id":null,"evidence_quote":"Provides the physics-based pitch-control model from which the Space-Creation Index is computed."}],"review_version":1}