{"id":"4389bac0-5070-4b04-8d2f-2318a6f536b6","arxiv_id":"2607.27125","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A taxonomy-guided, rally-level, video-anchored review system elicits more frequent, concrete, actionable, and appropriate tactical reflections from amateur badminton players than a report-and-statistics baseline.","lead":"TactiPlay helps amateur badminton players review their own match videos with taxonomy-guided tactical feedback tied to rallies and clips. A 16-player study found more concrete, actionable reflections than a stats-and-report baseline.","discovery_kind":"new_application","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The primary outcome (expert-rated reflection quality) appears to be analyzed on pooled entries (n=79 vs 44) rather than on the 16 participants, violating independence and inflating the reported p-values and power; unequal entry counts across conditions further conflate quality with quantity.","rationale":"The reader's weakest assumption targeted confounds the design does not isolate (different matches per condition, joint manipulation of workflow components, unverified traces, no diagnosis-correctness audit). Those are real and largely self-flagged in §7.3, and they justify CONDITIONAL. My concern is different and arguably more immediately load-bearing because it attacks the inference supporting the single strongest quantitative claim rather than the design's scope: even accepting the workflow comparison as valid, the entry-level analysis may overstate the evidence that reflection quality improved. I rate agreement as partial because the reader did not identify the unit-of-analysis issue. I do not recommend REJECT: the effect sizes are large (1.5–1.75 scale points), the qualitative strategy reversal (12/16 report-driven vs 6/16) and interview evidence provide independent converging support, and the fix is a reanalysis of existing data, not a new study. If the participant-level reanalysis holds, the paper's central claim stands essentially intact under the reader's existing CONDITIONAL framing; if it fails, CONDITIONAL should harden toward REJECT for the quality-difference claim while the system/design contribution remains. Hence UNCHANGED with a mandatory reanalysis condition. Note also the frequency result (79 vs 44) was reported without a stated test; a per-participant entry-count comparison should accompany the reanalysis.","tokens_in":21715,"tokens_out":1853,"duration_ms":47475,"concrete_test":"Recompute Figure 5 Part A with the participant as the unit of analysis: for each participant, average their entry-level scores per dimension per condition, then run Wilcoxon signed-rank tests (N=16 pairs) on overall Concreteness, Actionability, and Appropriateness; alternatively fit a linear mixed model on entries with a participant random intercept. If Actionability (5.44 vs 3.92) and Appropriateness (5.80 vs 4.05) remain significant at p<.05, the concern does not land; if significance collapses, the headline quality claims should be downgraded to trends. Secondary check: have a rater guess condition from de-labeled entries; above-chance guessing would confirm leakage.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Figure 5 Part A reports significance stars and observed power (0.97–1.00) on tests whose units are reflection entries: nT=79, nB=44, and per-domain subsets (e.g., 38 vs 29). But entries are nested within only 16 participants, each contributing a variable number of entries per condition (~5 vs ~3 on average). Entries from the same participant share style, skill, and session effects, so treating 79 vs 44 as independent observations overstates the effective sample size; the true denominator for a within-subjects claim is 16 pairs. The dramatic power values and *** markers therefore rest on a violated independence assumption, and the headline numbers in the abstract (e.g., Actionability M=5.44 vs 3.92) derive their inferential weight from this analysis. A second, related issue: the pools are unequal because TactiPlay elicited nearly twice as many entries. If the marginal entries are systematically different in quality (e.g., scaffolded restatements of report text), comparing means across unequal pools mixes a quantity effect with a quality effect. Third, raters scored entries \"while checking against the corresponding match footage\" after condition labels were removed, but TactiPlay-condition entries plausibly echo taxonomy descriptors, subtopic phrasing, and linked-clip references (cf. P7, P16 quotes), making the condition inferable; partial unblinding would bias exactly the three dimensions that carry the claim. None of this is alleged misconduct; it is a standard unit-of-analysis/clustering problem that is cheap to fix but load-bearing for the strongest claim.","agreement_with_reader":"partial"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The paper presents TactiPlay, a post-match review system for amateur badminton players. A formative interview study (N=8) yields three design requirements (video-linked feedback, tactical rather than stroke-only focus, rally-level entry points); free-form annotations of seven amateur matches by three national-level players are open-coded into a three-domain, 17-descriptor performance-issue taxonomy. A two-stage pipeline (CV/VLM event detection, then GPT-4.1-based issue generation, clustering, and rally summarization) feeds an interface with a four-topic structured report, subtopic-linked video clips, a Rally Explainer, and a dynamic court heatmap. A within-subjects study (N=16) compares TactiPlay against a report-and-statistics baseline using the same upstream detection pipeline, manual label verification, and report generator. The authors report that TactiPlay elicits nearly twice as many reflection entries, with higher expert-rated concreteness, actionability, and appropriateness, higher subjective ratings on 8 of 11 items, and a shift from video-driven manual search to a report-to-evidence review strategy.","tokens_in":22133,"tokens_out":3339,"duration_ms":143092,"significance":"If the main result holds, the paper makes a solid contribution to sports HCI: a worked example of how to organize AI-generated match analysis into a taxonomy-guided, rally-level, video-anchored reflection workflow for amateurs, with an evaluation showing that this organization changes both reflection quality and review strategy. The expert-derived taxonomy (304 annotations from national-level players, 17 descriptors) is a reusable artifact. The study is careful by field standards: within-subjects with Latin-square counterbalancing, identical upstream pipeline and verification protocol across conditions, blinded expert rating with strong ICCs, and unusually transparent non-claims about autonomous raw-video analysis. The detected strategy reversal (report-to-evidence loop) is a genuinely interesting interaction finding. Impact is moderate rather than transformative: N=16, single session, no behavioral outcome.","major_comments":[{"comment":"The primary outcome appears to be analyzed at the level of pooled reflection entries (nT=79, nB=44 overall; 38/29, 28/11, 23/9 per domain). Entries are nested within only 16 participants who contributed unequal numbers of entries per condition (~5 vs ~3 on average), so entry-level tests violate independence and the reported p-values and observed power (0.97-1.00) are not interpretable as stated. §6 says Wilcoxon signed-rank was used for paired data, but the figure's units are entries, not participants. Please re-analyze at the participant level (per-participant mean quality per condition, Wilcoxon, N=16, with effect sizes) or fit mixed-effects models with a random intercept per participant (and per match), and report per-participant entry counts. Also address the unequal-pool composition: with nearly twice as many TactiPlay entries, pooled means mix a quantity effect with a quality effec","section":"§6.1, Figure 5 Part A"},{"comment":"Condition labels were removed before expert rating, but TactiPlay-condition entries plausibly echo taxonomy descriptors, subtopic phrasing, and linked-clip references (cf. P7 and P16 quotes), making the condition inferable to raters who had also calibrated on taxonomy categories. Partial unblinding would bias exactly the three dimensions (concreteness, actionability, appropriateness) that carry the headline claim. Please either run a blinding audit (ask raters to guess condition per entry and test whether guessability correlates with scores), or add a dedicated limitation and a robustness check (e.g., restrict to entries with no taxonomy-specific vocabulary).","section":"§5.4.1 (rater blinding)"},{"comment":"The conditions differ jointly in prompt structure, taxonomy organization, clustering, video linking, interaction, and probably report length; each participant also reviewed two different matches. The integrated-workflow framing in §5.1 and §7.3 is honest, and I do not ask for a full ablation, but the causal language should be kept at workflow level throughout: the abstract and §6.1 phrases like 'elicits more frequent, concrete, actionable' reflections read as component-level attribution. Please also add one paragraph on the report-volume confound: higher helpfulness/efficiency ratings and more reflection entries are expected when one condition simply contains more content, and the interesting claim is about organization and linkage, not quantity.","section":"§5.1 vs §6-7 (attribution of effects)"},{"comment":"The pipeline's issue generation (GPT-4.1 with taxonomy reference) is central to the contribution, yet the correctness of generated diagnoses is never validated; §7.3 acknowledges this, and P3's quote ('the system said I was intercepted, but the shuttle was actually out') shows errors are visible to users. Since 'appropriateness' of reflections was rated against footage rather than against system output, the current evidence cannot rule out that participants reflect confidently on partly wrong diagnoses. Please add a small expert audit: sample k generated issues (e.g., 30-50) stratified by topic, have badminton experts judge correctness against the linked clips, and report the rate. This materially strengthens the taxonomy-grounding claim at modest cost.","section":"§3.3.2 and §7.3 (diagnosis correctness)"}],"minor_comments":[{"comment":"Reported detection metrics are modest (rally segmentation 75%, winner/win-reason 68%, hit detection precision 0.71/recall 0.72, hitter ID 77.9%), and stroke classification accuracy is only described as 'insufficient' without a number. Please give the number, and state the provenance/size of each test set and whether the 109-rally and 564-hit evaluations overlap with the seven taxonomy-annotation videos or the 32 study matches.","section":"§3.3.1"},{"comment":"Post-hoc observed power is generally uninformative and here especially misleading if computed on pooled entries (see Major Comment 1). Consider replacing the power column with effect sizes (e.g., rank-biserial r or Cliff's delta) and confidence intervals.","section":"Figure 5 (observed power column)"},{"comment":"Each participant reviewed two different matches with different opponents; match difficulty could affect reflection quality. The 2×2 Latin square handles order, but please report whether match identity (or opponent strength) showed any association with entry counts or quality, even descriptively.","section":"§5.3 (counterbalancing)"},{"comment":"Typesetting artifacts: 'presentTactiPlay' and 'inTactiPlay' (missing spaces) in Abstract/§1; 'Tactical A wareness Issues' in §4.1; 'R¯utenis Paulauskas' in ref [51]. Schön is cited as 2017 (a reprint); consider citing the original 1983 edition.","section":"Typos"},{"comment":"The frequency-based star rating on subtopic cards (Figure 3-F2) is mentioned only in the caption; one sentence on its rationale and scale would help, since frequency is not obviously a severity signal.","section":"Figure 3 / §4.1"},{"comment":"Figure 6 is descriptive over 16 participants; the report-driven vs video-driven shift (12/16 vs 6/16) is a nice result and could support a simple exact test (e.g., McNemar) rather than counts alone.","section":"§6.3 / Figure 6"}],"recommendation":"major_revision","confidential_remarks":"This is a competent, carefully bounded CHI-style systems paper; the design (counterbalancing, shared pipeline, blinded rating with ICCs) is above the norm for the venue. The one thing I would want verified before acceptance is the Figure 5 Part A analysis: if the stars were computed on pooled entries, that is a fixable but real statistical error, and the authors should redo the tests at participant level. The qualitative and interaction results stand regardless. No concerns about integrity; the limitations section is unusually honest, which I read as a strength rather than a weakness."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: this is a careful amateur-coaching HCI system, not a CV breakthrough. The useful new pieces are the expert-coded amateur issue taxonomy (304 annotations → 3 domains / 17 descriptors) and a full review workflow that ties those issues to rallies, clips, and court views. The N=16 finding that this package beats a commercial-style stats report on expert-rated reflection quality is credible in direction, weaker in the exact p-values they print.\n\nWhat they did well. Formative needfinding is tight and maps cleanly to DR1–DR3. They are unusually honest about the pipeline: stroke classification needed hand fix, ~25 min verification per 8-min match, trajectories not reviewed, no claim of autonomous raw-video coaching. Both conditions share the same upstream labels and GPT-4.1, so the comparison is about organization and linking, not magic detection. Reflection rating has strong ICCs, condition labels stripped, and the interaction-pattern shift (report-to-evidence vs video hunting) is a nice secondary result. Limitations section mostly names the real gaps.\n\nSoft spots, in proportion. The stress-test on unit of analysis lands. Figure 5A advertises nT=79 vs nB=44 with *** and power ~1.0. Entries nest inside 16 people; treating them as independent inflates significance and mixes “more entries” with “better entries.” Wilcoxon-for-paired is the right family only if they aggregated per participant—the figure’s denominators suggest they did not, or at least did not show it. Secondary issues they already flag: whole-workflow only (no component ablation), different matches per condition, no systematic audit that each generated diagnosis is factually right, and possible partial unblinding if TactiPlay phrasing leaks into written reflections. None of that kills the paper; it means the abstract’s quality deltas should be read as directional, not as precise effect sizes.\n\nCitations and related work look normal for the area (CoachAI, VIRD, Sporthesia, racket-sports viz). No math to break. No code/data, so reproducibility is weak.\n\nWho it’s for: sports HCI and people building amateur coaching products. Not for pure CV or anyone hunting a new learning theory. I’d send it to peer review; a referee should demand participant-level analyses (or mixed models) and a clearer separation of quantity vs quality. Worth engaging if you work adjacent; not mandatory reading otherwise.","headline":"Solid sports-HCI systems paper with a real taxonomy and a cleanly scoped study, but the headline stats likely overstate precision by treating nested reflection entries as independent.","tokens_in":23066,"tokens_out":606,"would_cite":false,"duration_ms":20715,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Organizing amateur badminton match evidence around an expert issue taxonomy, rallies, and linked video produces more frequent, concrete, actionable, and appropriate tactical reflections than a report-and-statistics baseline.","keywords":["Sports Analytics","Racket Sports","Tactical Awareness","Video-Based Review","Badminton","Reflective Learning","Interactive Systems","HCI"],"falsifier":"Run a component-controlled within-subjects study on the same matches with matched report length: if taxonomy organization plus video linking is removed while keeping the same verified labels and generator, expert-rated reflection frequency, concreteness, actionability, and appropriateness should fall back to baseline levels.","tokens_in":22818,"feed_emoji":"🏸","tokens_out":940,"duration_ms":22431,"temperature":0.7,"pith_summary":"Amateur badminton players now record their matches but usually get only aggregate stats or generic tips, so they struggle to turn footage into tactical insight without a coach. This paper argues that the missing piece is not more detection alone but a review workflow that organizes events with a coach-informed issue taxonomy, treats rallies as the natural unit of reflection, and anchors every claim to inspectable video and court views. The authors build that workflow in TactiPlay from a formative study and national-level athletes’ annotations, then compare it in a within-subjects study against a commercial-style report-and-statistics baseline that uses the same upstream event labels. Players produced nearly twice as many reflection entries under TactiPlay, with higher expert-rated concreteness, actionability, and appropriateness, and most shifted from hunting through video to a report-to-evidence verification loop. A sympathetic reader should care because the result suggests a practical path from raw recreational footage to structured reflection-on-action without requiring a human coach at every session.","feed_headline":"Taxonomy-linked match review sharpens amateur tactical reflection","feed_subtitle":"Rally-centered, video-anchored feedback beat stats reports on concreteness and actionability","key_machinery":"The expert-taxonomy-guided, rally-level, video-anchored review workflow: a hierarchical performance-issue taxonomy (Technical Execution, Positioning and Recovery, Tactical Choices) used to structure pipeline outputs into four report topics, rally explainers, subtopic-linked clips, and synchronized court visualizations so players can move between diagnosis and evidence.","core_discovery":"When match evidence is organized into an expert-taxonomy-guided, rally-level, video-anchored review workflow, amateur players produce more frequent reflections that experts rate as more concrete, actionable, and appropriate than those produced with a fixed report-and-statistics interface, even when both systems start from the same verified stroke- and rally-level labels.","pith_inferences":["If trajectory and position traces remain unverified, heatmap-backed spatial claims will be the first place trust breaks when the system is taken fully automatic.","The same scaffolding may matter most for players just above beginner technique, who have strokes but no shared language for recovery and shot choice.","Visualizing counterfactual footwork or return paths on the original clip is the natural next interaction after problem-linked replay.","Without longitudinal cross-match views, players may still confuse one-off errors with habits the taxonomy is meant to surface."],"forward_implications":["Amateur post-match tools should treat rallies as the primary entry point and link every issue card to the corresponding video snippets.","Feedback should separate terminal winning/losing patterns from non-terminal process issues using a coach-derived taxonomy rather than only court-zone technical stats.","Interface design that supports a report-to-evidence loop can replace predominantly manual video search for many recreational players.","Similar taxonomy-plus-video workflows could transfer to other rally-based racket sports after sport-specific taxonomy validation.","Coaches could use such systems as assistants that pre-compile key rallies and issue clusters from underused match recordings."],"fun_headline_variants":["Taxonomy-guided rally review boosts amateur tactical reflection","Video-anchored match evidence yields more actionable player insights","Expert taxonomy links rallies to concrete amateur reflections","Rally-level video workflow outperforms stats-only match reports","Amateurs reflect more appropriately with taxonomy-linked video review"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The better reflections come from the integrated review workflow itself, not from the fact that each player saw two different matches or from joint differences in prompt structure, clustering, linking, and report packaging that the study did not separate.","fun_headline_variants_meta":{"raw":{"variants":["Taxonomy-guided rally review boosts amateur tactical reflection","Video-anchored match evidence yields more actionable player insights","Expert taxonomy links rallies to concrete amateur reflections","Rally-level video workflow outperforms stats-only match reports","Amateurs reflect more appropriately with taxonomy-linked video review"]},"model":"grok-4.5","effort":"low","cost_usd":0.002105,"raw_usage":{"total_tokens":906,"prompt_tokens":724,"num_sources_used":0,"completion_tokens":62,"cost_in_usd_ticks":21048000,"prompt_tokens_details":{"text_tokens":724,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":120,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":724,"tokens_out":62,"duration_ms":4017,"temperature":1.0,"reasoning_tokens":120,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T11:52:43.437485+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Run a component-controlled within-subjects study on the same matches with matched report length: if taxonomy organization plus video linking is removed while keeping the same verified labels and generator, expert-rated reflection frequency, concreteness, actionability, and appropriateness should fall back to baseline levels.","supporting_citations":[],"review_version":1}