{"id":"f7d85b28-ade4-4749-9a12-4a1913c05918","arxiv_id":"2607.03633","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"Identity-specific free-throw motion signatures exist and are learnable by video models, but models prefer static appearance shortcuts unless silhouettes or skeletons suppress them.","lead":"Modern video models recognize basketball players accurately from free-throw clips but mostly use faces and jerseys, not motion, unless appearance is stripped away. The controlled BALLER120 dataset shows identity-specific shooting kinematics are learnable and more robust once static shortcuts are removed.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Residual non-dynamic identity cues (body proportions after normalization, pose-estimator idiosyncrasies) remain the softest point in claiming pure motion signatures drive silhouette/skeleton success.","rationale":"The reader correctly isolates the residual-acquisition-bias assumption as the weakest link for the motion-signature claim. The paper’s design (controlled free-throws, multi-regime inputs, appearance-disjoint and open-set splits, contour degradation, temporal shuffle, phase-aligned CAM, multi-view stability) supplies converging internal evidence that makes wholesale collapse of the claim unlikely; the residual concern is real but already partially stress-tested and does not overturn the diagnostic contrast between appearance shortcuts and suppressed regimes. Public release of BALLER120 + code remains the practical gate for community verification, so the CONDITIONAL verdict stands without adjustment. No stronger internal inconsistency or hidden assumption was found.","tokens_in":25584,"tokens_out":540,"duration_ms":20526,"concrete_test":"Warp every silhouette (and skeleton) sequence to a single canonical body template via thin-plate-spline on RTMPose keypoints, preserving only relative joint-angle trajectories and velocities; retrain MViTv2 under the appearance-disjoint protocol and re-measure Table 3 Top-1 / mAP. A drop >15 points relative to the original silhouette numbers would confirm residual anatomy leakage; stability would strengthen the pure-motion claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that competitive accuracy and robustness under silhouette/skeleton regimes (Tables 2–3, Fig. 6) plus phase-aligned saliency (Figs. 8–11) arise from identity-specific kinematics rather than residual execution-independent signals that survive the pipeline in §4. Spatial union-crop + common-scale resize removes absolute height/camera distance, and contour degradation (Table 4, Fig. 7) blurs fine boundaries, yet limb proportions, torso aspect, and absolute joint configurations (or RTMPose estimation biases correlated with body type) can still leak identity. Multi-frame static DINOv3 already recovers high accuracy (Table 5: ~90% on silhouettes), so ordered dynamics are helpful but not uniquely necessary. Temporal shuffle (Supp. Table F1) and multi-view consistency help, but do not fully isolate pure motion from multi-pose anatomy. If residual static shape or estimator artifacts dominate, the “motion signatures are learnable yet overlooked” conclusion weakens to “any remaining identity cue after appearance suppression is learnable.”","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"This paper asks whether modern video models use identity-specific motion when such cues are clearly available, or whether they prefer static appearance shortcuts. It introduces BALLER120, a controlled diagnostic set of 4,583 free-throw clips from 120 NBA players, with paired appearance, silhouette, and skeleton regimes. Using MViTv2, VideoMAEv2, and UniFormerV2 fine-tuned for closed-set identity recognition, the authors show near-ceiling accuracy under a standard split for all regimes (Table 2), sharp collapse of appearance models under an appearance-disjoint split while silhouette/skeleton models remain robust (Table 3), competitive open-set transfer for appearance-suppressed inputs (Fig. 6), limited sensitivity to contour degradation (Table 4), and CAM/pose-region analyses indicating phase-aligned, identity-distinct attention under suppression (Figs. 8–11). Complementary probes (static DINOv3 frames, temporal shuffle, action-recognition control) support a task-induced appearance bias: models exploit the easiest predictive signal unless appearance is suppressed.","tokens_in":25895,"tokens_out":1370,"duration_ms":19427,"significance":"If the result holds, the paper makes a clear diagnostic contribution rather than another unconstrained Re-ID leaderboard entry. BALLER120 is carefully designed to reduce action-level and acquisition confounds while pairing appearance with appearance-suppressed regimes, enabling cue attribution that gait and clothing-change Re-ID datasets do not isolate as cleanly. The multi-backbone, multi-split, ablation, and saliency package is unusually thorough for a diagnostic study and yields a falsifiable, practically relevant message: identity-specific execution cues are learnable and more robust under appearance shift, but standard video fine-tuning will often ignore them. Strengths include explicit multi-regime construction, appearance-disjoint and open-set protocols, contour-degradation and temporal-shuffle controls, and transparent positioning as a probe rather than a general benchmark. These make the work useful for biometrics, sports analytics, and video representation research even if residual non-dynamic cues remain partially entangled.","major_comments":[{"comment":"The central claim that silhouette/skeleton success reflects identity-specific motion signatures is only partially isolated from residual execution-independent or multi-pose static anatomy. §4’s union-crop + common-scale resize removes absolute height/camera distance, and Table 4/Fig. 7 show contour degradation barely hurts accuracy, but limb proportions, torso aspect, and phase-specific joint configurations (or RTMPose biases correlated with body type) can still identify players without continuous dynamics. Table 5 is the load-bearing control: multi-frame DINOv3 on silhouettes already reaches ~90% Top-1, so ordered video is helpful but not uniquely necessary. Temporal shuffle (Supp. Table F1) and multi-view consistency help, yet the abstract and §7 still state the result primarily as “motion micro-patterns” / “motion signatures.” Please either (i) reframe claims around execution-dependen","section":"§3–4, Tables 3–5, abstract/§7"},{"comment":"Skeleton results are reported as supporting appearance-suppressed motion learning, but Table 2 and Supp. Table C1 show strong backbone dependence: only MViTv2 remains near ceiling; VideoMAEv2 collapses (~25% Top-1). The paper notes this and supplies saliency hypotheses in Supp. §C–D, but the main-text claim that “skeleton-only inputs” induce a shift toward motion micro-patterns is overstated relative to the architecture-specific evidence. Either restrict the skeleton claim to MViTv2 throughout the abstract/§6.2, or provide a controlled diagnosis (e.g., denser keypoint rendering, temporal joint features, or architecture ablations) showing when sparse skeletons preserve identity-linked dynamics rather than architecture-specific compatibility.","section":"§6.2, Table 2, Supp. Table C1"}],"minor_comments":[{"comment":"Fig. 1 and the teaser caption are effective, but several later figures (e.g., Fig. 9 regional saliency curves, Fig. 6 open-set bars) would benefit from explicit axis units, error bars or seed variance, and a short note on how many clips/identities underlie each bar.","section":"Figs. 6, 9"},{"comment":"The loss combination Lid + Ltri (Eqs. 1–3) is reasonable, but the paper does not report an ablation of classification-only vs. triplet-only training. A short note or table would clarify whether the cue-shift findings depend on the metric-learning term.","section":"§5, Eqs. (1)–(3)"},{"comment":"Table 1’s “Explain. Support” column is informative but slightly overloaded; a one-sentence definition of what counts as phase-level interpretability support would help readers compare BALLER120 to gait sets.","section":"Table 1"},{"comment":"Minor wording: “execution-independent anatomy” vs. “fixed anatomy” is used somewhat interchangeably in §3–4; aligning terminology would reduce ambiguity with residual shape concerns.","section":"§3–4"},{"comment":"The action-recognition control (three-pointers vs. free-throws, Fig. 12) is valuable; stating the number of three-point clips and whether the same players/views are used would improve reproducibility.","section":"§6.4, Fig. 12"}],"recommendation":"major_revision","confidential_remarks":"Solid diagnostic paper with unusually careful experimental design for cs.CV. The residual multi-pose anatomy issue is real but addressable by reframing and one stronger control; I do not see a fatal construction circularity. Fit is good for a vision journal that values diagnostic benchmarks and representation analysis. No concerns about citation pattern or undisclosed novelty."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is simple: when free-throw motion is held constant and appearance is available, modern video backbones take the static shortcut; when you force silhouettes or skeletons, the same architectures pick up player-specific execution and stay robust under jersey shifts. That is a useful, well-supported diagnostic result, not a new recognition SOTA.\n\nWhat is new is BALLER120 itself—120 NBA players, ~38 free-throws each, three aligned regimes (appearance / silhouette / skeleton), plus the appearance-disjoint and open-set probes. The experimental design is careful: multiple Kinetics-pretrained backbones, contour degradation, temporal shuffle, static-frame DINOv3 controls, and pose-region CAM quantification all point the same way. The task-induced appearance-bias contrast (identity vs. free-throw/three-point discrimination) is a nice extra observation. Tables 2–3 and the open-set scaling figure are the load-bearing evidence; the saliency figures make the cue shift readable.\n\nThe softest point is exactly the residual non-dynamic leakage the stress-test flags. Union-box + common-scale resize and contour blur do not erase limb proportions or RTMPose idiosyncrasies, and multi-frame static silhouettes already hit high accuracy. So “pure kinematics” is stronger language than the pipeline strictly justifies. Still, temporal shuffle hurts silhouettes far more than appearance, multi-view consistency holds, and the appearance-disjoint collapse of RGB models is hard to explain by residual shape alone. The claim that motion signatures are present, learnable, and overlooked unless appearance is suppressed survives; the claim that they are the only remaining cue is softer.\n\nCitation pattern is appropriate (Re-ID, gait, Kinetics, CAM). Free parameters are ordinary training knobs. Dataset/code release is not stated in the manuscript I saw—that is the main practical caveat for reuse.\n\nThis is for people who care about what video models actually use in re-ID, gait, and behavioral biometrics, and for anyone building diagnostic probes rather than leaderboards. It deserves a serious referee. I would bring it to reading group and would cite the dataset and the cue-shift finding.","headline":"Clean diagnostic study with a useful new free-throw dataset: models ignore available motion until appearance is stripped; residual shape/pose leakage is real but does not sink the main result.","tokens_in":26524,"tokens_out":538,"would_cite":true,"duration_ms":5784,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Identity-specific motion is learnable from free-throws, but video models skip it for faces and jerseys unless appearance is stripped away.","keywords":["identity recognition","motion signatures","video models","appearance shortcuts","silhouette","skeleton","free-throw","diagnostic benchmark"],"falsifier":"Train and test the same silhouette or skeleton models after deliberately injecting controlled residual cues (fixed camera-angle clusters per player, or pose-estimator noise correlated with identity) and check whether accuracy and saliency maps still track player-specific execution phases; collapse under those controls would falsify the claim that the recovered signal is pure motion signature.","tokens_in":26496,"feed_emoji":"🏀","tokens_out":620,"duration_ms":5419,"temperature":0.7,"pith_summary":"This paper asks a diagnostic question: when identity-specific motion is clearly available, do modern video models actually use it to recognize people? The authors build BALLER120, a controlled set of free-throw clips from 120 professional basketball players, plus matched silhouette and skeleton versions that remove faces, jerseys, and texture. With full RGB input, action-recognition backbones reach near-perfect closed-set accuracy, but saliency and appearance-shift tests show they lean on static cues. When the same architectures are trained only on silhouettes or skeletons, they still achieve competitive accuracy, become far more robust when jerseys change, and attend to phase-aligned micro-patterns such as foot placement, elbow trajectory, and torso timing that differ across players. The study therefore establishes that individual motion signatures in a shared skilled action are present and learnable, yet are easily overshadowed by easier appearance shortcuts unless those shortcuts are deliberately removed.","feed_headline":"Video models ignore free-throw motion until faces are erased","feed_subtitle":"Silhouette and skeleton inputs recover player-specific kinematics that full RGB skips for jersey shortcuts","key_machinery":"BALLER120: a controlled diagnostic dataset of free-throw sequences from 120 NBA players, paired with appearance, silhouette, and skeleton regimes that hold the multi-phase action fixed while suppressing execution-independent cues, allowing direct comparison of what evidence identity classifiers actually use.","core_discovery":"Identity-specific motion signatures in free-throw execution are present, informative, and learnable by modern video backbones, but those models preferentially exploit static appearance shortcuts (faces, jersey regions) when both are available; only explicit appearance suppression forces the same architectures onto distinctive, phase-aligned kinematic patterns that remain competitive and more robust under appearance shift.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Video models skip free-throw motion for face and jersey cues","Silhouettes force free-throw models onto kinematic micro-patterns","Appearance suppression reveals free-throw identity signatures","Skeleton inputs make free-throw kinematics competitive and robust","Models favor static free-throw shortcuts until appearance is erased"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That free-throw differences among professional players, after cropping, scale normalization, and mask or skeleton abstraction, mainly reflect stable individual movement habits rather than leftover capture biases still tied to identity.","fun_headline_variants_meta":{"raw":{"variants":["Video models skip free-throw motion for face and jersey cues","Silhouettes force free-throw models onto kinematic micro-patterns","Appearance suppression reveals free-throw identity signatures","Skeleton inputs make free-throw kinematics competitive and robust","Models favor static free-throw shortcuts until appearance is erased"]},"model":"grok-4.5","effort":"low","cost_usd":0.005694,"raw_usage":{"total_tokens":1559,"prompt_tokens":819,"num_sources_used":0,"completion_tokens":82,"cost_in_usd_ticks":56940000,"prompt_tokens_details":{"text_tokens":819,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":658,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":819,"tokens_out":82,"duration_ms":5923,"temperature":1.0,"reasoning_tokens":658,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T01:02:25.667355+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Train and test the same silhouette or skeleton models after deliberately injecting controlled residual cues (fixed camera-angle clusters per player, or pose-estimator noise correlated with identity) and check whether accuracy and saliency maps still track player-specific execution phases; collapse under those controls would falsify the claim that the recovered signal is pure motion signature.","supporting_citations":[],"review_version":1}