Pith. sign in

SoccerMaster: A Vision Foundation Model for Soccer Understanding

2 Pith papers cite this work. Polarity classification is still indexing.

2 Pith papers citing it
abstract

Soccer understanding has recently garnered growing research interest due to its domain-specific complexity and unique challenges. Unlike prior works that typically rely on isolated, task-specific expert models, this work aims to propose a unified model to handle diverse soccer visual understanding tasks, ranging from fine-grained perception (e.g., athlete detection and identification) to high-level semantic reasoning (e.g., event classification). Concretely, our contributions are threefold: (i) we present SoccerMaster, the first soccer-specific vision foundation model that unifies diverse tasks within a single framework via supervised multi-task pretraining; (ii) we develop an automated data curation pipeline, SoccerFactory, to generate scalable spatial annotations, and integrate multiple existing soccer video datasets as a comprehensive pretraining data resource for multi-task pretraining; and (iii) we conduct extensive evaluations demonstrating that SoccerMaster consistently outperforms task-specific expert models across diverse downstream tasks, highlighting its breadth and superiority. The data, code, and model will be publicly available.

citation-role summary

background 2

citation-polarity summary

fields

cs.CV 2

years

2026 2

roles

background 2

polarities

background 2

representative citing papers

citing papers explorer

Showing 2 of 2 citing papers.

  • SoccerLens: Grounded Soccer Video Understanding Beyond Accuracy cs.CV · 2026-05-10 · unverdicted · none · ref 32 · 2 links · internal anchor

    SoccerLens benchmark shows state-of-the-art soccer VLMs achieve high classification accuracy yet fail to exceed 50% visual grounding performance and underutilize temporal information.

  • Towards Temporal Compositional Reasoning in Long-Form Sports Videos cs.CV · 2026-04-24 · conditional · none · ref 44 · internal anchor

    SportsTime plus Chain-of-Time Reasoning (temporal-reward GRPO and anchor-observe-infer) modestly lifts open-ended sports VideoQA and step-wise temporal grounding over 4B–8B MLLM baselines.