A three-stage curriculum (concept alignment, instruction tuning, downstream fine-tuning) adapted LLaVA-NeXT-Video to soccer, raising action classification accuracy from 11.8% to 63.5% and the VQA relative score from about 60 to 83.
Soccernet-tracking: Multiple object tracking dataset and benchmark in soccer videos
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Domain Adaptation of VLM for Soccer Video Understanding
A three-stage curriculum (concept alignment, instruction tuning, downstream fine-tuning) adapted LLaVA-NeXT-Video to soccer, raising action classification accuracy from 11.8% to 63.5% and the VQA relative score from about 60 to 83.