Pith. sign in

REVIEW 10 cited by

SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.08474 v4 pith:Y7QI4Z7O submitted 2024-10-11 cs.CV cs.CL

classification cs.CVcs.CL
keywords sportsmodelsreasoningsportuunderstandingevaluatelargemllms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal Large Language Models (MLLMs) are advancing the ability to reason about complex sports scenarios by integrating textual and visual information. To comprehensively evaluate their capabilities, we introduce SPORTU, a benchmark designed to assess MLLMs across multi-level sports reasoning tasks. SPORTU comprises two key components: SPORTU-text, featuring 900 multiple-choice questions with human-annotated explanations for rule comprehension and strategy understanding. This component focuses on testing models' ability to reason about sports solely through question-answering (QA), without requiring visual inputs; SPORTU-video, consisting of 1,701 slow-motion video clips across 7 different sports and 12,048 QA pairs, designed to assess multi-level reasoning, from simple sports recognition to complex tasks like foul detection and rule application. We evaluate four prevalent LLMs mainly utilizing few-shot learning paradigms supplemented by chain-of-thought (CoT) prompting on the SPORTU-text part. We evaluate four LLMs using few-shot learning and chain-of-thought (CoT) prompting on SPORTU-text. GPT-4o achieves the highest accuracy of 71%, but still falls short of human-level performance, highlighting room for improvement in rule comprehension and reasoning. The evaluation for the SPORTU-video part includes 7 proprietary and 6 open-source MLLMs. Experiments show that models fall short on hard tasks that require deep reasoning and rule-based understanding. Claude-3.5-Sonnet performs the best with only 52.6% accuracy on the hard task, showing large room for improvement. We hope that SPORTU will serve as a critical step toward evaluating models' capabilities in sports understanding and reasoning.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. RefereeBench: Are Video MLLMs Ready to be Multi-Sport Referees

    cs.CV 2026-04 unverdicted novelty 8.0 of 10

    RefereeBench shows that even the strongest video MLLMs reach only around 60% accuracy on multi-sport refereeing tasks and struggle with rule application and temporal grounding.

  2. Towards Temporal Compositional Reasoning in Long-Form Sports Videos

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    SportsTime benchmark and CoTR method improve multimodal AI's temporal compositional reasoning and evidence grounding in long-form sports videos.

  3. BoxComm: Benchmarking Category-Aware Commentary Generation and Narration Rhythm in Boxing

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    BoxComm is the first large-scale benchmark for category-aware commentary generation and rhythm assessment in boxing, showing state-of-the-art multimodal models struggle with tactical analysis and temporal pacing.

  4. TennisTV: Do Multimodal Large Language Models Understand Tennis Rallies?

    cs.CV 2025-09 unverdicted novelty 7.0 of 10

    Introduces TennisTV benchmark for evaluating 17 MLLMs on tennis video understanding from stroke-level to rally-level tasks with automated pipelines and human verification.

  5. Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    CREDiT applies counterfactual reasoning via structural causal models to decompose video representations into causal and non-causal parts for more reliable VideoQA on datasets like NExT-GQA and SportsQA.

  6. SoccerRef-Agents: Multi-Agent System for Automated Soccer Refereeing

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    SoccerRef-Agents is a multi-agent framework using MLLMs, cross-modal RAG, and a custom knowledge base that outperforms general MLLMs on soccer foul decisions and explanations.

  7. Towards Temporal Compositional Reasoning in Long-Form Sports Videos

    cs.CV 2026-04 conditional novelty 5.5 of 10

    SportsTime plus Chain-of-Time Reasoning (temporal-reward GRPO and anchor-observe-infer) modestly lifts open-ended sports VideoQA and step-wise temporal grounding over 4B–8B MLLM baselines.

  8. Watch, Remember, Reason: Human-View Video Understanding with MLLMs

    cs.CV 2026-06 unverdicted novelty 4.0 of 10

    This is a survey that frames video MLLM research via a human-view formulation of perceptual representations, memory states, reasoning traces, and predictions, then reviews methods, datasets, benchmarks, and open problems.

  9. SV3.3B: A Sports Video Understanding Model for Action Recognition

    cs.CV 2025-07 reject novelty 4.0 of 10

    A fine-tuned 3.3B video description model using DWT-VGG16-LDA keyframe sampling reports 29.2% higher validation scores than GPT-4o on a 1,315-clip NBA play-by-play subset.

  10. E-FreeM2: Efficient Training-Free Multi-Scale and Cross-Modal News Verification via MLLMs

    cs.MM 2025-06 conditional novelty 4.0 of 10

    A training-free pipeline using image and text retrieval plus two-stage Gemini and GPT-4o mini reasoning reaches 90.0% accuracy on NewsCLIPpings out-of-context detection, but code, prompts, and error bars are missing.

Pith tools