VLMs reach only 42.1% exact accuracy on counting pushups in videos, with weaker models exploiting modal counts, and 1k-sample fine-tuning transfers gains to MVBench, PerceptionTest, and TVBench.
Ovr: A dataset for open vocabulary temporal repetition counting in videos
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.CV 2years
2026 2representative citing papers
A motion-only embedding trained on synthetic point tracks matches or beats large appearance-based video models on temporal tasks and improves them when combined.
citing papers explorer
-
PushupBench: Your VLM is not good at counting pushups
VLMs reach only 42.1% exact accuracy on counting pushups in videos, with weaker models exploiting modal counts, and 1k-sample fine-tuning transfers gains to MVBench, PerceptionTest, and TVBench.
-
The TIME Machine: On The Power of Motion for Efficient Perception
A motion-only embedding trained on synthetic point tracks matches or beats large appearance-based video models on temporal tasks and improves them when combined.