A tiny COCO benchmark reports a 90% binary ViT accuracy, but the headline comparison is confounded by task difficulty and by missing code and data.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Pose Matters: Evaluating Vision Transformers and CNNs for Human Action Recognition on Small COCO Subsets
A tiny COCO benchmark reports a 90% binary ViT accuracy, but the headline comparison is confounded by task difficulty and by missing code and data.