A token-based autoregressive video detector represents objects as discrete-token 3D tracklets and achieves 91.14 mAP on UA-DETRAC, but its improvement over the static baseline mostly reflects redundant sliding windows, not temporal fusion.
Girshick, ‘‘Fast R-CNN,’’ in ICCV, Dec 2015, pp
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Improving Token-based Object Detection with Video
A token-based autoregressive video detector represents objects as discrete-token 3D tracklets and achieves 91.14 mAP on UA-DETRAC, but its improvement over the static baseline mostly reflects redundant sliding windows, not temporal fusion.