TAMT pre-trains on a source video dataset and fine-tunes small temporal-aware adapters plus a global moment-pooling head on the target, reporting new state-of-the-art cross-domain few-shot action recognition results.
ViViT: A video vision transformer
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
TAMT: Temporal-Aware Model Tuning for Cross-Domain Few-Shot Action Recognition
TAMT pre-trains on a source video dataset and fine-tunes small temporal-aware adapters plus a global moment-pooling head on the target, reporting new state-of-the-art cross-domain few-shot action recognition results.