AdaVid trains video-language encoders whose hidden dimensions can be stripped down at inference time, matching a standard model at half the FLOPs on EgoMCQ.
Frozen in time: A joint video and image encoder for end-to- end retrieval
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
AdaVid: Adaptive Video-Language Pretraining
AdaVid trains video-language encoders whose hidden dimensions can be stripped down at inference time, matching a standard model at half the FLOPs on EgoMCQ.