ACE fine-tunes video-language models with stochastically sampled action synonyms and shadow negatives, improving zero-shot classification of unseen procedural actions by up to 16 percent harmonic mean on ATA, IKEA, and GTEA.
Regen: A good generative zero-shot video classifier should be rewarded
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos
ACE fine-tunes video-language models with stochastically sampled action synonyms and shadow negatives, improving zero-shot classification of unseen procedural actions by up to 16 percent harmonic mean on ATA, IKEA, and GTEA.