A unified query-based model that alternately refines object and relationship representations achieves new best mAP on VidVRD and VidOR open-vocabulary video relationship detection.
Align and prompt: Video-and-language pre-training with entity prompts
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection
A unified query-based model that alternately refines object and relationship representations achieves new best mAP on VidVRD and VidOR open-vocabulary video relationship detection.