VLM-HOI distills BLIP image-text matching scores into an HOI detector via a contrastive loss, achieving 34.25 mAP on HICO-DET and 67.7 AP on V-COCO with a ResNet-50 backbone.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis
VLM-HOI distills BLIP image-text matching scores into an HOI detector via a contrastive loss, achieving 34.25 mAP on HICO-DET and 67.7 AP on V-COCO with a ResNet-50 backbone.