MObyGaze provides 6072 expert-annotated segments across 43 hours of film, with a multimodal thesaurus, and shows that current vision and text models can detect objectification above chance, while audio models cannot.
Best of Many
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
MObyGaze: a film dataset of multimodal objectification densely annotated by experts
MObyGaze provides 6072 expert-annotated segments across 43 hours of film, with a multimodal thesaurus, and shows that current vision and text models can detect objectification above chance, while audio models cannot.