DisenQ disentangles biometric, motion, and non-biometric features in videos via VLM-generated text supervision, achieving state-of-the-art activity-biometrics identification.
Flamingo: A visual language model for few-shot learning
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
DisenQ: Disentangling Q-Former for Activity-Biometrics
DisenQ disentangles biometric, motion, and non-biometric features in videos via VLM-generated text supervision, achieving state-of-the-art activity-biometrics identification.