Removing class names from LLM-generated descriptions collapses CLIP's zero-shot accuracy, and fine-tuning on synthetic attribute descriptions with a multi-resolution vision layer recovers much of that performance.
Ovarnet: Towards open- vocabulary object attribute recognition
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition
Removing class names from LLM-generated descriptions collapses CLIP's zero-shot accuracy, and fine-tuning on synthetic attribute descriptions with a multi-resolution vision layer recovers much of that performance.