Using Florence-2's multi-depth, multi-prompt visual features in an MLLM yields consistent benchmark gains over CLIP-based counterparts, especially on text-heavy tasks.
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
dataset 1
citation-polarity summary
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1roles
dataset 1polarities
use dataset 1representative citing papers
citing papers explorer
-
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion
Using Florence-2's multi-depth, multi-prompt visual features in an MLLM yields consistent benchmark gains over CLIP-based counterparts, especially on text-heavy tasks.