Trident, a training-free framework combining CLIP, DINO, and SAM, raises state-of-the-art open-vocabulary segmentation mIoU from 44.4 to 48.6 by splicing sub-image features and aggregating them with a SAM affinity matrix.
Segnet: A deep convolutional encoder-decoder architecture for image segmentation
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Harnessing Vision Foundation Models for High-Performance, Training-Free Open Vocabulary Segmentation
Trident, a training-free framework combining CLIP, DINO, and SAM, raises state-of-the-art open-vocabulary segmentation mIoU from 44.4 to 48.6 by splicing sub-image features and aggregating them with a SAM affinity matrix.