CASS achieves 44.4 average mIoU across eight open-vocabulary segmentation benchmarks by injecting DINO's low-rank spectral attention structure into CLIP and adjusting text embeddings with an object presence prior.
Single-stage semantic segmentation from image labels
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Distilling Spectral Graph for Object-Context Aware Open-Vocabulary Semantic Segmentation
CASS achieves 44.4 average mIoU across eight open-vocabulary segmentation benchmarks by injecting DINO's low-rank spectral attention structure into CLIP and adjusting text embeddings with an object presence prior.