A fovea-like patch sampling mechanism lets a 0.51M-parameter model outperform the 641M-parameter SAM-H on six segmentation benchmarks.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation
A fovea-like patch sampling mechanism lets a 0.51M-parameter model outperform the 641M-parameter SAM-H on six segmentation benchmarks.