A text-driven BLIP-GroundingDINO-SAM pipeline produces pseudo-labels and a new 260k-image dataset for SOD, with claimed SOTA results that are weakened by a likely PASCAL-S overlap and missing artifacts.
Re-caption: Saliency-enhanced image captioning through two-phase learning,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
REJECT 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Boosting Salient Object Detection with Knowledge Distillated from Large Foundation Models
A text-driven BLIP-GroundingDINO-SAM pipeline produces pseudo-labels and a new 260k-image dataset for SOD, with claimed SOTA results that are weakened by a likely PASCAL-S overlap and missing artifacts.