ViTs exhibit lazy aggregation by relying on irrelevant background patches for global semantics, and selectively integrating patch features into the CLS token reduces this effect and improves results across label-, text-, and self-supervision.
Visual instruction tuning
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2verdicts
UNVERDICTED 2representative citing papers
MACL deploys image, text, name, and coordination agents with message passing and adaptive balancing to achieve 1-5% precision gains on VISTA-Beyond for few-shot and zero-shot OOD alignment.
citing papers explorer
-
Vision Transformers Need More Than Registers
ViTs exhibit lazy aggregation by relying on irrelevant background patches for global semantics, and selectively integrating patch features into the CLS token reduces this effect and improves results across label-, text-, and self-supervision.
-
Multi-Agent Cooperative Learning for Robust Vision-Language Alignment under OOD Concepts
MACL deploys image, text, name, and coordination agents with message passing and adaptive balancing to achieve 1-5% precision gains on VISTA-Beyond for few-shot and zero-shot OOD alignment.