Pith. sign in

Visual instruction tuning

2 Pith papers cite this work. Polarity classification is still indexing.

2 Pith papers citing it

fields

cs.CV 1 cs.MA 1

years

2026 2

verdicts

UNVERDICTED 2

representative citing papers

Vision Transformers Need More Than Registers

cs.CV · 2026-02-25 · unverdicted · novelty 6.0

ViTs exhibit lazy aggregation by relying on irrelevant background patches for global semantics, and selectively integrating patch features into the CLS token reduces this effect and improves results across label-, text-, and self-supervision.

citing papers explorer

Showing 2 of 2 citing papers.

  • Vision Transformers Need More Than Registers cs.CV · 2026-02-25 · unverdicted · none · ref 20

    ViTs exhibit lazy aggregation by relying on irrelevant background patches for global semantics, and selectively integrating patch features into the CLS token reduces this effect and improves results across label-, text-, and self-supervision.

  • Multi-Agent Cooperative Learning for Robust Vision-Language Alignment under OOD Concepts cs.MA · 2026-01-11 · unverdicted · none · ref 22

    MACL deploys image, text, name, and coordination agents with message passing and adaptive balancing to achieve 1-5% precision gains on VISTA-Beyond for few-shot and zero-shot OOD alignment.