Pith. sign in

Optimizing Relevance Maps of Vision Transformers Improves Robustness

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

It has been observed that visual classification models often rely mostly on the image background, neglecting the foreground, which hurts their robustness to distribution changes. To alleviate this shortcoming, we propose to monitor the model's relevancy signal and manipulate it such that the model is focused on the foreground object. This is done as a finetuning step, involving relatively few samples consisting of pairs of images and their associated foreground masks. Specifically, we encourage the model's relevancy map (i) to assign lower relevance to background regions, (ii) to consider as much information as possible from the foreground, and (iii) we encourage the decisions to have high confidence. When applied to Vision Transformer (ViT) models, a marked improvement in robustness to domain shifts is observed. Moreover, the foreground masks can be obtained automatically, from a self-supervised variant of the ViT model itself; therefore no additional supervision is required.

fields

cs.LG 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Machine Learning from Explanations

cs.LG · 2025-07-07 · conditional · novelty 6.0

A two-stage optimization pipeline that alternates label loss with a KL divergence between feature maps of masked and unmasked inputs improves accuracy and robustness in small-data classification.

citing papers explorer

Showing 1 of 1 citing paper.

  • Machine Learning from Explanations cs.LG · 2025-07-07 · conditional · none · ref 2 · internal anchor

    A two-stage optimization pipeline that alternates label loss with a KL divergence between feature maps of masked and unmasked inputs improves accuracy and robustness in small-data classification.