Pith. sign in

Generalized Relation Modeling for Transformer Tracking

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Compared with previous two-stream trackers, the recent one-stream tracking pipeline, which allows earlier interaction between the template and search region, has achieved a remarkable performance gain. However, existing one-stream trackers always let the template interact with all parts inside the search region throughout all the encoder layers. This could potentially lead to target-background confusion when the extracted feature representations are not sufficiently discriminative. To alleviate this issue, we propose a generalized relation modeling method based on adaptive token division. The proposed method is a generalized formulation of attention-based relation modeling for Transformer tracking, which inherits the merits of both previous two-stream and one-stream pipelines whilst enabling more flexible relation modeling by selecting appropriate search tokens to interact with template tokens. An attention masking strategy and the Gumbel-Softmax technique are introduced to facilitate the parallel computation and end-to-end learning of the token division module. Extensive experiments show that our method is superior to the two-stream and one-stream pipelines and achieves state-of-the-art performance on six challenging benchmarks with a real-time running speed.

fields

cs.CV 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Decoupled Spatio-Temporal Consistency Learning for Self-Supervised Tracking

cs.CV · 2025-07-29 · conditional · novelty 6.0

SSTrack trains a Vision Transformer tracker without frame-wise box labels by combining forward global search, backward local association, and instance contrastive learning, and reports state-of-the-art self-supervised results on nine tracking benchmarks.

citing papers explorer

Showing 1 of 1 citing paper.

  • Decoupled Spatio-Temporal Consistency Learning for Self-Supervised Tracking cs.CV · 2025-07-29 · conditional · none · ref 20 · internal anchor

    SSTrack trains a Vision Transformer tracker without frame-wise box labels by combining forward global search, backward local association, and instance contrastive learning, and reports state-of-the-art self-supervised results on nine tracking benchmarks.