Pith. sign in

REVIEW 1 cited by

TrTr: Visual Tracking with Transformer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2105.03817 v1 pith:VD2WXOGS submitted 2021-05-09 cs.CV

classification cs.CV
keywords featurestrackingtransformertrtrarchitecturecross-correlationimagemethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Template-based discriminative trackers are currently the dominant tracking methods due to their robustness and accuracy, and the Siamese-network-based methods that depend on cross-correlation operation between features extracted from template and search images show the state-of-the-art tracking performance. However, general cross-correlation operation can only obtain relationship between local patches in two feature maps. In this paper, we propose a novel tracker network based on a powerful attention mechanism called Transformer encoder-decoder architecture to gain global and rich contextual interdependencies. In this new architecture, features of the template image is processed by a self-attention module in the encoder part to learn strong context information, which is then sent to the decoder part to compute cross-attention with the search image features processed by another self-attention module. In addition, we design the classification and regression heads using the output of Transformer to localize target based on shape-agnostic anchor. We extensively evaluate our tracker TrTr, on VOT2018, VOT2019, OTB-100, UAV, NfS, TrackingNet, and LaSOT benchmarks and our method performs favorably against state-of-the-art algorithms. Training code and pretrained models are available at https://github.com/tongtybj/TrTr.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. An Angular-Temporal Interaction Network for Light Field Object Tracking in Low-Light Scenes

    cs.CV 2025-07 conditional novelty 5.0 of 10

    ATINet, an angular-temporal interaction network with a novel epipolar-plane structure image representation, achieves the highest reported single and multiple object tracking scores on a new low-light light field dataset.

Pith tools