Pith. sign in

REVIEW 3 cited by

SparseTT: Visual Tracking with Sparse Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.03776 v1 pith:ZIK6EJKU submitted 2022-05-08 cs.CV

classification cs.CV
keywords trackingtransformersfocusinginformationmechanismmethodperformanceregions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Transformers have been successfully applied to the visual tracking task and significantly promote tracking performance. The self-attention mechanism designed to model long-range dependencies is the key to the success of Transformers. However, self-attention lacks focusing on the most relevant information in the search regions, making it easy to be distracted by background. In this paper, we relieve this issue with a sparse attention mechanism by focusing the most relevant information in the search regions, which enables a much accurate tracking. Furthermore, we introduce a double-head predictor to boost the accuracy of foreground-background classification and regression of target bounding boxes, which further improve the tracking performance. Extensive experiments show that, without bells and whistles, our method significantly outperforms the state-of-the-art approaches on LaSOT, GOT-10k, TrackingNet, and UAV123, while running at 40 FPS. Notably, the training time of our method is reduced by 75% compared to that of TransT. The source code and models are available at https://github.com/fzh0917/SparseTT.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Robust Tracking via Mamba-based Context-aware Token Learning

    cs.CV 2024-12 conditional novelty 6.0 of 10

    TemTrack tracks objects in video by learning temporal context from per-frame 'track tokens' with a Mamba and cross-attention module, achieving competitive accuracy at real-time speed.

  2. Visual Object Tracking across Diverse Data Modalities: A Review

    cs.CV 2024-12 conditional novelty 4.0 of 10

    A survey that taxonomizes deep-learning visual trackers across RGB, thermal, LiDAR, and four multi-modal combinations, with benchmark tables and future directions.

  3. Large Language models for Time Series Analysis: Techniques, Applications, and Challenges

    cs.LG 2025-05 reject novelty 3.0 of 10

    A review of LLM-based time series analysis that proposes several taxonomies, but is undermined by citation errors and a lack of systematic methodology.

Pith tools