REVIEW 2 cited by
SUTrack: Towards Simple and Unified Single Object Tracking
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this paper, we propose a simple yet unified single object tracking (SOT) framework, dubbed SUTrack. It consolidates five SOT tasks (RGB-based, RGB-Depth, RGB-Thermal, RGB-Event, RGB-Language Tracking) into a unified model trained in a single session. Due to the distinct nature of the data, current methods typically design individual architectures and train separate models for each task. This fragmentation results in redundant training processes, repetitive technological innovations, and limited cross-modal knowledge sharing. In contrast, SUTrack demonstrates that a single model with a unified input representation can effectively handle various common SOT tasks, eliminating the need for task-specific designs and separate training sessions. Additionally, we introduce a task-recognition auxiliary training strategy and a soft token type embedding to further enhance SUTrack's performance with minimal overhead. Experiments show that SUTrack outperforms previous task-specific counterparts across 11 datasets spanning five SOT tasks. Moreover, we provide a range of models catering edge devices as well as high-performance GPUs, striking a good trade-off between speed and accuracy. We hope SUTrack could serve as a strong foundation for further compelling research into unified tracking models. Code and models are available at github.com/chenxin-dlut/SUTrack.
Forward citations
Cited by 2 Pith papers
-
What You Have is What You Track: Adaptive and Robust Multimodal Tracking
FlexTrack claims state-of-the-art multimodal tracking on complete and simulated missing-modality benchmarks, using heterogeneous mixture-of-experts fusion and a video-level masking training strategy.
-
Towards Compact Unified Multimodal Tracking: Synergizing Knowledge Distillation with Structural Pruning
A pruned-head student with dual spatial and semantic distillation reaches 54 FPS and near-teacher accuracy on RGB-T and RGB-E tracking.
Discussion (0). Continue with ORCID to comment.