REVIEW 4 major objections 6 minor 132 references
A Deep Dive into Generic Object Tracking: A Survey
T0 review · 4 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This survey claims the entire generic object tracking literature can be organized into three design paradigms—Siamese matching, discriminative online learning, and transformer-based architectures—and that a unified taxonomy with…
desk verdict A genuinely useful three-paradigm survey whose unsourced Figure 29 comparison needs a data table and protocol before it can be trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing device is a two-axis categorization: paradigm (discriminative, Siamese, hybrid transformer, fully transformer) crossed with functional dimensions such as appearance model, backbone, template update, and architectural level of contribution. The comparison engine is Figure 29, an AUC-versus-FPS scatter on LaSOT that maps each tracker's accuracy against its speed. The standardized diagrams serve as the qualitative counterpart, letting readers compare design flow across trackers without re-reading each paper.
What would settle it
Recompute LaSOT AUC and FPS for a dozen representative trackers under one common protocol and hardware; if the paradigm-level ordering of fully transformer above hybrid above Siamese changes materially, the survey's central performance claim is disproved.
Extended reading notes
Core claim
The central claim is that generic object trackers can be organized into four groups—discriminative, Siamese, hybrid transformer, and fully transformer—and that doing so reveals a clear accuracy-versus-speed landscape. On LaSOT, fully transformer-based trackers occupy the high-accuracy region at moderate speed, hybrid transformer models sit in the middle, Siamese trackers are fast but less accurate, and discriminative trackers are comparatively slow. The paper's contribution is the organization itself: a unified taxonomy, standardized reconstructed architecture diagrams, and multi-dimensional tables comparing appearance model, backbone, update strategy, novelty, and drawbacks.
Load-bearing premise
The accuracy and speed numbers plotted in Figure 29 are taken as comparable across papers, even though those papers likely differ in training data, evaluation settings, and hardware.
Editorial extensions
If this is right
- If the taxonomy holds, newcomers can locate any tracker by paradigm and immediately know its update mechanism, backbone, and likely accuracy-speed profile.
- The LaSOT scatter implies that fully transformer trackers are the current accuracy ceiling, so future work aiming at state-of-the-art performance should start from attention-based architectures.
- The standardized diagrams make cross-paradigm architectural borrowing, such as adding transformer fusion to Siamese or discriminative trackers, easier to identify and replicate.
- The gap analysis suggests that hybrid and fully transformer trackers still pay a speed penalty, making efficiency a concrete open target for the next generation of trackers.
Reading between the lines
- A stricter criterion for the pure-attention category would reclassify some entries: STARK uses a CNN backbone but is labeled convolution-attention based on its head choice, so the boundary between convolution-attention and pure attention is drawn by prediction head rather than backbone.
- The LaSOT-only comparison may not generalize to long-term benchmarks like LTB-35 or OxUvA, where re-detection ability matters more than per-frame AUC; extending the same scatter to long-term datasets would test whether transformer dominance persists.
- The paper's own drawback tables could be mined into a per-challenge failure map covering occlusion, distractors, and long-term disappearance, which would be a testable extension of the taxonomy.
- Because the survey reports no unified re-implementation protocol, a reader should treat the plotted AUC and FPS values as approximate literature-reported numbers rather than measurements from a single controlled run.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a survey of generic object tracking (GOT), organized into discriminative-based, Siamese-based, hybrid transformer-based, and fully transformer-based trackers. It proposes a unified taxonomy, offers standardized architectural diagrams for representative methods, summarizes benchmarks and metrics, and provides a Section 4.3 performance comparison of AUC versus FPS on LaSOT. The paper claims to be among the first comprehensive surveys that jointly review all three major paradigms, with particular emphasis on transformer-based methods.
Significance. If its comparisons were properly sourced and its references corrected, the survey would be a useful reference for practitioners and researchers entering the field, as it covers a broad span of methods from MOSSE (2010) through prompt-based trackers such as PiVOT (2025). The paper's strengths include its breadth, the unified architectural diagrams, the functional categorization in Table 8, and the attempt to compare accuracy and efficiency across paradigms. However, the central quantitative contribution, Figure 29, is presented without any provenance or protocol, and several reference errors undermine the reliability of the survey as a reference work. There are no derivations or fitted parameters to check, so the main burden falls on completeness, accuracy of citations, and empirical transparency.
major comments (4)
- [Section 4.3, Figure 29] The AUC-versus-FPS scatter plot on LaSOT has no point-by-point provenance: there is no mapping from each plotted point to a paper or model variant, no statement of the evaluation split or LaSOT protocol used, no hardware or FPS measurement details, and no error bars. Because the original papers report AUC under different training settings and FPS on different GPUs, the plot may combine incomparable numbers, and the paradigm-level conclusions drawn from it are not verifiable or reproducible. A supplementary table listing, for every point, the tracker name, paper reference, model variant, training data, evaluation subset, reported AUC, reported FPS, and hardware configuration is needed; alternatively, the quantitative comparison should be removed or substantially qualified.
- [Section 3.3.2 / Reference [25]] The text describes SimTrack as a simplified one-branch transformer tracker, but reference [25] points to 'Exploring Simple 3D Multi-Object Tracking for Autonomous Driving' (arXiv:2108.10312), which is an unrelated paper about 3D multi-object tracking. This is not a formatting issue: a reader cannot locate the described method from the bibliography. The reference must be replaced with the correct SimTrack paper, and the in-text description should be checked against it.
- [Table 3 / Figure 5] Table 3 labels SiamFC++ as [12], but reference [12] is SiamFC; the correct reference for SiamFC++ is [16] (Xu et al., AAAI 2020). Figure 5's caption uses [16] for SiamFC++, so the table is internally inconsistent with the figure. This citation mismatch directly affects the reliability of the tabular comparison.
- [References / Section 3.3.1] TransT [45] is cited as 'Transtrack: Multiple Object Tracking with Transformer' (arXiv:2012.15460), which is a different paper. TransT should be cited as Chen et al., 'Transformer Tracking', CVPR 2021. Since the survey's contribution is partly bibliographic, this kind of substitution is a substantive error that requires correction.
minor comments (6)
- [Section 1 / Section 7] The introduction states that the survey reviews 'three major families' of trackers, while the concluding remarks refer to 'four major paradigms.' This inconsistency should be resolved by consistently counting hybrid and fully transformer-based trackers as separate categories or as subcategories of one transformer family.
- [Tables 2-6] Several table entries contain formatting artifacts such as leading spaces in 'F eature', 'F ocus', 'T rack', and 'T emplate', and Table 5 contains the phrase 'Introduction of NOTU largescake benchmark,' which appears to be a typo. A careful proofreading pass over the tables is needed.
- [Section 3.2 / Figure 5] The caption of Figure 5 refers to 'SimaRPN++' instead of 'SiamRPN++', and the text describes 'SiamFC++ [16]' while Table 3 uses '[12]'; these inconsistencies should be corrected together.
- [Section 3.1 / Reference [60]] The text credits 'previous adaptive CF methods such as ASEF [60]' to reference [60], but [60] is the same Bolme et al. MOSSE paper as [1]; the ASEF paper is a different work and should be cited separately.
- [References / Section 4.1] Reference [93] for VOT2016 lists 'G. Roffo, S. Melzi, et al.' as authors; the VOT2016 challenge paper is by Kristan et al. This reference should be corrected.
- [Section 3.3.2 / Figure 1] STMTrack [24] appears in the timeline (Figure 1) and in Table 8, but it is not described in any section of the survey. If it is intended to be a reviewed method, a description and architectural analysis should be added; otherwise it should be removed from the timeline and tables.
Circularity Check
No significant circularity: the survey organizes external results and derives none of its claims from its own inputs.
full rationale
This is a literature survey with no formal derivation chain, no fitted parameters, and no quantity that is predicted from its own inputs. The central contributions are a taxonomy, qualitative method descriptions, reconstructed architectural diagrams, and a performance comparison. The taxonomy is an organizational scheme, not a derived result: the paper classifies trackers into paradigms based on stated architectural principles and explicitly contrasts its categorization with existing surveys, so the classification is not circular. The architectural diagrams are descriptive reconstructions of published methods and do not feed back into any claim that depends on them by definition. The only self-citation, [71], is used as a general reference for vision transformer robustness in one sentence about attention mechanisms and is not load-bearing for any contribution. The AUC-vs-FPS scatter plot in Figure 29 does lack a disclosed per-point provenance table, evaluation protocol, and hardware details, which is a legitimate reproducibility and verifiability weakness, but it is not a circularity failure: the plotted numbers are taken from external benchmark results and original papers rather than generated by the survey's own model or fitted from a subset of data and then re-predicted. No step in the paper reduces, by construction or self-citation, to its own inputs. The honest finding is therefore no significant circularity, score 0.
Assumptions & free parameters
assumptions (2)
- domain assumption The cited papers are accurately represented and their reported results are trustworthy.
- domain assumption Performance numbers from different papers are directly comparable despite different evaluation protocols.
Cite this review
Pith. "Pith review of A Deep Dive into Generic Object Tracking: A Survey." pith.science (2026). https://pith.science/paper/H7GENZA6
@misc{pith2026250723251,
author = {Pith},
title = {Pith review of: A Deep Dive into Generic Object Tracking: A Survey},
year = {2026},
howpublished = {\url{https://pith.science/paper/H7GENZA6}},
note = {Machine review of arXiv:2507.23251}
}
read the original abstract
Generic object tracking remains an important yet challenging task in computer vision due to complex spatio-temporal dynamics, especially in the presence of occlusions, similar distractors, and appearance variations. Over the past two decades, a wide range of tracking paradigms, including Siamese-based trackers, discriminative trackers, and, more recently, prominent transformer-based approaches, have been introduced to address these challenges. While a few existing survey papers in this field have either concentrated on a single category or widely covered multiple ones to capture progress, our paper presents a comprehensive review of all three categories, with particular emphasis on the rapidly evolving transformer-based methods. We analyze the core design principles, innovations, and limitations of each approach through both qualitative and quantitative comparisons. Our study introduces a novel categorization and offers a unified visual and tabular comparison of representative methods. Additionally, we organize existing trackers from multiple perspectives and summarize the major evaluation benchmarks, highlighting the fast-paced advancements in transformer-based tracking driven by their robust spatio-temporal modeling capabilities.
Figures
Figures from the paper (26 more)
Reference graph
Works this paper leans on
-
[25]
Exploring Simple 3D Multi-Object Tracking for Autonomous Driving
Y. Zhang, Y. Wang, X. Wang, H. Li, Exploring simple 3d multi-object tracking for autonomous driving, arXiv preprint arXiv:2108.10312 (2021)
work page Pith review arXiv 2021
-
[12]
Bertinetto, J
L. Bertinetto, J. Valmadre, J. F. Henriques, A. Vedaldi, P. H. Torr, Fully-convolutional siamese networks for object tracking, in: ECCV Workshops, 2016, pp. 850–865
2016
-
[16]
Y. Xu, Z. Wang, Z. Li, Y. Yuan, Siamfc++: Towards robust and accurate visual tracking with target estimation guidelines, in: AAAI, volume 34, 2020, pp. 12549–12556
2020
- [45]
- [1]
-
[2]
J. F. Henriques, R. Caseiro, P. Martins, J. Batista, High-speed tracking with kernelized cor- relation filters, IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) 37 (2015) 583–596. doi: 10.1109/TPAMI.2014.2345390
arXiv 2015
-
[3]
M. Danelljan, G. Hager, F. Shahbaz Khan, M. Felsberg, Learning spatially regularized correlation filters for visual tracking, in: Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2015, pp. 4310–4318. doi: 10.1109/ICCV.2015.490
-
[4]
H. K. Galoogahi, A. Fagg, S. Lucey, Learning background-aware correlation filters for visual tracking, in: Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2017, pp. 1135–1143. doi: 10.1109/ICCV.2017.127
Show all 132 references
-
[5]
H. Nam, B. Han, Learning multi-domain convolutional neural networks for visual tracking, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 4293–4302. doi: 10.1109/CVPR.2016.464
2016 doi
-
[6]
Y. Li, J. Zhu, S. C. Hoi, Deep discriminative correlation filter learning for visual tracking, Pattern Recognition 94 (2019) 322–332. doi: 10.1016/j.patcog.2019.05.014
2019 doi
-
[7]
Valmadre, L
J. Valmadre, L. Bertinetto, J. F. Henriques, A. Vedaldi, P. H. Torr, End-to-end representation learning for correlation filter based tracking, in: Proceedings of the IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR), 2017, pp. 2805–2813. doi: 10.1109/CVPR.2017. 299
2017 doi
-
[8]
Danelljan, G
M. Danelljan, G. Bhat, F. S. Khan, M. Felsberg, Atom: Accurate tracking by overlap maxi- mization, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 4660–4669. doi: 10.1109/CVPR.2019.00479
2019
-
[9]
G. Bhat, M. Danelljan, L. Van Gool, R. Timofte, Learning discriminative model prediction for tracking, in: Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2019, pp. 6182–6191. doi: 10.1109/ICCV.2019.00628
2019
-
[10]
Danelljan, L
M. Danelljan, L. Van Gool, R. Timofte, Probabilistic regression for visual tracking, in: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 7183–7192. doi: 10.1109/CVPR42600.2020.00721
2020
-
[11]
Mayer, M
C. Mayer, M. Danelljan, L. Van Gool, R. Timofte, Learning to track multiple objects with a single tracker, in: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 6297–6307. doi: 10.1109/ICCV48922.2021.00623
2021
-
[13]
B. Li, J. Yan, W. Wu, Z. Zhu, X. Hu, High performance visual tracking with siamese region proposal network, in: CVPR, 2018, pp. 8971–8980
2018
-
[14]
Z. Chen, B. Zhong, G. Li, S. Zhang, R. Ji, Siamban: Siamese box adaptive network for visual tracking, in: CVPR, 2020, pp. 6668–6677. 47
2020
-
[15]
A. He, C. Luo, X. Tian, W. Zeng, Siamese network with spatial attention for visual tracking, in: CVPR, 2018, pp. 9351–9360
2018
-
[17]
B. Li, W. Wu, Q. Wang, F. Zhang, J. Xing, J. Yan, Siamrpn++: Evolution of siamese visual tracking with very deep networks, in: CVPR, 2019, pp. 4282–4291
2019
-
[18]
Y. Yu, Y. Xiong, W. Huang, M. R. Scott, Deformable siamese attention networks for visual object tracking, in: CVPR, 2020, pp. 6727–6736
2020
-
[19]
Voigtlaender, J
P. Voigtlaender, J. Luiten, P. H. Torr, B. Leibe, Siam r-cnn: Visual tracking by re-detection, in: CVPR, 2020, pp. 6577–6587
2020
-
[20]
Z. Zhu, Q. Wang, L. Bo, W. Wu, J. Yan, X. Hu, Distractor-aware siamese networks for visual object tracking, in: ECCV, 2018, pp. 103–119
2018
-
[21]
Y. Su, X. Yang, C. Ma, Siamdmu: Dual mask update for template adaptation in siamese trackers, IEEE Transactions on Emerging Topics in Computational Intelligence 8 (2024) 1658–1668
2024
-
[22]
B. Yan, H. Peng, J. Fu, D. Wang, H. Lu, Learning spatio-temporal transformer for visual tracking, arXiv preprint arXiv:2103.17154 (2021)
2021 arXiv
-
[23]
L. Lin, H. Fan, Z. Zhang, Y. Xu, H. Ling, Swintrack: A simple and strong baseline for transformer tracking, in: Advances in Neural Information Processing Systems (NeurIPS), 2022
2022
-
[24]
Z. Fu, Q. Liu, Z. Fu, Y. Wang, Stmtrack: Template-free visual tracking with space-time memory networks, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 13774–13783
2021
-
[26]
B. Ye, H. Chang, B. Ma, S. Shan, X. Chen, Joint feature learning and relation modeling for tracking: A one-stream framework, in: Proceedings of the European Conference on Computer Vision (ECCV), 2022
2022
-
[27]
K. Song, Y. Wang, M. Li, Y. Zhang, Transformer tracking with cyclic shifting window attention, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 12345–12354
2022
-
[28]
S. Gao, C. Zhou, C. Ma, X. Wang, J. Yuan, Aiatrack: Attention in attention for transformer visual tracking, in: Proceedings of the European Conference on Computer Vision (ECCV), 2022
2022
-
[29]
N. Wang, W. Zhou, J. Wang, H. Li, Correlation-embedded transformer tracking: A single-branch architecture, arXiv preprint arXiv:2401.12743 (2023)
2023 arXiv
-
[30]
Y. Cui, C. Jiang, L. Wang, G. Wu, Mixformer: End-to-end tracking with iterative mixed atten- tion, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 13608–13618
2022
-
[31]
Q. Wu, T. Yang, Z. Liu, B. Wu, Y. Shan, A. B. Chan, Dropmae: Masked autoencoders with spatial-attention dropout for tracking tasks, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 14561–14571. 48
2023
-
[32]
D. Yang, J. He, Y. Ma, Q. Yu, T. Zhang, Foreground-background distribution modeling trans- former for visual object tracking, in: Proceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 10117–10127
2023
-
[33]
H. Zhao, D. Wang, H. Lu, Representation learning for visual object tracking by masked appear- ance transfer, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 18696–18705
2023
-
[34]
X. Wei, Y. Bai, Y. Zheng, D. Shi, Y. Gong, Autoregressive visual tracking, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 9697–9706
2023
-
[35]
Y. Cui, T. Song, G. Wu, L. Wang, Mixformerv2: Efficient fully transformer tracking, Advances in neural information processing systems 36 (2023) 58736–58751
2023
-
[36]
X. Chen, H. Peng, D. Wang, H. Lu, H. Hu, Seqtrack: Sequence to sequence learning for visual object tracking, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 14572–14581
2023
-
[37]
S. Gao, C. Zhou, J. Zhang, Generalized relation modeling for transformer tracking, in: Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 18686–18695
2023
-
[38]
Y. Cai, J. Liu, J. Tang, G. Wu, Robust object modeling for visual tracking, in: Proceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 9589–9600
2023
-
[39]
F. Xie, L. Chu, J. Li, Y. Lu, C. Ma, Videotrack: Learning to track objects via video transformer, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 22826–22835
2023
-
[40]
J. Xie, B. Zhong, Z. Mo, S. Zhang, L. Shi, S. Song, R. Ji, Autoregressive queries for adaptive tracking with spatio-temporal transformers, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 19300–19309
2024
-
[41]
Zheng, B
Y. Zheng, B. Zhong, Q. Liang, Z. Mo, S. Zhang, X. Li, Odtrack: Online dense temporal token learning for visual tracking, in: Proceedings of the AAAI conference on artificial intelligence, volume 38, 2024, pp. 7588–7596
2024
-
[42]
L. Hong, S. Yan, R. Zhang, W. Li, X. Zhou, P. Guo, K. Jiang, Y. Chen, J. Li, Z. Chen, et al., Onetracker: Unifying visual object tracking with foundation models and efficient tuning, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, p...
2024
-
[43]
L. Gao, L. Chen, P. Liu, Y. Jiang, Y. Li, J. Ning, Transformer-based visual object tracking via fine–coarse concatenated attention and cross concatenated mlp, Pattern Recognition 146 (2024) 109964
2024
-
[44]
Chen, J.-C
S.-F. Chen, J.-C. Chen, I.-H. Jhuo, Y.-Y. Lin, Improving visual object tracking through visual prompting, IEEE Transactions on Multimedia (2025)
2025
-
[46]
N. Wang, W. Zhou, J. Wang, H. Li, Transformer meets tracker: Exploiting temporal context for robust visual tracking, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 8124–8133. 49
2021
-
[47]
Mayer, M
C. Mayer, M. Danelljan, G. Bhat, M. Paul, D. P. Paudel, F. Yu, L. Van Gool, Transforming model prediction for tracking, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 8731–8740
2022
-
[48]
Mayer, M
C. Mayer, M. Danelljan, G. Bhat, D. P. Paudel, L. Van Gool, Beyond sot: Tracking multiple generic objects at once, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (W ACV), 2024, pp. 1234–1243
2024
-
[49]
Zhang, Z
Y. Zhang, Z. Wang, M. Li, W. Liu, X. Wang, Cmat: Integrating convolution mixer and self- attention for visual tracking, IEEE Transactions on Multimedia 25 (2023) 1234–1245
2023
-
[50]
J. Li, W. Chen, M. Zhao, L. Wang, Reading relevant feature from global representation memory for visual object tracking, in: Advances in Neural Information Processing Systems (NeurIPS), 2023
2023
-
[51]
S. M. Marvasti-Zadeh, L. Cheng, H. Ghanei-Yakhdan, S. Kasaei, Deep learning for visual track- ing: A comprehensive survey, IEEE Transactions on Intelligent Transportation Systems 22 (2021) 3782–3804. doi:10.1109/TITS.2020.3046478
2021
-
[52]
Y. Li, J. Zhu, S. C. Hoi, Recent advances of single-object tracking methods: A brief survey, Neurocomputing 492 (2022) 318–329. doi: 10.1016/j.neucom.2021.05.011
2022 doi
-
[53]
C. Li, B. Yang, C. Li, Deep learning based visual tracking: A review, Neurocomputing 275 (2018) 2471–2480. doi: 10.1016/j.neucom.2017.10.070
2018 doi
-
[54]
M. Y. Abbass, K.-C. Kwon, N. Kim, S. A. Abdelwahab, F. E. Abd El-Samie, A. A. M. Khalaf, A survey on online learning for visual tracking, The Visual Computer 36 (2020) 993–1014. doi:10.1007/s00371-020-01848-y
2020 doi
-
[55]
Javed, M
S. Javed, M. Danelljan, F. S. Khan, M. H. Khan, M. Felsberg, J. Matas, Visual object tracking with discriminative filters and siamese networks: A survey and outlook, IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (2023) 1–1. doi: 10.1109/TPAMI.2022.3212594
2023
-
[56]
Ondraˇ soviˇ c, P
M. Ondraˇ soviˇ c, P. Taraba, Siamese visual object tracking: A survey, Electronics 10 (2021) 1876. doi:10.3390/electronics10151876
2021 doi
-
[57]
Zhang, Y
Y. Zhang, Y. Wang, Y. Wang, Y. Wang, Y. Wang, Visual object tracking: A survey, Computer Vision and Image Understanding 210 (2022) 103508. doi: 10.1016/j.cviu.2021.103508
2022
-
[58]
Thangavel, T
J. Thangavel, T. Kokul, A. Ramanan, S. Fernando, Transformers in single object tracking: An experimental survey, IEEE Access 11 (2023) 80297–80326. doi: 10.1109/ACCESS.2023.3237614
2023
-
[59]
Abdelaziz, M
O. Abdelaziz, M. Shehata, M. Mohamed, Beyond traditional single object tracking: A survey, arXiv preprint arXiv:2405.10439 (2024). URL: https://arxiv.org/abs/2405.10439
2024 arXiv
-
[60]
D. S. Bolme, J. R. Beveridge, B. A. Draper, Y. M. Lui, Visual object tracking using adaptive correlation filters, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, 2010, pp. 2544–2550
2010
-
[61]
Simonyan, A
K. Simonyan, A. Zisserman, Very deep convolutional networks for large-scale image recognition, in: International Conference on Learning Representations, 2015. URL: https://arxiv.org/ abs/1409.1556
2015 arXiv
-
[62]
K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, 2016, pp. 770–778. 50
2016
-
[63]
D. Held, S. Thrun, S. Savarese, Learning to track at 100 fps with deep regression networks, in: European Conference on Computer Vision (ECCV), Springer, 2016, pp. 749–765
2016
-
[64]
D. Guo, J. Xu, H. Zhu, Z. Huang, Learning dynamic siamese network for visual object tracking, in: ICCV, 2017
2017
-
[65]
Zhang, H
Z. Zhang, H. Peng, J. Fu, B. Li, W. Hu, Ocean: Object-aware anchor-free tracking, in: ECCV, 2020, pp. 771–787
2020
-
[66]
H. Chen, L. Zhang, et al., Enhanced correlation information mixer for siamese visual tracking, Knowledge-Based Systems 285 (2024) 111368
2024
-
[67]
Krizhevsky, I
A. Krizhevsky, I. Sutskever, G. E. Hinton, Imagenet classification with deep convolutional neural networks, Advances in neural information processing systems 25 (2012)
2012
-
[68]
A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, H. Adam, Mobilenets: Efficient convolutional neural networks for mobile vision applications, arXiv preprint arXiv:1704.04861 (2017)
2017 arXiv
-
[69]
Szegedy, W
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, A. Rabinovich, Going deeper with convolutions, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 1–9
2015
-
[70]
Szegedy, V
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, Z. Wojna, Rethinking the inception architecture for computer vision, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 2818–2826
2016
-
[71]
Alijani, J
S. Alijani, J. Fayyad, H. Najjaran, Vision transformers in domain adaptation and domain generalization: a study of robustness, Neural Computing and Applications 36 (2024) 17979– 18007
2024
-
[72]
Dosovitskiy, L
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, N. Houlsby, An image is worth 16x16 words: Transformers for image recognition at scale, in: International Conference on Learning Repr...
2021
-
[73]
Carion, F
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, S. Zagoruyko, End-to-end object detection with transformers, in: Proceedings of the European Conference on Computer Vision (ECCV), 2020, pp. 213–229
2020
-
[74]
Huang, X
L. Huang, X. Zhao, K. Huang, Got-10k: A large high-diversity benchmark for generic object tracking in the wild, IEEE Transactions on Pattern Analysis and Machine Intelligence 43 (2021) 1562–1577
2021
-
[75]
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, B. Guo, Swin transformer: Hierarchi- cal vision transformer using shifted windows, in: Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 10012–10022
2021
-
[76]
H. Wu, B. Xiao, N. Codella, M. Liu, X. Dai, L. Yuan, L. Zhang, Cvt: Introducing convolutions to vision transformers, in: Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 22–31
2021
-
[77]
Zhang, Y
X. Zhang, Y. Tian, L. Xie, W. Huang, Q. Dai, Q. Ye, Q. Tian, Hivit: A simpler and more efficient design of hierarchical vision transformer, in: The Eleventh International Conference on Learning Representations, 2023. 51
2023
-
[78]
Radford, J
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al., Learning transferable visual models from natural language supervi- sion, in: International conference on machine learning, PmLR, 2021, pp. 8748–8763
2021
-
[79]
Dosovitskiy, L
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, N. Houlsby, An image is worth 16x16 words: Transformers for image recognition at scale, in: International Conference on Learning Repr...
2021
-
[80]
Y. Wu, J. Lim, M.-H. Yang, Online object tracking: A benchmark, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2013, pp. 2411–2418
2013
-
[81]
Y. Wu, J. Lim, M.-H. Yang, Object tracking benchmark, IEEE Transactions on Pattern Analysis and Machine Intelligence 37 (2015) 1834–1848
2015
-
[82]
Liang, E
P. Liang, E. Blasch, H. Ling, Encoding color information for visual tracking: Algorithms and benchmark, IEEE transactions on image processing 24 (2015) 5630–5644
2015
-
[83]
A. W. Smeulders, D. M. Chu, R. Cucchiara, S. Calderara, A. Dehghan, M. Shah, Visual tracking: An experimental survey, IEEE transactions on pattern analysis and machine intelligence 36 (2013) 1442–1468
2013
-
[84]
Muller, A
M. Muller, A. Bibi, S. Giancola, S. Alsubaihi, B. Ghanem, Trackingnet: A large-scale dataset and benchmark for object tracking in the wild, in: Proceedings of the European conference on computer vision (ECCV), 2018, pp. 300–317
2018
-
[85]
A. Li, M. Lin, Y. Wu, M.-H. Yang, S. Yan, Nus-pro: A new visual tracking challenge, IEEE transactions on pattern analysis and machine intelligence 38 (2015) 335–349
2015
-
[86]
Kiani Galoogahi, A
H. Kiani Galoogahi, A. Fagg, S. Lucey, Need for speed: A benchmark for higher frame rate object tracking, Proceedings of the IEEE International Conference on Computer Vision (ICCV) (2017) 1125–1134
2017
-
[87]
H. Fan, F. Yang, P. Chu, Y. Lin, L. Yuan, H. Ling, Tracklinic: Diagnosis of challenge factors in visual tracking, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2021, pp. 970–979
2021
-
[88]
Valmadre, L
J. Valmadre, L. Bertinetto, J. F. Henriques, A. Vedaldi, P. H. Torr, Long-term tracking in the wild: A benchmark, in: Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 650–666
2018
-
[89]
Moudgil, V
A. Moudgil, V. Gandhi, Long-term visual object tracking benchmark, Proceedings of the Asian Conference on Computer Vision (ACCV) (2018)
2018
-
[90]
Lukeˇ ziˇ c, L.ˇC
A. Lukeˇ ziˇ c, L.ˇC. Zajc, T. Voj ´ ıˇ r, J. Matas, M. Kristan, Now you see me: evaluating performance in long-term visual tracking, arXiv preprint arXiv:1804.07056 (2018)
2018 arXiv
-
[91]
H. Fan, L. Lin, F. Yang, P. Chu, G. Deng, H. Yu, H. Bai, Y. Xu, C. Liao, H. Ling, Lasot: A high-quality benchmark for large-scale single object tracking, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2019) 5374–5383
2019
-
[92]
Kristan, J
M. Kristan, J. Matas, A. Leonardis, M. Felsberg, L. Cehovin, G. Fernandez, T. Vojir, G. Hager, G. Nebehay, R. Pflugfelder, The visual object tracking vot2015 challenge results, in: Proceedings of the IEEE international conference on computer vision workshops, 2015, pp. 1–23. 52
2015
-
[93]
Roffo, S
G. Roffo, S. Melzi, et al., The visual object tracking vot2016 challenge results, in: Computer Vision–ECCV 2016 Workshops: Amsterdam, The Netherlands, October 8-10 and 15-16, 2016, Proceedings, Part II, Springer International Publishing, 2016, pp. 777–823
2016
-
[94]
Kristan, J
M. Kristan, J. Matas, A. Leonardis, M. Felsberg, L. Cehovin Zajc, T. Vojir, G. D. Hager, et al., The sixth visual object tracking vot2018 challenge results, in: Proceedings of the European Conference on Computer Vision (ECCV) Workshops, 2018
2018
-
[95]
Mueller, N
M. Mueller, N. Smith, B. Ghanem, A benchmark and simulator for uav tracking, European Conference on Computer Vision (ECCV) (2016) 445–461
2016
-
[96]
W. Hu, Q. Wang, L. Zhang, L. Bertinetto, P. H. Torr, Siammask: A framework for fast on- line object tracking and segmentation, IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (2023) 3072–3089
2023
-
[97]
Lukezic, J
A. Lukezic, J. Matas, M. Kristan, D3s-a discriminative single shot segmentation tracker, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 7133–7142
2020
-
[98]
M. Paul, M. Danelljan, C. Mayer, L. Van Gool, Robust visual tracking by segmentation, in: European conference on computer vision, Springer, 2022, pp. 571–588
2022
-
[99]
A. Ali, A. Jalil, J. Niu, X. Zhao, S. Rathore, J. Ahmed, M. Aksam Iftikhar, Visual object tracking—classical and contemporary approaches, Frontiers of Computer Science 10 (2016) 167– 188
2016
-
[100]
S. Abba, A. M. Bizi, J.-A. Lee, S. Bakouri, M. L. Crespo, Real-time object detection, tracking, and monitoring framework for security surveillance systems, Heliyon 10 (2024)
2024
-
[101]
Sivalingam, A
R. Sivalingam, A. Cherian, J. Fasching, N. Walczak, N. Bird, V. Morellas, B. Murphy, K. Cullen, K. Lim, G. Sapiro, et al., A multi-sensor visual tracking system for behavior monitoring of at-risk children, in: 2012 IEEE International Conference on Robotics and Automation, IEEE...
2012
-
[102]
W. Hu, T. Tan, L. Wang, S. Maybank, A survey on visual surveillance of object motion and behaviors, IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews) 34 (2004) 334–352
2004
-
[103]
Polat, M
E. Polat, M. Yeasin, R. Sharma, Robust tracking of human body parts for collaborative human computer interaction, Computer Vision and Image Understanding 89 (2003) 44–69
2003
-
[104]
A. W. N. Ibrahim, P. W. Ching, G. G. Seet, W. M. Lau, W. Czajewski, Moving objects detection and tracking framework for uav-based surveillance, in: 2010 Fourth Pacific-Rim Symposium on Image and Video Technology, IEEE, 2010, pp. 456–461
2010
-
[105]
D. Du, Y. Qi, H. Yu, Y. Yang, K. Duan, G. Li, W. Zhang, Q. Huang, Q. Tian, The unmanned aerial vehicle benchmark: Object detection and tracking, in: Proceedings of the European conference on computer vision (ECCV), 2018, pp. 370–386
2018
-
[106]
L.-Y. Lo, C. H. Yiu, Y. Tang, A.-S. Yang, B. Li, C.-Y. Wen, Dynamic object tracking on autonomous uav system for surveillance applications, Sensors 21 (2021) 7888
2021
-
[107]
Fernandez-Sanjurjo, B
M. Fernandez-Sanjurjo, B. Bosquet, M. Mucientes, V. M. Brea, Real-time visual detection and tracking system for traffic monitoring, Engineering Applications of Artificial Intelligence 85 (2019) 410–420. 53
2019
-
[108]
Khemmar, M
R. Khemmar, M. Gouveia, B. Decoux, J.-Y. y Ertaud, Real time pedestrian and object detection and tracking-based deep learning. application to drone visual tracking, in: WSCG’2019-27. International Conference in Central Europe on Computer Graphics, Visualization and Computer Vi...
2019
-
[109]
Makhmutova, I
A. Makhmutova, I. V. Anikin, M. Dagaeva, Object tracking method for videomonitoring in intel- ligent transport systems, in: 2020 International Russian Automation Conference (RusAutoCon), IEEE, 2020, pp. 535–540
2020
-
[110]
D. M. Jim´ enez-Bravo, ´A. L. Murciego, A. S. Mendes, H. S. San Bl´ as, J. Bajo, Multi-object tracking in traffic environments: A systematic literature review, Neurocomputing 494 (2022) 43–55
2022
-
[111]
Bisio, C
I. Bisio, C. Garibotto, H. Haleem, F. Lavagetto, A. Sciarrone, A systematic review of drone based road traffic monitoring system, Ieee Access 10 (2022) 101537–101555
2022
-
[112]
Markiewicz, M
P. Markiewicz, M. D lugosz, P. Skruch, Review of tracking and object detection systems for advanced driver assistance and autonomous driving applications with focus on vulnerable road users sensing, in: Polish Control Conference, Springer, 2017, pp. 224–237
2017
-
[113]
Premachandra, S
C. Premachandra, S. Ueda, Y. Suzuki, Detection and tracking of moving objects at road inter- sections using a 360-degree camera for driver assistance and automated driving, IEEE Access 8 (2020) 135652–135660
2020
-
[114]
K. Cho, D. Cho, Autonomous driving assistance with dynamic objects using traffic surveillance cameras, Applied Sciences 12 (2022) 6247
2022
-
[115]
Petrovskaya, S
A. Petrovskaya, S. Thrun, Model based vehicle detection and tracking for autonomous urban driving, Autonomous Robots 26 (2009) 123–139
2009
-
[116]
Muller, M
A. Muller, M. Manz, M. Himmelsbach, H. Wunsche, A model-based object following system, in: 2009 IEEE Intelligent Vehicles Symposium, IEEE, 2009, pp. 242–249
2009
-
[117]
Rangesh, M
A. Rangesh, M. M. Trivedi, No blind spots: Full-surround multi-object tracking for autonomous vehicles using cameras and lidars, IEEE Transactions on Intelligent Vehicles 4 (2019) 588–599
2019
-
[118]
G´ omez-Hu´ elamo, L
C. G´ omez-Hu´ elamo, L. M. Bergasa, R. Guti´ errez, J. F. Arango, A. D ´ ıaz, Smartmot: Exploiting the fusion of hdmaps and multi-object tracking for real-time scene understanding in intelligent vehicles applications, in: 2021 IEEE Intelligent Vehicles Symposium (IV), IEEE, 2...
2021
-
[119]
C. A. Richards, N. P. Papanikolopoulos, Detection and tracking for robotic visual servoing systems, Robotics and Computer-Integrated Manufacturing 13 (1997) 101–120
1997
-
[120]
D. J. Jacques, R. Rodrigo, K. A. McIsaac, J. Samarabandu, An object tracking and visual servoing system for the visually impaired, in: Proceedings of the 2005 IEEE International Conference on Robotics and Automation, IEEE, 2005, pp. 3510–3515
2005
-
[121]
Chaumette, S
F. Chaumette, S. Hutchinson, Visual servoing and visual tracking, Handbook of Robotics (2008) 563–583
2008
-
[122]
W. E. Dixon, E. Zergeroglu, Y. Fang, D. M. Dawson, Object tracking by a robot manipulator: a robust cooperative visual servoing approach, in: Proceedings 2002 IEEE International Conference on Robotics and Automation (Cat. No. 02CH37292), volume 1, IEEE, 2002, pp. 211–216. 54
2002
-
[123]
S. Xu, K. Chen, Y. Ou, Z. Wang, C. Yang, A learning-based object tracking strategy using visual sensors and intelligent robot arm, IEEE Transactions on Automation Science and Engineering 20 (2022) 2280–2293
2022
-
[124]
Ortenzi, A
V. Ortenzi, A. Cosgun, T. Pardi, W. P. Chan, E. Croft, D. Kuli´ c, Object handovers: a review for robotics, IEEE Transactions on Robotics 37 (2021) 1855–1873
2021
-
[125]
Costanzo, G
M. Costanzo, G. De Maria, C. Natale, Handover control for human-robot and robot-robot collaboration, Frontiers in Robotics and AI 8 (2021) 672995
2021
-
[126]
Bouget, M
D. Bouget, M. Allan, D. Stoyanov, P. Jannin, Vision-based and marker-less surgical tool detection and tracking: a review of the literature, Medical Image Analysis 35 (2017) 633–654. doi:10.1016/ j.media.2016.09.003
2017
-
[127]
C. I. Nwoye, N. Padoy, Surgitrack: Fine-grained multi-class multi-tool tracking in surgical videos, Medical Image Analysis 101 (2025) 103438
2025
-
[128]
Z. Li, H. Shu, R. Liang, A. Goodridge, M. Sahu, F. X. Creighton, R. H. Taylor, M. Unberath, Tatoo: vision-based joint tracking of anatomy and tool for skull-base surgery, International journal of computer assisted radiology and surgery 18 (2023) 1303–1310
2023
-
[129]
Martin-Gomez, H
A. Martin-Gomez, H. Li, T. Song, S. Yang, G. Wang, H. Ding, N. Navab, Z. Zhao, M. Armand, Sttar: surgical tool tracking using off-the-shelf augmented reality head-mounted displays, IEEE Transactions on Visualization and Computer Graphics (2023)
2023
-
[130]
Singh, S
A. Singh, S. S. M. Salehi, A. Gholipour, Deep predictive motion tracking in magnetic resonance imaging: application to fetal imaging, IEEE transactions on medical imaging 39 (2020) 3523– 3534
2020
-
[131]
Koniar, L
D. Koniar, L. Hargaˇ s, Z. Loncova, A. Simonova, F. Duchoˇ n, P. Beˇ no, Visual system-based object tracking using image segmentation for biomedical applications, Electrical Engineering 99 (2017) 1349–1366
2017
-
[132]
Hayashida, K
J. Hayashida, K. Nishimura, R. Bise, Consistent cell tracking in multi-frames with spatio- temporal context by object-level warping loss, in: Proceedings of the IEEE/CVF Winter Con- ference on Applications of Computer Vision, 2022, pp. 1727–1736. 55
2022
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.