Pith. sign in

REVIEW 3 major objections 6 minor 293 references

Visual Object Tracking across Diverse Data Modalities: A Review

T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read A survey maps visual object tracking across seven data modalities, from RGB to LiDAR and language.

desk verdict A genuinely useful modality-organized survey whose reference tables need a primary-source audit before the paper can be trusted as a citation. read the letter →

arxiv 2412.09991 v1 pith:FEZ2K7FD submitted 2024-12-13 cs.CV cs.AI

classification cs.CVcs.AI
keywords visualobjecttrackingmulti-modalRGB-thermalLiDARRGB-Languagedeeplearningbenchmarkssurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This review aims to give a complete map of visual object tracking (VOT) methods organized by data modality rather than by algorithm family alone. It covers three single-modality tracks — RGB video, thermal infrared, and LiDAR point clouds — and four multi-modal combinations — RGB-Depth, RGB-Thermal, RGB-LiDAR, and RGB-Language. The authors claim this is the first survey to include LiDAR-based, RGB-LiDAR, and RGB-Language tracking, and they support the map with abstracted pipeline schemas, benchmark comparisons, and dataset statistics. A reader who wants to know which tracking paradigms exist, which methods inherit them, and what numbers they achieve on standard benchmarks would use this as a reference.

What carries the argument

The central organizing device is a modality-by-modality taxonomy with abstracted pipeline diagrams and named schemas. For RGB trackers the four schemas are DCF, Siamese, ICD, and OST; for multi-modal trackers the three fusion strategies are early, middle, and late fusion. These schemas do the argumentative work: they turn a list of several hundred methods into a small set of inheritance relations, and they let the survey transfer paradigm knowledge across modalities — for example, TIR trackers building on DCF and Siamese, and LiDAR trackers borrowing the Siamese schema before moving to motion-modeling and one-stream Transformers.

What would settle it

Pick any row in Tables 1 through 11, locate the cited paper's reported score, and check whether the numbers match; one confirmed mismatch, such as the MLSSNet entry described in Section 5.2 as a 500-sequence dataset versus the 430 videos listed in Table 12, would show that the reference value of the tables depends on verification against primary sources.

Watch

Extended reading notes

Core claim

On its own terms, the paper's central claim is that VOT is best understood from the perspective of data modalities, and that a survey organized that way can be both comprehensive and novel. For single-modality RGB tracking, it abstracts four paradigms: discriminative correlation filters (online-trained filters convolved with search features), Siamese trackers (a shared network matching a template to a search region), instance classification/detection (a network specialized to one target instance), and one-stream Transformers (a single Transformer that jointly extracts features and relates template to search). Thermal trackers are shown to inherit the DCF and Siamese schemas, and LiDAR trackers are shown to follow Siamese, motion-modeling, and one-stream Transformer designs. For multi-modal tracking, the organizing distinction is fusion stage: early fusion at the input, middle fusion at the feature level, or late fusion at the result level. The survey concludes that these taxonomies, together with its benchmark tables, constitute the first systematic reference for the newly emerged LiDAR-based, RGB-LiDAR, and RGB-Language tracking directions.

Load-bearing premise

The benchmark numbers and dataset statistics in the tables are faithful copies of the cited papers, and the modality and fusion taxonomies assign every method to exactly one correct box.

Editorial extensions

If this is right

  • A newcomer can identify the paradigm of any RGB tracker by matching its pipeline to one of four schemas rather than reading each paper in full.
  • Multi-modal trackers can be classified by fusion stage, which predicts whether the method requires aligned inputs, learns cross-modal feature interactions, or fuses final predictions.
  • The benchmark tables provide a single place to compare trackers across LaSOT, TrackingNet, GOT-10k, VOT, KITTI, nuScenes, Waymo, PTB, DepthTrack, RGBT234, LasHeR, and TNL2K, with numbers transcribed from the cited papers.
  • Identifying OST as the emerging RGB paradigm points to one-stream Transformers as the schema most likely to be transferred to TIR and LiDAR tracking.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the modality-first organization is right, the missing next piece is a cross-modal evaluation protocol that reuses the same target categories and metrics across RGB, TIR, and LiDAR, so that paradigm-transfer claims can be tested quantitatively rather than by inspection.
  • The fusion-stage taxonomy suggests a testable conjecture: middle fusion will keep dominating RGB-Thermal tracking because it offers learnable cross-modal parameters without requiring the strict input alignment that early fusion demands; a meta-analysis of the table entries could check whether late-fusion methods ever surpass middle-fusion ones at similar speed.
  • The inclusion of RGB-Language tracking implies the field may treat natural-language descriptions as a first-class query channel alongside boxes and point clouds, which would connect VOT to open-vocabulary and referring-expression benchmarks beyond those listed.
  • Because the survey records FPS alongside accuracy, a reader could use its tables to test whether the OST paradigm's accuracy gains come at a speed cost, a question the paper raises but does not resolve.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper is a survey of visual object tracking (VOT) organized by data modality. It reviews three single-modal families (RGB, thermal infrared, LiDAR) and four multi-modal combinations (RGB-Depth, RGB-Thermal, RGB-LiDAR, RGB-Language). For RGB trackers it proposes a taxonomy of four deep-learning paradigms (discriminative correlation filters, Siamese trackers, instance classification/detection, and one-stream transformers); for TIR and LiDAR it provides modality-specific taxonomies; and for multi-modal methods it applies an early/middle/late fusion categorization. The survey compiles comparison results on many benchmarks in Tables 1-11 and dataset statistics in Table 12, and closes with ten short discussions of future directions such as parameter-efficient transfer learning, online learning, and multi-modal tracking. The paper claims to be the first review covering LiDAR-based, RGB-LiDAR, and RGB-Language VOT methods.

Significance. If the benchmark tables and dataset statistics are reliable, this survey would be a useful entry point for researchers, especially for the less-covered LiDAR, RGB-LiDAR, and RGB-Language areas. The four-RGB-paradigm taxonomy and the early/middle/late fusion categorization are clear organizational principles, and the breadth of coverage (300+ papers, seven modality families) is a genuine strength. Because the survey's main contribution is reference value rather than new methods or derivations, the fidelity of the tables is load-bearing: a reader consulting this paper will use the tables to compare trackers and to select datasets. The paper ships no code or proofs, but as a survey that is not expected; its value rests on accurate transcription of the cited literature.

major comments (3)
  1. [Section 5.2 and Table 12 (Thermal block)] The dataset statistics for MLSSNet are internally contradictory: Section 5.2 states that MLSSNet [293] is a large-scale TIR dataset with 500 video sequences and 228k frames, while Table 12 reports 430 videos, 200k boxes, 20 classes, and an average duration of 15.5s for the same reference. Since the survey's reference value depends on faithful transcription of primary sources, the authors must verify the original paper and make the text and table agree.
  2. [Table 12 and Section 5.2 vs. Section 3.2 and Table 4] References [293] (MLSSNet), [294] (MMNet), and [30] (ECO-MM) are presented as trackers with benchmark results in Table 4 and Section 3.2, but the same references are listed as Thermal datasets in Table 12 and described as datasets in Section 5.2. If these papers indeed introduce both a tracker and a dataset, the survey should explicitly say so and clearly separate the two roles; as it stands, a reader cannot tell whether the rows in Table 12 are datasets, methods, or both, which undermines the dataset table.
  3. [Table 11] The caption of Table 11 states that results are evaluated by Precision/AUC, but several TNL2K cells contain three values (e.g., Feng et al. 0.27/0.34/0.25, Wang et al. 0.06/0.11/0.11 and 0.42/0.50/0.42, VLTTT 0.53/0.53). The legend must be expanded to explain what the third number represents, or the cells must be corrected, because the table is not interpretable as presented.
minor comments (6)
  1. [Section 2.1] The sentence claiming the survey covers "eight multiple modalities" should read "four multiple modalities," since Section 4 reviews exactly four multi-modal combinations (RGB-Depth, RGB-Thermal, RGB-LiDAR, RGB-Language).
  2. [Table 12] LaSOT appears in the RGB block with year 2019 and in the RGB-La block with year 2018; the year should be made consistent, and the double listing (the same dataset in two modality groups) should be explicitly justified.
  3. [Table 1] The last column header appears as "FPSSR(%)", which seems to merge the FPS and SR(%) columns; please split the header into separate columns for FPS and SR(%) or correct the label.
  4. [Section 4.4] The citation for VLTTT appears as "VLTT T[411]" in the text but as "V LTT T[41]" in Table 11; the reference number should be consistent (the reference list entry is [41]).
  5. [Section 5.2] The phrase "most of them are shotted at night" contains a typo; "shotted" should be "shot."
  6. [Section 1] The claim of being the first review to cover LiDAR-based, RGB-LiDAR, and RGB-Language VOT is plausible but should be substantiated by a more explicit comparison with the related surveys listed in Section 2, since the current discussion does not fully rule out partial coverage in prior works.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the survey's claims are descriptive summaries of external literature, with no derivation whose output depends on its own inputs.

full rationale

This is a survey paper whose central claims are taxonomic organization, method summaries, and transcribed benchmark/dataset statistics. There is no derived quantity, no fitted parameter, and no predictive claim that could reduce by construction to its own inputs. The few self-citations (e.g., Wang et al. [6], [7], [195], [263], [275], [277], [238]) appear in contextual passing remarks such as applications of tracking, lists of DCF variants, video object segmentation, PEFT, and future directions; none of these citations is load-bearing for the survey's stated contributions of organizing VOT methods by modality or reporting comparison results. The taxonomy distinctions (four RGB paradigms; early/middle/late fusion for multi-modal methods) are editorial classifications of external work, not outputs derived from the cited papers in a way that makes the classification equivalent to an input. The internal inconsistencies flagged by a skeptical reading, such as the MLSSNet numbers in Section 5.2 versus Table 12, the dataset-vs-method labeling of MMNet and ECO-MM in Table 12, and the TNL2K metric-legend mismatch in Table 11, are transcription or presentation defects that would affect the survey's reference reliability; they are correctness risks, not circular reasoning. Under the rule that non-consensus or factual errors are not circularity, these do not raise the circularity score. The paper is therefore self-contained as a review and exhibits no significant circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No new quantities, models, or entities are introduced; the paper's contribution is organizational. Its load-bearing assumptions are about the fidelity of transcribed benchmark numbers, the completeness of taxonomies, and the precedence of its coverage.

assumptions (3)
  • domain assumption The performance numbers in Tables 1 through 11 are accurately transcribed from the cited papers.
    The survey's comparative value depends on the fidelity of these reported benchmark results; Section 3 and Tables 1 through 11 compile numbers from external papers without independent reproduction.
  • domain assumption The proposed taxonomies, four RGB paradigms and three fusion strategies for RGB-D and RGB-T methods, are faithful, complete, and non-overlapping.
    Sections 3.1 and 4 organize methods into categories; misclassification would mislead readers about paradigm boundaries and the relationships between methods.
  • ad hoc to paper The claim of being the first review for LiDAR-based, RGB-LiDAR, and RGB-Language VOT is correct.
    Sections 1 and 2 assert novelty of coverage without an exhaustive comparison against all earlier surveys, so this precedence claim is unverified.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Visual Object Tracking across Diverse Data Modalities: A Review." pith.science (2026). https://pith.science/paper/FEZ2K7FD

@misc{pith2026241209991,
  author       = {Pith},
  title        = {Pith review of: Visual Object Tracking across Diverse Data Modalities: A Review},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FEZ2K7FD}},
  note         = {Machine review of arXiv:2412.09991}
}
read the original abstract

Visual Object Tracking (VOT) is an attractive and significant research area in computer vision, which aims to recognize and track specific targets in video sequences where the target objects are arbitrary and class-agnostic. The VOT technology could be applied in various scenarios, processing data of diverse modalities such as RGB, thermal infrared and point cloud. Besides, since no one sensor could handle all the dynamic and varying environments, multi-modal VOT is also investigated. This paper presents a comprehensive survey of the recent progress of both single-modal and multi-modal VOT, especially the deep learning methods. Specifically, we first review three types of mainstream single-modal VOT, including RGB, thermal infrared and point cloud tracking. In particular, we conclude four widely-used single-modal frameworks, abstracting their schemas and categorizing the existing inheritors. Then we summarize four kinds of multi-modal VOT, including RGB-Depth, RGB-Thermal, RGB-LiDAR and RGB-Language. Moreover, the comparison results in plenty of VOT benchmarks of the discussed modalities are presented. Finally, we provide recommendations and insightful observations, inspiring the future development of this fast-growing literature.

Figures

Figures reproduced from arXiv: 2412.09991 by the authors.

Figure 1
Figure 1. The advantages, disadvantages and applications of the dis￾cussed three single-modal VOT and four multi-modal VOT. [14], [15], [16], [17] and large-scale datasets [18], [19], [20]. We mainly focus on the methods of the past decade, especially the methods based on Deep Neural Network (DNN). According to their pipelines, we classify the mainstream RGB trackers into four categories, namely Discriminative Correlation Fil… view at source ↗
Figure 2
Figure 2. Develop lineage of representative DCF-based RGB trackers. The methods are divided according to their employed features, including [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Visualization of the four schemas of RGB-based VOT, which [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Develop lineage of representative Siamese-based RGB trackers. [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: Develop lineage of representative ICD-based RGB trackers. [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 7
Figure 7. Figure 7: Develop lineage of Thermal-based trackers. The methods are split into Discriminative Correlation Filter, Siamese and Others. HIPTrack [94], RFGM [88], ROMTrack [86], SeqTrack [92], AR￾Track [91], VideoTrack [93], ODTrack [95] and EVPTrack [96]. Specifically, leveraging…
Figure 8
Figure 8. Figure 8: Develop lineage of LiDAR-based trackers. The methods are split into Siamese, Motion-Modeling and Single-Branch framework, where the first one has dominated till now. Then, we make fine-grained cat￾egorization to conclude the common labels of related methods, like using…
Figure 9
Figure 9. Figure 9: Develop lineage of RGB-Depth based trackers. We categorize the methods based on the feature fusion manners, which can be sum￾marized into Early Fusion, Middle Fusion, and Late Fusion specifically. and has fewer model parameters and computational overheads. Moreover, wi…
Figure 10
Figure 10. Figure 10: Develop lineage of RGB-Thermal based trackers. We categorize the methods based on the feature fusion manners, which can be sum￾marized into Early Fusion, Middle Fusion, and Late Fusion specifically. is of practical value and significance. 4.2 RGB-Thermal Trackers RGB …

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

293 extracted references · 65 canonical work pages

  1. [293]

    Learning deep multi-level similarity for thermal infrared object tracking,

    Q. Liu, X. Li, Z. He, N. Fan, D. Yuan, and H. Wang, “Learning deep multi-level similarity for thermal infrared object tracking,” IEEE Transactions on Multimedia, vol. 23, pp. 2114–2126, 2020

  2. [294]

    Multi-task driven feature models for thermal infrared tracking,

    Q. Liu, X. Li, Z. He, N. Fan, D. Yuan, W. Liu, and Y . Liang, “Multi-task driven feature models for thermal infrared tracking,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 07, 2020, pp. 11 604–11 611

  3. [30]

    Learning dual-level deep representation for thermal infrared tracking,

    Q. Liu, D. Yuan, N. Fan, P. Gao, X. Li, and Z. He, “Learning dual-level deep representation for thermal infrared tracking,” IEEE Transactions on Multimedia, 2022

  4. [1]

    Backbone is all your need: A simplified architecture for visual object tracking,

    B. Chen, P. Li, L. Bai, L. Qiao, Q. Shen, B. Li, W. Gan, W. Wu, and W. Ouyang, “Backbone is all your need: A simplified architecture for visual object tracking,” arXiv preprint arXiv:2203.05328, 2022

  5. [2]

    High-speed tracking with kernelized correlation filters,

    J. F. Henriques, R. Caseiro, P. Martins, and J. Batista, “High-speed tracking with kernelized correlation filters,” IEEE transactions on pattern analysis and machine intelligence , vol. 37, no. 3, pp. 583–596, 2014

  6. [3]

    High performance visual tracking with siamese region proposal network,

    B. Li, J. Yan, W. Wu, Z. Zhu, and X. Hu, “High performance visual tracking with siamese region proposal network,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 8971–8980

  7. [4]

    3d-siamrpn: An end-to-end learning method for real-time 3d single object tracking using raw point cloud,

    Z. Fang, S. Zhou, Y . Cui, and S. Scherer, “3d-siamrpn: An end-to-end learning method for real-time 3d single object tracking using raw point cloud,” IEEE Sensors Journal, vol. 21, no. 4, pp. 4995–5011, 2020

  8. [5]

    3d siamese voxel- to-bev tracker for sparse point clouds,

    L. Hui, L. Wang, M. Cheng, J. Xie, and J. Yang, “3d siamese voxel- to-bev tracker for sparse point clouds,” Advances in Neural Information Processing Systems, vol. 34, pp. 28 714–28 727, 2021

Show all 293 references
  1. [6]

    Real-time 3d human tracking for mobile robots with multisensors,

    M. Wang, D. Su, L. Shi, Y . Liu, and J. V . Miro, “Real-time 3d human tracking for mobile robots with multisensors,” in 2017 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2017, pp. 5081–5087

  2. [7]

    Accurate and real-time 3-d tracking for the following robots by fusing vision and ultrasonar information,

    M. Wang, Y . Liu, D. Su, Y . Liao, L. Shi, J. Xu, and J. V . Miro, “Accurate and real-time 3-d tracking for the following robots by fusing vision and ultrasonar information,” IEEE/ASME Transactions On Mechatronics , vol. 23, no. 3, pp. 997–1006, 2018

  3. [8]

    Grounding-tracking- integration,

    Z. Yang, T. Kumar, T. Chen, J. Su, and J. Luo, “Grounding-tracking- integration,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 31, no. 9, pp. 3433–3443, 2020

  4. [9]

    Siamese natural lan- guage tracker: Tracking by natural language descriptions with siamese trackers,

    Q. Feng, V . Ablavsky, Q. Bai, and S. Sclaroff, “Siamese natural lan- guage tracker: Tracking by natural language descriptions with siamese trackers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 5851–5860

  5. [10]

    Tracking-learning-detection,

    Z. Kalal, K. Mikolajczyk, and J. Matas, “Tracking-learning-detection,” IEEE transactions on pattern analysis and machine intelligence, vol. 34, no. 7, pp. 1409–1422, 2011

  6. [11]

    Struck: Structured output tracking with kernels,

    S. Hare, S. Golodetz, A. Saffari, V . Vineet, M.-M. Cheng, S. L. Hicks, and P. H. Torr, “Struck: Structured output tracking with kernels,” IEEE transactions on pattern analysis and machine intelligence , vol. 38, no. 10, pp. 2096–2109, 2015

  7. [12]

    Discrimina- tive scale space tracking,

    M. Danelljan, G. H ¨ager, F. S. Khan, and M. Felsberg, “Discrimina- tive scale space tracking,” IEEE transactions on pattern analysis and machine intelligence, vol. 39, no. 8, pp. 1561–1575, 2016

  8. [13]

    Joint feature learning and relation modeling for tracking: A one-stream framework,

    B. Ye, H. Chang, B. Ma, S. Shan, and X. Chen, “Joint feature learning and relation modeling for tracking: A one-stream framework,” in Euro- pean Conference on Computer Vision. Springer, 2022, pp. 341–357

  9. [14]

    Siamese box adaptive network for visual tracking,

    Z. Chen, B. Zhong, G. Li, S. Zhang, and R. Ji, “Siamese box adaptive network for visual tracking,” in Proceedings of the IEEE/CVF confer- ence on computer vision and pattern recognition, 2020, pp. 6668–6677

  10. [15]

    Fear: Fast, efficient, accurate and robust visual tracker,

    V . Borsuk, R. Vei, O. Kupyn, T. Martyniuk, I. Krashenyi, and J. Matas, “Fear: Fast, efficient, accurate and robust visual tracker,” in European Conference on Computer Vision. Springer, 2022, pp. 644–663

  11. [16]

    Fully-convolutional siamese networks for object tracking,

    L. Bertinetto, J. Valmadre, J. F. Henriques, A. Vedaldi, and P. H. Torr, “Fully-convolutional siamese networks for object tracking,” in European conference on computer vision . Springer, 2016, pp. 850– 865

  12. [17]

    Learning discrimi- native model prediction for tracking,

    G. Bhat, M. Danelljan, L. V . Gool, and R. Timofte, “Learning discrimi- native model prediction for tracking,” in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 6182–6191

  13. [18]

    Lasot: A high-quality benchmark for large-scale single object tracking,

    H. Fan, L. Lin, F. Yang, P. Chu, G. Deng, S. Yu, H. Bai, Y . Xu, C. Liao, and H. Ling, “Lasot: A high-quality benchmark for large-scale single object tracking,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 5374–5383

  14. [19]

    Got-10k: A large high-diversity benchmark for generic object tracking in the wild,

    L. Huang, X. Zhao, and K. Huang, “Got-10k: A large high-diversity benchmark for generic object tracking in the wild,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 43, no. 5, pp. 1562– 1577, 2019

  15. [20]

    Track- ingnet: A large-scale dataset and benchmark for object tracking in the wild,

    M. Muller, A. Bibi, S. Giancola, S. Alsubaihi, and B. Ghanem, “Track- ingnet: A large-scale dataset and benchmark for object tracking in the wild,” in Proceedings of the European conference on computer vision (ECCV), 2018, pp. 300–317

  16. [21]

    Transforming model prediction for tracking,

    C. Mayer, M. Danelljan, G. Bhat, M. Paul, D. P. Paudel, F. Yu, and L. Van Gool, “Transforming model prediction for tracking,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 8731–8740

  17. [22]

    Siamrpn++: Evolution of siamese visual tracking with very deep networks,

    B. Li, W. Wu, Q. Wang, F. Zhang, J. Xing, and J. Yan, “Siamrpn++: Evolution of siamese visual tracking with very deep networks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 4282–4291

  18. [23]

    Siamcar: Siamese fully convolutional classification and regression for visual tracking,

    D. Guo, J. Wang, Y . Cui, Z. Wang, and S. Chen, “Siamcar: Siamese fully convolutional classification and regression for visual tracking,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 6269–6277

  19. [24]

    Siamban: Target-aware tracking with siamese box adaptive network,

    Z. Chen, B. Zhong, G. Li, S. Zhang, R. Ji, Z. Tang, and X. Li, “Siamban: Target-aware tracking with siamese box adaptive network,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022

  20. [26]

    Uct: Learning unified convolutional networks for real-time visual tracking,

    Z. Zhu, G. Huang, W. Zou, D. Du, and C. Huang, “Uct: Learning unified convolutional networks for real-time visual tracking,” in Proceedings of JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 24 the IEEE international conference on computer vision workshops, 2017, pp....

  21. [27]

    Branchout: Regularization for online ensemble tracking with convolutional neural networks,

    B. Han, J. Sim, and H. Adam, “Branchout: Regularization for online ensemble tracking with convolutional neural networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 3356–3365

  22. [28]

    Correlation- aware deep tracking,

    F. Xie, C. Wang, G. Wang, Y . Cao, W. Yang, and W. Zeng, “Correlation- aware deep tracking,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 8751–8760

  23. [29]

    Mixformer: End-to-end track- ing with iterative mixed attention,

    Y . Cui, C. Jiang, L. Wang, and G. Wu, “Mixformer: End-to-end track- ing with iterative mixed attention,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 13 608–13 618

  24. [31]

    Gfsnet: Generalization-friendly siamese network for thermal infrared object tracking,

    R. Chen, S. Liu, Z. Miao, and F. Li, “Gfsnet: Generalization-friendly siamese network for thermal infrared object tracking,” Infrared Physics & Technology, p. 104190, 2022

  25. [32]

    Thermal infrared object tracking using correlation filters improved by level set,

    H. Zhang, Z. Yin, and H. Zhang, “Thermal infrared object tracking using correlation filters improved by level set,”Signal, Image and Video Processing, pp. 1–7, 2022

  26. [33]

    P2b: Point-to-box network for 3d object tracking in point clouds,

    H. Qi, C. Feng, Z. Cao, F. Zhao, and Y . Xiao, “P2b: Point-to-box network for 3d object tracking in point clouds,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 6329–6338

  27. [34]

    Box- aware feature enhancement for single object tracking on point clouds,

    C. Zheng, X. Yan, J. Gao, W. Zhao, W. Zhang, Z. Li, and S. Cui, “Box- aware feature enhancement for single object tracking on point clouds,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 13 199–13 208

  28. [35]

    Beyond 3d siamese tracking: A motion-centric paradigm for 3d single object tracking in point clouds,

    C. Zheng, X. Yan, H. Zhang, B. Wang, S. Cheng, S. Cui, and Z. Li, “Beyond 3d siamese tracking: A motion-centric paradigm for 3d single object tracking in point clouds,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 8111–8120

  29. [36]

    Sowp: Spatially or- dered and weighted patch descriptor for visual tracking,

    H.-U. Kim, D.-Y . Lee, J.-Y . Sim, and C.-S. Kim, “Sowp: Spatially or- dered and weighted patch descriptor for visual tracking,” inProceedings of the IEEE International Conference on Computer Vision , 2015, pp. 3011–3019

  30. [37]

    Weighted sparse representa- tion regularized graph learning for rgb-t object tracking,

    C. Li, N. Zhao, Y . Lu, C. Zhu, and J. Tang, “Weighted sparse representa- tion regularized graph learning for rgb-t object tracking,” inProceedings of the 25th ACM international conference on Multimedia , 2017, pp. 1856–1864

  31. [38]

    Learning adaptive attribute- driven representation for real-time rgb-t tracking,

    P. Zhang, D. Wang, H. Lu, and X. Yang, “Learning adaptive attribute- driven representation for real-time rgb-t tracking,”International Journal of Computer Vision, vol. 129, no. 9, pp. 2714–2729, 2021

  32. [39]

    F-siamese tracker: A frustum-based double siamese network for 3d single object tracking,

    H. Zou, J. Cui, X. Kong, C. Zhang, Y . Liu, F. Wen, and W. Li, “F-siamese tracker: A frustum-based double siamese network for 3d single object tracking,” in 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2020, pp. 8133–8139

  33. [40]

    3d object tracking using rgb and lidar data,

    A. Asvadi, P. Girao, P. Peixoto, and U. Nunes, “3d object tracking using rgb and lidar data,” in 2016 IEEE 19th International Conference on Intelligent Transportation Systems (ITSC). IEEE, 2016, pp. 1255– 1260

  34. [41]

    Divert more attention to vision- language tracking,

    M. Guo, Z. Zhang, H. Fan, and L. Jing, “Divert more attention to vision- language tracking,” arXiv preprint arXiv:2207.01076, 2022

  35. [42]

    Handcrafted and deep trackers: Recent visual object tracking approaches and trends,

    M. Fiaz, A. Mahmood, S. Javed, and S. K. Jung, “Handcrafted and deep trackers: Recent visual object tracking approaches and trends,” ACM Computing Surveys (CSUR), vol. 52, no. 2, pp. 1–44, 2019

  36. [43]

    Visual object tracking with discriminative filters and siamese networks: a survey and outlook,

    S. Javed, M. Danelljan, F. S. Khan, M. H. Khan, M. Felsberg, and J. Matas, “Visual object tracking with discriminative filters and siamese networks: a survey and outlook,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 5, pp. 6552–6574, 2022

  37. [44]

    Siamese visual object tracking: A survey,

    M. Ondra ˇsoviˇc and P. Tar ´abek, “Siamese visual object tracking: A survey,”IEEE Access, vol. 9, pp. 110 149–110 172, 2021

  38. [45]

    Recent advances of single-object tracking methods: A brief survey,

    Y . Zhang, T. Wang, K. Liu, B. Zhang, and L. Chen, “Recent advances of single-object tracking methods: A brief survey,” Neurocomputing, vol. 455, pp. 1–11, 2021

  39. [46]

    Visual object tracking: A survey,

    F. Chen, X. Wang, Y . Zhao, S. Lv, and X. Niu, “Visual object tracking: A survey,” Computer Vision and Image Understanding , vol. 222, p. 103508, 2022

  40. [47]

    Deep learning for visual tracking: A comprehensive survey,

    S. M. Marvasti-Zadeh, L. Cheng, H. Ghanei-Yakhdan, and S. Kasaei, “Deep learning for visual tracking: A comprehensive survey,” IEEE Transactions on Intelligent Transportation Systems, 2021

  41. [48]

    Deep visual tracking: Review and experimental comparison,

    P. Li, D. Wang, L. Wang, and H. Lu, “Deep visual tracking: Review and experimental comparison,” Pattern Recognition, vol. 76, pp. 323–338, 2018

  42. [49]

    Single object tracking: A survey of methods, datasets, and evaluation metrics,

    Z. Soleimanitaleb and M. A. Keyvanrad, “Single object tracking: A survey of methods, datasets, and evaluation metrics,” arXiv preprint arXiv:2201.13066, 2022

  43. [50]

    Single object tracking research: A survey,

    R. Han, W. Feng, Q. Guo, and Q. Hu, “Single object tracking research: A survey,” arXiv preprint arXiv:2204.11410, 2022

  44. [51]

    Multi-modal visual tracking: Review and experimental comparison,

    P. Zhang, D. Wang, and H. Lu, “Multi-modal visual tracking: Review and experimental comparison,” arXiv preprint arXiv:2012.04176, 2020

  45. [52]

    Recent advances on multicue object tracking: a survey,

    G. S. Walia and R. Kapoor, “Recent advances on multicue object tracking: a survey,” Artificial Intelligence Review , vol. 46, no. 1, pp. 1–39, 2016

  46. [53]

    Recent trends in multicue based visual tracking: A review,

    A. Kumar, G. S. Walia, and K. Sharma, “Recent trends in multicue based visual tracking: A review,” Expert Systems with Applications , vol. 162, p. 113711, 2020

  47. [54]

    Recent advances and trends in visual tracking: A review,

    H. Yang, L. Shao, F. Zheng, L. Wang, and Z. Song, “Recent advances and trends in visual tracking: A review,” Neurocomputing, vol. 74, no. 18, pp. 3823–3831, 2011

  48. [55]

    A survey of appearance models in visual object tracking,

    X. Li, W. Hu, C. Shen, Z. Zhang, A. Dick, and A. V . D. Hengel, “A survey of appearance models in visual object tracking,” ACM transactions on Intelligent Systems and Technology (TIST), vol. 4, no. 4, pp. 1–48, 2013

  49. [56]

    Visual tracking: An experimental survey,

    A. W. Smeulders, D. M. Chu, R. Cucchiara, S. Calderara, A. Dehghan, and M. Shah, “Visual tracking: An experimental survey,” IEEE trans- actions on pattern analysis and machine intelligence, vol. 36, no. 7, pp. 1442–1468, 2013

  50. [57]

    A comparison of correlation filter-based trackers and struck trackers,

    J. Wang, L. Zheng, M. Tang, and J. Feng, “A comparison of correlation filter-based trackers and struck trackers,”IEEE Transactions on Circuits and Systems for Video Technology, vol. 30, no. 9, pp. 3106–3118, 2019

  51. [58]

    Sparse coding based visual tracking: Review and experimental comparison,

    S. Zhang, H. Yao, X. Sun, and X. Lu, “Sparse coding based visual tracking: Review and experimental comparison,” Pattern Recognition, vol. 46, no. 7, pp. 1772–1788, 2013

  52. [59]

    An in-depth analysis of visual tracking with siamese neural networks,

    R. Pflugfelder, “An in-depth analysis of visual tracking with siamese neural networks,” arXiv preprint arXiv:1707.00569, 2017

  53. [60]

    Object tracking: A survey,

    A. Yilmaz, O. Javed, and M. Shah, “Object tracking: A survey,” Acm computing surveys (CSUR), vol. 38, no. 4, pp. 13–es, 2006

  54. [61]

    Rgbd object tracking: An in-depth review,

    J. Yang, Z. Li, S. Yan, F. Zheng, A. Leonardis, J.-K. K ¨am¨ar¨ainen, and L. Shao, “Rgbd object tracking: An in-depth review,” arXiv preprint arXiv:2203.14134, 2022

  55. [62]

    Object fusion track- ing based on visible and infrared images: A comprehensive review,

    X. Zhang, P. Ye, H. Leung, K. Gong, and G. Xiao, “Object fusion track- ing based on visible and infrared images: A comprehensive review,” Information Fusion, vol. 63, pp. 166–187, 2020

  56. [63]

    The sixth visual object tracking vot2018 challenge results,

    M. Kristan, A. Leonardis, J. Matas, M. Felsberg, R. Pflugfelder, L. ˇCehovin Zajc, T. V ojir, G. Bhat, A. Lukezic, A. Eldesokey et al., “The sixth visual object tracking vot2018 challenge results,” in Pro- ceedings of the European Conference on Computer Vision (ECCV) Workshops...

  57. [64]

    A benchmark and simulator for uav tracking,

    M. Mueller, N. Smith, and B. Ghanem, “A benchmark and simulator for uav tracking,” in European conference on computer vision . Springer, 2016, pp. 445–461

  58. [65]

    Need for speed: A benchmark for higher frame rate object tracking,

    H. Kiani Galoogahi, A. Fagg, C. Huang, D. Ramanan, and S. Lucey, “Need for speed: A benchmark for higher frame rate object tracking,” in Proceedings of the IEEE International Conference on Computer Vision, 2017, pp. 1125–1134

  59. [66]

    Object tracking benchmark,

    Y . Wu, J. Lim, and M.-H. Yang, “Object tracking benchmark,” IEEE transactions on pattern analysis and machine intelligence, vol. 37, no. 9, pp. 1834—-1848, 2015

  60. [67]

    Robust visual object tracking via adaptive attribute-aware discriminative correlation filters,

    X.-F. Zhu, X.-J. Wu, T. Xu, Z.-H. Feng, and J. Kittler, “Robust visual object tracking via adaptive attribute-aware discriminative correlation filters,” IEEE Transactions on Multimedia, vol. 24, pp. 301–312, 2021

  61. [68]

    Bilateral weighted regression ranking model with spatial-temporal correlation filter for visual tracking,

    H. Zhu, H. Peng, G. Xu, L. Deng, Y . Cheng, and A. Song, “Bilateral weighted regression ranking model with spatial-temporal correlation filter for visual tracking,” IEEE Transactions on Multimedia , vol. 24, pp. 2098–2111, 2021

  62. [69]

    Siamcorners: Siamese corner networks for visual tracking,

    K. Yang, Z. He, W. Pei, Z. Zhou, X. Li, D. Yuan, and H. Zhang, “Siamcorners: Siamese corner networks for visual tracking,” IEEE Transactions on Multimedia, vol. 24, pp. 1956–1967, 2021

  63. [70]

    Stgl: Spatial- temporal graph representation and learning for visual tracking,

    B. Jiang, Y . Zhang, B. Luo, X. Cao, and J. Tang, “Stgl: Spatial- temporal graph representation and learning for visual tracking,” IEEE Transactions on Multimedia, vol. 23, pp. 2162–2171, 2020

  64. [71]

    Discriminative siamese complementary tracker with flexible update,

    B. Fan, J. Tian, Y . Peng, and Y . Tang, “Discriminative siamese complementary tracker with flexible update,” IEEE Transactions on Multimedia, 2021

  65. [72]

    Siamese implicit region proposal network with compound attention for visual tracking,

    S. Chan, J. Tao, X. Zhou, C. Bai, and X. Zhang, “Siamese implicit region proposal network with compound attention for visual tracking,” IEEE Transactions on Image Processing, vol. 31, pp. 1882–1894, 2022

  66. [73]

    Learning feature channel weighting for real-time visual tracking,

    Z. Li, J. Zhang, Y . Li, J. Zhu, S. Long, D. Xue, and L. Fan, “Learning feature channel weighting for real-time visual tracking,” IEEE Transac- tions on Image Processing, vol. 31, pp. 2190–2200, 2022. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 25

  67. [74]

    Global tracking via ensemble of local trackers,

    Z. Zhou, J. Chen, W. Pei, K. Mao, H. Wang, and Z. He, “Global tracking via ensemble of local trackers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 8761–8770

  68. [75]

    Unified transformer tracker for object tracking,

    F. Ma, M. Z. Shou, L. Zhu, H. Fan, Y . Xu, Y . Yang, and Z. Yan, “Unified transformer tracker for object tracking,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 8781–8790

  69. [76]

    Ranking-based siamese visual tracking,

    F. Tang and Q. Ling, “Ranking-based siamese visual tracking,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 8741–8750

  70. [77]

    Transformer tracking with cyclic shifting window attention,

    Z. Song, J. Yu, Y .-P. P. Chen, and W. Yang, “Transformer tracking with cyclic shifting window attention,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 8791–8800

  71. [78]

    Learning target-aware representation for visual tracking via informative interac- tions,

    M. Guo, Z. Zhang, H. Fan, L. Jing, Y . Lyu, B. Li, and W. Hu, “Learning target-aware representation for visual tracking via informative interac- tions,” arXiv preprint arXiv:2201.02526, 2022

  72. [79]

    Sparsett: Visual tracking with sparse transformers,

    Z. Fu, Z. Fu, Q. Liu, W. Cai, and Y . Wang, “Sparsett: Visual tracking with sparse transformers,” arXiv preprint arXiv:2205.03776, 2022

  73. [80]

    Online hybrid lightweight rep- resentations learning: Its application to visual tracking,

    I. Jung, M. Kim, E. Park, and B. Han, “Online hybrid lightweight rep- resentations learning: Its application to visual tracking,” arXiv preprint arXiv:2205.11179, 2022

  74. [81]

    Towards sequence-level training for visual tracking,

    M. Kim, S. Lee, J. Ok, B. Han, and M. Cho, “Towards sequence-level training for visual tracking,” in European Conference on Computer Vision. Springer, 2022, pp. 534–551

  75. [82]

    Towards grand unification of object tracking,

    B. Yan, Y . Jiang, P. Sun, D. Wang, Z. Yuan, P. Luo, and H. Lu, “Towards grand unification of object tracking,” inEuropean Conference on Computer Vision. Springer, 2022, pp. 733–751

  76. [83]

    Robust visual tracking by segmentation,

    M. Paul, M. Danelljan, C. Mayer, and L. Van Gool, “Robust visual tracking by segmentation,” in European Conference on Computer Vi- sion. Springer, 2022, pp. 571–588

  77. [84]

    Aiatrack: Attention in attention for transformer visual tracking,

    S. Gao, C. Zhou, C. Ma, X. Wang, and J. Yuan, “Aiatrack: Attention in attention for transformer visual tracking,” in European Conference on Computer Vision. Springer, 2022, pp. 146–164

  78. [85]

    Foreground-background distribution modeling transformer for visual object tracking,

    D. Yang, J. He, Y . Ma, Q. Yu, and T. Zhang, “Foreground-background distribution modeling transformer for visual object tracking,” in Pro- ceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 10 117–10 127

  79. [86]

    Robust object modeling for visual tracking,

    Y . Cai, J. Liu, J. Tang, and G. Wu, “Robust object modeling for visual tracking,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 9589–9600

  80. [87]

    Mixformerv2: Efficient fully transformer tracking,

    Y . Cui, T. Song, G. Wu, and L. Wang, “Mixformerv2: Efficient fully transformer tracking,” Advances in Neural Information Processing Sys- tems, vol. 36, 2024

  81. [88]

    Reading relevant feature from global representation memory for visual object tracking,

    X. Zhou, P. Guo, L. Hong, J. Li, W. Zhang, W. Ge, and W. Zhang, “Reading relevant feature from global representation memory for visual object tracking,” Advances in Neural Information Processing Systems , vol. 36, 2023

  82. [89]

    Dropmae: Masked autoencoders with spatial-attention dropout for tracking tasks,

    Q. Wu, T. Yang, Z. Liu, B. Wu, Y . Shan, and A. B. Chan, “Dropmae: Masked autoencoders with spatial-attention dropout for tracking tasks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 14 561–14 571

  83. [90]

    Generalized relation modeling for transformer tracking,

    S. Gao, C. Zhou, and J. Zhang, “Generalized relation modeling for transformer tracking,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 18 686–18 695

  84. [91]

    Autoregressive visual tracking,

    X. Wei, Y . Bai, Y . Zheng, D. Shi, and Y . Gong, “Autoregressive visual tracking,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 9697–9706

  85. [92]

    Seqtrack: Sequence to sequence learning for visual object tracking,

    X. Chen, H. Peng, D. Wang, H. Lu, and H. Hu, “Seqtrack: Sequence to sequence learning for visual object tracking,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 14 572–14 581

  86. [93]

    Videotrack: Learning to track objects via video transformer,

    F. Xie, L. Chu, J. Li, Y . Lu, and C. Ma, “Videotrack: Learning to track objects via video transformer,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 22 826–22 835

  87. [94]

    Learning historical status prompt for accurate and robust visual tracking,

    W. Cai, Q. Liu, and Y . Wang, “Learning historical status prompt for accurate and robust visual tracking,” arXiv preprint arXiv:2311.02072, 2023

  88. [95]

    Odtrack: Online dense temporal token learning for visual tracking,

    Y . Zheng, B. Zhong, Q. Liang, Z. Mo, S. Zhang, and X. Li, “Odtrack: Online dense temporal token learning for visual tracking,” arXiv preprint arXiv:2401.01686, 2024

  89. [96]

    Explicit visual prompts for visual object tracking,

    L. Shi, B. Zhong, Q. Liang, N. Li, S. Zhang, and X. Li, “Explicit visual prompts for visual object tracking,” arXiv preprint arXiv:2401.03142 , 2024

  90. [97]

    Dynamic saliency- aware regularization for correlation filter-based object tracking,

    W. Feng, R. Han, Q. Guo, J. Zhu, and S. Wang, “Dynamic saliency- aware regularization for correlation filter-based object tracking,” IEEE Transactions on Image Processing, vol. 28, no. 7, pp. 3232–3245, 2019

  91. [98]

    Learning adaptive discrim- inative correlation filters via temporal consistency preserving spatial feature selection for robust visual object tracking,

    T. Xu, Z.-H. Feng, X.-J. Wu, and J. Kittler, “Learning adaptive discrim- inative correlation filters via temporal consistency preserving spatial feature selection for robust visual object tracking,” IEEE Transactions on Image Processing, vol. 28, no. 11, pp. 5596–5609, 2019

  92. [99]

    Quadruplet network with one-shot learning for fast visual object tracking,

    X. Dong, J. Shen, D. Wu, K. Guo, X. Jin, and F. Porikli, “Quadruplet network with one-shot learning for fast visual object tracking,” IEEE Transactions on Image Processing, vol. 28, no. 7, pp. 3516–3527, 2019

  93. [100]

    Visual tracking via dynamic graph learning,

    C. Li, L. Lin, W. Zuo, J. Tang, and M.-H. Yang, “Visual tracking via dynamic graph learning,” IEEE transactions on pattern analysis and machine intelligence, vol. 41, no. 11, pp. 2770–2782, 2018

  94. [101]

    Robust estimation of similarity transformation for visual object tracking,

    Y . Li, J. Zhu, S. C. Hoi, W. Song, Z. Wang, and H. Liu, “Robust estimation of similarity transformation for visual object tracking,” in Proceedings of the AAAI conference on artificial intelligence , vol. 33, no. 01, 2019, pp. 8666–8673

  95. [102]

    Spm-tracker: Series-parallel matching for real-time visual object tracking,

    G. Wang, C. Luo, Z. Xiong, and W. Zeng, “Spm-tracker: Series-parallel matching for real-time visual object tracking,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 3643–3652

  96. [103]

    Roi pooled correlation filters for visual tracking,

    Y . Sun, C. Sun, D. Wang, Y . He, and H. Lu, “Roi pooled correlation filters for visual tracking,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 5783–5791

  97. [104]

    Fast online object tracking and segmentation: A unifying approach,

    Q. Wang, L. Zhang, L. Bertinetto, W. Hu, and P. H. Torr, “Fast online object tracking and segmentation: A unifying approach,” in Proceedings of the IEEE/CVF conference on Computer Vision and Pattern Recognition, 2019, pp. 1328–1338

  98. [105]

    Target-aware deep tracking,

    X. Li, C. Ma, B. Wu, Z. He, and M.-H. Yang, “Target-aware deep tracking,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 1369–1378

  99. [106]

    Siamese cascaded region proposal networks for real-time visual tracking,

    H. Fan and H. Ling, “Siamese cascaded region proposal networks for real-time visual tracking,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 7952–7961

  100. [107]

    Visual tracking via adaptive spatially-regularized correlation filters,

    K. Dai, D. Wang, H. Lu, C. Sun, and J. Li, “Visual tracking via adaptive spatially-regularized correlation filters,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 4670–4679

  101. [108]

    Deeper and wider siamese networks for real- time visual tracking,

    Z. Zhang and H. Peng, “Deeper and wider siamese networks for real- time visual tracking,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 4591–4600

  102. [109]

    Graph convolutional tracking,

    J. Gao, T. Zhang, and C. Xu, “Graph convolutional tracking,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 4649–4659

  103. [110]

    Atom: Accurate tracking by overlap maximization,

    M. Danelljan, G. Bhat, F. S. Khan, and M. Felsberg, “Atom: Accurate tracking by overlap maximization,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 4660–4669

  104. [111]

    Bridging the gap between detection and tracking: A unified approach,

    L. Huang, X. Zhao, and K. Huang, “Bridging the gap between detection and tracking: A unified approach,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 3999–4009

  105. [112]

    Deep meta learning for real- time target-aware visual tracking,

    J. Choi, J. Kwon, and K. M. Lee, “Deep meta learning for real- time target-aware visual tracking,” in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 911–920

  106. [113]

    Gradnet: Gradient-guided network for visual object tracking,

    P. Li, B. Chen, W. Ouyang, D. Wang, X. Yang, and H. Lu, “Gradnet: Gradient-guided network for visual object tracking,” in Proceedings of the IEEE/CVF International conference on computer vision , 2019, pp. 6162–6171

  107. [114]

    Joint group feature selection and discriminative filter learning for robust visual object tracking,

    T. Xu, Z.-H. Feng, X.-J. Wu, and J. Kittler, “Joint group feature selection and discriminative filter learning for robust visual object tracking,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 7950–7960

  108. [115]

    Learning the model update for siamese trackers,

    L. Zhang, A. Gonzalez-Garcia, J. v. d. Weijer, M. Danelljan, and F. S. Khan, “Learning the model update for siamese trackers,” inProceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 4010–4019

  109. [116]

    ’skimming- perusal’tracking: A framework for real-time and robust long-term tracking,

    B. Yan, H. Zhao, D. Wang, H. Lu, and X. Yang, “’skimming- perusal’tracking: A framework for real-time and robust long-term tracking,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 2385–2393

  110. [117]

    Situp: Scale invariant tracking using average peak-to-correlation energy,

    H. Ma, S. T. Acton, and Z. Lin, “Situp: Scale invariant tracking using average peak-to-correlation energy,”IEEE Transactions on Image Processing, vol. 29, pp. 3546–3557, 2020

  111. [118]

    Robust visual tracking via constrained multi-kernel correlation filters,

    B. Huang, T. Xu, S. Jiang, Y . Chen, and Y . Bai, “Robust visual tracking via constrained multi-kernel correlation filters,” IEEE Transactions on Multimedia, vol. 22, no. 11, pp. 2820–2832, 2020. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 26

  112. [119]

    Real-time correlation tracking via joint model compression and transfer,

    N. Wang, W. Zhou, Y . Song, C. Ma, and H. Li, “Real-time correlation tracking via joint model compression and transfer,” IEEE Transactions on Image Processing, vol. 29, pp. 6123–6135, 2020

  113. [120]

    Mining spatial-temporal similarity for visual tracking,

    Y . Zhang, X. Gao, Z. Chen, H. Zhong, H. Xie, and C. Yan, “Mining spatial-temporal similarity for visual tracking,” IEEE Transactions on Image Processing, vol. 29, pp. 8107–8119, 2020

  114. [121]

    Fast learning of spatially regularized and content aware correlation filter for visual tracking,

    R. Han, W. Feng, and S. Wang, “Fast learning of spatially regularized and content aware correlation filter for visual tracking,” IEEE Transac- tions on Image Processing, vol. 29, pp. 7128–7140, 2020

  115. [122]

    Ensemble tracking based on diverse collaborative framework with multi-cue dynamic fusion,

    Y . Han, P. Zhang, T. Zhuo, W. Huang, Y . Zha, and Y . Zhang, “Ensemble tracking based on diverse collaborative framework with multi-cue dynamic fusion,” IEEE Transactions on Multimedia , vol. 22, no. 10, pp. 2698–2710, 2019

  116. [123]

    Deep object tracking with shrinkage loss,

    X. Lu, C. Ma, J. Shen, X. Yang, I. Reid, and M.-H. Yang, “Deep object tracking with shrinkage loss,” IEEE transactions on pattern analysis and machine intelligence, 2020

  117. [124]

    Discriminative and robust online learning for siamese visual tracking,

    J. Zhou, P. Wang, and H. Sun, “Discriminative and robust online learning for siamese visual tracking,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, no. 07, 2020, pp. 13 017– 13 024

  118. [125]

    Siamfc++: Towards robust and accurate visual tracking with target estimation guidelines,

    Y . Xu, Z. Wang, Z. Li, Y . Yuan, and G. Yu, “Siamfc++: Towards robust and accurate visual tracking with target estimation guidelines,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 07, 2020, pp. 12 549–12 556

  119. [126]

    Globaltrack: A simple and strong baseline for long-term tracking,

    L. Huang, X. Zhao, and K. Huang, “Globaltrack: A simple and strong baseline for long-term tracking,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 34, no. 07, 2020, pp. 11 037–11 044

  120. [127]

    Online decision based visual tracking via reinforcement learning,

    W. Zhang, R. Song, Y . Li et al., “Online decision based visual tracking via reinforcement learning,” Advances in Neural Information Process- ing Systems, vol. 33, pp. 11 778–11 788, 2020

  121. [128]

    Tlpg-tracker: Joint learning of target localization and proposal generation for visual tracking,

    S. Li, Z. Zhang, Z. Liu, A. Wang, L. Qiu, and F. Du, “Tlpg-tracker: Joint learning of target localization and proposal generation for visual tracking,” in Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelligence , 2...

  122. [129]

    Know your surroundings: Exploiting scene information for object tracking,

    G. Bhat, M. Danelljan, L. V . Gool, and R. Timofte, “Know your surroundings: Exploiting scene information for object tracking,” in European Conference on Computer Vision . Springer, 2020, pp. 205– 221

  123. [130]

    Pg-net: Pixel to global matching network for visual tracking,

    B. Liao, C. Wang, Y . Wang, Y . Wang, and J. Yin, “Pg-net: Pixel to global matching network for visual tracking,” in European Conference on Computer Vision. Springer, 2020, pp. 429–444

  124. [131]

    Object tracking using spatio-temporal networks for future prediction location,

    Y . Liu, R. Li, Y . Cheng, R. T. Tan, and X. Sui, “Object tracking using spatio-temporal networks for future prediction location,” in European Conference on Computer Vision. Springer, 2020, pp. 1–17

  125. [132]

    Ocean: Object-aware anchor-free tracking,

    Z. Zhang, H. Peng, J. Fu, B. Li, and W. Hu, “Ocean: Object-aware anchor-free tracking,” in European Conference on Computer Vision . Springer, 2020, pp. 771–787

  126. [133]

    Tracking by instance detection: A meta-learning approach,

    G. Wang, C. Luo, X. Sun, Z. Xiong, and W. Zeng, “Tracking by instance detection: A meta-learning approach,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 6288– 6297

  127. [134]

    Siam r-cnn: Visual tracking by re-detection,

    P. V oigtlaender, J. Luiten, P. H. Torr, and B. Leibe, “Siam r-cnn: Visual tracking by re-detection,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 6578–6588

  128. [135]

    Roam: Recurrently op- timizing tracking model,

    T. Yang, P. Xu, R. Hu, H. Chai, and A. B. Chan, “Roam: Recurrently op- timizing tracking model,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 6718–6727

  129. [136]

    Recursive least-squares estimator-aided online learning for visual tracking,

    J. Gao, W. Hu, and Y . Lu, “Recursive least-squares estimator-aided online learning for visual tracking,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 7386–7395

  130. [137]

    Probabilistic regression for visual tracking,

    M. Danelljan, L. V . Gool, and R. Timofte, “Probabilistic regression for visual tracking,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 7183–7192

  131. [139]

    Deformable siamese attention networks for visual object tracking,

    Y . Yu, Y . Xiong, W. Huang, and M. R. Scott, “Deformable siamese attention networks for visual object tracking,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2020, pp. 6728–6737

  132. [140]

    Correlation-guided attention for corner detection based visual tracking,

    F. Du, P. Liu, W. Zhao, and X. Tang, “Correlation-guided attention for corner detection based visual tracking,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 6836–6845

  133. [141]

    Learning target candidate association to keep track of what not to track,

    C. Mayer, M. Danelljan, D. P. Paudel, and L. Van Gool, “Learning target candidate association to keep track of what not to track,” inProceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 13 444–13 454

  134. [143]

    Learn to match: Automatic matching network design for visual tracking,

    Z. Zhang, Y . Liu, X. Wang, B. Li, and W. Hu, “Learn to match: Automatic matching network design for visual tracking,” inProceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 13 339–13 348

  135. [144]

    Saliency- associated object tracking,

    Z. Zhou, W. Pei, X. Li, H. Wang, F. Zheng, and Z. He, “Saliency- associated object tracking,” in Proceedings of the IEEE/CVF Interna- tional Conference on Computer Vision, 2021, pp. 9866–9875

  136. [145]

    Learning spatio-temporal transformer for visual tracking,

    B. Yan, H. Peng, J. Fu, D. Wang, and H. Lu, “Learning spatio-temporal transformer for visual tracking,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 10 448–10 457

  137. [146]

    Adaptive multi-feature reliability re- determinative correlation filter for visual tracking,

    M. Guan and C. Wen, “Adaptive multi-feature reliability re- determinative correlation filter for visual tracking,” IEEE Transactions on Multimedia, vol. 23, pp. 3841–3852, 2020

  138. [147]

    Siamese tracking network with informative enhanced loss,

    S. Tian, X. Liu, M. Liu, S. Li, and B. Yin, “Siamese tracking network with informative enhanced loss,” IEEE Transactions on Multimedia , vol. 23, pp. 120–132, 2020

  139. [148]

    Cat: Corner aided tracking with deep regression network,

    S. Zhang, X. Zhao, and L. Fang, “Cat: Corner aided tracking with deep regression network,” IEEE Transactions on Multimedia , vol. 23, pp. 859–870, 2020

  140. [149]

    Learning recurrent memory activation networks for visual tracking,

    S. Pu, Y . Song, C. Ma, H. Zhang, and M.-H. Yang, “Learning recurrent memory activation networks for visual tracking,” IEEE Transactions on Image Processing, vol. 30, pp. 725–738, 2020

  141. [150]

    Visual tracking via dynamic memory networks,

    T. Yang and A. B. Chan, “Visual tracking via dynamic memory networks,” IEEE transactions on pattern analysis and machine intel- ligence, vol. 43, no. 1, pp. 360–374, 2019

  142. [151]

    Nocal- siam: Refining visual features and response with advanced non-local blocks for real-time siamese tracking,

    H. Tan, X. Zhang, Z. Zhang, L. Lan, W. Zhang, and Z. Luo, “Nocal- siam: Refining visual features and response with advanced non-local blocks for real-time siamese tracking,” IEEE Transactions on Image Processing, vol. 30, pp. 2656–2668, 2021

  143. [152]

    Siamcan: Real- time visual tracking based on siamese center-aware network,

    W. Zhou, L. Wen, L. Zhang, D. Du, T. Luo, and Y . Wu, “Siamcan: Real- time visual tracking based on siamese center-aware network,” IEEE Transactions on Image Processing, vol. 30, pp. 3597–3609, 2021

  144. [153]

    Dy- namical hyperparameter optimization via deep reinforcement learning in tracking,

    X. Dong, J. Shen, W. Wang, L. Shao, H. Ling, and F. Porikli, “Dy- namical hyperparameter optimization via deep reinforcement learning in tracking,” IEEE transactions on pattern analysis and machine intel- ligence, vol. 43, no. 5, pp. 1515–1529, 2019

  145. [154]

    Toward accurate pixelwise object tracking via attention retrieval,

    Z. Zhang, Y . Liu, B. Li, W. Hu, and H. Peng, “Toward accurate pixelwise object tracking via attention retrieval,” IEEE Transactions on Image Processing, vol. 30, pp. 8553–8566, 2021

  146. [155]

    Transformer meets tracker: Exploiting temporal context for robust visual tracking,

    N. Wang, W. Zhou, J. Wang, and H. Li, “Transformer meets tracker: Exploiting temporal context for robust visual tracking,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2021, pp. 1571–1580

  147. [156]

    Transformer tracking,

    X. Chen, B. Yan, J. Zhu, D. Wang, X. Yang, and H. Lu, “Transformer tracking,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 8126–8135

  148. [157]

    Distractor-aware fast tracking via dynamic convolutions and mot philosophy,

    Z. Zhang, B. Zhong, S. Zhang, Z. Tang, X. Liu, and Z. Zhang, “Distractor-aware fast tracking via dynamic convolutions and mot philosophy,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 1024–1033

  149. [158]

    Rotation equivariant siamese networks for tracking,

    D. K. Gupta, D. Arya, and E. Gavves, “Rotation equivariant siamese networks for tracking,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 12 362–12 371

  150. [159]

    Learning to fuse asymmetric feature maps in siamese trackers,

    W. Han, X. Dong, F. S. Khan, L. Shao, and J. Shen, “Learning to fuse asymmetric feature maps in siamese trackers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 16 570–16 580

  151. [160]

    Learning to filter: Siamese relation network for robust tracking,

    S. Cheng, B. Zhong, G. Li, X. Liu, Z. Tang, X. Li, and J. Wang, “Learning to filter: Siamese relation network for robust tracking,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 4421–4431

  152. [161]

    Graph attention tracking,

    D. Guo, Y . Shao, Y . Cui, Z. Wang, L. Zhang, and C. Shen, “Graph attention tracking,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 9543–9552

  153. [162]

    Capsulerrt: Relationships-aware regression tracking via capsules,

    D. Ma and X. Wu, “Capsulerrt: Relationships-aware regression tracking via capsules,” in Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, 2021, pp. 10 948–10 957

  154. [163]

    Alpha-refine: Boosting tracking performance by precise bounding box estimation,

    B. Yan, X. Zhang, D. Wang, H. Lu, and X. Yang, “Alpha-refine: Boosting tracking performance by precise bounding box estimation,” JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 27 in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2...

  155. [165]

    Model uncertainty guides visual object tracking,

    L. Zhou, A. Ledent, Q. Hu, T. Liu, J. Zhang, and M. Kloft, “Model uncertainty guides visual object tracking,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 35, no. 4, 2021, pp. 3581– 3589

  156. [166]

    Visual tracking via hierarchical deep reinforcement learning,

    D. Zhang, Z. Zheng, R. Jia, and M. Li, “Visual tracking via hierarchical deep reinforcement learning,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, no. 4, 2021, pp. 3315–3323

  157. [167]

    Multi-kernel correlation filter for visual tracking,

    M. Tang and J. Feng, “Multi-kernel correlation filter for visual tracking,” in Proceedings of the IEEE international conference on computer vision, 2015, pp. 3038–3046

  158. [168]

    Joint scale-spatial correlation tracking with adaptive rotation estimation,

    M. Zhang, J. Xing, J. Gao, X. Shi, Q. Wang, and W. Hu, “Joint scale-spatial correlation tracking with adaptive rotation estimation,” in Proceedings of the IEEE international conference on computer vision workshops, 2015, pp. 32–40

  159. [169]

    Learning spatially regularized correlation filters for visual tracking,

    M. Danelljan, G. Hager, F. Shahbaz Khan, and M. Felsberg, “Learning spatially regularized correlation filters for visual tracking,” in Proceed- ings of the IEEE international conference on computer vision, 2015, pp. 4310–4318

  160. [170]

    Hierarchical con- volutional features for visual tracking,

    C. Ma, J.-B. Huang, X. Yang, and M.-H. Yang, “Hierarchical con- volutional features for visual tracking,” in Proceedings of the IEEE international conference on computer vision, 2015, pp. 3074–3082

  161. [171]

    Long-term correlation tracking,

    C. Ma, X. Yang, C. Zhang, and M.-H. Yang, “Long-term correlation tracking,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 5388–5396

  162. [172]

    Multi- store tracker (muster): A cognitive psychology inspired approach to object tracking,

    Z. Hong, Z. Chen, C. Wang, X. Mei, D. Prokhorov, and D. Tao, “Multi- store tracker (muster): A cognitive psychology inspired approach to object tracking,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 749–758

  163. [173]

    Con- volutional features for correlation filter based visual tracking,

    M. Danelljan, G. Hager, F. Shahbaz Khan, and M. Felsberg, “Con- volutional features for correlation filter based visual tracking,” in Proceedings of the IEEE international conference on computer vision workshops, 2015, pp. 58–66

  164. [174]

    Online tracking by learning discriminative saliency map with convolutional neural network,

    S. Hong, T. You, S. Kwak, and B. Han, “Online tracking by learning discriminative saliency map with convolutional neural network,” in International conference on machine learning. PMLR, 2015, pp. 597– 606

  165. [175]

    Visual tracking with fully convolutional networks,

    L. Wang, W. Ouyang, X. Wang, and H. Lu, “Visual tracking with fully convolutional networks,” inProceedings of the IEEE international conference on computer vision, 2015, pp. 3119–3127

  166. [176]

    Once for all: a two-flow convolutional neural net- work for visual tracking,

    K. Chen and W. Tao, “Once for all: a two-flow convolutional neural net- work for visual tracking,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 28, no. 12, pp. 3377–3386, 2017

  167. [177]

    Learning feed-forward one-shot learners,

    L. Bertinetto, J. F. Henriques, J. Valmadre, P. Torr, and A. Vedaldi, “Learning feed-forward one-shot learners,” Advances in neural infor- mation processing systems, vol. 29, 2016

  168. [178]

    Siamese instance search for tracking,

    R. Tao, E. Gavves, and A. W. Smeulders, “Siamese instance search for tracking,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 1420–1429

  169. [179]

    Stct: Sequentially training convolutional networks for visual tracking,

    L. Wang, W. Ouyang, X. Wang, and H. Lu, “Stct: Sequentially training convolutional networks for visual tracking,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 1373– 1381

  170. [180]

    Learning to track at 100 fps with deep regression networks,

    D. Held, S. Thrun, and S. Savarese, “Learning to track at 100 fps with deep regression networks,” in European conference on computer vision. Springer, 2016, pp. 749–765

  171. [181]

    Robust visual tracking via convolutional networks without training,

    K. Zhang, Q. Liu, Y . Wu, and M.-H. Yang, “Robust visual tracking via convolutional networks without training,” IEEE Transactions on Image Processing, vol. 25, no. 4, pp. 1779–1792, 2016

  172. [182]

    Staple: Complementary learners for real-time tracking,

    L. Bertinetto, J. Valmadre, S. Golodetz, O. Miksik, and P. H. Torr, “Staple: Complementary learners for real-time tracking,” inProceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 1401–1409

  173. [183]

    Hedged deep tracking,

    Y . Qi, S. Zhang, L. Qin, H. Yao, Q. Huang, J. Lim, and M.-H. Yang, “Hedged deep tracking,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 4303–4311

  174. [184]

    Deep motion features for visual tracking,

    S. Gladh, M. Danelljan, F. S. Khan, and M. Felsberg, “Deep motion features for visual tracking,” in 2016 23rd international conference on pattern recognition (ICPR). IEEE, 2016, pp. 1243–1248

  175. [185]

    Adaptive decontamination of the training set: A unified formulation for discrim- inative visual tracking,

    M. Danelljan, G. Hager, F. Shahbaz Khan, and M. Felsberg, “Adaptive decontamination of the training set: A unified formulation for discrim- inative visual tracking,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 1430–1438

  176. [187]

    Sanet: Structure-aware network for visual track- ing,

    H. Fan and H. Ling, “Sanet: Structure-aware network for visual track- ing,” in Proceedings of the IEEE conference on computer vision and pattern recognition workshops, 2017, pp. 42–49

  177. [188]

    Action-decision networks for visual tracking with deep reinforcement learning,

    S. Yun, J. Choi, Y . Yoo, K. Yun, and J. Young Choi, “Action-decision networks for visual tracking with deep reinforcement learning,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 2711–2720

  178. [189]

    Context-aware correlation filter tracking,

    M. Mueller, N. Smith, and B. Ghanem, “Context-aware correlation filter tracking,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 1396–1404

  179. [190]

    Dcco: Towards deformable continuous convolution operators for visual track- ing,

    J. Johnander, M. Danelljan, F. S. Khan, and M. Felsberg, “Dcco: Towards deformable continuous convolution operators for visual track- ing,” in International Conference on Computer Analysis of Images and Patterns. Springer, 2017, pp. 55–67

  180. [191]

    Learning background- aware correlation filters for visual tracking,

    H. Kiani Galoogahi, A. Fagg, and S. Lucey, “Learning background- aware correlation filters for visual tracking,” inProceedings of the IEEE international conference on computer vision, 2017, pp. 1135–1143

  181. [192]

    Discriminative correlation filter with channel and spatial reliability,

    A. Lukezic, T. V ojir, L. ˇCehovin Zajc, J. Matas, and M. Kristan, “Discriminative correlation filter with channel and spatial reliability,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 6309–6318

  182. [193]

    Multi-task correlation particle filter for robust object tracking,

    T. Zhang, C. Xu, and M.-H. Yang, “Multi-task correlation particle filter for robust object tracking,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 4335–4343

  183. [194]

    End-to-end representation learning for correlation filter based track- ing,

    J. Valmadre, L. Bertinetto, J. Henriques, A. Vedaldi, and P. H. Torr, “End-to-end representation learning for correlation filter based track- ing,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 2805–2813

  184. [195]

    Large margin object tracking with circulant feature maps,

    M. Wang, Y . Liu, and Z. Huang, “Large margin object tracking with circulant feature maps,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 4021–4029

  185. [196]

    Robust visual tracking using oblique random forests,

    L. Zhang, J. Varadarajan, P. Nagaratnam Suganthan, N. Ahuja, and P. Moulin, “Robust visual tracking using oblique random forests,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 5589–5598

  186. [198]

    Robust object tracking based on temporal and spatial deep networks,

    Z. Teng, J. Xing, Q. Wang, C. Lang, S. Feng, and Y . Jin, “Robust object tracking based on temporal and spatial deep networks,” in Proceedings of the IEEE International Conference on Computer Vision , 2017, pp. 1144–1153

  187. [199]

    Good features to correlate for visual tracking,

    E. Gundogdu and A. A. Alatan, “Good features to correlate for visual tracking,” IEEE Transactions on Image Processing , vol. 27, no. 5, pp. 2526–2540, 2018

  188. [200]

    Recurrent filter learning for visual tracking,

    T. Yang and A. B. Chan, “Recurrent filter learning for visual tracking,” in Proceedings of the IEEE International Conference on Computer Vision Workshops, 2017, pp. 2010–2019

  189. [201]

    Dual deep network for visual tracking,

    Z. Chi, H. Li, H. Lu, and M.-H. Yang, “Dual deep network for visual tracking,” IEEE Transactions on Image Processing , vol. 26, no. 4, pp. 2005–2015, 2017

  190. [202]

    Robust structural sparse tracking,

    T. Zhang, C. Xu, and M.-H. Yang, “Robust structural sparse tracking,” IEEE transactions on pattern analysis and machine intelligence, vol. 41, no. 2, pp. 473–486, 2018

  191. [203]

    Online object tracking, learning and parsing with and-or graphs,

    Y . Lu, T. Wu, and S. Chun Zhu, “Online object tracking, learning and parsing with and-or graphs,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2014, pp. 3462–3469

  192. [204]

    Integrating boundary and center correlation filters for visual tracking with aspect ratio variation,

    F. Li, Y . Yao, P. Li, D. Zhang, W. Zuo, and M.-H. Yang, “Integrating boundary and center correlation filters for visual tracking with aspect ratio variation,” in Proceedings of the IEEE International Conference on Computer Vision Workshops, 2017, pp. 2001–2009

  193. [205]

    Correlation filters with weighted convolution responses,

    Z. He, Y . Fan, J. Zhuang, Y . Dong, and H. Bai, “Correlation filters with weighted convolution responses,” in Proceedings of the IEEE International Conference on Computer Vision Workshops , 2017, pp. 1992–2000

  194. [206]

    Learning dynamic siamese network for visual object tracking,

    Q. Guo, W. Feng, C. Zhou, R. Huang, L. Wan, and S. Wang, “Learning dynamic siamese network for visual object tracking,” in Proceedings of the IEEE international conference on computer vision, 2017, pp. 1763– 1771

  195. [207]

    Learning policies for adaptive tracking with deep feature cascades,

    C. Huang, S. Lucey, and D. Ramanan, “Learning policies for adaptive tracking with deep feature cascades,” in Proceedings of the IEEE international conference on computer vision, 2017, pp. 105–114. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 28

  196. [208]

    Crest: Convolutional residual learning for visual tracking,

    Y . Song, C. Ma, L. Gong, J. Zhang, R. W. Lau, and M.-H. Yang, “Crest: Convolutional residual learning for visual tracking,” in Proceedings of the IEEE international conference on computer vision, 2017, pp. 2555– 2564

  197. [209]

    End-to-end flow correlation tracking with spatial-temporal attention,

    Z. Zhu, W. Wu, W. Zou, and J. Yan, “End-to-end flow correlation tracking with spatial-temporal attention,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 548– 557

  198. [210]

    Vital: Visual tracking via adversarial learning,

    Y . Song, C. Ma, X. Wu, L. Gong, L. Bao, W. Zuo, C. Shen, R. W. Lau, and M.-H. Yang, “Vital: Visual tracking via adversarial learning,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 8990–8999

  199. [211]

    P2t: Part-to-target tracking via deep regression learning,

    J. Gao, T. Zhang, X. Yang, and C. Xu, “P2t: Part-to-target tracking via deep regression learning,” IEEE Transactions on Image Processing, vol. 27, no. 6, pp. 3074–3086, 2018

  200. [212]

    Iterative graph seeking for object tracking,

    D. Du, L. Wen, H. Qi, Q. Huang, Q. Tian, and S. Lyu, “Iterative graph seeking for object tracking,” IEEE Transactions on Image Processing , vol. 27, no. 4, pp. 1809–1821, 2017

  201. [213]

    Learning to update for object tracking with recurrent meta-learner,

    B. Li, W. Xie, W. Zeng, and W. Liu, “Learning to update for object tracking with recurrent meta-learner,” IEEE Transactions on Image Processing, vol. 28, no. 7, pp. 3624–3635, 2019

  202. [214]

    Visual tracking with weighted adaptive local sparse appearance model via spatio-temporal context learning,

    Z. Li, J. Zhang, K. Zhang, and Z. Li, “Visual tracking with weighted adaptive local sparse appearance model via spatio-temporal context learning,” IEEE Transactions on Image Processing , vol. 27, no. 9, pp. 4478–4489, 2018

  203. [215]

    Deformable object tracking with gated fusion,

    W. Liu, Y . Song, D. Chen, S. He, Y . Yu, T. Yan, G. P. Hancke, and R. W. Lau, “Deformable object tracking with gated fusion,” IEEE Transactions on Image Processing, vol. 28, no. 8, pp. 3766–3777, 2019

  204. [216]

    Learning attentions: residual attentional siamese network for high performance online visual tracking,

    Q. Wang, Z. Teng, J. Xing, J. Gao, W. Hu, and S. Maybank, “Learning attentions: residual attentional siamese network for high performance online visual tracking,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 4854–4863

  205. [217]

    Robust visual tracking revisited: From correlation filter to template matching,

    F. Liu, C. Gong, X. Huang, T. Zhou, J. Yang, and D. Tao, “Robust visual tracking revisited: From correlation filter to template matching,” IEEE Transactions on Image Processing, vol. 27, no. 6, pp. 2777–2790, 2018

  206. [218]

    Efficient correlation tracking via center-biased spatial regularization,

    Y . Zhou, J. Han, F. Yang, K. Zhang, and R. Hong, “Efficient correlation tracking via center-biased spatial regularization,” IEEE Transactions on Image Processing, vol. 27, no. 12, pp. 6159–6173, 2018

  207. [219]

    Correlation tracking via joint discrimination and reliability learning,

    C. Sun, D. Wang, H. Lu, and M.-H. Yang, “Correlation tracking via joint discrimination and reliability learning,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 489– 497

  208. [220]

    Hyper- parameter optimization for tracking with continuous deep q-learning,

    X. Dong, J. Shen, W. Wang, Y . Liu, L. Shao, and F. Porikli, “Hyper- parameter optimization for tracking with continuous deep q-learning,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 518–527

  209. [221]

    A twofold siamese network for real-time object tracking,

    A. He, C. Luo, X. Tian, and W. Zeng, “A twofold siamese network for real-time object tracking,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 4834–4843

  210. [223]

    Unveiling the power of deep tracking,

    G. Bhat, J. Johnander, M. Danelljan, F. S. Khan, and M. Felsberg, “Unveiling the power of deep tracking,” inProceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 483–498

  211. [224]

    Learning spatial- temporal regularized correlation filters for visual tracking,

    F. Li, C. Tian, W. Zuo, L. Zhang, and M.-H. Yang, “Learning spatial- temporal regularized correlation filters for visual tracking,” in Proceed- ings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 4904–4913

  212. [225]

    Multi- cue correlation filters for robust visual tracking,

    N. Wang, W. Zhou, Q. Tian, R. Hong, M. Wang, and H. Li, “Multi- cue correlation filters for robust visual tracking,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 4844–4853

  213. [226]

    Sint++: Robust visual tracking via adversarial positive instance generation,

    X. Wang, C. Li, B. Luo, and J. Tang, “Sint++: Robust visual tracking via adversarial positive instance generation,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 4864– 4873

  214. [227]

    Deep attentive tracking via reciprocative learning,

    S. Pu, Y . Song, C. Ma, H. Zhang, and M.-H. Yang, “Deep attentive tracking via reciprocative learning,” Advances in neural information processing systems, vol. 31, 2018

  215. [228]

    Learning dynamic memory networks for ob- ject tracking,

    T. Yang and A. B. Chan, “Learning dynamic memory networks for ob- ject tracking,” in Proceedings of the European conference on computer vision (ECCV), 2018, pp. 152–167

  216. [229]

    Distractor-aware siamese networks for visual object tracking,

    Z. Zhu, Q. Wang, B. Li, W. Wu, J. Yan, and W. Hu, “Distractor-aware siamese networks for visual object tracking,” in Proceedings of the European conference on computer vision (ECCV), 2018, pp. 101–117

  217. [230]

    Deep reinforcement learning with iterative shift for visual tracking,

    L. Ren, X. Yuan, J. Lu, M. Yang, and J. Zhou, “Deep reinforcement learning with iterative shift for visual tracking,” in Proceedings of the European conference on computer vision (ECCV), 2018, pp. 684–700

  218. [231]

    Deep regression tracking with shrinkage loss,

    X. Lu, C. Ma, B. Ni, X. Yang, I. Reid, and M.-H. Yang, “Deep regression tracking with shrinkage loss,” inProceedings of the European conference on computer vision (ECCV), 2018, pp. 353–369

  219. [232]

    Real-time’actor- critic’tracking,

    B. Chen, D. Wang, P. Li, S. Wang, and H. Lu, “Real-time’actor- critic’tracking,” inProceedings of the European conference on computer vision (ECCV), 2018, pp. 318–334

  220. [233]

    Real-time mdnet,

    I. Jung, J. Son, M. Baek, and B. Han, “Real-time mdnet,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 83– 98

  221. [234]

    Structured siamese network for real-time visual tracking,

    Y . Zhang, L. Wang, J. Qi, D. Wang, M. Feng, and H. Lu, “Structured siamese network for real-time visual tracking,” in Proceedings of the European conference on computer vision (ECCV), 2018, pp. 351–366

  222. [235]

    Visual tracking via spatially aligned correlation filters network,

    M. Zhang, Q. Wang, J. Xing, J. Gao, P. Peng, W. Hu, and S. Maybank, “Visual tracking via spatially aligned correlation filters network,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 469–485

  223. [236]

    Triplet loss in siamese network for object tracking,

    X. Dong and J. Shen, “Triplet loss in siamese network for object tracking,” in Proceedings of the European conference on computer vision (ECCV), 2018, pp. 459–474

  224. [237]

    Real-time compressive track- ing,

    K. Zhang, L. Zhang, and M.-H. Yang, “Real-time compressive track- ing,” in European conference on computer vision. Springer, 2012, pp. 864–877

  225. [238]

    Robust object tracking with a hierarchical ensemble framework,

    M. Wang, Y . Liu, and R. Xiong, “Robust object tracking with a hierarchical ensemble framework,” in 2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2016, pp. 438–445

  226. [239]

    Adaptive color attributes for real-time visual tracking,

    M. Danelljan, F. Shahbaz Khan, M. Felsberg, and J. Van de Weijer, “Adaptive color attributes for real-time visual tracking,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2014, pp. 1090–1097

  227. [240]

    Adaptive color attributes for real-time visual tracking,

    ——, “Adaptive color attributes for real-time visual tracking,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2014, pp. 1090–1097

  228. [241]

    Robust object tracking with online multiple instance learning,

    B. Babenko, M.-H. Yang, and S. Belongie, “Robust object tracking with online multiple instance learning,” IEEE transactions on pattern analysis and machine intelligence, vol. 33, no. 8, pp. 1619–1632, 2010

  229. [242]

    Accurate scale estimation for robust visual tracking,

    M. Danelljan, G. H ¨ager, F. Khan, and M. Felsberg, “Accurate scale estimation for robust visual tracking,” in British Machine Vision Con- ference, Nottingham, September 1-5, 2014. Bmva Press, 2014

  230. [243]

    Learning a deep compact image represen- tation for visual tracking,

    N. Wang and D.-Y . Yeung, “Learning a deep compact image represen- tation for visual tracking,” Advances in neural information processing systems, vol. 26, 2013

  231. [244]

    Visual ob- ject tracking using adaptive correlation filters,

    D. S. Bolme, J. R. Beveridge, B. A. Draper, and Y . M. Lui, “Visual ob- ject tracking using adaptive correlation filters,” in 2010 IEEE computer society conference on computer vision and pattern recognition. IEEE, 2010, pp. 2544–2550

  232. [245]

    Exploiting the circulant structure of tracking-by-detection with kernels,

    J. F. Henriques, R. Caseiro, P. Martins, and J. Batista, “Exploiting the circulant structure of tracking-by-detection with kernels,” in European conference on computer vision. Springer, 2012, pp. 702–715

  233. [246]

    Fast visual tracking via dense spatio-temporal context learning,

    K. Zhang, L. Zhang, Q. Liu, D. Zhang, and M.-H. Yang, “Fast visual tracking via dense spatio-temporal context learning,” in European conference on computer vision. Springer, 2014, pp. 127–141

  234. [247]

    A scale adaptive kernel correlation filter tracker with feature integration,

    Y . Li and J. Zhu, “A scale adaptive kernel correlation filter tracker with feature integration,” in European conference on computer vision . Springer, 2014, pp. 254–265

  235. [248]

    Reliable patch trackers: Robust visual tracking by exploiting reliable patches,

    Y . Li, J. Zhu, and S. C. Hoi, “Reliable patch trackers: Robust visual tracking by exploiting reliable patches,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2015, pp. 353–361

  236. [249]

    Correlation filters with lim- ited boundaries,

    H. Kiani Galoogahi, T. Sim, and S. Lucey, “Correlation filters with lim- ited boundaries,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 4630–4638

  237. [250]

    Target response adaptation for correlation filter tracking,

    A. Bibi, M. Mueller, and B. Ghanem, “Target response adaptation for correlation filter tracking,” in European conference on computer vision. Springer, 2016, pp. 419–433

  238. [251]

    Real-time visual tracking: Promoting the robustness of correlation filter learning,

    Y . Sui, Z. Zhang, G. Wang, Y . Tang, and L. Zhang, “Real-time visual tracking: Promoting the robustness of correlation filter learning,” in European conference on computer vision . Springer, 2016, pp. 662– 678

  239. [252]

    Structural correlation filter for robust visual tracking,

    S. Liu, T. Zhang, X. Cao, and C. Xu, “Structural correlation filter for robust visual tracking,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 4312–4320. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 29

  240. [253]

    Parallel tracking and verifying: A framework for real-time and high accuracy visual tracking,

    H. Fan and H. Ling, “Parallel tracking and verifying: A framework for real-time and high accuracy visual tracking,” inProceedings of the IEEE international conference on computer vision, 2017, pp. 5486–5494

  241. [254]

    Learning spatial-aware regressions for visual tracking,

    C. Sun, D. Wang, H. Lu, and M.-H. Yang, “Learning spatial-aware regressions for visual tracking,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 8962–8970

  242. [255]

    Be- yond correlation filters: Learning continuous convolution operators for visual tracking,

    M. Danelljan, A. Robinson, F. Shahbaz Khan, and M. Felsberg, “Be- yond correlation filters: Learning continuous convolution operators for visual tracking,” in European conference on computer vision. Springer, 2016, pp. 472–488

  243. [256]

    Eco: Efficient convolution operators for tracking,

    M. Danelljan, G. Bhat, F. Shahbaz Khan, and M. Felsberg, “Eco: Efficient convolution operators for tracking,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 6638– 6646

  244. [257]

    High- performance long-term tracking with meta-updater,

    K. Dai, Y . Zhang, D. Wang, J. Li, H. Lu, and X. Yang, “High- performance long-term tracking with meta-updater,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2020, pp. 6298–6307

  245. [258]

    Signature verification using a

    J. Bromley, I. Guyon, Y . LeCun, E. S ¨ackinger, and R. Shah, “Signature verification using a” siamese” time delay neural network,” Advances in neural information processing systems, vol. 6, 1993

  246. [259]

    Facenet: A unified embed- ding for face recognition and clustering,

    F. Schroff, D. Kalenichenko, and J. Philbin, “Facenet: A unified embed- ding for face recognition and clustering,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 815– 823

  247. [260]

    Deep face recognition,

    O. M. Parkhi, A. Vedaldi, and A. Zisserman, “Deep face recognition,” 2015

  248. [261]

    Exploring simple siamese representation learning,

    X. Chen and K. He, “Exploring simple siamese representation learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 15 750–15 758

  249. [262]

    Video object segmentation using space-time memory networks,

    S. W. Oh, J.-Y . Lee, N. Xu, and S. J. Kim, “Video object segmentation using space-time memory networks,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 9226–9235

  250. [263]

    Delving deeper into mask utilization in video object segmentation,

    M. Wang, J. Mei, L. Liu, G. Tian, Y . Liu, and Z. Pan, “Delving deeper into mask utilization in video object segmentation,” IEEE Transactions on Image Processing, 2022

  251. [264]

    Faster r-cnn: Towards real-time object detection with region proposal networks,

    S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” Advances in neural information processing systems, vol. 28, 2015

  252. [265]

    Cract: Cascaded regression-align-classification for robust tracking,

    H. Fan and H. Ling, “Cract: Cascaded regression-align-classification for robust tracking,” in 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2021, pp. 7013–7020

  253. [266]

    High-performance discriminative tracking with transformers,

    B. Yu, M. Tang, L. Zheng, G. Zhu, J. Wang, H. Feng, X. Feng, and H. Lu, “High-performance discriminative tracking with transformers,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 9856–9865

  254. [267]

    Human tracking using convolu- tional neural networks,

    J. Fan, W. Xu, Y . Wu, and Y . Gong, “Human tracking using convolu- tional neural networks,”IEEE transactions on Neural Networks, vol. 21, no. 10, pp. 1610–1623, 2010

  255. [268]

    Robust online visual tracking with a single convolutional neural network,

    H. Li, Y . Li, and F. Porikli, “Robust online visual tracking with a single convolutional neural network,” inAsian Conference on Computer Vision. Springer, 2014, pp. 194–209

  256. [269]

    Transferring rich feature hierarchies for robust visual tracking,

    N. Wang, S. Li, A. Gupta, and D.-Y . Yeung, “Transferring rich feature hierarchies for robust visual tracking,” arXiv preprint arXiv:1501.04587, 2015

  257. [270]

    Fcos: Fully convolutional one- stage object detection,

    Z. Tian, C. Shen, H. Chen, and T. He, “Fcos: Fully convolutional one- stage object detection,” in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 9627–9636

  258. [271]

    Focal loss for dense object detection,

    T.-Y . Lin, P. Goyal, R. Girshick, K. He, and P. Doll ´ar, “Focal loss for dense object detection,” in Proceedings of the IEEE international conference on computer vision, 2017, pp. 2980–2988

  259. [272]

    Learning multi-domain convolutional neural networks for visual tracking,

    H. Nam and B. Han, “Learning multi-domain convolutional neural networks for visual tracking,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 4293–4302

  260. [273]

    Meta-tracker: Fast and robust online adaptation for visual object trackers,

    E. Park and A. C. Berg, “Meta-tracker: Fast and robust online adaptation for visual object trackers,” in Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 569–585

  261. [274]

    Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers,

    S. Zheng, J. Lu, H. Zhao, X. Zhu, Z. Luo, Y . Wang, Y . Fu, J. Feng, T. Xiang, P. H. Torr et al., “Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition ...

  262. [275]

    Mail: A unified mask-image- language trimodal network for referring image segmentation,

    Z. Li, M. Wang, J. Mei, and Y . Liu, “Mail: A unified mask-image- language trimodal network for referring image segmentation,” arXiv preprint arXiv:2111.10747, 2021

  263. [276]

    Pix2seq: A language modeling framework for object detection,

    T. Chen, S. Saxena, L. Li, D. J. Fleet, and G. Hinton, “Pix2seq: A language modeling framework for object detection,” arXiv preprint arXiv:2109.10852, 2021

  264. [277]

    Actionclip: Adapting language-image pretrained models for video action recognition,

    M. Wang, J. Xing, J. Mei, Y . Liu, and Y . Jiang, “Actionclip: Adapting language-image pretrained models for video action recognition,” IEEE Transactions on Neural Networks and Learning Systems, 2023

  265. [278]

    Multiscale vision transformers,

    H. Fan, B. Xiong, K. Mangalam, Y . Li, Z. Yan, J. Malik, and C. Fe- ichtenhofer, “Multiscale vision transformers,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 6824–6835

  266. [279]

    Vilt: Vision-and-language transformer without convolution or region supervision,

    W. Kim, B. Son, and I. Kim, “Vilt: Vision-and-language transformer without convolution or region supervision,” inInternational Conference on Machine Learning. PMLR, 2021, pp. 5583–5594

  267. [280]

    Infrared target tracking using multiple instance learning with adaptive motion prediction and spatially template weighting,

    X. Shi, W. Hu, Y . Cheng, G. Chen, J. Ji, and H. Ling, “Infrared target tracking using multiple instance learning with adaptive motion prediction and spatially template weighting,” in Sensors and Systems for Space Applications VI, vol. 8739. SPIE, 2013, pp. 351–357

  268. [281]

    Infrared target tracking in multiple feature pseudo- color image with kernel density estimation,

    R. Liu and Y . Lu, “Infrared target tracking in multiple feature pseudo- color image with kernel density estimation,” Infrared Physics & Tech- nology, vol. 55, no. 6, pp. 505–512, 2012

  269. [282]

    Channel coded distribution field tracking for thermal infrared imagery,

    A. Berg, J. Ahlberg, and M. Felsberg, “Channel coded distribution field tracking for thermal infrared imagery,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops , 2016, pp. 9–17

  270. [283]

    Infrared target tracking based on robust low-rank sparse learning,

    Y . He, M. Li, J. Zhang, and J. Yao, “Infrared target tracking based on robust low-rank sparse learning,”IEEE Geoscience and Remote Sensing Letters, vol. 13, no. 2, pp. 232–236, 2015

  271. [284]

    Dense structural learning for infrared object tracking at 200+ frames per second,

    X. Yu, Q. Yu, Y . Shang, and H. Zhang, “Dense structural learning for infrared object tracking at 200+ frames per second,”Pattern Recognition Letters, vol. 100, pp. 152–159, 2017

  272. [285]

    Part-based co-difference object tracking algorithm for infrared videos,

    H. S. Demir and O. F. Adil, “Part-based co-difference object tracking algorithm for infrared videos,” in 2018 25th IEEE International Con- ference on Image Processing (ICIP). IEEE, 2018, pp. 3723–3727

  273. [286]

    Mask sparse representation based on semantic features for thermal infrared target tracking,

    M. Li, L. Peng, Y . Chen, S. Huang, F. Qin, and Z. Peng, “Mask sparse representation based on semantic features for thermal infrared target tracking,” Remote Sensing, vol. 11, no. 17, p. 1967, 2019

  274. [287]

    Deep convolutional neural networks for thermal infrared object tracking,

    Q. Liu, X. Lu, Z. He, C. Zhang, and W.-S. Chen, “Deep convolutional neural networks for thermal infrared object tracking,”Knowledge-Based Systems, vol. 134, pp. 189–198, 2017

  275. [288]

    Synthetic data generation for end-to-end thermal infrared tracking,

    L. Zhang, A. Gonzalez-Garcia, J. Van De Weijer, M. Danelljan, and F. S. Khan, “Synthetic data generation for end-to-end thermal infrared tracking,” IEEE Transactions on Image Processing , vol. 28, no. 4, pp. 1837–1850, 2018

  276. [289]

    Infrared pedestrian tracking with graph memory features,

    L. Jin, J. Cheng, and C. Zhang, “Infrared pedestrian tracking with graph memory features,” IEEE Signal Processing Letters , vol. 28, pp. 1933– 1937, 2021

  277. [290]

    Exploring reliable infrared object tracking with spatio-temporal fusion transformer,

    M. Qi, Q. Wang, S. Zhuang, K. Zhang, K. Li, Y . Liu, and Y . Yang, “Exploring reliable infrared object tracking with spatio-temporal fusion transformer,” Knowledge-Based Systems, vol. 284, p. 111234, 2024

  278. [291]

    Two-stage spatio-temporal feature correlation network for infrared ground target tracking,

    S. Li, G. Fu, X. Yang, X. Cao, S. Niu, and Z. Meng, “Two-stage spatio-temporal feature correlation network for infrared ground target tracking,” IEEE Transactions on Geoscience and Remote Sensing, 2024

  279. [292]

    Hierarchical spatial-aware siamese network for thermal infrared object tracking,

    X. Li, Q. Liu, N. Fan, Z. He, and H. Wang, “Hierarchical spatial-aware siamese network for thermal infrared object tracking,” Knowledge- Based Systems, vol. 166, pp. 71–81, 2019

  280. [295]

    Thermal infrared object tracking based on adaptive feature fusion,

    Y . Wang, J. Ma, J. Lv, and Z. Zhao, “Thermal infrared object tracking based on adaptive feature fusion,” in 2021 11th International Confer- ence on Information Technology in Medicine and Education (ITME) . IEEE, 2021, pp. 71–75

  281. [296]

    Hierarchical convolution fusion-based adaptive siamese network for infrared target tracking,

    Y . Xu, M. Wan, Q. Chen, W. Qian, K. Ren, and G. Gu, “Hierarchical convolution fusion-based adaptive siamese network for infrared target tracking,” IEEE Transactions on Instrumentation and Measurement , vol. 70, pp. 1–12, 2021

  282. [297]

    Scale and appearance variation enhanced siamese network for thermal infrared target tracking,

    T. Yao, J. Hu, B. Zhang, Y . Gao, P. Li, and Q. Hu, “Scale and appearance variation enhanced siamese network for thermal infrared target tracking,” Infrared Physics & Technology , vol. 117, p. 103825, 2021

  283. [298]

    Structural target-aware model for thermal infrared tracking,

    D. Yuan, X. Shu, Q. Liu, and Z. He, “Structural target-aware model for thermal infrared tracking,” Neurocomputing, vol. 491, pp. 44–56, 2022. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 30

  284. [299]

    Infrared target tracking via weighted correlation filter,

    Y .-J. He, M. Li, J. Zhang, and J.-P. Yao, “Infrared target tracking via weighted correlation filter,” Infrared Physics & Technology, vol. 73, pp. 103–114, 2015

  285. [300]

    Comparison of infrared and visible imagery for object tracking: Toward trackers with superior ir performance,

    E. Gundogdu, H. Ozkan, H. Seckin Demir, H. Ergezer, E. Akagunduz, and S. Kubilay Pakin, “Comparison of infrared and visible imagery for object tracking: Toward trackers with superior ir performance,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognit...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.