REVIEW 5 major objections 5 minor 82 references
MVTD: A Benchmark Dataset for Maritime Visual Object Tracking
T0 review · 5 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper introduces MVTD, a maritime visual object tracking benchmark of 182 video sequences and roughly 150,000 annotated frames, and reports that 14 state-of-the-art trackers perform substantially worse on it than on general-purpose…
desk verdict MVTD is a useful new maritime tracking dataset, but the paper's fine-tuning gains are not supported because Protocol I and II are evaluated on different splits. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the MVTD dataset itself, assembled from onshore static cameras and a camera mounted on an unmanned surface vehicle, with bounding boxes annotated in an annotation tool, then manually checked and adjusted by student annotators and averaged across corrections. The argument is carried by the contrast between two evaluation protocols: Protocol I runs pretrained trackers directly on the full test set, and Protocol II fine-tunes the five best-performing trackers on the training split and re-evaluates them on the test split. The delta between these protocols is the evidence for performance degradation and for the value of domain adaptation. Nine attribute labels, covering occlusion, illumination change, scale variation, motion blur, appearance variation, partial visibility, low resolution, background clutter, and low contrast, are the machinery for the paper's attribute-level analysis of where trackers fail.
What would settle it
Re-annotate a random sample of MVTD sequences with a second independent team and compute per-frame inter-annotator IoU; if mean IoU falls well below typical benchmark agreement (say, below 0.7) or if tracker rankings change materially when re-evaluated on the second annotation set, the reported degradation and fine-tuning gains would not be stable.
Extended reading notes
Core claim
The paper's central claim is that maritime visual object tracking is a distinct domain in which state-of-the-art trackers underperform, and that a dedicated dataset can both expose and remedy this gap. Concretely, MVTD contains 182 sequences totaling 150,058 frames with four object classes and nine attributes spanning maritime-specific and generic challenges. On Protocol I, all 14 pretrained trackers score lower on MVTD than on general-purpose benchmarks; HIPTrack leads with 68.65 percent AUC, while ETTrack falls to 50.00 percent. On Protocol II, fine-tuning the five best trackers on MVTD's training split raises AUC to 71.90 to 75.14 percent, with SimTrack improving from 67.01 to 75.14 percent. The paper takes this as evidence that transfer learning on domain-specific data is an effective route to robust maritime tracking.
Load-bearing premise
The benchmark's conclusions rest on the assumption that its manually reviewed bounding-box annotations are accurate and consistent enough to support the reported tracker scores, yet the paper reports no inter-annotator agreement or quantitative quality check.
Editorial extensions
If this is right
- MVTD provides a public benchmark so that future trackers can be compared on maritime-specific challenges rather than only on open-air datasets.
- The observed degradation across all 14 trackers indicates that strong performance on LaSOT, GOT-10K, or TrackingNet does not automatically transfer to maritime scenes.
- Fine-tuning on MVTD's training split produces substantial gains, up to about 8 AUC points, supporting domain adaptation as a practical step before maritime deployment.
- Attribute-level results point to low-contrast objects, motion blur, and scale variation as persistent failure modes, suggesting where tracker design should focus.
Reading between the lines
- A direct comparison that controls for sequence difficulty or object class, rather than comparing across different benchmarks, would tell whether the degradation is caused by maritime-specific factors or simply by harder sequences; the paper does not run this control.
- The fine-tuning gains could partly reflect additional training data or longer training rather than domain adaptation per se; fine-tuning the same trackers on a same-sized generic dataset would isolate the maritime contribution.
- Because the paper reports no inter-annotator statistics, a small re-annotation study of sampled sequences would test whether the benchmark's rankings and absolute scores are stable under annotation noise.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces MVTD, a maritime visual object tracking dataset with 182 high-resolution sequences, about 150,000 annotated frames, four object classes, and nine tracking attributes. The authors evaluate 14 recent trackers under two protocols: Protocol I uses pretrained trackers, and Protocol II fine-tunes the top five trackers on a training split and evaluates on a test split. They report performance degradation relative to general-purpose datasets and substantial fine-tuning gains, concluding that domain adaptation is important for maritime tracking.
Significance. If the central claims are supported, MVTD would fill a clear gap: existing VOT benchmarks are predominantly terrestrial, and a public, large-scale maritime benchmark with modern trackers would be useful. The paper's strengths include the size and scope of the dataset, the inclusion of maritime-specific attributes such as low contrast and water reflections, the use of recent 2024/2025 trackers, and a public release link. However, the two headline empirical claims—cross-dataset degradation and fine-tuning gains—are not currently backed by the evidence shown in the manuscript, and the annotation quality is not quantified. These issues are fixable with additional analysis and reporting, but they are load-bearing for the paper's conclusions.
major comments (5)
- [Abstract and Section VI.A] The claim of 'substantial performance degradation compared to their performance on general-purpose datasets' is not directly supported by any quantitative comparison in the manuscript. Fig. 1 is referenced as showing AUC comparisons with LaSOT, GOT-10K, and TrackingNet, but no table or text reports the per-tracker scores on those datasets. The reader cannot verify the magnitude of the degradation, nor whether the same evaluation conditions were used. The authors should add a table reporting per-tracker AUC (and precision) on MVTD and on the general-purpose datasets, or temper the claim to what the shown data actually supports.
- [Section IV.A, V.B, Tables III and IV] The fine-tuning gains are not a controlled comparison. Protocol I states that results are reported for the entire dataset and also for the testing split, yet only one Protocol I table (Table III) and Fig. 5 are given, and Section V.A analyzes 'the overall performance' without a test-split table. Table IV reports fine-tuned results on the testing split. If Table III is the full-dataset result, then the improvements in Table IV could reflect that the testing split is easier than the full dataset, not the effect of fine-tuning. The authors must report the Protocol I results on the identical testing split for the five fine-tuned trackers and, ideally, per-sequence gains, before claiming that fine-tuning on MVTD yields substantial improvements.
- [Section III.C] The bounding-box annotations are the foundation of every reported metric, but no quantitative quality assessment is provided. The text describes manual checking by students and averaging of corrected boxes, but there is no inter-annotator agreement score, no validation against a second annotation pass, and no analysis of how averaging affects box tightness. For a dataset paper, reporting e.g., IoU between annotators on a sample of frames and per-class box-size distributions would substantiate the reliability of the benchmark results.
- [Section VI.C] The discussion makes attribute-level claims, stating that 'attribute-specific representations and results reveal consistent weaknesses' for low-contrast objects, motion blur, occlusion, and scale variation. However, the manuscript contains no attribute-level performance breakdown. Since the dataset defines nine attributes, the authors should provide per-attribute AUC/precision tables for Protocol I (and ideally Protocol II) to support these claims.
- [Section IV.A and IV.C] The training/testing split is not described with enough detail to interpret Protocol II or to reproduce it. The paper says the dataset is divided into training and testing splits with all object categories present, but it does not state the number of sequences or frames in each split, nor how the split was performed. This information is essential for evaluating the fine-tuning experiments and for anyone wishing to use the benchmark.
minor comments (5)
- [Section III.A and Table I] The text states an average sequence length of 863 frames, while Table I reports 824; the total of 150,058 frames divided by 182 sequences gives approximately 824. These numbers should be reconciled.
- [Section III.B] The sentence describing the offshore camera contains an unresolved citation placeholder: 'mounted on an Unmanned Surface Vessel (USV) [?]'. This reference should be filled in or removed.
- [Section I] There is a typo in the sentence beginning 'T This highlights the pressing need...' where 'T' is a stray character. Additionally, 'Occlussion' in Section III.D should be 'Occlusion'.
- [Section V.A] The text uses 'success rate' interchangeably with AUC values reported in Table III; for clarity, the authors should distinguish between the AUC metric and the success rate at a fixed overlap threshold (e.g., OP50).
- [Section IV.A] The protocol description says pretrained trackers are 'tested on the complete test set' and then 'we report the results for the entire dataset'; this wording is contradictory. Please clarify which set was used for the numbers in Table III and Fig. 5.
Circularity Check
No circularity: the dataset, annotations, and tracker evaluations are empirical measurements, not results derived from their own inputs.
full rationale
MVTD is an empirical dataset paper. The central contributions are the dataset itself, manual annotations, and benchmark scores, with no analytical derivation whose output is equivalent to its input. Protocol I runs pretrained trackers with default parameters, and Protocol II fine-tunes on a training split and evaluates on a disjoint testing split, which is standard supervised evaluation rather than circular prediction. The reported gains are measured results, not fitted parameters renamed as predictions. The only self-citations (e.g., reference [35]) appear in related-work descriptions of maritime challenges and are not load-bearing for the benchmark's conclusions. The skeptical concern that Protocol II gains are compared against full-dataset Protocol I numbers is a methodological control issue: the paper states in Section IV.A that results are reported for the entire dataset with test-split results also intended for comparison, but no separate Protocol I test-split table appears. This affects whether the 10-15% improvements are an apples-to-apples comparison, but it is not circularity because no result is defined in terms of another result or forced by construction. Similarly, selecting the top-5 trackers from Protocol I introduces selection-on-test-set concerns, but that weakens the evidence without making the derivation circular. Therefore, the circularity score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption The 182 sequences and bounding-box annotations are accurate and consistent enough for benchmarking.
- ad hoc to paper The dataset is representative of maritime visual tracking environments.
- domain assumption Standard VOT metrics (precision, success rate, OP50, OP75) are appropriate for maritime evaluation.
- domain assumption The training and testing splits prevent leakage and contain all object categories.
Cite this review
Pith. "Pith review of MVTD: A Benchmark Dataset for Maritime Visual Object Tracking." pith.science (2026). https://pith.science/paper/EA4ES4R4
@misc{pith2026250602866,
author = {Pith},
title = {Pith review of: MVTD: A Benchmark Dataset for Maritime Visual Object Tracking},
year = {2026},
howpublished = {\url{https://pith.science/paper/EA4ES4R4}},
note = {Machine review of arXiv:2506.02866}
}
read the original abstract
Visual Object Tracking (VOT) is a fundamental task with widespread applications in autonomous navigation, surveillance, and maritime robotics. Despite significant advances in generic object tracking, maritime environments continue to present unique challenges, including specular water reflections, low-contrast targets, dynamically changing backgrounds, and frequent occlusions. These complexities significantly degrade the performance of state-of-the-art tracking algorithms, highlighting the need for domain-specific datasets. To address this gap, we introduce the Maritime Visual Tracking Dataset (MVTD), a comprehensive and publicly available benchmark specifically designed for maritime VOT. MVTD comprises 182 high-resolution video sequences, totaling approximately 150,000 frames, and includes four representative object classes: boat, ship, sailboat, and unmanned surface vehicle (USV). The dataset captures a diverse range of operational conditions and maritime scenarios, reflecting the real-world complexities of maritime environments. We evaluated 14 recent SOTA tracking algorithms on the MVTD benchmark and observed substantial performance degradation compared to their performance on general-purpose datasets. However, when fine-tuned on MVTD, these models demonstrate significant performance gains, underscoring the effectiveness of domain adaptation and the importance of transfer learning in specialized tracking contexts. The MVTD dataset fills a critical gap in the visual tracking community by providing a realistic and challenging benchmark for maritime scenarios. Dataset and Source Code can be accessed here "https://github.com/AhsanBaidar/MVTD".
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[35]
Benchmarking vision-based object tracking for usvs in complex maritime environments,
M. U. Din, A. B. Bakht, W. Akram, Y . Dong, L. Seneviratne, and I. Hussain, “Benchmarking vision-based object tracking for usvs in complex maritime environments,” IEEE Access, 2025
2025
-
[1]
Rgbt tracking: A comprehensive review,
M. Feng and J. Su, “Rgbt tracking: A comprehensive review,” Inf. Fusion, p. 102492, 2024
2024
-
[2]
Visual object tracking with discriminative filters and siamese networks: a survey and outlook,
S. Javed, M. Danelljan, F. S. Khan, M. H. Khan, M. Felsberg, and J. Matas, “Visual object tracking with discriminative filters and siamese networks: a survey and outlook,” IEEE TPAMI, vol. 45, no. 5, pp. 6552– 6574, 2022
2022
-
[3]
A robust method for multi object tracking in autonomous ship navigation systems,
Z. Shao, Y . Yin, H. Lyu, and C. Guedes Soares, “A robust method for multi object tracking in autonomous ship navigation systems,” Ocean Eng., vol. 311, p. 118560, 2024
work page 2024
-
[4]
M. Ud Din, X. He, H. Kuang, S. Yang, W. Akram, D. Lin, L. Senevi- ratne, D. Yihao, S. He, and I. Hussain, “A multi-stage heading control approach for autonomous usv navigation in gnss-denied environments with uav cooperation,” Xiaoyu and Kuang, Hailiang and Yang, Siyuan and Akram, Waseem and Lin, Defu and Seneviratne, Lakmal and Yihao, Dong and He, Shaomi...
work page 2024
-
[5]
J. Ding, W. Li, L. Pei, M. Yang, A. Tian, and B. Yuan, “Novel pipeline integrating cross-modality and motion model for nearshore multi-object tracking in optical video surveillance,” IEEE T-ITS, 2024
work page 2024
-
[6]
Foundationpose: Unified 6d pose estimation and tracking of novel objects,
B. Wen, W. Yang, J. Kautz, and S. Birchfield, “Foundationpose: Unified 6d pose estimation and tracking of novel objects,” in CVPR, 2024, pp. 17 868–17 879
work page 2024
-
[7]
Q. Nguyen, R. Jiang, M. Ellingwood, and R. Yurko, “Fractional tackles: leveraging player tracking data for within-play tackling evaluation in american football,” Sci. Rep., vol. 15, no. 1, p. 2148, 2025
work page 2025
Show all 82 references
-
[8]
Tracking and navigation of a microswarm under laser speckle contrast imaging for targeted delivery,
Q. Wang, Q. Wang, Z. Ning, K. F. Chan, J. Jiang, Y . Wang, L. Su, S. Jiang, B. Wang, B. Y . M. Ip, H. Ko, T. W. H. Leung, P. W. Y . Chiu, S. C. H. Yu, and L. Zhang, “Tracking and navigation of a microswarm under laser speckle contrast imaging for targeted delivery,” Sci. Robot...
2024
-
[9]
Aquaculture defects recognition via multi-scale semantic segmentation,
W. Akram, T. Hassan, H. Toubar, M. Ahmed, N. Mi ˇskovic, L. Senevi- ratne, and I. Hussain, “Aquaculture defects recognition via multi-scale semantic segmentation,” Expert systems with applications , vol. 237, p. 121197, 2024
2024
-
[10]
Enhancing aquaculture net pen inspection: A benchmark study on detection and semantic segmentation,
W. Akram, A. Bakht, M. U. Din, L. Seneviratne, and I. Hussain, “Enhancing aquaculture net pen inspection: A benchmark study on detection and semantic segmentation,” IEEE Access, vol. 13, pp. 3453– 3474, 2025
2025
-
[11]
Mula-gan: Multi-level attention gan for enhanced underwater visibility,
A. B. Bakht, Z. Jia, M. U. Din, W. Akram, L. S. Saoud, L. Seneviratne, D. Lin, S. He, and I. Hussain, “Mula-gan: Multi-level attention gan for enhanced underwater visibility,” Ecological Informatics, vol. 81, p. 102631, 2024
2024
-
[12]
Fishtrack: Multi- object tracking method for fish using spatiotemporal information fusion,
Y . Liu, B. Li, X. Zhou, D. Li, and Q. Duan, “Fishtrack: Multi- object tracking method for fish using spatiotemporal information fusion,” Expert Syst Appl. , vol. 238, p. 122194, 2024
2024
-
[13]
A guide to image- and video-based small object detection using deep learning: Case study of maritime surveillance,
A. Miri Rekavandi, L. Xu, F. Boussaid, A.-K. Seghouane, S. Hoefs, and M. Bennamoun, “A guide to image- and video-based small object detection using deep learning: Case study of maritime surveillance,” EEE T-ITS, vol. 26, no. 3, pp. 2851–2879, 2025
2025
-
[14]
Fully-convolutional siamese networks for object tracking,
L. Bertinetto, J. Valmadre, J. F. Henriques, A. Vedaldi, and P. H. Torr, “Fully-convolutional siamese networks for object tracking,” in ECCV Workshops. Springer, 2016, pp. 850–865
2016
-
[15]
Transformer tracking,
X. Chen, B. Yan, J. Zhu, D. Wang, X. Yang, and H. Lu, “Transformer tracking,” in CVPR, 2021
2021
-
[16]
Adaptive text feature updating for visual- language tracking,
X. Liu, Z. Zou, and J. Hao, “Adaptive text feature updating for visual- language tracking,” in ICPR. Springer, 2024, pp. 366–381
2024
-
[17]
Fast online object tracking and segmentation: A unifying approach,
Q. Wang, L. Zhang, L. Bertinetto, W. Hu, and P. H. Torr, “Fast online object tracking and segmentation: A unifying approach,” in CVPR, 2019, pp. 1328–1338
2019
-
[18]
Siamese box adaptive network for visual tracking,
Z. Chen, B. Zhong, G. Li, S. Zhang, and R. Ji, “Siamese box adaptive network for visual tracking,” in CVPR, 2020, pp. 6668–6677
2020
-
[19]
High performance visual tracking with siamese region proposal network,
B. Li, J. Yan, W. Wu, Z. Zhu, and X. Hu, “High performance visual tracking with siamese region proposal network,” in CVPR, 2018, pp. 8971–8980
2018
-
[20]
Seqtrack: Sequence to sequence learning for visual object tracking,
X. Chen, H. Peng, D. Wang, H. Lu, and H. Hu, “Seqtrack: Sequence to sequence learning for visual object tracking,” in CVPR, June 2023, pp. 14 572–14 581
2023
-
[21]
Exploring enhanced contextual information for video-level object tracking,
B. Kang, X. Chen, S. Lai, Y . Liu, Y . Liu, and D. Wang, “Exploring enhanced contextual information for video-level object tracking,” in AAAI, 2025
2025
-
[22]
Aittrack: Attention-based image-text align- ment for visual tracking,
B. Alawode and S. Javed, “Aittrack: Attention-based image-text align- ment for visual tracking,” IEEE Access, 2025
2025
-
[23]
Lasot: A high-quality benchmark for large-scale single object tracking,
H. Fan, L. Lin, F. Yang, P. Chu, G. Deng, S. Yu, H. Bai, Y . Xu, C. Liao, and H. Ling, “Lasot: A high-quality benchmark for large-scale single object tracking,” in CVPR, 2019, pp. 5374–5383
2019
-
[24]
The seventh visual object tracking vot2019 challenge results,
M. Kristan, J. Matas, A. Leonardis, M. Felsberg, R. Pflugfelder, J.-K. Kamarainen, L. Cehovin Zajc, O. Drbohlav, A. Lukezic, A. Berg et al., “The seventh visual object tracking vot2019 challenge results,” in CVPR workshops, 2019, pp. 0–0
2019
-
[25]
The sixth visual object tracking vot2018 challenge results,
M. Kristan, A. Leonardis, J. Matas, M. Felsberg, R. Pflugfelder, L. Ce- hovin Zajc, T. V ojir, G. Bhat, A. Lukezic, A. Eldesokeyet al., “The sixth visual object tracking vot2018 challenge results,” in ECCV Workshops, 2018, pp. 0–0
2018
-
[26]
The eighth visual object tracking vot2020 challenge results,
M. Kristan, A. Leonardis, J. Matas, M. Felsberg, R. Pflugfelder, J.-K. K¨am¨ar¨ainen, M. Danelljan, L. ˇC. Zajc, A. Luke ˇziˇc, O. Drbohlav et al., “The eighth visual object tracking vot2020 challenge results,” in ECCV Workshops. Springer, 2020, pp. 547–601
2020
-
[27]
The ninth visual object tracking vot2021 challenge results,
M. Kristan, J. Matas, A. Leonardis, M. Felsberg, R. Pflugfelder, J.-K. K¨am¨ar¨ainen, H. J. Chang, M. Danelljan, L. Cehovin, A. Luke ˇziˇc et al., “The ninth visual object tracking vot2021 challenge results,” in CVPR, 2021, pp. 2711–2738
2021
-
[28]
The tenth visual object tracking vot2022 challenge results,
M. Kristan, A. Leonardis, J. Matas, M. Felsberg, R. Pflugfelder, J.-K. K¨am¨ar¨ainen, H. J. Chang, M. Danelljan, L. ˇC. Zajc, A. Luke ˇziˇc et al., “The tenth visual object tracking vot2022 challenge results,” in ECCV. Springer, 2022, pp. 431–460
2022
-
[29]
The first visual object tracking segmentation vots2023 challenge results,
M. Kristan, J. Matas, M. Danelljan, M. Felsberg, H. J. Chang, L. ˇC. Zajc, A. Luke ˇziˇc, O. Drbohlav, Z. Zhang, K.-T. Tran et al. , “The first visual object tracking segmentation vots2023 challenge results,” in CVPR, 2023, pp. 1796–1818
2023
-
[30]
Got-10k: A large high-diversity benchmark for generic object tracking in the wild,
L. Huang, X. Zhao, and K. Huang, “Got-10k: A large high-diversity benchmark for generic object tracking in the wild,” IEEE TPAMI , vol. 43, no. 5, pp. 1562–1577, 2019
2019
-
[31]
Track- ingnet: A large-scale dataset and benchmark for object tracking in the wild,
M. Muller, A. Bibi, S. Giancola, S. Alsubaihi, and B. Ghanem, “Track- ingnet: A large-scale dataset and benchmark for object tracking in the wild,” in ECCV, 2018, pp. 300–317
2018
-
[32]
Lasot: A high-quality large-scale single object tracking benchmark,
H. Fan, H. Bai, L. Lin, F. Yang, P. Chu, G. Deng, S. Yu, Harshit, M. Huang, J. Liu et al., “Lasot: A high-quality large-scale single object tracking benchmark,” IJCV, vol. 129, pp. 439–461, 2021
2021
-
[33]
To- wards more flexible and accurate object tracking with natural language: Algorithms and benchmark,
X. Wang, X. Shu, Z. Zhang, B. Jiang, Y . Wang, Y . Tian, and F. Wu, “To- wards more flexible and accurate object tracking with natural language: Algorithms and benchmark,” in CVPR, 2021, pp. 13 763–13 773
2021
-
[34]
Maritime mission planning for unmanned surface vessel using large language model,
M. U. Din, W. Akram, A. B. Bakht, Y . Dong, and I. Hussain, “Maritime mission planning for unmanned surface vessel using large language model,” in 2025 IEEE International Conference on Simulation, Mod- eling, and Programming for Autonomous Robots (SIMPAR) . IEEE, 2025, pp. 1–6
2025
-
[36]
Video processing from electro-optical sensors for object detection and tracking in a maritime environment: A survey,
D. K. Prasad, D. Rajan, L. Rachmawati, E. Rajabally, and C. Quek, “Video processing from electro-optical sensors for object detection and tracking in a maritime environment: A survey,” IEEE T-ITS , vol. 18, no. 8, pp. 1993–2016, 2017
1993
-
[37]
Vision-based autonomous navi- gation for unmanned surface vessel in extreme marine conditions,
M. Ahmed, A. B. Bakht, T. Hassan, W. Akram, A. Humais, L. Senevi- ratne, S. He, D. Lin, and I. Hussain, “Vision-based autonomous navi- gation for unmanned surface vessel in extreme marine conditions,” in IROS. IEEE, 2023, pp. 7097–7103
2023
-
[38]
A review of intelligent ship marine object detection based on rgb camera,
D. Yang, M. I. Solihin, Y . Zhao, B. Yao, C. Chen, B. Cai, and A. Machmudah, “A review of intelligent ship marine object detection based on rgb camera,” IET IP, vol. 18, no. 2, pp. 281–297, 2024
2024
-
[39]
Exploring the behavior feature of complex trajectories of ships with fourier transform processing: a case from fishing vessels,
Q. Zhu, Y . Xi, S. Hu, and Y . Chen, “Exploring the behavior feature of complex trajectories of ships with fourier transform processing: a case from fishing vessels,” Front. Mar. Sci., vol. 10, p. 1271930, 2023
2023
-
[40]
Visual object tracking: A survey,
F. Chen, X. Wang, Y . Zhao, S. Lv, and X. Niu, “Visual object tracking: A survey,” CVIU, vol. 222, p. 103508, 2022
2022
-
[41]
Online object tracking: A benchmark,
Y . Wu, J. Lim, and M.-H. Yang, “Online object tracking: A benchmark,” in CVPR, 2013, pp. 2411–2418
2013
-
[42]
Object tracking benchmark,
——, “Object tracking benchmark,” IEEE TPAMI, vol. 37, no. 9, pp. 1834–1848, 2015
2015
-
[43]
Encoding color information for visual tracking: Algorithms and benchmark,
P. Liang, E. Blasch, and H. Ling, “Encoding color information for visual tracking: Algorithms and benchmark,” IEEE TIP , vol. 24, no. 12, pp. 5630–5644, 2015
2015
-
[44]
Nus-pro: A new visual tracking challenge,
A. Li, M. Lin, Y . Wu, M.-H. Yang, and S. Yan, “Nus-pro: A new visual tracking challenge,” IEEE TPAMI, vol. 38, no. 2, pp. 335–349, 2015
2015
-
[45]
A benchmark and simulator for uav tracking,
U. Benchmark, “A benchmark and simulator for uav tracking,” in ECCV, 2016
2016
-
[46]
Need for speed: A benchmark for higher frame rate object tracking,
H. Kiani Galoogahi, A. Fagg, C. Huang, D. Ramanan, and S. Lucey, “Need for speed: A benchmark for higher frame rate object tracking,” in CVPR, 2017, pp. 1125–1134
2017
-
[47]
Cdtb: A color and depth visual object tracking dataset and benchmark,
A. Lukezic, U. Kart, J. Kapyla, A. Durmush, J.-K. Kamarainen, J. Matas, and M. Kristan, “Cdtb: A color and depth visual object tracking dataset and benchmark,” in CVPR, 2019, pp. 10 013–10 022
2019
-
[48]
Seadronessee: A maritime benchmark for detecting humans in open water,
L. A. Varga, B. Kiefer, M. Messmer, and A. Zell, “Seadronessee: A maritime benchmark for detecting humans in open water,” in WACV, 2022, pp. 2260–2270
2022
-
[49]
Polaris dataset: A maritime object detection and tracking dataset in pohang canal,
J. Choi, D. Cho, G. Lee, H. Kim, G. Yang, J. Kim, and Y . Cho, “Polaris dataset: A maritime object detection and tracking dataset in pohang canal,” arXiv preprint arXiv:2412.06192 , 2024
2024 arXiv
-
[50]
Learning multi-domain convolutional neural networks for visual tracking,
H. Nam and B. Han, “Learning multi-domain convolutional neural networks for visual tracking,” in CVPR, 2016, pp. 4293–4302
2016
-
[51]
Learning spatio-temporal discrimina- tive model for affine subspace based visual object tracking,
T. Xu, X.-F. Zhu, and X.-J. Wu, “Learning spatio-temporal discrimina- tive model for affine subspace based visual object tracking,” Vis. Intell., vol. 1, no. 1, p. 4, 2023
2023
-
[52]
Siamese instance search for tracking,
R. Tao, E. Gavves, and A. W. Smeulders, “Siamese instance search for tracking,” in CVPR, 2016, pp. 1420–1429
2016
-
[53]
Topology-aware universal adversarial attack on 3d object tracking,
R. Cheng, X. Wang, F. Sohel, and H. Lei, “Topology-aware universal adversarial attack on 3d object tracking,” Vis. Intell., vol. 1, no. 1, p. 31, 2023
2023
-
[54]
Ocean: Object-aware anchor-free tracking,
Z. Zhang, H. Peng, J. Fu, B. Li, and W. Hu, “Ocean: Object-aware anchor-free tracking,” in ECCV. Springer, 2020, pp. 771–787
2020
-
[55]
Learn to match: Automatic matching network design for visual tracking,
Z. Zhang, Y . Liu, X. Wang, B. Li, and W. Hu, “Learn to match: Automatic matching network design for visual tracking,” inCVPR, 2021, pp. 13 339–13 348
2021
-
[56]
Atom: Accurate tracking by overlap maximization,
M. Danelljan, G. Bhat, F. S. Khan, and M. Felsberg, “Atom: Accurate tracking by overlap maximization,” in CVPR, 2019, pp. 4660–4669
2019
-
[57]
Learning discrim- inative model prediction for tracking,
G. Bhat, M. Danelljan, L. V . Gool, and R. Timofte, “Learning discrim- inative model prediction for tracking,” in CVPR, 2019, pp. 6182–6191
2019
-
[58]
Probabilistic regression for visual tracking,
M. Danelljan, L. V . Gool, and R. Timofte, “Probabilistic regression for visual tracking,” in CVPR, 2020, pp. 7183–7192
2020
-
[59]
Transformer meets tracker: Exploiting temporal context for robust visual tracking,
N. Wang, W. Zhou, J. Wang, and H. Li, “Transformer meets tracker: Exploiting temporal context for robust visual tracking,” in CVPR, 2021, pp. 1571–1580
2021
-
[60]
Transformer tracking,
X. Chen, B. Yan, J. Zhu, D. Wang, X. Yang, and H. Lu, “Transformer tracking,” in CVPR, 2021, pp. 8126–8135. 13
2021
-
[61]
Transforming model prediction for tracking,
C. Mayer, M. Danelljan, G. Bhat, M. Paul, D. P. Paudel, F. Yu, and L. Van Gool, “Transforming model prediction for tracking,” in CVPR, 2022, pp. 8731–8740
2022
-
[62]
Aiatrack: Attention in attention for transformer visual tracking,
S. Gao, C. Zhou, C. Ma, X. Wang, and J. Yuan, “Aiatrack: Attention in attention for transformer visual tracking,” in ECCV. Springer, 2022, pp. 146–164
2022
-
[63]
Backbone is all your need: A simplified architecture for visual object tracking,
B. Chen, P. Li, L. Bai, L. Qiao, Q. Shen, B. Li, W. Gan, W. Wu, and W. Ouyang, “Backbone is all your need: A simplified architecture for visual object tracking,” in ECCV. Springer, 2022, pp. 375–392
2022
-
[64]
Joint feature learning and relation modeling for tracking: A one-stream framework,
B. Ye, H. Chang, B. Ma, S. Shan, and X. Chen, “Joint feature learning and relation modeling for tracking: A one-stream framework,” in ECCV. Springer, 2022, pp. 341–357
2022
-
[65]
Learning from images: A distillation learning framework for event cameras,
Y . Deng, H. Chen, H. Chen, and Y . Li, “Learning from images: A distillation learning framework for event cameras,” IEEE TIP , vol. 30, pp. 4919–4931, 2021
2021
-
[66]
Distilled siamese networks for visual tracking,
J. Shen, Y . Liu, X. Dong, X. Lu, F. S. Khan, and S. Hoi, “Distilled siamese networks for visual tracking,” IEEE TPAMI, vol. 44, no. 12, pp. 8896–8909, 2021
2021
-
[67]
Teacher-student knowledge distillation for real-time correlation tracking,
Q. Chen, B. Zhong, Q. Liang, Q. Deng, and X. Li, “Teacher-student knowledge distillation for real-time correlation tracking,” Neurocomput- ing, vol. 500, pp. 537–546, 2022
2022
-
[68]
Ensemble learning with siamese networks for visual tracking,
J. Zhuang, Y . Dong, and H. Bai, “Ensemble learning with siamese networks for visual tracking,” Neurocomputing, vol. 464, pp. 497–506, 2021
2021
-
[69]
Unsupervised cross-modal distillation for thermal infrared tracking,
J. Sun, L. Zhang, Y . Zha, A. Gonzalez-Garcia, P. Zhang, W. Huang, and Y . Zhang, “Unsupervised cross-modal distillation for thermal infrared tracking,” in ACM Multimedia, 2021, pp. 2262–2270
2021
-
[70]
Real-time correlation tracking via joint model compression and transfer,
N. Wang, W. Zhou, Y . Song, C. Ma, and H. Li, “Real-time correlation tracking via joint model compression and transfer,” IEEE TIP, vol. 29, pp. 6123–6135, 2020
2020
-
[71]
Distillation, ensemble and selection for building a better and faster siamese based tracker,
S. Zhao, T. Xu, X.-J. Wu, and J. Kittler, “Distillation, ensemble and selection for building a better and faster siamese based tracker,” IEEE TCSVT, vol. 34, no. 1, pp. 182–194, 2022
2022
-
[72]
Distilling channels for efficient deep tracking,
S. Ge, Z. Luo, C. Zhang, Y . Hua, and D. Tao, “Distilling channels for efficient deep tracking,” IEEE TIP, vol. 29, pp. 2610–2621, 2019
2019
-
[73]
Hiptrack: Visual tracking with historical prompts,
W. Cai, Q. Liu, and Y . Wang, “Hiptrack: Visual tracking with historical prompts,” in CVPR, 2024, pp. 19 258–19 267
2024
-
[74]
Autoregressive queries for adaptive tracking with spatio-temporal trans- formers,
J. Xie, B. Zhong, Z. Mo, S. Zhang, L. Shi, S. Song, and R. Ji, “Autoregressive queries for adaptive tracking with spatio-temporal trans- formers,” in CVPR, 2024, pp. 19 300–19 309
2024
-
[75]
Correlation-embedded transformer tracking: A single-branch frame- work,
F. Xie, W. Yang, C. Wang, L. Chu, Y . Cao, C. Ma, and W. Zeng, “Correlation-embedded transformer tracking: A single-branch frame- work,” arXiv preprint arXiv:2401.12743 , 2024
2024 arXiv
-
[76]
Odtrack: Online dense temporal token learning for visual tracking,
Y . Zheng, B. Zhong, Q. Liang, Z. Mo, S. Zhang, and X. Li, “Odtrack: Online dense temporal token learning for visual tracking,” in AAAI, 2024
2024
-
[77]
Artrackv2: Prompting autore- gressive tracker where to look and how to describe,
Y . Bai, Z. Zhao, Y . Gong, and X. Wei, “Artrackv2: Prompting autore- gressive tracker where to look and how to describe,” in CVPR, 2024, pp. 19 048–19 057
2024
-
[78]
Autoregressive visual tracking,
X. Wei, Y . Bai, Y . Zheng, D. Shi, and Y . Gong, “Autoregressive visual tracking,” in CVPR, 2023, pp. 9697–9706
2023
-
[79]
Efficient visual tracking with exemplar transformers,
P. Blatter, M. Kanakis, M. Danelljan, and L. Van Gool, “Efficient visual tracking with exemplar transformers,” in WACV, 2023, pp. 1571–1581
2023
-
[80]
Generalized relation modeling for transformer tracking,
S. Gao, C. Zhou, and J. Zhang, “Generalized relation modeling for transformer tracking,” in CVPR, 2023, pp. 18 686–18 695
2023
-
[81]
Towards sequence-level training for visual tracking,
M. Kim, S. Lee, J. Ok, B. Han, and M. Cho, “Towards sequence-level training for visual tracking,” in ECCV, 2022
2022
-
[82]
Learning spatio-temporal transformer for visual tracking,
B. Yan, H. Peng, J. Fu, D. Wang, and H. Lu, “Learning spatio-temporal transformer for visual tracking,” in CVPR, 2021, pp. 10 448–10 457
2021
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.