REVIEW 4 major objections 6 minor 74 references
High Performance Visual Object Tracking with Unified Convolutional Networks
T0 review · 4 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Training features and tracking together as one convolution produces top-ranked accuracy at 58 frames per second.
desk verdict Competent real-time tracker paper with strong ablations, but the headline VOT ranking rests on an unverified comparability assumption. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is the unified convolutional tracker (UCT): a fully convolutional network in which the feature extractor and the tracking filter are both convolution operations trained jointly. The filter is learned by minimizing the $L_2$ distance between its response map $R(x_k)$ and a Gaussian label $y_k$, with a weight-decay term, using gradient descent; this keeps the filtering step differentiable, so gradients flow back into the feature extractor. During online tracking the whole patch is scored in one forward pass, model updates are gated by a peak-versus-noise ratio $\mathrm{PNR} = (R_{\max}-R_{\min})/\mathrm{mean}(R\setminus R_{\max})$ against historical thresholds, and a separate one-dimensional convolutional filter branch estimates scale changes.
What would settle it
Re-run UCT with the official VOT2016 evaluation toolkit and the same initializations and compare the resulting EAO to the reported 0.342; if the official number is materially lower, the ranking claim fails. A second, more targeted test: keep everything identical but freeze the pretrained backbone during offline training, so only the two tracking convolution layers learn; if this frozen version matches UCT's accuracy, the paper's central claim that joint feature learning delivers the performance is not supported.
Extended reading notes
Core claim
The paper's central claim is that the usual separation in deep trackers—a feature extractor pretrained for another task and a separate tracking module—is a fixable cause of suboptimal performance. UCT replaces that separation with a single fully convolutional network in which the tracking filter is a convolution layer and the feature extractor is a convolutional network, both optimized together by gradient descent on an $L_2$ loss between the response map and a Gaussian label. Online, one forward pass produces the whole foreground response map, an update is triggered only when a peak-versus-noise ratio and the peak value stay above historical thresholds, and a one-dimensional scale filter handles size changes. The paper reports that this system scores an AUC of 0.693 on OTB2013 and 0.670 on OTB2015, and an EAO of 0.3576 and 0.342 on VOT2015 and VOT2016, ranking first on both VOT challenges' EAO lists at 58 FPS, while a lighter version runs at 154 FPS with a smaller accuracy margin.
Load-bearing premise
The whole comparison rests on the assumption that benchmark scores quoted from other trackers were produced under the same evaluation conditions as UCT's own runs, so the rankings the paper draws from them are fair.
Editorial extensions
If this is right
- Accuracy and speed need not be opposed: a single forward pass through one network can produce both a full response map and a scale estimate, so a tracker can hold its own against much slower champions.
- The feature extractor's learned features become tailored to tracking in general rather than inherited from a classification task, which is why the paper attributes strong results on rotation and deformation attributes to end-to-end training.
- The peak-versus-noise-ratio update rule means the model tends to skip learning from occluded or low-confidence frames, which should reduce drift and also saves computation by updating infrequently.
- Because UCT-lite keeps most of the accuracy at 154 FPS, the same architecture can be traded off along a speed-accuracy curve to fit different hardware constraints.
- A tracker that runs beyond real time while keeping top accuracy makes deep tracking practical for robots, automated driving, and other latency-sensitive applications.
Reading between the lines
- My inference: beyond the paper, the idea of treating the tracking filter as a trainable convolution suggests a unifying view in which Siamese trackers and correlation-filter trackers differ mainly in when and how the filter layer is updated.
- My inference: the peak-versus-noise-ratio criterion is a generic confidence measure that could be lifted into other online-learning systems, such as detectors or re-identifiers, whenever model updates risk being poisoned by bad frames.
- My inference: the paper reports one-step SGD updates when confident frames occur; varying the number of update steps per confident frame is a natural speed-accuracy trade-off the paper leaves untested.
- My inference: the evaluations use benchmarks from 2015-2016, so whether the advantage persists on later and harder tracking benchmarks is an open question the paper does not address.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a unified convolutional tracker (UCT) that represents both feature extraction and the correlation-filter tracking operation as convolutional layers in a single fully convolutional network. The filter is learned by minimizing an L2 ridge-regression loss (Eq. 2) via gradient descent, which permits end-to-end offline training on ImageNet VID and first-frame adaptation. Online, a peak-versus-noise ratio (PNR, Eq. 4) gates model updates, and a 1-D scale-filter branch handles scale changes. The authors evaluate UCT and a lighter variant, UCT-lite, on OTB2013, OTB2015, VOT2015, and VOT2016, and report leading performance among the compared trackers at 58 FPS and 154 FPS respectively. The central claim is that UCT achieves leading accuracy while running beyond real-time speed.
Significance. The paper offers a clean, unified formulation: converting the correlation-filter ridge regression of Eq. (2) into a differentiable convolution layer and training it jointly with a feature extractor is a sensible and timely idea, and the ablations in Table II are internally consistent (each component improves AUC). The PNR-based adaptive update is a simple and practical mechanism, and the reported speeds (58 FPS for UCT, 154 FPS for UCT-lite) would be valuable for real-time applications. If the VOT protocol issues are resolved, the work could serve as a strong real-time baseline. However, the paper currently does not provide code, per-sequence results, or uncertainty estimates, so the empirical contribution is not yet fully verifiable.
major comments (4)
- [Section V-D, Section V-E, Figs. 8 and 10, Tables IV and V] The claim that 'UCT ranks 1st in 70 trackers according to EAO criterion' on VOT2016 (Section V-E) and 'ranks 1st in 61 trackers' on VOT2015 (Section V-D) is constructed by inserting the authors' own EAO numbers into the official challenge leaderboards. The manuscript does not state that UCT was run through the official VOT evaluation toolkit; that is, the reset-after-failure protocol, the exact overlap computation, the sequence-length weighting, and the VOT release used are not specified. The single sentence 'All the tracking results use the reported results to ensure a fair comparison' (Section V) addresses the provenance of the competing trackers' numbers, not the protocol used for UCT's own numbers. Since the VOT2016 margin over CCOT is 0.342 vs 0.331 (about 3.2% relative), a modest difference in evaluation protocol could alter the ranking. Please provide the exact VOT version, confirm use of the official evaluation scripts, and release per-sequence overlap curves and raw results.
- [Section V-D] The exclusion of MDNet from the VOT2015 comparison relies on the statement that 'MDNet [50] is not compatible with the latest VOT rules because of OTB training data.' However, the VOT2016 results in Figure 10 and Table V include the entry 'MDNet_N,' whose relationship to reference [50] is not clarified in the paper. The paper does not explain why OTB-trained MDNet is ineligible for VOT2015 but appears in VOT2016. If the eligibility rules differ between the two challenges, this should be stated explicitly; as written, the VOT2015 'ranks 1st' claim is not established against the full participant field that would be expected under a consistent set of rules.
- [Sections V-B through V-E] The performance differences that support the 'leading performance' claims are reported without any measure of uncertainty. For instance, on OTB2013 the success AUC advantage over BACF is 0.693 vs 0.657 (a 5.5% relative difference), on OTB2015 the precision advantage over HDT is 0.899 vs 0.848, and on VOT2016 the EAO advantage over CCOT is 0.342 vs 0.331. These margins are small relative to the known variance of tracking benchmarks, yet no per-sequence results, confidence intervals, or paired statistical tests are given. The paper should report per-sequence overlap values and perform appropriate significance testing before concluding that UCT 'outperforms all the other trackers' (Section V-C).
- [Abstract, Section V-B, Section V-C] The abstract claims 'leading performance on these benchmarks' without qualification, but on OTB2013 (Section V-B) and OTB2015 (Section V-C) the comparison set is restricted to a selected list of real-time and near-real-time trackers; offline top performers such as MDNet, CCOT, and DeepSRDCF are not included in the OTB plots. Thus the OTB results demonstrate leadership within a subset, not across the full state of the art. The claim should be narrowed to 'leading performance among real-time trackers on OTB' or the OTB experiments should be extended to include the broader set of top-performing trackers.
minor comments (6)
- [Section III-B, Eq. (4)] Please define the set R\Rmax explicitly; in particular, state whether all entries equal to Rmax are removed and whether the mean is over the remaining spatial responses.
- [Section IV-A, Eq. (6) and Algorithm 1] Clarify whether the thresholds in Eq. (6) are accumulated over all frames up to the current frame or only over frames where the model was updated, and whether the current frame's PNR and Rmax are included in the threshold before the comparison. The current pseudocode order makes this ambiguous.
- [Section V-E, Figure 10] The text says UCT 'ranks 1st in 70 trackers,' but the horizontal axis in Figure 10 extends to 71; please reconcile the reported number of trackers.
- [Section V] The sentence 'All the tracking results use the reported results to ensure a fair comparison' should be rewritten to specify which numbers are taken from prior publications and which are measured by the authors, and to describe the evaluation protocol used for the authors' own numbers.
- [Section V-A] The paper states that code and results 'will be made publicly available,' but no link is provided; for a system paper, including code and per-sequence raw results at submission is important for reproducibility.
- [Section IV-B, Eq. (7)] The scale branch is described as 'inspired by [29]' but the details of how the 1-D scale filter is trained and how its output is combined with the translation response are not fully specified; please add the training objective and the rule for updating the scale model.
Circularity Check
No circularity: UCT's central claims are empirical benchmark results, not derivations from their own inputs; the only self-citation is a normal extension note and is not load-bearing.
full rationale
The paper's central claims are empirical: UCT achieves leading performance on OTB2013, OTB2015, VOT2015, and VOT2016 while running in real time. There is no derivation chain in which a predicted quantity is equivalent by construction to a fitted input. The PNR-based model update uses an online adaptive threshold computed from the historical average of PNR and Rmax (Eq. 4-6); it is a decision rule, not a parameter fitted to benchmark outcomes, so it does not constitute a fitted input called a prediction. The scale branch is a separate 1D convolutional filter trained on first-frame samples, following standard scale-estimation practice. The paper's self-citation to its ICCVW 2017 work [42] is used to state that the present paper extends that work, but the current paper changes the backbone, training data, and training strategy and reports new experiments on four benchmarks; the performance claims rest on those new measurements, not on the cited paper. The VOT rankings are constructed by comparing self-reported EAO numbers with reported results from challenge participants, and the paper states it follows the latest VOT rules; any concern about protocol comparability is a correctness or verification issue, not a circular reduction. No step in the paper reduces to its own inputs, and no load-bearing claim is justified solely by a self-citation. Therefore the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (10)
- Offline learning rate =
1e-5
- First-frame learning rate =
5e-7
- Online update learning rate =
1e-7
- Weight decay lambda offline =
0.005
- Weight decay lambda first frame =
0.01
- Scale filter size S =
33
- Scale factor a =
1.02
- Jittering translation and scale =
0.05 and 0.02
- Hann window padding =
1.8
- Fine-tuned layer counts =
last 6 of VGG, last 3 of ZF
assumptions (4)
- domain assumption ImageNet VID training data is representative of generic visual tracking targets
- standard math Convexity of the L2 loss in equation (2) allows gradient descent to converge to a good optimum
- ad hoc to paper The PNR statistic is a reliable indicator of tracking confidence
- domain assumption The scale branch from fDSST transfers to the convolutional framework
invented entities (1)
-
Peak-versus-noise ratio (PNR)
Cite this review
Pith. "Pith review of High Performance Visual Object Tracking with Unified Convolutional Networks." pith.science (2026). https://pith.science/paper/3MHCVLMC
@misc{pith2026190809445,
author = {Pith},
title = {Pith review of: High Performance Visual Object Tracking with Unified Convolutional Networks},
year = {2026},
howpublished = {\url{https://pith.science/paper/3MHCVLMC}},
note = {Machine review of arXiv:1908.09445}
}
read the original abstract
Convolutional neural networks (CNN) based tracking approaches have shown favorable performance in recent benchmarks. Nonetheless, the chosen CNN features are always pre-trained in different tasks and individual components in tracking systems are learned separately, thus the achieved tracking performance may be suboptimal. Besides, most of these trackers are not designed towards real-time applications because of their time-consuming feature extraction and complex optimization details. In this paper, we propose an end-to-end framework to learn the convolutional features and perform the tracking process simultaneously, namely, a unified convolutional tracker (UCT). Specifically, the UCT treats feature extractor and tracking process both as convolution operation and trains them jointly, which enables learned CNN features are tightly coupled with tracking process. During online tracking, an efficient model updating method is proposed by introducing peak-versus-noise ratio (PNR) criterion, and scale changes are handled efficiently by incorporating a scale branch into network. Experiments are performed on four challenging tracking datasets: OTB2013, OTB2015, VOT2015 and VOT2016. Our method achieves leading performance on these benchmarks while maintaining beyond real-time speed.
Figures
Figures from the paper (9 more)
Reference graph
Works this paper leans on
-
[50]
Learning multi-domain convolu- tional neural networks for visual tracking,
H. Nam and B. Han, “Learning multi-domain convolu- tional neural networks for visual tracking,” in Proceed- ings of the IEEE Conference on Computer Vision and Pattern Recognition, June 2016
work page 2016
-
[1]
Sparse representation for crowd attributes recognition,
A. N. Shuaibu, I. Faye, Y . S. Ali, N. Kamel, M. N. Saad, and A. S. Malik, “Sparse representation for crowd attributes recognition,” IEEE Access , vol. 5, pp. 10 422– 10 433, 2017
work page 2017
-
[2]
Research on automatic parking systems based on parking scene recognition,
S. Ma, H. Jiang, M. Han, J. Xie, and C. Li, “Research on automatic parking systems based on parking scene recognition,” IEEE Access , vol. 5, pp. 21 901–21 917, 2017
work page 2017
-
[3]
Fastpose: Towards real-time pose estimation and tracking via scale-normalized multi-task networks,
J. Zhang, Z. Zhu, W. Zou, P. Li, Y . Li, H. Su, and G. Huang, “Fastpose: Towards real-time pose estimation and tracking via scale-normalized multi-task networks,” arXiv preprint arXiv:1908.05593 , 2019
arXiv 1908
-
[4]
Exploiting Offset-guided Network for Pose Estimation and Tracking
R. Zhang, Z. Zhu, P. Li, R. Wu, C. Guo, G. Huang, and H. Xia, “Exploiting offset-guided network for pose esti- mation and tracking,” arXiv preprint arXiv:1906.01344 , 2019
work page Pith review arXiv 1906
-
[5]
State-aware re-identification feature for multi-target multi-camera tracking,
P. Li, J. Zhang, Z. Zhu, Y . Li, L. Jiang, and G. Huang, “State-aware re-identification feature for multi-target multi-camera tracking,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, 2019, pp. 0–0
work page 2019
-
[6]
H.-X. Ma, W. Zou, Z. Zhu, C. Zhang, and Z.-B. Kang, “Selection of observation position and orientation in visual servoing with eye-in-vehicle configuration for manipulator,” International Journal of Automation and Computing, pp. 1–14
-
[7]
Adaptive tra- jectory tracking of wheeled mobile robots based on a fish-eye camera,
Z. Kang, W. Zou, H. Ma, and Z. Zhu, “Adaptive tra- jectory tracking of wheeled mobile robots based on a fish-eye camera,” International Journal of Control, Automation and Systems , pp. 1–13
Show all 74 references
-
[8]
A velocity compensation visual servo method for oculomotor con- trol of bionic eyes,
Z. Zhu, W. Zou, Q. Wang, F. Zhang et al. , “A velocity compensation visual servo method for oculomotor con- trol of bionic eyes,” 2018
2018
-
[9]
Motion control in saccade and smooth pursuit for bionic eye based on three- dimensional coordinates,
Q. Wang, W. Zou, D. Xu, and Z. Zhu, “Motion control in saccade and smooth pursuit for bionic eye based on three- dimensional coordinates,” Journal of Bionic Engineering, vol. 14, no. 2, pp. 336–347, 2017
2017
-
[10]
Optical flow based real-time moving object detection in unconstrained scenes,
J. Huang, W. Zou, J. Zhu, and Z. Zhu, “Optical flow based real-time moving object detection in unconstrained scenes,” arXiv preprint arXiv:1807.04890 , 2018
2018 arXiv
-
[11]
Optical flow based online moving foreground analysis,
J. Huang, W. Zou, and Z. Zhu, “Optical flow based online moving foreground analysis,” arXiv preprint arXiv:1811.07256, 2018. 11 UCT(Ours) SiamFCCCOT CFNet PTAV Fig. 13: Comparisons of our approach with four state-of-the-art trackers in the changing scenario skiing, singer2, iro...
2018 arXiv
-
[12]
An effi- cient optical flow based motion detection method for non-stationary scenes,
J. Huang, W. Zou, Z. Zhu, and J. Zhu, “An effi- cient optical flow based motion detection method for non-stationary scenes,” arXiv preprint arXiv:1811.08290, 2018
2018 arXiv
-
[13]
Motion cue based instance-level moving object detection,
J. Huang, W. Zou, Z. Zhu, J. Zhu et al. , “Motion cue based instance-level moving object detection,” 2019
2019
-
[14]
Std: A stereo tracking dataset for evaluating binocular tracking algorithms,
Z. Zhu, W. Zou, Q. Wang, and F. Zhang, “Std: A stereo tracking dataset for evaluating binocular tracking algorithms,” in 2016 IEEE International Conference on Robotics and Biomimetics (ROBIO) . IEEE, 2016, pp. 2215–2220
2016
-
[15]
Object tracking benchmark,
Y . Wu, J. Lim, and M. H. Yang, “Object tracking benchmark,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 37, no. 9, pp. 1834–1848, 2015
2015
-
[16]
Visual tracking: An exper- imental survey,
A. W. Smeulders, D. M. Chu, R. Cucchiara, S. Calderara, A. Dehghan, and M. Shah, “Visual tracking: An exper- imental survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 36, no. 7, pp. 1442–1468, 2014
2014
-
[17]
The sixth visual object tracking vot2018 challenge results,
M. Kristan, A. Leonardis, J. Matas, M. Felsberg, R. Pflugfelder, L. ˇC. Zajc, T. V oj´ır, G. Bhat, A. Lukeˇziˇc, A. Eldesokey et al. , “The sixth visual object tracking vot2018 challenge results,” in European Conference on Computer Vision. Springer, Cham, 2018, pp. 3–53
2018
-
[18]
The visual object tracking vot2015 challenge results,
K. Matej, J. Matas, A. Leonardis, M. Felsberg, L. ˇCehovin, G. Fern ´andez, T. V oj´ıˇr, G. H ¨ager, G. Nebe- hay, R. Pflugfelder et al. , “The visual object tracking vot2015 challenge results,” in Workshop on the Visual Object Tracking Challenge (VOT, in conjunction with 12 IC...
2015
-
[19]
Online object tracking with sparse prototypes,
D. Wang, H. Lu, and M.-H. Yang, “Online object tracking with sparse prototypes,” IEEE Transactions on Image Processing, vol. 22, no. 1, pp. 314–325, 2013
2013
-
[20]
Inverse sparse tracker with a locally weighted distance metric,
D. Wang, H. Lu, Z. Xiao, and M.-H. Yang, “Inverse sparse tracker with a locally weighted distance metric,” IEEE Transactions on Image Processing , vol. 24, no. 9, pp. 2646–2657, 2015
2015
-
[21]
Struck: Structured output tracking with kernels,
S. Hare, S. Golodetz, A. Saffari, V . Vineet, M.-M. Cheng, S. L. Hicks, and P. H. Torr, “Struck: Structured output tracking with kernels,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 38, no. 10, pp. 2096–2109, 2016
2016
-
[22]
Real-time tracking via on-line boosting,
H. Grabner, M. Grabner, and H. Bischof, “Real-time tracking via on-line boosting,” in Proceedings of the British Machine Vision Conference , 2006
2006
-
[23]
Robust object tracking with online multiple instance learning,
B. Babenko, M.-H. Yang, and S. Belongie, “Robust object tracking with online multiple instance learning,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 33, no. 8, pp. 1619–1632, 2011
2011
-
[24]
High-speed tracking with kernelized correlation filters,
J. F. Henriques, R. Caseiro, P. Martins, and J. Batista, “High-speed tracking with kernelized correlation filters,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 37, no. 3, p. 583, 2015
2015
-
[25]
Learning spatially regularized correlation filters for vi- sual tracking,
M. Danelljan, G. Hager, F. S. Khan, and M. Felsberg, “Learning spatially regularized correlation filters for vi- sual tracking,” in Proceedings of the IEEE International Conference on Computer Vision , 2015, pp. 4310–4318
2015
-
[26]
Long- term correlation tracking,
C. Ma, X. Yang, C. Zhang, and M.-H. Yang, “Long- term correlation tracking,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 5388–5396
2015
-
[27]
A scale adaptive kernel correlation filter tracker with feature integration,
Y . Li and J. Zhu, “A scale adaptive kernel correlation filter tracker with feature integration,” in Proceedings of the European Conference on Computer Vision Workshop , 2014, pp. 254–265
2014
-
[28]
Ex- ploiting the circulant structure of tracking-by-detection with kernels,
J. F. Henriques, C. Rui, P. Martins, and J. Batista, “Ex- ploiting the circulant structure of tracking-by-detection with kernels,” in Proceedings of the European Confer- ence on Computer Vision , 2012, pp. 702–715
2012
-
[29]
Discriminative scale space tracking,
M. Danelljan, G. H ¨ager, F. S. Khan, and M. Felsberg, “Discriminative scale space tracking,”IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 39, no. 8, pp. 1561–1575, 2017
2017
-
[30]
Learning background-aware correlation filters for visual tracking,
H. Kiani Galoogahi, A. Fagg, and S. Lucey, “Learning background-aware correlation filters for visual tracking,” in Proceedings of the IEEE International Conference on Computer Vision, Oct 2017
2017
-
[31]
Imagenet classification with deep convolutional neural networks,
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Proceedings of the Advances in Neural Information Processing Systems, 2012, pp. 1097–1105
2012
-
[32]
Deep resid- ual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep resid- ual learning for image recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 770–778
2016
-
[33]
Faster r-cnn: towards real-time object detection with region proposal networks,
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: towards real-time object detection with region proposal networks,” in Proceedings of the Advances in Neural Information Processing Systems , 2015, pp. 91–99
2015
-
[34]
Fully convolu- tional networks for semantic segmentation,
J. Long, E. Shelhamer, and T. Darrell, “Fully convolu- tional networks for semantic segmentation,” in Proceed- ings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 3431–3440
2015
-
[35]
Hi- erarchical convolutional features for visual tracking,
C. Ma, J.-B. Huang, X. Yang, and M.-H. Yang, “Hi- erarchical convolutional features for visual tracking,” in Proceedings of the IEEE International Conference on Computer Vision, December 2015
2015
-
[36]
Hedged deep tracking,
Y . Qi, S. Zhang, L. Qin, H. Yao, Q. Huang, J. Lim, and M.-H. Yang, “Hedged deep tracking,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, June 2016
2016
-
[37]
Beyond correlation filters: Learning continuous convo- lution operators for visual tracking,
M. Danelljan, A. Robinson, F. S. Khan, and M. Felsberg, “Beyond correlation filters: Learning continuous convo- lution operators for visual tracking,” in Proceedings of the European Conference on Computer Vision , 2016, pp. 472–488
2016
-
[38]
Convolutional features for correlation filter based visual tracking,
M. Danelljan, G. Hager, F. S. Khan, and M. Felsberg, “Convolutional features for correlation filter based visual tracking,” in Proceedings of the IEEE International Con- ference on Computer Vision Workshop , 2015, pp. 621– 629
2015
-
[39]
Fully-convolutional siamese networks for object tracking,
L. Bertinetto, J. Valmadre, J. F. Henriques, A. Vedaldi, and P. H. S. Torr, “Fully-convolutional siamese networks for object tracking,” in Proceedings of the European Conference on Computer Vision Workshop , 2016, pp. 850–865
2016
-
[40]
Visual track- ing with fully convolutional networks,
L. Wang, W. Ouyang, X. Wang, and H. Lu, “Visual track- ing with fully convolutional networks,” in Proceedings of the IEEE International Conference on Computer Vision , 2015, pp. 3119–3127
2015
-
[41]
The visual object tracking vot2016 challenge results,
M. Kristan, A. Leonardis, J. Matas, M. Felsberg, R. Pflugfelder, L. ehovin, T. V ojr, G. Hger, A. Lukei, and G. Fernndez, “The visual object tracking vot2016 challenge results,” in Proceedings of the European Con- ference on Computer Vision Workshop . Springer Inter- national P...
2016
-
[42]
UCT: Learning unified convolutional networks for real-time visual tracking,
Z. Zhu, G. Huang, W. Zou, D. Du, and C. Huang, “UCT: Learning unified convolutional networks for real-time visual tracking,” in 2017 IEEE International Conference on Computer Vision Workshops (ICCVW) , Oct 2017, pp. 1973–1982
2017
-
[43]
Two-stream gated fusion convnets for action recognition,
J. Zhu, W. Zou, and Z. Zhu, “Two-stream gated fusion convnets for action recognition,” in 2018 24th Interna- tional Conference on Pattern Recognition (ICPR). IEEE, 2018, pp. 597–602
2018
-
[44]
Attention-guided unified network for panoptic segmentation,
Y . Li, X. Chen, Z. Zhu, L. Xie, G. Huang, D. Du, and X. Wang, “Attention-guided unified network for panoptic segmentation,” in CVPR, 2019
2019
-
[45]
Action machine: Rethinking action recognition in trimmed videos,
J. Zhu, W. Zou, L. Xu, Y . Hu, Z. Zhu, M. Chang, J. Huang, G. Huang, and D. Du, “Action machine: Rethinking action recognition in trimmed videos,” arXiv preprint arXiv:1812.05770, 2018
2018 arXiv
-
[46]
Learning gating convnet for two-stream based methods in action recognition,
J. Zhu, W. Zou, and Z. Zhu, “Learning gating convnet for two-stream based methods in action recognition,” arXiv preprint arXiv:1709.03655, vol. 1, no. 2, p. 6, 2017. 13
2017 arXiv
-
[47]
Transferring rich feature hierarchies for robust visual tracking,
N. Wang, S. Li, A. Gupta, and D.-Y . Yeung, “Transferring rich feature hierarchies for robust visual tracking,” arXiv preprint arXiv:1501.04587, 2015
2015 arXiv
-
[48]
Deeptrack: Learning dis- criminative feature representations online for robust vi- sual tracking,
H. Li, Y . Li, and F. Porikli, “Deeptrack: Learning dis- criminative feature representations online for robust vi- sual tracking,” IEEE Transactions on Image Processing , vol. 25, no. 4, pp. 1834–1848, 2016
2016
-
[49]
Multi-hierarchical independent correlation filters for vi- sual tracking,
S. Bai, Z. He, T.-B. Xu, Z. Zhu, Y . Dong, and H. Bai, “Multi-hierarchical independent correlation filters for vi- sual tracking,” arXiv preprint arXiv:1811.10302 , 2018
2018 arXiv
-
[51]
Template matching using fast normalized cross correlation,
K. Briechle and U. D. Hanebeck, “Template matching using fast normalized cross correlation,” in Optical Pat- tern Recognition XII , vol. 4387, 2001, pp. 95–103
2001
-
[52]
Real-time track- ing of non-rigid objects using mean shift,
D. Comaniciu, V . Ramesh, and P. Meer, “Real-time track- ing of non-rigid objects using mean shift,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, vol. 2, 2000, pp. 142–149
2000
-
[53]
Visual object tracking using adaptive correlation filters,
D. S. Bolme, J. R. Beveridge, B. A. Draper, and Y . M. Lui, “Visual object tracking using adaptive correlation filters,” in Proceedings of the IEEE Conference on Com- puter Vision and Pattern Recognition , 2010, pp. 2544– 2550
2010
-
[54]
Staple: Complementary learners for real- time tracking,
L. Bertinetto, J. Valmadre, S. Golodetz, O. Miksik, and P. H. S. Torr, “Staple: Complementary learners for real- time tracking,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , June 2016
2016
-
[55]
Correla- tion filters with limited boundaries,
H. Kiani Galoogahi, T. Sim, and S. Lucey, “Correla- tion filters with limited boundaries,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 4630–4638
2015
-
[56]
Learning to track at 100 fps with deep regression networks,
D. Held, S. Thrun, and S. Savarese, “Learning to track at 100 fps with deep regression networks,” in Proceedings of the European Conference on Computer Vision , 2016, pp. 749–765
2016
-
[57]
End-to-end representation learning for correlation filter based tracking,
J. Valmadre, L. Bertinetto, J. F. Henriques, A. Vedaldi, and P. H. S. Torr, “End-to-end representation learning for correlation filter based tracking,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017
2017
-
[58]
Dcfnet: Discriminant correlation filters network for visual track- ing,
Q. Wang, J. Gao, J. Xing, M. Zhang, and W. Hu, “Dcfnet: Discriminant correlation filters network for visual track- ing,” arXiv preprint arXiv:1704.04057 , 2017
2017 arXiv
-
[59]
End-to-end flow correlation tracking with spatial-temporal attention,
Z. Zhu, W. Wu, W. Zou, and J. Yan, “End-to-end flow correlation tracking with spatial-temporal attention,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018
2018
-
[60]
End-to-end video-level representation learning for action recognition,
J. Zhu, Z. Zhu, and W. Zou, “End-to-end video-level representation learning for action recognition,” in 2018 24th International Conference on Pattern Recognition (ICPR). IEEE, 2018, pp. 645–650
2018
-
[61]
High performance visual tracking with siamese region proposal network,
B. Li, J. Yan, W. Wu, Z. Zhu, and X. Hu, “High performance visual tracking with siamese region proposal network,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 8971–8980
2018
-
[62]
Distractor-aware siamese networks for visual object tracking,
Z. Zhu, Q. Wang, B. Li, W. Wu, J. Yan, and W. Hu, “Distractor-aware siamese networks for visual object tracking,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 101–117
2018
-
[63]
Densebox: Unifying landmark localization with end to end object detection,
L. Huang, Y . Yang, Y . Deng, and Y . Yu, “Densebox: Unifying landmark localization with end to end object detection,” arXiv preprint arXiv:1509.04874 , 2015
2015 arXiv
-
[64]
Learning a deep compact im- age representation for visual tracking,
N. Wang and D.-Y . Yeung, “Learning a deep compact im- age representation for visual tracking,” in Proceedings of the Advances in Neural Information Processing Systems , 2013, pp. 809–817
2013
-
[65]
Online object tracking: A benchmark,
Y . Wu, J. Lim, and M. H. Yang, “Online object tracking: A benchmark,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2013, pp. 2411–2418
2013
-
[66]
The visual object tracking vot2015 challenge results,
M. Kristan, J. Matas, A. Leonardis, and M. Felsberg, “The visual object tracking vot2015 challenge results,” in Proceedings of the IEEE International Conference on Computer Vision Workshop , 2015, pp. 564–586
2015
-
[67]
Very deep convolu- tional networks for large-scale image recognition,
K. Simonyan and A. Zisserman, “Very deep convolu- tional networks for large-scale image recognition,” arXiv preprint arXiv:1409.1556, 2014
2014 arXiv
-
[68]
Visualizing and under- standing convolutional networks,
M. D. Zeiler and R. Fergus, “Visualizing and under- standing convolutional networks,” in Proceedings of the European Conference on Computer Vision . Springer, 2014, pp. 818–833
2014
-
[69]
ImageNet Large Scale Visual Recognition Challenge,
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei, “ImageNet Large Scale Visual Recognition Challenge,” International Journal of Computer Vision, vol. 115, no. 3, pp. 211–252, 2015
2015
-
[70]
Caffe: Convolutional architecture for fast feature embedding,
Y . Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell, “Caffe: Convolutional architecture for fast feature embedding,” in Proceedings of the 22nd ACM international conference on Multimedia . ACM, 2014, pp. 675–678
2014
-
[71]
Parallel tracking and verifying: A framework for real-time and high accuracy visual tracking,
H. Fan and H. Ling, “Parallel tracking and verifying: A framework for real-time and high accuracy visual tracking,” in Proceedings of the IEEE International Con- ference on Computer Vision , 2017
2017
-
[72]
Context-aware correlation filter tracking,
M. Mueller, N. Smith, and B. Ghanem, “Context-aware correlation filter tracking,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 1396–1404
2017
-
[73]
Discriminative correlation filter with channel and spatial reliability,
A. Lukei, T. V oj, L. ehovin, J. Matas, and M. Kristan, “Discriminative correlation filter with channel and spatial reliability,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017
2017
-
[74]
Visual tracking using attention- modulated disintegration and integration,
J. Choi, H. Jin Chang, J. Jeong, Y . Demiris, and J. Young Choi, “Visual tracking using attention- modulated disintegration and integration,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 4321–4330
2016
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.