REVIEW 3 major objections 5 minor 44 references
Real Time Visual Tracking using Spatial-Aware Temporal Aggregation Network
T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read The paper proposes SATA, a correlation-filter tracker that aligns and aggregates historical frame features with deformable convolutions, and claims real-time leading results on five public tracking benchmarks.
desk verdict Solid and promising tracking architecture, but the headline accuracy and speed claims don't survive contact with the paper's own tables. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the Alignment Module: a deformable-convolution block that concatenates search and history features, predicts a per-pixel offset field, and samples the history feature at the aligned locations. It is paired with an Aggregation Module that assigns each aligned history feature a per-location weight equal to normalized cosine similarity between embedded history and search features, so unreliable history is down-weighted. A feature pyramid separates shallow, detail-rich layers from deep, semantic layers, allowing offsets to be learned at suitable resolutions, and the differentiable correlation filter layer carries the tracking loss that supervises the whole pipeline.
What would settle it
Take a synthetic video with a precisely known constant translation and compare the Alignment Module's predicted per-pixel offsets to the true displacement; if the tracked output stays accurate even when the predicted offsets are wrong, temporal alignment is not what drives the gain. Alternatively, ablate the Alignment Module by fixing all offsets to zero and check whether the aggregation gain over the No Agg baseline survives; if it does, the claimed spatial-alignment mechanism is not load-bearing.
Extended reading notes
Core claim
The central claim is that spatial-aware temporal aggregation is an effective upgrade for correlation-filter trackers. Historical frame features are warped to the search frame by an Alignment Module that predicts pixel-level offsets with deformable convolutions, then combined with the search frame features by per-location adaptive weights in an embedding space. A feature pyramid aligns deep semantic layers and shallow detail layers at their own resolutions before fusing them top-down, and the aggregated features feed a differentiable correlation filter layer so the whole pipeline trains end-to-end from the tracking loss. The paper reports that this design outperforms the no-aggregation version and reaches competitive or leading results on OTB2013, OTB2015, VOT2015, VOT2016, and LaSOT.
Load-bearing premise
The alignment offsets are learned only from the final tracking loss with no direct supervision, so the whole benefit of aggregation rests on the network spontaneously learning offsets that move history features onto the right pixels; if the offsets are wrong, aggregation blurs instead of sharpens.
Editorial extensions
If this is right
- On OTB2013, adding temporal aggregation raises success AUC from 0.646 (No Agg) to 0.698 (full model), so history frames measurably improve accuracy.
- Aggregating the two deepest pyramid layers ($P^l$ and $P^{l-1}$) reaches 0.692 AUC at 19 FPS, so most of the gain is available at practical speed.
- The middle-layer aggregation alone (0.671 AUC) beats the shallowest alone (0.662 AUC), while adding the deepest layer gives the best result, supporting the paper's coarse-to-fine fusion order.
- Aggregation weights decrease as history frames get farther from the search frame, so a moderate history window balances information gain against appearance drift.
- The model needs no fine-tuning on VOT or LaSOT to be competitive there, suggesting the learned alignment and aggregation transfer across benchmarks.
Reading between the lines
- The align-and-aggregate module is not tied to a correlation filter; the same deformable alignment and cosine-similarity weighting could be inserted into Siamese trackers or other online learners, a transfer the paper does not test.
- Because the offsets are trained only through the final tracking loss, adding a light optical-flow supervision signal to the Alignment Module might improve long-range aggregation on fast-motion sequences; this is a testable extension.
- The per-location aggregation weights behave like an online confidence map, so a future tracker could use a sharp drop in history weights as a signal for occlusion or drift.
- The reported speed-accuracy table suggests the two-level aggregation configuration is the practical operating point, since the three-level version buys only a small AUC gain while slowing the tracker noticeably.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes SATA, a correlation-filter-based visual tracker that aggregates features from historical frames into the current search frame using a pixel-level alignment module based on deformable convolutions, combined with a feature pyramid for multi-scale aggregation and an end-to-end trainable correlation filter layer. The method is evaluated on OTB-2013, OTB-2015, VOT-2015, VOT-2016, and LaSOT, with ablations on OTB-2013. The authors claim leading performance on these benchmarks and a real-time speed of 26 FPS. The central technical idea is that spatial-aware temporal aggregation of historical features improves tracking accuracy over using a single static frame.
Significance. If the technical claims were fully substantiated, the work would make a useful contribution by showing that video-level temporal aggregation, previously successful in video object detection, can be adapted to visual tracking with deformable-convolution alignment and multi-scale fusion. The paper includes end-to-end training, ablations that consistently show improvement of the full aggregation over the no-aggregation baseline, and evaluation on several standard benchmarks. However, the headline claims of 'leading performance' and 'real-time speed of 26 FPS' are internally inconsistent with the paper's own reported tables and figures, and the alignment module's behavior is not directly validated. These issues are load-bearing for the paper's central claims and require correction before the contributions can be accepted as stated.
major comments (3)
- [Abstract, §5.1, §5.5, Table 3] The real-time speed claim is not supported by the paper's own ablation table. The abstract and §5.1 state an average speed of 26 FPS, and §5.5 says 'PlPl−1 Agg operate on real time speed 26 FPS and AUC 0.677 when set T = 3 (last line in Table 3)'. However, Table 3 shows that the configuration achieving the headline OTB-2013 AUC of 0.698 is PlPl−1Pl−2 Agg with T=3, which runs at 14 FPS. The only row reporting 26 FPS has T=2 and all three feature layers (Pl, Pl−1, Pl−2) checked, and its AUC is 0.677, not 0.698. The text's reference to 'last line' is also inconsistent with the table's actual last line. No single configuration presented supports both the top reported accuracy and the claimed real-time speed; please correct the text and either report the speed of the configuration used for the headline accuracy or clearly specify which configuration produces 26 FPS.
- [Abstract, §5.2, §5.3, Table 1, Figures 3-4] The claim of 'leading performance' on OTB2013, OTB2015, VOT2015, and VOT2016 is contradicted by the paper's own results. On OTB-2015, Figure 4 shows CCOT with a higher success score (0.671 vs 0.661) and higher precision (0.898 vs 0.872) than SATA. In VOT-2015, §5.3 states that 'the performance of SATA ranks 2nd after MDNet'. In VOT-2016, Table 1 reports CCOT with EAO 0.331 versus SATA's 0.319. The abstract and conclusion should be revised to state accurately that SATA is competitive or superior on specific benchmarks where the evidence supports it, rather than claiming leading performance across all listed benchmarks.
- [§3.2, §4.1] The Alignment Module is the mechanism that makes temporal aggregation safe, but the paper provides no direct evaluation or supervision of the predicted offsets. Offsets are learned only through the final tracking loss, and the paper does not report offset accuracy, visualizations of alignment quality, or any ablation comparing the learned deformable alignment against an oracle or flow-based alignment. If the learned offsets are inaccurate, aggregation could blur rather than sharpen features, and the claimed benefit would not be attributable to the proposed mechanism. Please add an analysis of the learned offsets, or at minimum an ablation that isolates the effect of the alignment module, to support the claim that spatial alignment is what drives the improvement.
minor comments (5)
- [§5.5] The text says that PlPl−1 Agg 'gains the performance with more than 0.05' compared to the No Agg baseline, but the numbers in Table 3 show an improvement of 0.692 − 0.646 = 0.046, which is less than 0.05.
- [Table 3] The last row of Table 3 is labeled 'PlPl−1 Agg' but has all three layers checked and T=2; this appears to be a copy-and-paste error, as the row should be labeled consistently with its configuration, e.g., 'PlPl−1Pl−2 Agg (T=2)'.
- [§5.1, Figure 2, §5.5] There are several typos and minor wording issues, including 'seted', 'comoared', 'trakcers', and 'Temporal infromation' in Figure 2; these should be corrected.
- [Figure 8] The temporal aggregation analysis in Figure 8 does not state which pyramid layers are used for the varying number of historical frames, nor does it provide error bars or statistical significance; please clarify the configuration used.
- [Table 2] The LaSOT comparison reports only a single run without statistical significance or protocol details beyond 'protocol I'; please state whether these are mean results over multiple runs and cite the protocol definition.
Circularity Check
No circular derivation found; reported benchmark numbers are external, though the headline 26 FPS and 0.698 AUC come from different Table 3 configurations.
full rationale
The paper's central claim, that spatial-aware temporal aggregation improves tracking, is supported by ablation studies and by evaluation on external benchmarks (OTB, VOT, LaSOT) rather than being derived from a fitted parameter or from the paper's own definitions. The model is trained on ILSVRC-2015 and tested on other benchmarks, so no reported benchmark result is used as a training target. The Alignment Module is learned end-to-end with the final CF loss, which is an internal design assumption and a potential correctness risk, but not circularity because the claimed improvement is not asserted by construction. Citations to Refs. [16,36] for CF-layer backpropagation are external prior work, not self-citations, and they do not carry the central claim. One in-scope issue is flagged here only for completeness: Section 5.5 states 'PlPl−1 Agg operate on real time speed 26 FPS and AUC 0.677 when set T = 3 (last line in Table 3)', but the last row of Table 3 has T=2 and checks all three pyramid layers, while the headline OTB-2013 AUC of 0.698 corresponds to the 14 FPS PlPl−1Pl−2 row. That is an internal reporting inconsistency, not a circular step, so it does not affect the circularity score.
Assumptions & free parameters
free parameters (6)
- Historical frame window size T =
3 (training), 2 to 5 evaluated
- Number of aggregated feature pyramid layers =
3 (Pl, Pl-1, Pl-2) for full SATA
- CF regularization coefficient lambda =
1e-4
- Gaussian response bandwidth =
0.1
- Scale estimation hyperparameters =
scale step 1.03, number of scales 3, scale penalty 0.993, model update rate 0.01
- Training schedule hyperparameters =
learning rate 1e-5, weight decay 5e-4, momentum 0.9, batch 32, epochs 50
assumptions (4)
- standard math The closed-form DCF solution in Eq. (4) follows from minimizing the ridge loss in Eq. (3) in the Fourier domain.
- domain assumption ILSVRC-2015 video detection clips provide a training distribution sufficient for single-object tracking.
- domain assumption Deformable-convolution offsets, trained only by the tracking loss, encode meaningful inter-frame motion.
- domain assumption Weighted aggregation of aligned historical features improves the response map compared to the single-frame feature.
Cite this review
Pith. "Pith review of Real Time Visual Tracking using Spatial-Aware Temporal Aggregation Network." pith.science (2026). https://pith.science/paper/NO4MLUYU
@misc{pith2026190800692,
author = {Pith},
title = {Pith review of: Real Time Visual Tracking using Spatial-Aware Temporal Aggregation Network},
year = {2026},
howpublished = {\url{https://pith.science/paper/NO4MLUYU}},
note = {Machine review of arXiv:1908.00692}
}
read the original abstract
More powerful feature representations derived from deep neural networks benefit visual tracking algorithms widely. However, the lack of exploitation on temporal information prevents tracking algorithms from adapting to appearances changing or resisting to drift. This paper proposes a correlation filter based tracking method which aggregates historical features in a spatial-aligned and scale-aware paradigm. The features of historical frames are sampled and aggregated to search frame according to a pixel-level alignment module based on deformable convolutions. In addition, we also use a feature pyramid structure to handle motion estimation at different scales, and address the different demands on feature granularity between tracking losses and deformation offset learning. By this design, the tracker, named as Spatial-Aware Temporal Aggregation network (SATA), is able to assemble appearances and motion contexts of various scales in a time period, resulting in better performance compared to a single static image. Our tracker achieves leading performance in OTB2013, OTB2015, VOT2015, VOT2016 and LaSOT, and operates at a real-time speed of 26 FPS, which indicates our method is effective and practical. Our code will be made publicly available at \href{https://github.com/ecart18/SATA}{https://github.com/ecart18/SATA}.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
N. Ballas, L. Yao, C. Pal, and A. C. Courville. Delving deeper into convolutional networks for learning video rep- resentations. In 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings, 2016. 2
work page 2016
-
[2]
G. Bertasius, L. Torresani, and J. Shi. Object detection in video with spatiotemporal sampling networks. In Com- puter Vision - ECCV 2018 - 15th European Conference, Mu- nich, Germany, September 8-14, 2018, Proceedings, Part XII, pages 342–357, 2018. 1, 2, 3
work page 2018
-
[3]
L. Bertinetto, J. Valmadre, J. F. Henriques, A. Vedaldi, and P. H. S. Torr. Fully-convolutional siamese networks for ob- ject tracking. In Computer Vision - ECCV 2016 Workshops - Amsterdam, The Netherlands, October 8-10 and 15-16, 2016, Proceedings, Part II, pages 850–865, 2016. 2, 6
work page 2016
-
[4]
G. Bhat, J. Johnander, M. Danelljan, F. S. Khan, and M. Fels- berg. Unveiling the power of deep tracking. In Computer Vision - ECCV 2018 - 15th European Conference, Munich, Germany, September 8-14, 2018, Proceedings, Part II, pages 493–509, 2018. 1, 2
work page 2018
-
[5]
D. S. Bolme, J. R. Beveridge, B. A. Draper, and Y . M. Lui. Visual object tracking using adaptive correlation filters. In The Twenty-Third IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2010, San Francisco, CA, USA, 13-18 June 2010, pages 2544–2550, 2010. 1, 2
work page 2010
-
[6]
J. Dai, H. Qi, Y . Xiong, Y . Li, G. Zhang, H. Hu, and Y . Wei. Deformable convolutional networks. In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017, pages 764–773, 2017. 2, 4
work page 2017
-
[7]
M. Danelljan, G. Bhat, F. S. Khan, and M. Felsberg. ECO: efficient convolution operators for tracking. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, July 21-26, 2017 , pages 6931–6939, 2017. 6
work page 2017
-
[8]
M. Danelljan, G. H ¨ager, F. S. Khan, and M. Felsberg. Ac- curate scale estimation for robust visual tracking. In British Machine Vision Conference, BMVC 2014, Nottingham, UK, September 1-5, 2014, 2014. 6
work page 2014
Show all 44 references
-
[9]
Danelljan, G
M. Danelljan, G. H ¨ager, F. S. Khan, and M. Felsberg. Con- volutional features for correlation filter based visual tracking. In 2015 IEEE International Conference on Computer Vision Workshop, ICCV Workshops 2015, Santiago, Chile, Decem- ber 7-13, 2015, pages 621–629, 2015. 1, 2
2015
-
[10]
Danelljan, G
M. Danelljan, G. H ¨ager, F. S. Khan, and M. Felsberg. Learn- ing spatially regularized correlation filters for visual track- ing. In 2015 IEEE International Conference on Computer Vision, ICCV 2015, Santiago, Chile, December 7-13, 2015 , pages 4310–4318, 2015. 2, 3, 6
2015
-
[11]
Danelljan, F
M. Danelljan, F. S. Khan, M. Felsberg, and J. van de Wei- jer. Adaptive color attributes for real-time visual tracking. In 2014 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2014, Columbus, OH, USA, June 23-28, 2014, pages 1090–1097, 2014. 1, 2
2014
-
[12]
Danelljan, A
M. Danelljan, A. Robinson, F. S. Khan, and M. Felsberg. Beyond correlation filters: Learning continuous convolution operators for visual tracking. In Computer Vision - ECCV 2016 - 14th European Conference, Amsterdam, The Nether- lands, October 11-14, 2016, Proceedings, Part V ,...
2016
-
[13]
J. Deng, W. Dong, R. Socher, L. Li, K. Li, and F. Li. Im- agenet: A large-scale hierarchical image database. In 2009 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2009), 20-25 June 2009, Miami, Florida, USA, pages 248–255, 2009. 2, 6
2009
-
[14]
Dosovitskiy, P
A. Dosovitskiy, P. Fischer, E. Ilg, P. H ¨ausser, C. Hazirbas, V . Golkov, P. van der Smagt, D. Cremers, and T. Brox. Flownet: Learning optical flow with convolutional networks. In 2015 IEEE International Conference on Computer Vision, ICCV 2015, Santiago, Chile, December 7-13,...
2015
-
[15]
H. Fan, L. Lin, F. Yang, P. Chu, G. Deng, S. Yu, H. Bai, Y . Xu, C. Liao, and H. Ling. Lasot: A high-quality benchmark for large-scale single object tracking. CoRR, abs/1809.07845, 2018. 7
2018 arXiv
-
[16]
Gundogdu and A
E. Gundogdu and A. A. Alatan. Good features to corre- late for visual tracking. IEEE Trans. Image Processing , 27(5):2526–2540, 2018. 1, 2, 3, 5
2018
-
[17]
J. F. Henriques, R. Caseiro, P. Martins, and J. Batista. High- speed tracking with kernelized correlation filters. IEEE Trans. Pattern Anal. Mach. Intell., 37(3):583–596, 2015. 1, 2
2015
-
[18]
J. F. Henriques, R. Caseiro, P. Martins, and J. P. Batista. Ex- ploiting the circulant structure of tracking-by-detection with kernels. In Computer Vision - ECCV 2012 - 12th European Conference on Computer Vision, Florence, Italy, October 7- 13, 2012, Proceedings, Part IV, pag...
2012
-
[19]
P. Hu, G. Wang, X. Kong, J. Kuen, and Y . Tan. Motion- guided cascaded refinement network for video object seg- mentation. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, pages 1400–1409, 2018. 3
2018
-
[20]
E. Ilg, N. Mayer, T. Saikia, M. Keuper, A. Dosovitskiy, and T. Brox. Flownet 2.0: Evolution of optical flow estimation with deep networks. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, July 21-26, 2017, pages 1647–1655, 2017. 2, 3
2017
-
[21]
Kristan, J
M. Kristan, J. Matas, A. Leonardis, M. Felsberg, L. Cehovin, G. Fern ´andez, T. V oj´ır, G. H ¨ager, G. Nebehay, and R. P. Pflugfelder. The visual object tracking VOT2015 challenge results. In 2015 IEEE International Conference on Computer Vision Workshop, ICCV Workshops 2015, ...
2015
-
[22]
H. Li, Y . Li, and F. Porikli. Deeptrack: Learning discrimina- tive feature representations online for robust visual tracking. IEEE Trans. Image Processing, 25(4):1834–1848, 2016. 2
2016
-
[23]
T. Lin, P. Doll ´ar, R. B. Girshick, K. He, B. Hariharan, and S. J. Belongie. Feature pyramid networks for object detec- tion. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, July 21-26, 2017, pages 936–944, 2017. 3
2017
-
[24]
H. Lu, P. Li, and D. Wang. Visual object tracking: A survey. Pattern Recognition and Artificial Intelligence, 31(1):61–76,
-
[25]
C. Ma, J. Huang, X. Yang, and M. Yang. Hierarchical con- volutional features for visual tracking. In 2015 IEEE Inter- national Conference on Computer Vision, ICCV 2015, San- tiago, Chile, December 7-13, 2015, pages 3074–3082, 2015. 2
2015
-
[26]
Nam and B
H. Nam and B. Han. Learning multi-domain convolutional neural networks for visual tracking. In 2016 IEEE Confer- ence on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV , USA, June 27-30, 2016, pages 4293– 4302, 2016. 7
2016
-
[27]
Paszke, S
A. Paszke, S. Gross, S. Chintala, and G. Chanan. Pytorch: Tensors and dynamic neural networks in python with strong gpu acceleration. PyTorch: Tensors and dynamic neural net- works in Python with strong GPU acceleration, 2017. 6
2017
-
[28]
Ranjan and M
A. Ranjan and M. J. Black. Optical flow estimation using a spatial pyramid network. In 2017 IEEE Conference on Com- puter Vision and Pattern Recognition, CVPR 2017, Hon- olulu, HI, USA, July 21-26, 2017 , pages 2720–2729, 2017. 3
2017
-
[29]
Russakovsky, J
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. S. Bernstein, A. C. Berg, and F. Li. Imagenet large scale visual recog- nition challenge. International Journal of Computer Vision, 115(3):211–252, 2015. 2, 6
2015
-
[30]
Simonyan and A
K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. In 3rd Interna- tional Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Pro- ceedings, 2015. 6
2015
-
[31]
Y . Song, C. Ma, X. Wu, L. Gong, L. Bao, W. Zuo, C. Shen, R. W. H. Lau, and M. Yang. VITAL: visual tracking via adversarial learning. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, pages 8990–8999, 2018. 1
2018
-
[32]
Tripathi, Z
S. Tripathi, Z. C. Lipton, S. J. Belongie, and T. Q. Nguyen. Context matters: Refining object detection in video with recurrent neural networks. In Proceedings of the British Machine Vision Conference 2016, BMVC 2016, York, UK, September 19-22, 2016, 2016. 2
2016
-
[33]
Z. Tu, W. Xie, D. Zhang, R. Poppe, R. C. Veltkamp, B. Li, and J. Yuan. A survey of variational and cnn-based optical flow techniques. Sig. Proc.: Image Comm. , 72:9–24, 2019. 3
2019
-
[34]
Valmadre, L
J. Valmadre, L. Bertinetto, J. F. Henriques, A. Vedaldi, and P. H. S. Torr. End-to-end representation learning for correla- tion filter based tracking. In 2017 IEEE Conference on Com- puter Vision and Pattern Recognition, CVPR 2017, Hon- olulu, HI, USA, July 21-26, 2017 , pages...
2017
-
[35]
N. Wang, W. Zhou, Q. Tian, R. Hong, M. Wang, and H. Li. Multi-cue correlation filters for robust visual tracking. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18- 22, 2018, pages 4844–4853, 2018. 2
2018
-
[36]
Q. Wang, J. Gao, J. Xing, M. Zhang, and W. Hu. Dcfnet: Discriminant correlation filters network for visual tracking. CoRR, abs/1704.04057, 2017. 1, 2, 5
2017 arXiv
-
[37]
Y . Wu, J. Lim, and M. Yang. Online object tracking: A benchmark. In 2013 IEEE Conference on Computer Vision and Pattern Recognition, Portland, OR, USA, June 23-28, 2013, pages 2411–2418, 2013. 6
2013
-
[38]
Y . Wu, J. Lim, and M. Yang. Object tracking benchmark. IEEE Trans. Pattern Anal. Mach. Intell. , 37(9):1834–1848,
-
[39]
Xiao and Y
F. Xiao and Y . J. Lee. Video object detection with an aligned spatial-temporal memory. In Computer Vision - ECCV 2018 - 15th European Conference, Munich, Germany, September 8-14, 2018, Proceedings, Part VIII, pages 494–510, 2018. 1, 2, 3
2018
-
[40]
H. Xiao, J. Feng, G. Lin, Y . Liu, and M. Zhang. Monet: Deep motion exploitation for video object segmentation. In 2018 IEEE Conference on Computer Vision and Pattern Recogni- tion, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, pages 1140–1148, 2018. 3
2018
-
[41]
L. Yang, Y . Wang, X. Xiong, J. Yang, and A. K. Katsaggelos. Efficient video object segmentation via network modulation. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18- 22, 2018, pages 6499–6507, 2018. 3
2018
-
[42]
Yosinski, J
J. Yosinski, J. Clune, A. M. Nguyen, T. J. Fuchs, and H. Lip- son. Understanding neural networks through deep visualiza- tion. CoRR, abs/1506.06579, 2015. 3
2015 arXiv
-
[43]
X. Zhu, Y . Wang, J. Dai, L. Yuan, and Y . Wei. Flow-guided feature aggregation for video object detection. In IEEE In- ternational Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017 , pages 408–417, 2017. 1, 2, 3, 5
2017
-
[44]
Z. Zhu, W. Wu, W. Zou, and J. Yan. End-to-end flow cor- relation tracking with spatial-temporal attention. In 2018 IEEE Conference on Computer Vision and Pattern Recog- nition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, pages 548–557, 2018. 2, 3
2018
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.