REVIEW 29 references
YO-CSA-T: A Real-time Badminton Tracking System Utilizing YOLO Based on Contextual and Spatial Attention
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The 3D trajectory of a shuttlecock required for a badminton rally robot for human-robot competition demands real-time performance with high accuracy. However, the fast flight speed of the shuttlecock, along with various visual effects, and its tendency to blend with environmental elements, such as court lines and lighting, present challenges for rapid and accurate 2D detection. In this paper, we first propose the YO-CSA detection network, which optimizes and reconfigures the YOLOv8s model's backbone, neck, and head by incorporating contextual and spatial attention mechanisms to enhance model's ability in extracting and integrating both global and local features. Next, we integrate three major subtasks, detection, prediction, and compensation, into a real-time 3D shuttlecock trajectory detection system. Specifically, our system maps the 2D coordinate sequence extracted by YO-CSA into 3D space using stereo vision, then predicts the future 3D coordinates based on historical information, and re-projects them onto the left and right views to update the position constraints for 2D detection. Additionally, our system includes a compensation module to fill in missing intermediate frames, ensuring a more complete trajectory. We conduct extensive experiments on our own dataset to evaluate both YO-CSA's performance and system effectiveness. Experimental results show that YO-CSA achieves a high accuracy of 90.43% mAP@0.75, surpassing both YOLOv8s and YOLO11s. Our system performs excellently, maintaining a speed of over 130 fps across 12 test sequences.
Reference graph
Works this paper leans on
-
[1]
Rich feature hierarchies for accurate object detection and semantic segmentation,
R. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation,” in Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 580–587, 2014
2014
-
[2]
The pascal visual object classes challenge: A retrospective,
M. Everingham, S. A. Eslami, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman, “The pascal visual object classes challenge: A retrospective,” International journal of computer vision , vol. 111, pp. 98–136, 2015
work page 2015
-
[3]
You only look once: Unified, real-time object detection,
J. Redmon, “You only look once: Unified, real-time object detection,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016
2016
-
[4]
Center- net: Keypoint triplets for object detection,
K. Duan, S. Bai, L. Xie, H. Qi, Q. Huang, and Q. Tian, “Center- net: Keypoint triplets for object detection,” in Proceedings of the IEEE/CVF international conference on computer vision , pp. 6569– 6578, 2019
work page 2019
-
[5]
Efficientnet: Rethinking model scaling for con- volutional neural networks,
M. Tan and Q. Le, “Efficientnet: Rethinking model scaling for con- volutional neural networks,” in International conference on machine learning, pp. 6105–6114, PMLR, 2019
2019
-
[6]
Attention is all you need,
A. Vaswani, “Attention is all you need,” Advances in Neural Informa- tion Processing Systems , 2017
2017
-
[7]
Deformable detr: Deformable transformers for end-to-end object detection,
X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai, “Deformable detr: Deformable transformers for end-to-end object detection,” arXiv preprint arXiv:2010.04159, 2020
arXiv 2010
-
[8]
Swin transformer: Hierarchical vision transformer using shifted windows,
Z. Liu, Y . Lin, Y . Cao, H. Hu, Y . Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF international conference on computer vision , pp. 10012–10022, 2021
work page 2021
Show all 29 references
-
[9]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929 , 2020
2010 arXiv
-
[10]
Imagenet: A large-scale hierarchical image database,
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE conference on computer vision and pattern recognition , pp. 248–255, Ieee, 2009
2009
-
[11]
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , pp. 770–778, 2016
2016
-
[12]
Td-net: Trans- deformer network for automatic pancreas segmentation,
S. Dai, Y . Zhu, X. Jiang, F. Yu, J. Lin, and D. Yang, “Td-net: Trans- deformer network for automatic pancreas segmentation,” Neurocom- puting, vol. 517, pp. 279–293, 2023
2023
-
[13]
Tse deeplab: An efficient visual transformer for medical image segmentation,
J. Yang, J. Tu, X. Zhang, S. Yu, and X. Zheng, “Tse deeplab: An efficient visual transformer for medical image segmentation,” Biomedical Signal Processing and Control , vol. 80, p. 104376, 2023
2023
-
[14]
Swin- t-nfc crfs: An encoder–decoder neural model for high-precision uav positioning via point cloud super resolution and image semantic segmentation,
S. Wang, H. Wang, S. She, Y . Zhang, Q. Qiu, and Z. Xiao, “Swin- t-nfc crfs: An encoder–decoder neural model for high-precision uav positioning via point cloud super resolution and image semantic segmentation,” Computer Communications, vol. 197, pp. 52–60, 2023
2023
-
[15]
Ttnet: Real-time temporal and spatial video analysis of table tennis,
R. V oeikov, N. Falaleev, and R. Baikulov, “Ttnet: Real-time temporal and spatial video analysis of table tennis,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pp. 884–885, 2020
2020
-
[16]
Tracknetv2: Efficient shuttlecock tracking network,
N.-E. Sun, Y .-C. Lin, S.-P. Chuang, T.-H. Hsu, D.-R. Yu, H.-Y . Chung, and T.-U. ˙Ik, “Tracknetv2: Efficient shuttlecock tracking network,” in 2020 International Conference on Pervasive Artificial Intelligence (ICPAI), pp. 86–91, IEEE, 2020
2020
-
[17]
Widely applicable strong baseline for sports ball detection and tracking,
S. Tarashima, M. A. Haq, Y . Wang, and N. Tagawa, “Widely applicable strong baseline for sports ball detection and tracking,” arXiv preprint arXiv:2311.05237, 2023
2023 arXiv
-
[18]
U-net: Convolutional networks for biomedical image segmentation,
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, pro- ceedings, part III 18 ...
2015
-
[19]
Deep high-resolution representation learning for visual recognition,
J. Wang, K. Sun, T. Cheng, B. Jiang, C. Deng, Y . Zhao, D. Liu, Y . Mu, M. Tan, X. Wang, et al., “Deep high-resolution representation learning for visual recognition,” IEEE transactions on pattern analysis and machine intelligence , vol. 43, no. 10, pp. 3349–3364, 2020
2020
-
[20]
Deep learning-based algorithm for recognizing tennis balls,
D. Wu and A. Xiao, “Deep learning-based algorithm for recognizing tennis balls,” Applied Sciences, vol. 12, no. 23, p. 12116, 2022
2022
-
[21]
Efficient golf ball detection and tracking based on convolutional neural networks and kalman filter,
T. Zhang, X. Zhang, Y . Yang, Z. Wang, and G. Wang, “Efficient golf ball detection and tracking based on convolutional neural networks and kalman filter,” arXiv preprint arXiv:2012.09393 , 2020
2012 arXiv
-
[22]
Yolov3: An incremental improvement,
J. Redmon, “Yolov3: An incremental improvement,” arXiv preprint arXiv:1804.02767, 2018
2018 arXiv
-
[23]
An introduction to the kalman filter,
G. Bishop, G. Welch, et al. , “An introduction to the kalman filter,” Proc of SIGGRAPH, Course , vol. 8, no. 27599-23175, p. 41, 2001
2001
-
[24]
Table tennis track detection based on temporal feature multiplexing network,
W. Li, X. Liu, K. An, C. Qin, and Y . Cheng, “Table tennis track detection based on temporal feature multiplexing network,” Sensors, vol. 23, no. 3, p. 1726, 2023
2023
-
[25]
Faster r-cnn: Towards real- time object detection with region proposal networks,
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real- time object detection with region proposal networks,” IEEE transac- tions on pattern analysis and machine intelligence , vol. 39, no. 6, pp. 1137–1149, 2016
2016
-
[26]
Contextual transformer networks for visual recognition,
Y . Li, T. Yao, Y . Pan, and T. Mei, “Contextual transformer networks for visual recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 2, pp. 1489–1500, 2022
2022
-
[27]
Spatial group-wise enhance: Improving semantic feature learning in convolutional networks,
X. Li, X. Hu, and J. Yang, “Spatial group-wise enhance: Improving semantic feature learning in convolutional networks,” arXiv preprint arXiv:1905.09646, 2019
1905 arXiv
-
[28]
Dynamic routing between capsules,
S. Sabour, N. Frosst, and G. E. Hinton, “Dynamic routing between capsules,” Advances in neural information processing systems, vol. 30, 2017
2017
-
[29]
Tracknetv3: Enhancing shuttlecock track- ing with augmentations and trajectory rectification,
Y .-J. Chen and Y .-S. Wang, “Tracknetv3: Enhancing shuttlecock track- ing with augmentations and trajectory rectification,” in Proceedings of the 5th ACM International Conference on Multimedia in Asia, pp. 1–7, 2023
2023
Discussion (0). Continue with ORCID to comment.