REVIEW 4 major objections 5 minor 1 cited by
Survey on Hand Gesture Recognition from Visual Input
T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read This survey claims that recent hand gesture recognition research from visual input can be organized into six research topics and a five-axis method taxonomy, and that this organization reveals systematic associations—box/filter capture…
desk verdict A genuinely useful survey of visual-input HGR with a data-quality problem: the taxonomy is solid, but the paper's own counts don't add up and the retrieval pipeline misses depth-only work. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the survey's classification framework itself: six topic clusters discovered by non-negative matrix factorization of paper titles, keywords, and abstracts—hand gesture classification, hand gesture estimation, sign language recognition, hand/body reconstruction, multimodal fusion, and real-time recognition—combined with a five-axis methodological table (input type, camera setup, capture method, task goal, recognition method). The framework is used to tabulate all 125 reviewed papers and, through Bayes' theorem, to compute conditional probabilities such as P(classification | box/filter), the quantitative evidence for the paper's associations. The same machinery structures the dataset and challenge sections, making the taxonomy the device that connects selection, analysis, and conclusions.
What would settle it
Re-run the selection pipeline with an additional broad query such as 'egocentric hand' or 'hand-object interaction' for the same period and same venues, then recompute the reported percentages for input type and recognition method; if the added papers shift those percentages substantially or introduce new frequent topics, the 125-paper corpus is not representative and the associations the survey draws would need re-examination.
Extended reading notes
Core claim
The paper's central claim is that a systematic review of visual-input hand gesture recognition, built from top-venue publications and targeted database queries, yields a coherent taxonomy that previous surveys lacked. The taxonomy separates gesture classification from gesture estimation, then cross-cuts those tasks by input modality (RGB, RGB-D, video), camera setup (monocular, multi-view), hand capture representation (skeleton-based versus box/filter-based), and recognition method (neural network, non-neural, or hybrid). Within this corpus the authors report that video input accounts for 53% of studies, hybrid methods for 68%, box/filter capture is about twice as likely to be associated with classification than estimation, and multi-view setups are more strongly associated with estimation. They further provide a dataset inventory showing ASL and HO3D as the most-used benchmarks for classification and estimation respectively and argue that lack of standardized benchmarks is a central limitation of current research.
Load-bearing premise
The survey's claims about trends and associations depend on its literature retrieval pipeline—a crawl of top-venue publication lists plus four database queries, filtered by titles, abstracts, and topic modeling—returning a representative sample of hand gesture recognition research from 2018 to 2025.
Editorial extensions
If this is right
- Researchers entering the field can use the taxonomy to position a new method against the dominant video-and-hybrid baseline rather than searching across hundreds of papers.
- The reported associations give testable expectations: a new box/filter method is more likely aimed at classification, and a multi-view system at hand pose estimation.
- Datasets ASL and HO3D function as de facto benchmarks for classification and estimation, so new methods will be compared against the headline numbers the survey compiles.
- Because accuracy on some benchmark datasets is already very high (100% on ASL, 98.53% on AUTSL), progress signals will increasingly come from harder, more naturalistic datasets such as WLASL and isoGD, where top accuracies remain below 85%.
- The survey's call for unified evaluation frameworks, if heeded, would make its own cross-paper performance table reproducible and comparable.
Reading between the lines
- A consequence the authors leave implicit: if the field adopted their taxonomy as a reporting standard, meta-analyses could track shifts in method prevalence over time and test whether the 2024 spike in publications continues.
- The Bayes associations are corpus-relative; a broader corpus that included more hand-object interaction and egocentric work would likely raise the estimation share and strengthen the multiview-estimation link.
- A testable extension would be to run the same selection pipeline on the 2025-2026 literature and check whether transformer-based methods displace CNN+LSTM hybrids as the dominant approach, a trend the paper identifies as emerging.
- The benchmarking gap they identify suggests a concrete next step: a reproducibility study that fixes data splits and metrics, re-evaluating the leading methods on ASL, AUTSL, JESTER, WLASL, HO3D, and FreiHAND under one protocol.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a structured survey of hand gesture recognition (HGR) and 3D hand pose estimation from visual input (RGB, RGB-D, and video), covering work published between 2018 and 2025. It proposes a taxonomy based on task objective, input type, capture method, and recognition technique; applies non-negative matrix factorization (NNMF) to organize retrieved papers into six topics; tabulates benchmark datasets and state-of-the-art results; and discusses challenges, edge deployment, explainability, bias, and privacy. The stated goal is to provide a comprehensive yet focused alternative to broader recent surveys that include non-visual modalities.
Significance. If its selection process is representative and its counts are correct, the survey would be a useful entry point for researchers: it offers a clear classification scheme, a compact dataset table, an algorithmic overview with common formulations, and welcome sections on deployment and ethical considerations. The use of topic modeling to structure the literature is a constructive addition to typical survey methodology. Its value as a reference, however, depends on reproducible retrieval and internally consistent statistics, both of which currently need repair; those issues affect the headline claims about trends and method prevalence.
major comments (4)
- [Table I, Section III-A, Section III-C] The reported paper counts are internally inconsistent: Table I states that the survey reviews 137 papers, Section III-C states 125 studies (37 from Scholar and 88 from Scopus), and Table II's Selected Papers column sums to 89 rather than 88. Since the quantitative claims (venue distribution, input-type percentages, topic timeline) are aggregates over this set, please reconcile the totals and provide a complete list of the included papers, or a public manifest, so the counts can be audited.
- [Table II, Section IV-B] The Scopus query design systematically excludes depth-only work: query 1 requires the term 'RGB', and query 3 permits 'RGB', 'video', 'skeleton', or 'multi modal' but not 'depth' or 'RGB-D'. Because the survey explicitly claims to cover depth images as input and reports that 19% of its selected papers use RGB-D input (Fig. 8a), any depth-based hand pose estimation or gesture recognition paper that does not mention RGB/video/skeleton is invisible to the selection pipeline. Please add explicit depth/RGB-D queries and rerun the selection, or qualify the comprehensiveness claim accordingly.
- [Section III-A, Section III-D] The NNMF-based selection step is not reproducible as reported: Section III-A says NNMF was applied to the titles, keywords, and abstracts of the papers, while Section III-D says it was applied only to titles and keywords; no relevance threshold, per-paper topic assignment rule, or list of papers rejected after topic modeling is provided. Please specify the exact text fields used, define the relevance criterion, and release the topic assignments or the selection script so that another group can reproduce the 125/137-paper set.
- [Table VII] The state-of-the-art table requires verification before it can be trusted: the reported best MPJPE of 1.1 on HO3D and 1.18 on FreiHAND are well below typical published results on these benchmarks (which are usually in the millimeter-to-centimeter range depending on the protocol), and the 100% accuracy on ASL is presented without any dataset split or evaluation details. Please state the exact metric unit, evaluation protocol, and data split for each row, or remove values that cannot be substantiated.
minor comments (5)
- [Section I] The sentence 'a "true" vision-based approach reported by in 1993 [161]' is missing the author name and should be reworded.
- [Section III-C] The sentence 'The methodology for extracting the topics from the collection of articles is detailed in Section II-D that follows' should refer to Section III-D, not Section II-D.
- [Table V] The 'American Sign Language Digits' row cites the same reference [17] as the MUGD row, reports a different sample count, and lists 36 classes for a digit dataset; this appears to be a dataset mislabeling and should be checked.
- [Section V-F] The notation 'WQ.WK.WV' should be written as 'WQ, WK, WV' or an equivalent list, since the periods are ambiguous.
- [Section VII-B] The citation for Transformer models in the sentence 'Vision Transformers (ViT) [16]' points to a gesture-recognition application paper rather than the original ViT paper (Dosovitskiy et al.); please update the reference.
Circularity Check
No significant circularity: the survey's claims are descriptive syntheses of external literature, with no derivation chain that reduces to its own inputs.
full rationale
This paper is a literature survey with no original derivations, no fitted parameters, and no predictive claims that could be reducible to its inputs by construction. The central claim is that the survey comprehensively synthesizes recent hand gesture recognition research from RGB, depth, and video input (Abstract, Section I). That claim is supported by the retrieval pipeline of Section III, which selects papers from Google Scholar top venues and Scopus queries and then organizes them into a taxonomy. The taxonomy categories (classification vs. estimation, RGB vs. RGB-D vs. video, monocular vs. multiview, skeleton vs. box/filter, NN vs. non-NN vs. hybrid) are applied to the selected papers rather than derived from them, so there is no Eq. X = Eq. Y by construction. The two self-citations by the authors ([2] in the Introduction and [170] in the Explainability section) are peripheral background references and are not load-bearing for any of the survey's conclusions. The manuscript's internal inconsistencies and potential retrieval blind spots (e.g., Table I says 137 papers while Section III-C says 125; the Scopus query strings omit depth-only search terms) are correctness and reproducibility risks in the sampling methodology, and are appropriately classified as such rather than as circularity. Because the survey is self-contained against external literature and makes no prediction that is statistically forced by its own inputs, no circular step can be exhibited.
Assumptions & free parameters
free parameters (1)
- Number of NNMF topics =
6
assumptions (3)
- domain assumption The Google Scholar top-venue crawl plus Scopus queries returns a representative sample of HGR research.
- domain assumption NNMF topic modeling with six manually refined components yields stable, meaningful categories.
- domain assumption Performance numbers from cited papers are correctly transcribed and comparable enough for Table VII.
Cite this review
Pith. "Pith review of Survey on Hand Gesture Recognition from Visual Input." pith.science (2026). https://pith.science/paper/SPKWQDRA
@misc{pith2026250111992,
author = {Pith},
title = {Pith review of: Survey on Hand Gesture Recognition from Visual Input},
year = {2026},
howpublished = {\url{https://pith.science/paper/SPKWQDRA}},
note = {Machine review of arXiv:2501.11992}
}
read the original abstract
Hand gesture recognition has become an important research area, driven by the growing demand for human-computer interaction in fields such as sign language recognition, virtual and augmented reality, and robotics. Despite the rapid growth of the field, there are few surveys that comprehensively cover recent research developments, available solutions, and benchmark datasets. This survey addresses this gap by examining the latest advancements in hand gesture and 3D hand pose recognition from various types of camera input data including RGB images, depth images, and videos from monocular or multiview cameras, examining the differing methodological requirements of each approach. Furthermore, an overview of widely used datasets is provided, detailing their main characteristics and application domains. Finally, open challenges such as achieving robust recognition in real-world environments, handling occlusions, ensuring generalization across diverse users, and addressing computational efficiency for real-time applications are highlighted to guide future research directions. By synthesizing the objectives, methodologies, and applications of recent studies, this survey offers valuable insights into current trends, challenges, and opportunities for future research in human hand gesture recognition.
Figures
Figures from the paper (12 more)
Forward citations
Cited by 1 Pith paper
-
TRIFFID: Autonomous Robotic Aid For Increasing First Responders Efficiency
The paper presents the TRIFFID system architecture for autonomous UAV and UGV disaster reconnaissance, as a design proposal without experimental validation.
Reference graph
Works this paper leans on
-
[1]
Spatial–temporal feature-based end-to-end fourier network for 3d sign language recognition
Sunusi Bala Abdullahi, Kosin Chamnongthai, Veronica Bolon-Canedo, and Brais Cancela. Spatial–temporal feature-based end-to-end fourier network for 3d sign language recognition. Expert Systems with Applications, 248:123258, 2024
2024
-
[2]
Papadopoulos, Vassia Zacharopoulou, George J
Nikolas Adaloglou, Theocharis Chatzis, Ilias Papastratis, Andreas Ster- gioulas, Georgios Th. Papadopoulos, Vassia Zacharopoulou, George J. Xydopoulos, Klimnis Atzakas, Dimitris Papazachariou, and Petros Daras. A comprehensive study on deep learning-based methods for sign language recognition. IEEE Transactions on Multimedia, 24:1750– 1762, 2022
2022
-
[3]
Enhancing hand gesture image recognition by integrating various feature groups
Ismail Taha Ahmed, Wisam Hazim Gwad, Baraa Tareq Hammad, and Entisar Alkayal. Enhancing hand gesture image recognition by integrating various feature groups. Technologies, 13(4), 2025
2025
-
[4]
Rgb arabic alphabets sign language dataset
Muhammad Al-Barham, Adham Alsharkawi, Musa Al-Yaman, Mo- hammad Al-Fetyani, Ashraf Elnagar, Ahmad Abu SaAleek, and Mo- hammad Al-Odat. Rgb arabic alphabets sign language dataset. arXiv preprint arXiv:2301.11932, 2023
arXiv 2023
-
[5]
A structured and methodological review on vision-based hand gesture recognition system
Fahmid Al Farid, Noramiza Hashim, Junaidi Abdullah, Md Roman Bhuiyan, Wan Noor Shahida Mohd Isa, Jia Uddin, Mohammad Ahsanul Haque, and Mohd Nizam Husen. A structured and methodological review on vision-based hand gesture recognition system. Journal of Imaging, 8(6):153, 2022
2022
-
[6]
Innovative hand pose based sign language recognition using hybrid metaheuristic optimization algorithms with deep learning model for hearing impaired persons
Bayan Alabduallah, Reham Al Dayil, Abdulwhab Alkharashi, and Amani A Alneil. Innovative hand pose based sign language recognition using hybrid metaheuristic optimization algorithms with deep learning model for hearing impaired persons. Scientific Reports , 15(1):9320, 2025
2025
-
[7]
Real-time sign language recognition based on yolo algorithm
Melek Alaftekin, Ishak Pacal, and Kenan Cicek. Real-time sign language recognition based on yolo algorithm. Neural Computing and Applications, 36(14):7609–7624, 2024
2024
-
[8]
A survey on sign language literature
Marie Alaghband, Hamid Reza Maghroor, and Ivan Garibay. A survey on sign language literature. Machine Learning with Applications , 14:100504, 2023
2023
Show all 242 references
-
[9]
Snapture—a novel neural architecture for combined static and dynamic hand gesture recognition
Hassan Ali, Doreen Jirak, and Stefan Wermter. Snapture—a novel neural architecture for combined static and dynamic hand gesture recognition. Cognitive Computation, 15(6):2014–2033, 2023
2014
-
[10]
Posetrack: A benchmark for human pose estimation and tracking
Mykhaylo Andriluka, Umar Iqbal, Eldar Insafutdinov, Leonid Pishchulin, Anton Milan, Juergen Gall, and Bernt Schiele. Posetrack: A benchmark for human pose estimation and tracking. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 5167–5176, 2018
2018
-
[11]
The american sign language lexicon video dataset
Vassilis Athitsos, Carol Neidle, Stan Sclaroff, Joan Nash, Alexandra Stefan, Quan Yuan, and Ashwin Thangali. The american sign language lexicon video dataset. In 2008 IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops, pages 1–8. IEEE, 2008
2008
-
[12]
3d hand pose and shape estimation from rgb images for keypoint-based hand gesture recogni- tion
Danilo Avola, Luigi Cinque, Alessio Fagioli, Gian Luca Foresti, Adriano Fragomeni, and Daniele Pannone. 3d hand pose and shape estimation from rgb images for keypoint-based hand gesture recogni- tion. Pattern Recognition, 129:108762, 2022
2022
-
[13]
Local extrema min-max pattern: A novel descriptor for extracting compact and discrete features for hand gesture recognition
Arti Bahuguna, Gopa Bhaumik, and Mahesh Chandra Govil. Local extrema min-max pattern: A novel descriptor for extracting compact and discrete features for hand gesture recognition. Biomedical Signal Processing and Control, 93:106203, 2024
2024
-
[14]
A hybrid approach for static hand gesture recognition with integrated bigru- bilstm and sequential self-attention mechanism
Arti Bahuguna, Mahesh Chandra Govil, and Gopa Bhaumik. A hybrid approach for static hand gesture recognition with integrated bigru- bilstm and sequential self-attention mechanism. Signal, Image and Video Processing, 19(6):1–19, 2025
2025
-
[15]
Multimodal fusion hierarchi- cal self-attention network for dynamic hand gesture recognition
Pranav Balaji and Manas Ranjan Prusty. Multimodal fusion hierarchi- cal self-attention network for dynamic hand gesture recognition. Jour- nal of Visual Communication and Image Representation , 98:104019, 2024
2024
-
[16]
Ultra-range gesture recognition using a web-camera in human–robot interaction
Eran Bamani, Eden Nissinman, Inbar Meir, Lisa Koenigsberg, and Avishai Sintov. Ultra-range gesture recognition using a web-camera in human–robot interaction. Engineering Applications of Artificial Intelligence, 132:108443, 2024
2024
-
[17]
A new 2d static hand gesture colour image dataset for asl gestures
Andre Barczak, Napoleon Reyes, M Abastillas, A Piccio, and Teo Susnjak. A new 2d static hand gesture colour image dataset for asl gestures. Res Lett Inf Math Sci , 15, 01 2011
2011
-
[18]
Improving real- time hand gesture recognition with semantic segmentation
Gibran Benitez-Garcia, Lidia Prudente-Tixteco, Luis Carlos Castro- Madrid, Rocio Toscano-Medina, Jesus Olivares-Mercado, Gabriel Sanchez-Perez, and Luis Javier Garcia Villalba. Improving real- time hand gesture recognition with semantic segmentation. Sensors, 21(2):356, 2021
2021
-
[19]
Video based hand gesture recognition dataset using thermal camera
Simen Birkeland, Lin Julie Fjeldvik, Nadia Noori, Sreenivasa Reddy Yeduri, and Linga Reddy Cenkeramaddi. Video based hand gesture recognition dataset using thermal camera. Data in Brief , 54:110299, 2024
2024
-
[20]
Weakly- supervised 3d hand pose estimation from monocular rgb images
Yujun Cai, Liuhao Ge, Jianfei Cai, and Junsong Yuan. Weakly- supervised 3d hand pose estimation from monocular rgb images. In Proceedings of the European conference on computer vision (ECCV) , pages 666–682, 2018
2018
-
[21]
Exploiting spatial-temporal re- lationships for 3d pose estimation via graph convolutional networks
Yujun Cai, Liuhao Ge, Jun Liu, Jianfei Cai, Tat-Jen Cham, Junsong Yuan, and Nadia Magnenat Thalmann. Exploiting spatial-temporal re- lationships for 3d pose estimation via graph convolutional networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision ...
2019
-
[22]
Neural sign language translation
Necati Cihan Camgoz, Simon Hadfield, Oscar Koller, Hermann Ney, and Richard Bowden. Neural sign language translation. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 7784–7793, 2018
2018
-
[23]
Sign language recognition for assisting the deaf in hospitals
Necati Cihan Camgöz, Ahmet Alp Kındıro ˘glu, and Lale Akarun. Sign language recognition for assisting the deaf in hospitals. In Human Behavior Understanding: 7th International Workshop, HBU 2016, Amsterdam, The Netherlands, October 16, 2016, Proceedings 7 , pages 89–101. Sprin...
2016
-
[24]
Emerging properties in self-supervised vision transformers, 2021
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers, 2021
2021
-
[25]
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6299– 6308, 2017
2017
-
[26]
Hand-gesture recogni- tion based on emg and event-based camera sensor fusion: A benchmark in neuromorphic computing
Enea Ceolini, Charlotte Frenkel, Sumit Bam Shrestha, Gemma Taverni, Lyes Khacef, Melika Payvand, and Elisa Donati. Hand-gesture recogni- tion based on emg and event-based camera sensor fusion: A benchmark in neuromorphic computing. Frontiers in neuroscience, 14:637, 2020
2020
-
[27]
Dexycb: A benchmark for capturing hand grasping of objects
Yu-Wei Chao, Wei Yang, Yu Xiang, Pavlo Molchanov, Ankur Handa, Jonathan Tremblay, Yashraj S Narang, Karl Van Wyk, Umar Iqbal, Stan Birchfield, et al. Dexycb: A benchmark for capturing hand grasping of objects. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pa...
2021
-
[28]
Convolutional neural network hand gesture recognition for american sign language
Shruti Chavan, Xinrui Yu, and Jafar Saniie. Convolutional neural network hand gesture recognition for american sign language. In 2021 IEEE International Conference on Electro Information Technology (EIT), pages 188–192, 2021
2021
-
[29]
Multi-scale attention 3d convolutional network for multimodal gesture recognition
Huizhou Chen, Yunan Li, Huijuan Fang, Wentian Xin, Zixiang Lu, and Qiguang Miao. Multi-scale attention 3d convolutional network for multimodal gesture recognition. Sensors, 22(6):2405–2405, Mar 2022
2022
-
[30]
Lisa: Learning implicit shape and appearance of hands, 2022
Enric Corona, Tomas Hodan, Minh V o, Francesc Moreno-Noguer, Chris Sweeney, Richard Newcombe, and Lingni Ma. Lisa: Learning implicit shape and appearance of hands, 2022
2022
-
[31]
Dabwan, Mukti E
Basel A. Dabwan, Mukti E. Jadhav, Mohammed Al Yami, Eman A. Hassan, Soad M. Almula, and Yahya A. Ali. Classifying hand gestures for people with disabilities utilizing the mobilenetv2 model. In 2024 1st International Conference on Innovative Sustainable Technologies for Energy,...
2024
-
[32]
Histograms of oriented gradients for human detection
Navneet Dalal and Bill Triggs. Histograms of oriented gradients for human detection. In 2005 IEEE computer society conference on computer vision and pattern recognition (CVPR’05) , volume 1, pages 886–893. Ieee, 2005
2005
-
[33]
Scaling egocentric vision: The epic-kitchens dataset
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al. Scaling egocentric vision: The epic-kitchens dataset. In Proceedings of the European conference on comput...
2018
-
[34]
Cnn based static hand gesture recognition using rgb-d data
N.C Dayananda Kumar, K.V Suresh, and R Dinesh. Cnn based static hand gesture recognition using rgb-d data. In 2022 2nd International Conference on Artificial Intelligence and Signal Processing (AISP) , pages 1–6, 2022
2022
-
[35]
Spatial-temporal graph convolutional networks for sign language recog- nition
Cleison Correia de Amorim, David Macêdo, and Cleber Zanchettin. Spatial-temporal graph convolutional networks for sign language recog- nition. In International Conference on Artificial Neural Networks , pages 646–657. Springer, 2019
2019
-
[36]
Automatic translation of sign language with multi-stream 3d cnn and generation of artificial depth maps
Giulia Zanon de Castro, Rúbia Reis Guerra, and Frederico Gadelha Guimarães. Automatic translation of sign language with multi-stream 3d cnn and generation of artificial depth maps. Expert Systems with Applications, 215:119394, 2023
2023
-
[37]
Isolated sign recognition from rgb video using pose flow and self-attention
Mathieu De Coster, Mieke Van Herreweghe, and Joni Dambre. Isolated sign recognition from rgb video using pose flow and self-attention. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pages 3436–3445, 2021
2021
-
[38]
Skeleton-based dynamic hand gesture recognition
Quentin De Smedt, Hazem Wannous, and Jean-Philippe Vandeborre. Skeleton-based dynamic hand gesture recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition workshops, pages 1–9, 2016
2016
-
[39]
Heterogeneous hand gesture recognition using 3d dynamic skeletal data
Quentin De Smedt, Hazem Wannous, and Jean-Philippe Vandeborre. Heterogeneous hand gesture recognition using 3d dynamic skeletal data. Computer Vision and Image Understanding , 181:60–72, 2019
2019
-
[40]
Shrec’17 track: 3d hand gesture recognition using a depth and skeletal dataset
Quentin De Smedt, Hazem Wannous, Jean-Philippe Vandeborre, Joris Guerry, Bertrand Le Saux, and David Filliat. Shrec’17 track: 3d hand gesture recognition using a depth and skeletal dataset. In 3DOR-10th Eurographics Workshop on 3D Object Retrieval , pages 1–6, 2017
2017
-
[41]
Tms-net: A multi-feature multi-stream multi-level infor- mation sharing network for skeleton-based sign language recognition
Zhiwen Deng, Yuquan Leng, Junkang Chen, Xiang Yu, Yang Zhang, and Qing Gao. Tms-net: A multi-feature multi-stream multi-level infor- mation sharing network for skeleton-based sign language recognition. Neurocomputing, 572:127194, 2024
2024
-
[42]
Deep learning for hand gesture recognition on skeletal data
Guillaume Devineau, Fabien Moutarde, Wang Xi, and Jie Yang. Deep learning for hand gesture recognition on skeletal data. In 2018 13th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2018) , pages 106–113. IEEE, 2018
2018
-
[43]
Modeling image variability in appearance-based gesture recognition
Philippe Dreuw, Thomas Deselaers, Daniel Keysers, and Hermann Ney. Modeling image variability in appearance-based gesture recognition. In ECCV workshop on statistical methods in multi-image and video processing, pages 7–18, 2006
2006
-
[44]
Beyond granularity: Enhancing continuous sign language recognition with granularity-aware feature fusion and attention optimization
Yao Du, Taiying Peng, and Xiaohui Hu. Beyond granularity: Enhancing continuous sign language recognition with granularity-aware feature fusion and attention optimization. Applied Sciences, 14(19), 2024
2024
-
[45]
Enes Duran, Muhammed Kocabas, Vasileios Choutas, Zicong Fan, and Michael J. Black. Hmp: Hand motion priors for pose and shape estimation from video. In 2024 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) , pages 6341–6351, 2024
2024
-
[46]
Lavrf: Sign language recognition via lightweight attentive vgg16 with random forest
Edmond Li Ren Ewe, Chin Poo Lee, Kian Ming Lim, Lee Chung Kwek, and Ali Alqahtani. Lavrf: Sign language recognition via lightweight attentive vgg16 with random forest. PLOS ONE, 19(4):1– 22, 04 2024
2024
-
[47]
Multi-task and multi-modal learning for rgb dynamic gesture recognition
Dinghao Fan, Hengjie Lu, Shugong Xu, and Shan Cao. Multi-task and multi-modal learning for rgb dynamic gesture recognition. IEEE Sensors Journal, 21(23):27026–27036, 2021
2021
-
[48]
Mdsi: Pluggable multi-strategy decoupling with semantic inte- gration for rgb-d gesture recognition
Fengyi Fang, Zihan Liao, Zhehan Kan, Guijin Wang, and Wenming Yang. Mdsi: Pluggable multi-strategy decoupling with semantic inte- gration for rgb-d gesture recognition. Pattern Recognition, 166:111653, 2025
2025
-
[49]
Yolov8-g2f: A portable gesture recognition optimization algorithm
Zhao Feng, Junjian Huang, Wei Zhang, Shiping Wen, Yangpeng Liu, and Tingwen Huang. Yolov8-g2f: A portable gesture recognition optimization algorithm. Neural Networks, 188:107469, 2025
2025
-
[50]
Hand gesture recognition on edge devices: Sensor technologies, algorithms, and processing hardware
Elfi Fertl, Encarnación Castillo, Georg Stettinger, Manuel P Cuéllar, and Diego P Morales. Hand gesture recognition on edge devices: Sensor technologies, algorithms, and processing hardware. Sensors, 25(6):1687, 2025
2025
-
[51]
An efficient rgb-d hand gesture detection framework for dexterous robot hand-arm teleoperation system
Qing Gao, Zhaojie Ju, Yongquan Chen, Qiwen Wang, and Chuliang Chi. An efficient rgb-d hand gesture detection framework for dexterous robot hand-arm teleoperation system. IEEE Transactions on Human- Machine Systems, 53(1):13–23, 2023
2023
-
[52]
First-person hand action benchmark with rgb-d videos and 3d hand pose annotations
Guillermo Garcia-Hernando, Shanxin Yuan, Seungryul Baek, and Tae- Kyun Kim. First-person hand action benchmark with rgb-d videos and 3d hand pose annotations. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 409–419, 2018
2018
-
[53]
Convmixformer-a resource-efficient convolution mixer for transformer-based dynamic hand gesture recognition
Mallika Garg, Debashis Ghosh, and Pyari Mohan Pradhan. Convmixformer-a resource-efficient convolution mixer for transformer-based dynamic hand gesture recognition. arXiv preprint arXiv:2411.07118, 2024
2024 arXiv
-
[54]
Gestformer: Multiscale wavelet pooling transformer network for dynamic hand gesture recognition
Mallika Garg, Debashis Ghosh, and Pyari Mohan Pradhan. Gestformer: Multiscale wavelet pooling transformer network for dynamic hand gesture recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 2473–2483, 2024
2024
-
[55]
Letizia Gionfrida, Wan M. R. Rusli, Angela E. Kedgley, and Anil A. Bharath. A 3dcnn-lstm multi-class temporal segmentation for hand gesture recognition. Electronics, 11(15), 2022
2022
-
[56]
The role of gesture in communication and thinking
Susan Goldin-Meadow. The role of gesture in communication and thinking. Trends in Cognitive Sciences , 3(11):419–429, 1999
1999
-
[57]
Towards end-to-end speech recog- nition with recurrent neural networks
Alex Graves and Navdeep Jaitly. Towards end-to-end speech recog- nition with recurrent neural networks. In International conference on machine learning, pages 1764–1772. PMLR, 2014
2014
-
[58]
Mska: Multi-stream keypoint attention network for sign language recognition and translation
Mo Guan, Yan Wang, Guangkun Ma, Jiarui Liu, and Mingzu Sun. Mska: Multi-stream keypoint attention network for sign language recognition and translation. Pattern Recognition, 165:111602, 2025
2025
-
[59]
Multi-view isolated sign language recognition based on cross-view and multi-level transformer
Zhong Guan, Yongli Hu, Huajie Jiang, Yanfeng Sun, and Baocai Yin. Multi-view isolated sign language recognition based on cross-view and multi-level transformer. Multimedia Systems, 31(3):1–15, 2025
2025
-
[60]
A hierarchical attention gcn network for body-hand gesture recognition
Xiaofeng Guo, Qing Zhu, Yaonan Wang, and Yang Mo. A hierarchical attention gcn network for body-hand gesture recognition. In 2023 China Automation Congress (CAC) , pages 799–804, 2023
2023
-
[61]
A rapid adaptation approach for dynamic air-writing recognition using wearable wristbands with self- supervised contrastive learning
Yunjian Guo, Kunpeng Li, Wei Yue, Nam-Young Kim, Yang Li, Guozhen Shen, and Jong-Chul Lee. A rapid adaptation approach for dynamic air-writing recognition using wearable wristbands with self- supervised contrastive learning. Nano-Micro Letters, 17(1):1–15, 2025
2025
-
[62]
Honnotate: A method for 3d annotation of hand and object poses
Shreyas Hampali, Mahdi Rad, Markus Oberweger, and Vincent Lepetit. Honnotate: A method for 3d annotation of hand and object poses. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3196–3206, 2020
2020
-
[63]
A systematic review of hand gesture recognition: An update from 2018 to 2024
Abdirahman Osman Hashi, Siti Zaiton Mohd Hashim, and Azurah Bte Asamah. A systematic review of hand gesture recognition: An update from 2018 to 2024. IEEE Access, 2024
2018
-
[64]
Developing a real-time hand exoskeleton system that controlled by a hand gesture recognition system via wireless sensors
Yunus Hazar and Ömer Faruk Ertu ˘grul. Developing a real-time hand exoskeleton system that controlled by a hand gesture recognition system via wireless sensors. Biomedical Signal Processing and Control, 99:106886, 2025
2025
-
[65]
Distilling the knowl- edge in a neural network, 2015
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowl- edge in a neural network, 2015
2015
-
[66]
Efficient multimodal fusion for hand pose estimation with hourglass network
Dinh-Cuong Hoang, Phan Xuan Tan, Duc-Long Pham, Hai-Nam Pham, Son-Anh Bui, Chi-Minh Nguyen, An-Binh Phi, Khanh-Duong Tran, Viet-Anh Trinh, van-Duc Tran, Duc-Thanh Tran, van-Hiep Duong, Khanh-Toan Phan, van-Thiep Nguyen, van-Duc Vu, and Thu-Uyen Nguyen. Efficient multimodal fus...
2024
-
[67]
Bdsl36: A dataset for bangladeshi sign letters recognition
Oishee Bintey Hoque, Mohammad Imrul Jubair, Al-Farabi Akash, and Saiful Islam. Bdsl36: A dataset for bangladeshi sign letters recognition. In Proceedings of the Asian Conference on Computer Vision , 2020
2020
-
[68]
Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam
Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. Mobilenets: Efficient convolutional neural networks for mobile vision applications, 2017
2017
-
[69]
Stfe-net: A spatial- temporal feature extraction network for continuous sign language translation
Jiwei Hu, Yunfei Liu, Kin-Man Lam, and Ping Lou. Stfe-net: A spatial- temporal feature extraction network for continuous sign language translation. IEEE Access, 11:46204–46217, 2023
2023
-
[70]
Hand gesture recognition algorithm using svm and hog model for control of robotic system
Phat Nguyen Huu and Tan Phung Ngoc. Hand gesture recognition algorithm using svm and hog model for control of robotic system. Journal of Robotics , 2021(1):3986497, 2021
2021
-
[71]
Iandola, Song Han, Matthew W
Forrest N. Iandola, Song Han, Matthew W. Moskewicz, Khalid Ashraf, William J. Dally, and Kurt Keutzer. Squeezenet: Alexnet-level accuracy with 50x fewer parameters and <0.5mb model size, 2016
2016
-
[72]
Tools and approaches for topic detection from twitter streams: survey
Rania Ibrahim, Ahmed Elbagoury, Mohamed S Kamel, and Fakhri Karray. Tools and approaches for topic detection from twitter streams: survey. Knowledge and Information Systems , 54:511–539, 2018
2018
-
[73]
Dataset of pakistan sign language and automatic recognition of hand configuration of urdu alphabet through machine learning
Ali Imran, Abdul Razzaq, Irfan Ahmad Baig, Aamir Hussain, Sharaiz Shahid, and Tausif-ur Rehman. Dataset of pakistan sign language and automatic recognition of hand configuration of urdu alphabet through machine learning. Data in Brief , 36:107021, 2021
2021
-
[74]
Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchis- escu. Human3. 6m: Large scale datasets and predictive methods for 3d human sensing in natural environments. IEEE transactions on pattern analysis and machine intelligence , 36(7):1325–1339, 2013
2013
-
[75]
Hand pose estimation via latent 2.5d heatmap regression
Umar Iqbal, Pavlo Molchanov, Thomas Breuel Juergen Gall, and Jan Kautz. Hand pose estimation via latent 2.5d heatmap regression. In Proceedings of the European Conference on Computer Vision (ECCV) , September 2018
2018
-
[76]
Quantization and training of neural networks for efficient integer- arithmetic-only inference, 2017
Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, and Dmitry Kalenichenko. Quantization and training of neural networks for efficient integer- arithmetic-only inference, 2017
2017
-
[77]
Latent dirichlet allocation (lda) and topic modeling: models, applications, a survey
Hamed Jelodar, Yongli Wang, Chi Yuan, Xia Feng, Xiahui Jiang, Yanchao Li, and Liang Zhao. Latent dirichlet allocation (lda) and topic modeling: models, applications, a survey. Multimedia tools and applications, 78:15169–15211, 2019
2019
-
[78]
Stm: Spatiotemporal and motion encoding for action recogni- tion
Boyuan Jiang, MengMeng Wang, Weihao Gan, Wei Wu, and Junjie Yan. Stm: Spatiotemporal and motion encoding for action recogni- tion. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , October 2019
2019
-
[79]
Skeleton aware multi-modal sign language recognition
Songyao Jiang, Bin Sun, Lichen Wang, Yue Bai, Kunpeng Li, and Yun Fu. Skeleton aware multi-modal sign language recognition. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pages 3408–3418, 2021
2021
-
[80]
Handformer: Hand pose reconstructing from a single rgb image
Zixun Jiao, Xihan Wang, Jingcao Li, Rongxin Gao, Miao He, Jiao Liang, Zhaoqiang Xia, and Quanli Gao. Handformer: Hand pose reconstructing from a single rgb image. Pattern Recognition Letters , 183:155–164, 2024
2024
-
[81]
Gesture recognition matching based on dynamic skeleton
Wang Jingyao, Yu Naigong, and Essaf Firdaous. Gesture recognition matching based on dynamic skeleton. In 2021 33rd Chinese Control and Decision Conference (CCDC) , pages 1680–1685, 2021
2021
-
[82]
Learning effective human pose estimation from inaccurate annotation
Sam Johnson and Mark Everingham. Learning effective human pose estimation from inaccurate annotation. In CVPR 2011 , pages 1465–
2011
-
[83]
Panoptic studio: A massively multiview system for social motion capture
Hanbyul Joo, Hao Liu, Lei Tan, Lin Gui, Bart Nabbe, Iain Matthews, Takeo Kanade, Shohei Nobuhara, and Yaser Sheikh. Panoptic studio: A massively multiview system for social motion capture. In Proceedings of the IEEE international conference on computer vision , pages 3334– 3342, 2015
2015
-
[84]
Total capture: A 3d deformation model for tracking faces, hands, and bodies
Hanbyul Joo, Tomas Simon, and Yaser Sheikh. Total capture: A 3d deformation model for tracking faces, hands, and bodies. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 8320–8329, 2018
2018
-
[85]
Ms-asl: A large-scale data set and benchmark for understanding american sign language
Hamid Reza Vaezi Joze and Oscar Koller. Ms-asl: A large-scale data set and benchmark for understanding american sign language. arXiv preprint arXiv:1812.01053, 2018
2018 arXiv
-
[86]
Dynamic japanese sign language recognition throw hand pose es- timation using effective feature extraction and classification approach
Manato Kakizaki, Abu Saleh Musa Miah, Koki Hirooka, and Jungpil Shin. Dynamic japanese sign language recognition throw hand pose es- timation using effective feature extraction and classification approach. Sensors, 24(3), 2024
2024
-
[87]
Temporal signed gestures segmentation in an image sequence using deep reinforcement learning
Dawid Kalandyk and Tomasz Kapu ´sci´nski. Temporal signed gestures segmentation in an image sequence using deep reinforcement learning. Engineering Applications of Artificial Intelligence , 131:107879, 2024
2024
-
[88]
Learning 3d human dynamics from video
Angjoo Kanazawa, Jason Y Zhang, Panna Felsen, and Jitendra Malik. Learning 3d human dynamics from video. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 5614–5623, 2019
2019
-
[89]
Sign language apprehension using convolution neural networks
Meghana Pai Kane, Sherwin Fernandes, Ricky Fonseca, Shika Desai, Akhil Shetye, and Ananya Sharma. Sign language apprehension using convolution neural networks. In 2022 13th International Conference on Computing Communication and Networking Technologies (ICCCNT) , pages 1–7, 2022
2022
-
[90]
Real-time sign language fingerspelling recognition using convolutional neural networks from depth map
Byeongkeun Kang, Subarna Tripathi, and Truong Q Nguyen. Real-time sign language fingerspelling recognition using convolutional neural networks from depth map. In 2015 3rd IAPR Asian Conference on Pattern Recognition (ACPR), pages 136–140. IEEE, 2015
2015
-
[91]
Black, Krikamol Muandet, and Siyu Tang
Korrawe Karunratanakul, Jinlong Yang, Yan Zhang, Michael J. Black, Krikamol Muandet, and Siyu Tang. Grasping field: Learning implicit representations for human grasps. In 2020 International Conference on 3D Vision (3DV) , pages 333–344, 2020
2020
-
[92]
Muhammed Kocabas, Nikos Athanasiou, and Michael J. Black. Vibe: Video inference for human body pose and shape estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020
2020
-
[93]
Weakly supervised learning with multi-stream cnn-lstm- hmms to discover sequential parallelism in sign language videos
Oscar Koller, Necati Cihan Camgoz, Hermann Ney, and Richard Bowden. Weakly supervised learning with multi-stream cnn-lstm- hmms to discover sequential parallelism in sign language videos. IEEE transactions on pattern analysis and machine intelligence, 42(9):2306– 2320, 2019
2019
-
[94]
Deep sign: Enabling robust statistical continuous sign language recog- nition via hybrid cnn-hmms
Oscar Koller, Sepehr Zargaran, Hermann Ney, and Richard Bowden. Deep sign: Enabling robust statistical continuous sign language recog- nition via hybrid cnn-hmms. International Journal of Computer Vision, 126:1311–1325, 2018
2018
-
[95]
Black, and Kostas Daniilidis
Nikos Kolotouros, Georgios Pavlakos, Michael J. Black, and Kostas Daniilidis. Learning to reconstruct 3d human pose and shape via model- fitting in the loop. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , October 2019
2019
-
[96]
Convo- lutional mesh regression for single-image human shape reconstruction
Nikos Kolotouros, Georgios Pavlakos, and Kostas Daniilidis. Convo- lutional mesh regression for single-image human shape reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2019
2019
-
[97]
Real-time hand gesture detection and classification using convolutional neural networks
Okan Köpüklü, Ahmet Gunduz, Neslihan Kose, and Gerhard Rigoll. Real-time hand gesture detection and classification using convolutional neural networks. In 2019 14th IEEE international conference on automatic face & gesture recognition (FG 2019) , pages 1–8. IEEE, 2019
2019
-
[98]
Motion fused frames: Data level fusion strategy for hand gesture recognition
Okan Kopuklu, Neslihan Kose, and Gerhard Rigoll. Motion fused frames: Data level fusion strategy for hand gesture recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, June 2018
2018
-
[99]
Kothadiya, Chintan M
Deep R. Kothadiya, Chintan M. Bhatt, Hena Kharwa, and Felix Albu. Hybrid inceptionnet based enhanced architecture for isolated sign language recognition. IEEE Access, 12:90889–90899, 2024
2024
-
[100]
Hmdb: a large video database for human motion recognition
Hildegard Kuehne, Hueihan Jhuang, Estíbaliz Garrote, Tomaso Poggio, and Thomas Serre. Hmdb: a large video database for human motion recognition. In 2011 International conference on computer vision , pages 2556–2563. IEEE, 2011
2011
-
[101]
Isolated video-based sign language recognition using a hybrid cnn-lstm framework based on attention mechanism
Diksha Kumari and Radhey Shyam Anand. Isolated video-based sign language recognition using a hybrid cnn-lstm framework based on attention mechanism. Electronics, 13(7), 2024
2024
-
[102]
Recognition of jsl fingerspelling using deep convolutional neural networks
Bogdan Kwolek, Wojciech Baczynski, and Shinji Sako. Recognition of jsl fingerspelling using deep convolutional neural networks. Neuro- computing, 456:586–598, 2021
2021
-
[103]
H2o: Two hands manipulating objects for first person interaction recognition
Taein Kwon, Bugra Tekin, Jan St ¨"uhmer, Federica Bogo, and Marc Pollefeys. H2o: Two hands manipulating objects for first person interaction recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 10138–10148, 2021
2021
-
[104]
Yanushkevich
Kenneth Lai and Svetlana N. Yanushkevich. Cnn+rnn depth and skeleton based dynamic hand gesture recognition. In 2018 24th International Conference on Pattern Recognition (ICPR) , pages 3451– 3456, 2018
2018
-
[105]
Sst-gcn: Structure aware spatial-temporal gcn for 3d hand pose estimation
Viet-Thanh Le, Thanh-Hai Tran, Van-Nam Hoang, Van-Hung Le, Thi- Lan Le, and Hai Vu. Sst-gcn: Structure aware spatial-temporal gcn for 3d hand pose estimation. In 2021 13th International Conference on Knowledge and Systems Engineering (KSE) , pages 1–6, 2021
2021
-
[106]
Video hand gestures recognition using depth camera and lightweight cnn
David González León, Jade Gröli, Sreenivasa Reddy Yeduri, Daniel Rossier, Romuald Mosqueron, Om Jee Pandey, and Linga Reddy Cenkeramaddi. Video hand gestures recognition using depth camera and lightweight cnn. IEEE Sensors Journal , 22(14):14610–14619, 2022
2022
-
[107]
Word-level deep sign language recognition from video: A new large- scale dataset and methods comparison
Dongxu Li, Cristian Rodriguez Opazo, Xin Yu, and Hongdong Li. Word-level deep sign language recognition from video: A new large- scale dataset and methods comparison. In 2020 IEEE Winter Confer- ence on Applications of Computer Vision (WACV) , pages 1448–1458, 2020
2020
-
[108]
Efficient hand gesture recognition using multi-task multi-modal learning and self-distillation
Jie-Ying Li, Herman Prawiro, Chia-Chen Chiang, Hsin-Yu Chang, Tse- Yu Pan, Chih-Tsun Huang, and Min-Chun Hu. Efficient hand gesture recognition using multi-task multi-modal learning and self-distillation. In Proceedings of the 5th ACM International Conference on Multimedia in ...
2024
-
[109]
First-person hand action recognition using multimodal data
Rui Li, Hongyu Wang, Zhenyu Liu, Na Cheng, and Hongye Xie. First-person hand action recognition using multimodal data. IEEE Transactions on Cognitive and Developmental Systems , 14(4):1449– 1464, 2022
2022
-
[110]
Action recognition based on a bag of 3d points
Wanqing Li, Zhengyou Zhang, and Zicheng Liu. Action recognition based on a bag of 3d points. In 2010 IEEE computer society conference on computer vision and pattern recognition-workshops , pages 9–14. IEEE, 2010
2010
-
[111]
Eva: Key values eclosion with space anchor used in hand pose estimation and shape reconstruction
Xuefeng Li and Xiangbo Lin. Eva: Key values eclosion with space anchor used in hand pose estimation and shape reconstruction. Infor- mation Sciences, 706:122003, 2025
2025
-
[112]
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Per- ona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedin...
2014
-
[113]
Han: An efficient hierarchical self-attention network for skeleton-based gesture recognition
Jianbo Liu, Ying Wang, Shiming Xiang, and Chunhong Pan. Han: An efficient hierarchical self-attention network for skeleton-based gesture recognition. Pattern Recognition, 162:111343, 2025
2025
-
[114]
Dynamic gesture recognition based on cnn-lstm-attention
Jinwei Liu, Baoguo Wei, Mingzhi Cai, and Yong Xu. Dynamic gesture recognition based on cnn-lstm-attention. In 2021 IEEE International Conference on Signal Processing, Communications and Computing (ICSPCC), pages 1–6. IEEE, 2021
2021
-
[115]
Keypoint fusion for rgb-d based 3d hand pose estimation
Xingyu Liu, Pengfei Ren, Yuanyuan Gao, Jingyu Wang, Haifeng Sun, Qi Qi, Zirui Zhuang, and Jianxin Liao. Keypoint fusion for rgb-d based 3d hand pose estimation. Proceedings of the AAAI Conference on Artificial Intelligence , 38(4):3756–3764, Mar 2024
2024
-
[116]
Gtignet: Global topology interaction graphormer network for 3d hand pose estimation
Yanjun Liu, Wanshu Fan, Cong Wang, Shixi Wen, Xin Yang, Qiang Zhang, Xiaopeng Wei, and Dongsheng Zhou. Gtignet: Global topology interaction graphormer network for 3d hand pose estimation. Neural Networks, 185:107221, 2025
2025
-
[117]
Spmhand: Segmentation- guided progressive multi-path 3d hand pose and shape estimation
Haofan Lu, Shuiping Gou, and Ruimin Li. Spmhand: Segmentation- guided progressive multi-path 3d hand pose and shape estimation. IEEE Transactions on Multimedia , 26:6822–6833, 2024
2024
-
[118]
A unified approach to interpreting model predictions, 2017
Scott Lundberg and Su-In Lee. A unified approach to interpreting model predictions, 2017
2017
-
[119]
Probabilistic non-negative matrix factor- ization and its robust extensions for topic modeling
Minnan Luo, Feiping Nie, Xiaojun Chang, Yi Yang, Alexander Haupt- mann, and Qinghua Zheng. Probabilistic non-negative matrix factor- ization and its robust extensions for topic modeling. Proceedings of the AAAI Conference on Artificial Intelligence , 31(1), Feb. 2017
2017
-
[120]
Two-stream mixed convolu- tional neural network for american sign language recognition
Ying Ma, Tianpei Xu, and Kangchul Kim. Two-stream mixed convolu- tional neural network for american sign language recognition. Sensors, 22(16), 2022
2022
-
[121]
Deephps: End- to-end estimation of 3d hand pose and shape by learning from synthetic depth
Jameel Malik, Ahmed Elhayek, Fabrizio Nunnari, Kiran Varanasi, Kiarash Tamaddon, Alexis Heloir, and Didier Stricker. Deephps: End- to-end estimation of 3d hand pose and shape by learning from synthetic depth. In 2018 International Conference on 3D Vision (3DV) , pages 110–119, 2018
2018
-
[122]
Hand gestures for the human-car interaction: The briareo dataset
Fabio Manganaro, Stefano Pini, Guido Borghi, Roberto Vezzani, and Rita Cucchiara. Hand gestures for the human-car interaction: The briareo dataset. In Image Analysis and Processing–ICIAP 2019: 20th International Conference, Trento, Italy, September 9–13, 2019, Proceedings, Par...
2019
-
[123]
Hand gesture recognition with leap motion and kinect devices
Giulio Marin, Fabio Dominio, and Pietro Zanuttigh. Hand gesture recognition with leap motion and kinect devices. In 2014 IEEE International conference on image processing (ICIP) , pages 1565–
2014
-
[124]
The jester dataset: A large-scale video dataset of human gestures
Joanna Materzynska, Guillaume Berger, Ingo Bax, and Roland Memi- sevic. The jester dataset: A large-scale video dataset of human gestures. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops , Oct 2019
2019
-
[125]
Ouhands database for hand detection and pose recognition
Matti Matilainen, Pekka Sangi, Jukka Holappa, and Olli Silvén. Ouhands database for hand detection and pose recognition. In 2016 Sixth international conference on image processing theory, tools and applications (IPTA), pages 1–5. IEEE, 2016
2016
-
[126]
A survey on bias and fairness in machine learning
Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. A survey on bias and fairness in machine learning. ACM Comput. Surv., 54(6), July 2021
2021
-
[127]
Monocular 3d human pose estimation in the wild using improved cnn supervision
Dushyant Mehta, Helge Rhodin, Dan Casas, Pascal Fua, Oleksandr Sotnychenko, Weipeng Xu, and Christian Theobalt. Monocular 3d human pose estimation in the wild using improved cnn supervision. In 2017 international conference on 3D vision (3DV) , pages 506–516. IEEE, 2017
2017
-
[128]
Using motion history images with 3d convolutional networks in isolated sign language recognition
Ozge Mercanoglu Sincan and Hacer Yalim Keles. Using motion history images with 3d convolutional networks in isolated sign language recognition. IEEE Access, 10:18608–18618, 2022
2022
-
[129]
Al Mehedi Hasan, Satoshi Nishimura, and Jungpil Shin
Abu Saleh Musa Miah, Md. Al Mehedi Hasan, Satoshi Nishimura, and Jungpil Shin. Sign language recognition using graph and general deep neural network based on large scale dataset. IEEE Access, 12:34553– 34569, 2024
2024
-
[130]
Al Mehedi Hasan, Yoichi Tomioka, and Jungpil Shin
Abu Saleh Musa Miah, Md. Al Mehedi Hasan, Yoichi Tomioka, and Jungpil Shin. Hand gesture recognition for multi-culture sign language using graph and general deep learning network. IEEE Open Journal of the Computer Society , 5:144–155, 2024
2024
-
[131]
Multiple-hand 2d pose estimation from a monocular rgb image
Purnendu Mishra and Kishor Sarawadekar. Multiple-hand 2d pose estimation from a monocular rgb image. IEEE Access , 12:40722– 40735, 2024
2024
-
[132]
A review of the hand gesture recognition system: Current progress and future directions
Noraini Mohamed, Mumtaz Begum Mustafa, and Nazean Jomhari. A review of the hand gesture recognition system: Current progress and future directions. IEEE Access, 9:157422–157436, 2021
2021
-
[133]
Online detection and classification of dynamic hand gestures with recurrent 3d convolutional neural network
Pavlo Molchanov, Xiaodong Yang, Shalini Gupta, Kihwan Kim, Stephen Tyree, and Jan Kautz. Online detection and classification of dynamic hand gestures with recurrent 3d convolutional neural network. In Proceedings of the IEEE conference on computer vision and pattern recognitio...
2016
-
[134]
Interhand2.6m: A dataset and baseline for 3d interacting hand pose estimation from a single rgb image
Gyeongsik Moon, Shoou-I Yu, He Wen, Takaaki Shiratori, and Ky- oung Mu Lee. Interhand2.6m: A dataset and baseline for 3d interacting hand pose estimation from a single rgb image. In Computer Vision– ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedi...
2020
-
[135]
In my perspective, in my hands: Accurate egocentric 2d hand pose and action recognition
Wiktor Mucha and Martin Kampel. In my perspective, in my hands: Accurate egocentric 2d hand pose and action recognition. In 2024 IEEE 18th International Conference on Automatic Face and Gesture Recognition (FG), pages 1–9, 2024
2024
-
[136]
Ganer- ated hands for real-time 3d hand tracking from monocular rgb
Franziska Mueller, Florian Bernard, Oleksandr Sotnychenko, Dushyant Mehta, Srinath Sridhar, Dan Casas, and Christian Theobalt. Ganer- ated hands for real-time 3d hand tracking from monocular rgb. In Proceedings of the IEEE conference on computer vision and pattern recognition,...
2018
-
[137]
Real-time hand tracking under occlusion from an egocentric rgb-d sensor
Franziska Mueller, Dushyant Mehta, Oleksandr Sotnychenko, Srinath Sridhar, Dan Casas, and Christian Theobalt. Real-time hand tracking under occlusion from an egocentric rgb-d sensor. In Proceedings of the IEEE international conference on computer vision , pages 1154–1163, 2017
2017
-
[138]
Real-time hand gesture recognition based on deep learning yolov3 model
Abdullah Mujahid, Mazhar Javed Awan, Awais Yasin, Mazin Abed Mohammed, Robertas Damaševi ˇcius, Rytis Maskeli ¯unas, and Kar- rar Hameed Abdulkareem. Real-time hand gesture recognition based on deep learning yolov3 model. Applied Sciences, 11(9), 2021
2021
-
[139]
Multi-modal domain adaptation for fine-grained action recognition
Jonathan Munro and Dima Damen. Multi-modal domain adaptation for fine-grained action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020
2020
-
[140]
Sabarimalai Manikandan, Ajit Jha, Jing Zhou, and Linga Reddy Cenkeramaddi
Sai Sree Nathala, Rakesh Reddy Yakkati, Daniel Skomedal Breland, Sreenivasa Reddy Yeduri, M. Sabarimalai Manikandan, Ajit Jha, Jing Zhou, and Linga Reddy Cenkeramaddi. A deep cnn-based hand gestures recognition using high-resolution thermal imaging. In 2024 IEEE 19th Conferenc...
2024
-
[141]
Signgraph: An efficient and accurate pose-based graph convolution approach toward sign language recognition
Neelma Naz, Hasan Sajid, Sara Ali, Osman Hasan, and Muham- mad Khurram Ehsan. Signgraph: An efficient and accurate pose-based graph convolution approach toward sign language recognition. IEEE Access, 11:19135–19147, 2023
2023
-
[142]
A decade of progress in human motion recognition: A comprehensive survey from 2010 to
Donghyeon Noh, Hojin Yoon, and Donghun Lee. A decade of progress in human motion recognition: A comprehensive survey from 2010 to
2010
-
[143]
Hand pose recognition using parallel multi stream cnn
Iram Noreen, Muhammad Hamid, Uzma Akram, Saadia Malik, and Muhammad Saleem. Hand pose recognition using parallel multi stream cnn. Sensors, 21(24):8469–8469, Dec 2021
2021
-
[144]
Núñez, Raúl Cabido, Juan J
Juan C. Núñez, Raúl Cabido, Juan J. Pantrigo, Antonio S. Montemayor, and José F. Vélez. Convolutional neural networks and long short- term memory for skeleton-based human activity and hand gesture recognition. Pattern Recognition, 76:80–94, 2018
2018
-
[145]
Real-time sign language fingerspelling recognition using convolutional neural network
Abiodun Oguntimilehin and Kolade Balogun. Real-time sign language fingerspelling recognition using convolutional neural network. Interna- tional Arab Journal of Information Technology, 21(1):158 – 165, 2024. Cited by: 1; All Open Access, Gold Open Access
2024
-
[146]
Efficient model-based 3d tracking of hand articulations using kinect
Iason Oikonomidis, Nikolaos Kyriazis, Antonis A Argyros, et al. Efficient model-based 3d tracking of hand articulations using kinect. In BmVC, volume 1, page 3, 2011
2011
-
[147]
Multiresolution gray-scale and rotation invariant texture classification with local bi- nary patterns
Timo Ojala, Matti Pietikainen, and Topi Maenpaa. Multiresolution gray-scale and rotation invariant texture classification with local bi- nary patterns. IEEE Transactions on pattern analysis and machine intelligence, 24(7):971–987, 2002
2002
-
[148]
Dhgd: Dynamic hand gesture dataset for skeleton-based gesture recog- nition and baseline evaluations
Masaya Okano, Jia-Qing Liu, Tomoko Tateyama, and Yen-Wei Chen. Dhgd: Dynamic hand gesture dataset for skeleton-based gesture recog- nition and baseline evaluations. In 2024 IEEE International Conference on Consumer Electronics (ICCE) , pages 1–4, 2024
2024
-
[149]
Hand gesture recognition based on computer vision: a review of techniques
Munir Oudah, Ali Al-Naji, and Javaan Chahl. Hand gesture recognition based on computer vision: a review of techniques. journal of Imaging, 6(8):73, 2020
2020
-
[150]
Ozdemir, Ahmet Alp Kındıro ˘glu, Necati Cihan Camg ¨
O ˘gulcan ¨"Ozdemir, Ahmet Alp Kındıro ˘glu, Necati Cihan Camg ¨"oz, and Lale Akarun. Bosphorussign22k sign language recognition dataset. arXiv preprint arXiv:2004.01283 , 2020
2004 arXiv
-
[151]
Back to rgb: 3d tracking of hands and hand-object interactions based on short-baseline stereo
Paschalis Panteleris and Antonis Argyros. Back to rgb: 3d tracking of hands and hand-object interactions based on short-baseline stereo. In Proceedings of the IEEE International Conference on Computer Vision Workshops, pages 575–584, 2017
2017
-
[152]
Using a single rgb frame for real time 3d hand pose estimation in the wild
Paschalis Panteleris, Iason Oikonomidis, and Antonis Argyros. Using a single rgb frame for real time 3d hand pose estimation in the wild. In 2018 IEEE Winter Conference on Applications of Computer Vision (WACV), pages 436–445, 2018
2018
-
[153]
Topic modeling of research fields: An interdisciplinary perspective
Michael Paul and Roxana Girju. Topic modeling of research fields: An interdisciplinary perspective. In Proceedings of the International Conference RANLP-2009, pages 337–342, 2009
2009
-
[154]
Abul Ala Walid, Rakhi Rani Paul, Md
Subrata Kumer Paul, Md. Abul Ala Walid, Rakhi Rani Paul, Md. Jamal Uddin, Md. Sohel Rana, Maloy Kumar Devnath, Ishaat Rahman Dipu, and Md. Momenul Haque. An adam based cnn and lstm approach for sign language recognition in real time for deaf people. Bulletin of Electrical Engi...
2024
-
[155]
Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black. Expressive body capture: 3d hands, face, and body from a single image. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogn...
2019
-
[156]
Learning to estimate 3d human pose and shape from a single color image
Georgios Pavlakos, Luyang Zhu, Xiaowei Zhou, and Kostas Daniilidis. Learning to estimate 3d human pose and shape from a single color image. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , June 2018
2018
-
[157]
Beyond temporal pooling: Recurrence and temporal convolutions for gesture recognition in video
Lionel Pigou, A ¨"aron van den Oord, Sander Dieleman, Mieke Van Herreweghe, and Joni Dambre. Beyond temporal pooling: Recurrence and temporal convolutions for gesture recognition in video. CoRR, abs/1506.01911, 2015
2015 arXiv
-
[158]
Gesture recognition performance score: A new metric to evaluate gesture recognition systems
Pramod Kumar Pisharady and Martin Saerbeck. Gesture recognition performance score: A new metric to evaluate gesture recognition systems. In Asian Conference on Computer Vision , pages 157–173. Springer, 2014
2014
-
[159]
Attention based detection and recognition of hand postures against complex backgrounds
Pramod Kumar Pisharady, Prahlad Vadakkepat, and Ai Poh Loh. Attention based detection and recognition of hand postures against complex backgrounds. International Journal of Computer Vision , 101:403–419, 2013
2013
-
[160]
Real-time multi- view bimanual gesture recognition
Geoffrey Poon, Kin Chung Kwan, and Wai-Man Pang. Real-time multi- view bimanual gesture recognition. In 2018 IEEE 3rd International Conference on Signal and Image Processing (ICSIP) , pages 19–23, 2018
2018
-
[161]
Historical development of hand gesture recognition
Prashan Premaratne and Prashan Premaratne. Historical development of hand gesture recognition. human computer interaction using hand gestures, pages 5–29, 2014
2014
-
[162]
Spelling it out: Real-time asl fingerspelling recognition
Nicolas Pugeault and Richard Bowden. Spelling it out: Real-time asl fingerspelling recognition. In 2011 IEEE International conference on computer vision workshops (ICCV workshops) , pages 1114–1119. IEEE, 2011
2011
-
[163]
Realtime and robust hand tracking from depth
Chen Qian, Xiao Sun, Yichen Wei, Xiaoou Tang, and Jian Sun. Realtime and robust hand tracking from depth. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 1106–1113, 2014
2014
-
[164]
Real time hand gesture recognition applied for flight simulator controls
Zhuang Qianzheng, Li Xiaodong, Ren Jie, and Qiao Yuanyuan. Real time hand gesture recognition applied for flight simulator controls. In 2021 IEEE 7th International Conference on Virtual Reality (ICVR) , pages 407–411, 2021
2021
-
[165]
Selvi Rajendran
Rajesh George Rajan and P. Selvi Rajendran. Gesture recognition of rgb-d and rgb static images using ensemble-based cnn architecture. In 2021 5th International Conference on Intelligent Computing and Control Systems (ICICCS) , pages 1579–1584, 2021
2021
-
[166]
Hand sign language recognition using multi-view hand skeleton
Razieh Rastgoo, Kourosh Kiani, and Sergio Escalera. Hand sign language recognition using multi-view hand skeleton. Expert Systems with Applications, 150:113336, 2020
2020
-
[167]
Multi-modal zero-shot dynamic hand gesture recognition
Razieh Rastgoo, Kourosh Kiani, Sergio Escalera, and Mohammad Sabokrou. Multi-modal zero-shot dynamic hand gesture recognition. Expert Systems with Applications , 247:123349, 2024
2024
-
[168]
Development and validation of a brazilian sign language database for human gesture recognition
Tamires Martins Rezende, Sílvia Grasiella Moreira Almeida, and Frederico Gadelha Guimarães. Development and validation of a brazilian sign language database for human gesture recognition. Neural Computing and Applications , 33(16):10449–10467, 2021
2021
-
[169]
why should i trust you?
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. "why should i trust you?": Explaining the predictions of any classifier, 2016
2016
-
[170]
Pa- padopoulos
Nikolaos Rodis, Christos Sardianos, Panagiotis Radoglou-Grammatikis, Panagiotis Sarigiannidis, Iraklis Varlamis, and Georgios Th. Pa- padopoulos. Multimodal explainable artificial intelligence: A com- prehensive review of methodological advances and future research directions, 2024
2024
-
[171]
3d hand pose detection in egocentric rgb-d images
Grégory Rogez, Maryam Khademi, JS Supan ˇciˇc III, Jose Maria Mar- tinez Montiel, and Deva Ramanan. 3d hand pose detection in egocentric rgb-d images. In Computer Vision-ECCV 2014 Workshops: Zurich, Switzerland, September 6-7 and 12, 2014, Proceedings, Part I 13, pages 356–371...
2014
-
[172]
Lsa64: A dataset of argentinian sign language
Franco Ronchetti, Facundo Quiroga, Cesar Estrebou, Laura Lanzarini, and Alejandro Rosete. Lsa64: A dataset of argentinian sign language. XX II Congreso Argentino de Ciencias de la Computación (CACIC) , 2016
2016
-
[173]
Frankmocap: A monocular 3d whole-body pose estimation system via regression and integration
Yu Rong, Takaaki Shiratori, and Hanbyul Joo. Frankmocap: A monocular 3d whole-body pose estimation system via regression and integration. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops , pages 1749–1759, October 2021
2021
-
[174]
Shahen Shah, and Md Baharul Islam
Arezoo Sadeghzadeh, A.F.M. Shahen Shah, and Md Baharul Islam. Mlmsign: Multi-lingual multi-modal illumination-invariant sign lan- guage recognition. Intelligent Systems with Applications , 22:200384, 2024
2024
-
[175]
Pifuhd: Multi-level pixel-aligned implicit function for high-resolution 3d human digitization
Shunsuke Saito, Tomas Simon, Jason Saragih, and Hanbyul Joo. Pifuhd: Multi-level pixel-aligned implicit function for high-resolution 3d human digitization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2020
2020
-
[176]
Methods, databases and recent advancement of vision-based hand gesture recognition for hci systems: A review
Debajit Sarma and Manas Kamal Bhuyan. Methods, databases and recent advancement of vision-based hand gesture recognition for hci systems: A review. SN Computer Science , 2(6):436, 2021
2021
-
[177]
Selvaraju, Michael Cogswell, Abhishek Das, Ramakr- ishna Vedantam, Devi Parikh, and Dhruv Batra
Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakr- ishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-cam: Visual explanations from deep networks via gradient-based localization. Inter- national Journal of Computer Vision , 128(2):336–359, October 2019
2019
-
[178]
Ntu rgb+ d: A large scale dataset for 3d human activity analysis
Amir Shahroudy, Jun Liu, Tian-Tsong Ng, and Gang Wang. Ntu rgb+ d: A large scale dataset for 3d human activity analysis. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 1010–1019, 2016
2016
-
[179]
An accurate estimation of hand gestures using optimal modified convolutional neural network
Subhashini Shanmugam and Revathi Sathya Narayanan. An accurate estimation of hand gestures using optimal modified convolutional neural network. Expert Systems with Applications , 249:123351, 2024
2024
-
[180]
Hand gesture recognition with YOLOv8 on OAK-D in near real-time
Aditya Sharma. Hand gesture recognition with YOLOv8 on OAK-D in near real-time. In Puneet Chugh, Aritra Roy Gosthipaty, Susan Huot, Kseniia Kidriavsteva, Ritwik Raha, and Abhishek Thanki, editors, PyImageSearch. 2023
2023
-
[181]
Vision-based hand gesture recognition using deep learning for the interpretation of sign language
Sakshi Sharma and Sukhwinder Singh. Vision-based hand gesture recognition using deep learning for the interpretation of sign language. Expert Systems with Applications , 182:115657, 2021
2021
-
[182]
Decoupled spatial-temporal attention network for skeleton-based action-gesture recognition
Lei Shi, Yifan Zhang, Jian Cheng, and Hanqing Lu. Decoupled spatial-temporal attention network for skeleton-based action-gesture recognition. In Proceedings of the Asian Conference on Computer Vision (ACCV), November 2020
2020
-
[183]
A methodological and structural review of hand gesture recognition across diverse data modalities
Jungpil Shin, Abu Saleh Musa Miah, Md Humaun Kabir, Md Abdur Rahim, and Abdullah Al Shiam. A methodological and structural review of hand gesture recognition across diverse data modalities. IEEE Access, 2024
2024
-
[184]
3d hand reconstruction via aggregating intra and inter graphs guided by prior knowledge for hand-object interaction scenario
Feng Shuang, Wenbo He, and Shaodong Li. 3d hand reconstruction via aggregating intra and inter graphs guided by prior knowledge for hand-object interaction scenario. Journal of Visual Communication and Image Representation, 100:104129, 2024
2024
-
[185]
Autsl: A large scale multi-modal turkish sign language dataset and baseline methods
Ozge Mercanoglu Sincan and Hacer Yalim Keles. Autsl: A large scale multi-modal turkish sign language dataset and baseline methods. IEEE access, 8:181340–181355, 2020
2020
-
[186]
Impact of colour image and skeleton plotting on sign language recognition using convolutional neural networks (cnn)
Anushka Singh, Fahad Eqbal Hashmi, Naman Tyagi, and Anant Kumar Jayswal. Impact of colour image and skeleton plotting on sign language recognition using convolutional neural networks (cnn). In 2024 14th International Conference on Cloud Computing, Data Science & Engineering (C...
2024
-
[187]
Video understanding-based random hand gesture authentication
Wenwei Song, Wenxiong Kang, Lu Wang, Zenan Lin, and Mengting Gan. Video understanding-based random hand gesture authentication. IEEE Transactions on Biometrics, Behavior, and Identity Science , 4(4):453–470, 2022
2022
-
[188]
Accurate hand contact detection from rgb images via image-to-image translation
Suzanne Sorli, Marc Comino-Trinidad, and Dan Casas. Accurate hand contact detection from rgb images via image-to-image translation. Computers & Graphics , page 104200, 2025
2025
-
[189]
Include: A large scale dataset for indian sign language recognition
Advaith Sridhar, Rohith Gandhi Ganesan, Pratyush Kumar, and Mitesh Khapra. Include: A large scale dataset for indian sign language recognition. In Proceedings of the 28th ACM international conference on multimedia, pages 1366–1375, 2020
2020
-
[190]
Real-time joint tracking of a hand manipulating an object from rgb-d input
Srinath Sridhar, Franziska Mueller, Michael Zollhöfer, Dan Casas, Antti Oulasvirta, and Christian Theobalt. Real-time joint tracking of a hand manipulating an object from rgb-d input. In Computer Vision– ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October ...
2016
-
[191]
Interactive markerless articulated hand motion tracking using rgb and depth data
Srinath Sridhar, Antti Oulasvirta, and Christian Theobalt. Interactive markerless articulated hand motion tracking using rgb and depth data. In Proceedings of the IEEE international conference on computer vision, pages 2456–2463, 2013
2013
-
[192]
Depth-based hand pose estimation: methods, data, and challenges
James Steven Supan ˇciˇc, Gregory Rogez, Yi Yang, Jamie Shotton, and Deva Ramanan. Depth-based hand pose estimation: methods, data, and challenges. International Journal of Computer Vision, 126:1180–1198, 2018
2018
-
[193]
Showme: Robust object- agnostic hand-object 3d reconstruction from rgb video
Anilkumar Swamy, Vincent Leroy, Philippe Weinzaepfel, Fabien Ba- radel, Salma Galaaoui, Romain Brégier, Matthieu Armando, Jean- Sebastien Franco, and Grégory Rogez. Showme: Robust object- agnostic hand-object 3d reconstruction from rgb video. Computer Vision and Image Understa...
2024
-
[194]
Talaat, Walid El-Shafai, Naglaa F
Fatma M. Talaat, Walid El-Shafai, Naglaa F. Soliman, Abeer D. Al- garni, Fathi E. Abd El-Samie, and Ali I. Siam. Real-time arabic avatar for deaf-mute communication enabled by deep learning sign language translation. Computers and Electrical Engineering , 119:109475, 2024
2024
-
[195]
Hgr-vit: hand gesture recognition with vision transformer
Chun Keat Tan, Kian Ming Lim, Roy Kwang Yang Chang, Chin Poo Lee, and Ali Alqahtani. Hgr-vit: hand gesture recognition with vision transformer. Sensors, 23(12):5555, 2023
2023
-
[196]
Latent regression forest: Structured estimation of 3d articulated hand posture
Danhang Tang, Hyung Jin Chang, Alykhan Tejani, and Tae-Kyun Kim. Latent regression forest: Structured estimation of 3d articulated hand posture. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 3786–3793, 2014
2014
-
[197]
A general skeleton-based action and gesture recognition framework for human–robot collaboration
Matteo Terreran, Leonardo Barcellona, and Stefano Ghidoni. A general skeleton-based action and gesture recognition framework for human–robot collaboration. Robotics and Autonomous Systems , 170:104523, 2023
2023
-
[198]
Real- time continuous pose recovery of human hands using convolutional networks
Jonathan Tompson, Murphy Stein, Yann Lecun, and Ken Perlin. Real- time continuous pose recovery of human hands using convolutional networks. ACM Transactions on Graphics (ToG) , 33(5):1–10, 2014
2014
-
[199]
Motion feature estimation using bi- directional gru for skeleton-based dynamic hand gesture recognition
Reena Tripathi and Bindu Verma. Motion feature estimation using bi- directional gru for skeleton-based dynamic hand gesture recognition. Signal, image and video processing , 18(Suppl 1):299–308, 2024
2024
-
[200]
Survey on vision-based dynamic hand gesture recognition
Reena Tripathi and Bindu Verma. Survey on vision-based dynamic hand gesture recognition. The Visual Computer , 40(9):6171–6199, 2024
2024
-
[201]
An analysis of convolutional long short-term memory recurrent neural networks for gesture recognition
Eleni Tsironi, Pablo Barros, Cornelius Weber, and Stefan Wermter. An analysis of convolutional long short-term memory recurrent neural networks for gesture recognition. Neurocomputing, 268:76–86, 2017. Advances in artificial neural networks, machine learning and compu- tationa...
2017
-
[202]
Consistent 3d hand reconstruction in video via self-supervised learning
Zhigang Tu, Zhisheng Huang, Yujin Chen, Di Kang, Linchao Bao, Bisheng Yang, and Junsong Yuan. Consistent 3d hand reconstruction in video via self-supervised learning. IEEE Transactions on Pattern Analysis and Machine Intelligence , 45(8):9469–9485, 2023
2023
-
[203]
Capturing hands in action using discriminative salient points and physics simulation
Dimitrios Tzionas, Luca Ballan, Abhilash Srikantha, Pablo Aponte, Marc Pollefeys, and Juergen Gall. Capturing hands in action using discriminative salient points and physics simulation. International Journal of Computer Vision , 118:172–193, 2016
2016
-
[204]
Representation learning with contrastive predictive coding, 2019
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding, 2019
2019
-
[205]
Bodynet: V olumetric inference of 3d human body shapes
Gul Varol, Duygu Ceylan, Bryan Russell, Jimei Yang, Ersin Yumer, Ivan Laptev, and Cordelia Schmid. Bodynet: V olumetric inference of 3d human body shapes. In Proceedings of the European Conference on Computer Vision (ECCV) , September 2018
2018
-
[206]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need, 2023
2023
-
[207]
Recovering accurate 3d human pose in the wild using imus and a moving camera
Timo V on Marcard, Roberto Henschel, Michael J Black, Bodo Rosen- hahn, and Gerard Pons-Moll. Recovering accurate 3d human pose in the wild using imus and a moving camera. In Proceedings of the European conference on computer vision (ECCV) , pages 601–617, 2018
2018
-
[208]
Chalearn looking at people rgb-d isolated and continuous datasets for gesture recognition
Jun Wan, Yibing Zhao, Shuai Zhou, Isabelle Guyon, Sergio Escalera, and Stan Z Li. Chalearn looking at people rgb-d isolated and continuous datasets for gesture recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition workshops , pages 56–64, 2016
2016
-
[209]
Al-mobilenet: a novel model for 2d gesture recognition in intelligent cockpit based on multi-modal data
Bin Wang, Liwen Yu, and Bo Zhang. Al-mobilenet: a novel model for 2d gesture recognition in intelligent cockpit based on multi-modal data. Artificial Intelligence Review , 57(10):282, 2024
2024
-
[210]
Real-time block-based embedded cnn for gesture classification on an fpga
Ching-Chen Wang, Yu-Chun Ding, Ching-Te Chiu, Chao-Tsung Huang, Yen-Yu Cheng, Shih-Yi Sun, Chih-Han Cheng, and Hsueh-Kai Kuo. Real-time block-based embedded cnn for gesture classification on an fpga. IEEE Transactions on Circuits and Systems I: Regular Papers, 68(10):4182–4193, 2021
2021
-
[211]
Region ensemble network: Towards good practices for deep 3d hand pose estimation
Guijin Wang, Xinghao Chen, Hengkai Guo, and Cairong Zhang. Region ensemble network: Towards good practices for deep 3d hand pose estimation. Journal of Visual Communication and Image Repre- sentation, 55:404–414, 2018
2018
-
[212]
Mining actionlet ensemble for action recognition with depth cameras
Jiang Wang, Zicheng Liu, Ying Wu, and Junsong Yuan. Mining actionlet ensemble for action recognition with depth cameras. In 2012 IEEE conference on computer vision and pattern recognition , pages 1290–1297. IEEE, 2012
2012
-
[213]
Mask-pose cascaded cnn for 2d hand pose estimation from single color image
Yangang Wang, Cong Peng, and Yebin Liu. Mask-pose cascaded cnn for 2d hand pose estimation from single color image. IEEE Transac- tions on Circuits and Systems for Video Technology, 29(11):3258–3268, 2018
2018
-
[214]
Accurate and real- time variant hand pose estimation based on gray code bounding box representation
Yangang Wang, Wenqian Sun, and Ruting Rao. Accurate and real- time variant hand pose estimation based on gray code bounding box representation. IEEE Sensors Journal , 24(11):18043–18053, 2024
2024
-
[215]
Handgcn- former: A novel topology-aware transformer network for 3d hand pose estimation
Yintong Wang, LiLi Chen, Jiamao Li, and Xiaolin Zhang. Handgcn- former: A novel topology-aware transformer network for 3d hand pose estimation. In 2023 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pages 5664–5673, 2023
2023
-
[216]
Nonnegative matrix factorization: A comprehensive review
Yu-Xiong Wang and Yu-Jin Zhang. Nonnegative matrix factorization: A comprehensive review. IEEE Transactions on Knowledge and Data Engineering, 25(6):1336–1353, 2013
2013
-
[217]
Yo3rl-net:a fusion of two-phase end-to-end deep net framework for hand detection and gesture recognition
Xiang Wu, Yuanhao Ma, Shijie Zhang, Tianfei Chen, and He Jiang. Yo3rl-net:a fusion of two-phase end-to-end deep net framework for hand detection and gesture recognition. Alexandria Engineering Journal, 121:77–89, 2025
2025
-
[218]
View invariant human action recognition using histograms of 3d joints
Lu Xia, Chia-Chih Chen, and Jake K Aggarwal. View invariant human action recognition using histograms of 3d joints. In 2012 IEEE computer society conference on computer vision and pattern recognition workshops, pages 20–27. IEEE, 2012
2012
-
[219]
Multi-modal sign language recognition with enhanced spatiotemporal representation
Shiwei Xiao, Yuchun Fang, and Lan Ni. Multi-modal sign language recognition with enhanced spatiotemporal representation. In 2021 International Joint Conference on Neural Networks (IJCNN) , pages 1–8, 2021
2021
-
[220]
Efficient hand pose estimation from a single depth image
Chi Xu and Li Cheng. Efficient hand pose estimation from a single depth image. In Proceedings of the IEEE international conference on computer vision, pages 3456–3462, 2013
2013
-
[221]
Continuous sign language recognition based on hierarchi- cal memory sequence network
Cuihong Xue, Jingli Jia, Ming Yu, Gang Yan, Yingchun Guo, and Yuehao Liu. Continuous sign language recognition based on hierarchi- cal memory sequence network. IET Computer Vision , 18(2):247–259, 2024
2024
-
[222]
Skeleton- based hand gesture recognition for assembly line operation
Chao-Lung Yang, Wen-Ting Li, and Shang-Che Hsu. Skeleton- based hand gesture recognition for assembly line operation. In 2020 International Conference on Advanced Robotics and Intelligent Systems (ARIS), pages 1–6, 2020
2020
-
[223]
Combination of semantic segmentation and skeleton estimation for human hands detection
Chao-Lung Yang, Yang-Hsiu Tung, Tzu-Ching Kao, Chao-Hung Huang, En Liou, Po-Ting Lin, and Kai-Lung Hua. Combination of semantic segmentation and skeleton estimation for human hands detection. In 2023 International Conference on Advanced Robotics and Intelligent Systems (ARIS) ...
2023
-
[224]
The korean sign language dataset for action recognition
Seunghan Yang, Seungjun Jung, Heekwang Kang, and Changick Kim. The korean sign language dataset for action recognition. In Interna- tional conference on multimedia modeling , pages 532–542. Springer, 2019
2019
-
[225]
A systematic review on hand gesture recognition techniques, challenges and applications
Mais Yasen and Shaidah Jusoh. A systematic review on hand gesture recognition techniques, challenges and applications. PeerJ Computer Science, 5:e218, 2019
2019
-
[226]
3d hand pose estimation and gesture recognition based on hand-object interaction information
Qingshan Yin, Qiaochu Zhao, Huijie Jia, Yan Gao, Luoluo Feng, Chaoming Li, Yao Cheng, and Bin Lin. 3d hand pose estimation and gesture recognition based on hand-object interaction information. In 2023 IEEE/CIC International Conference on Communications in China (ICCC), pages 1–6, 2023
2023
-
[227]
The 2017 hands in the million challenge on 3d hand pose estimation
Shanxin Yuan, Qi Ye, Guillermo Garcia-Hernando, and Tae-Kyun Kim. The 2017 hands in the million challenge on 3d hand pose estimation. arXiv preprint arXiv:1707.02237 , 2017
2017 arXiv
-
[228]
Development of a lightweight real- time application for dynamic hand gesture recognition
Oluwaleke Yusuf and Maki Habib. Development of a lightweight real- time application for dynamic hand gesture recognition. In 2023 IEEE International Conference on Mechatronics and Automation (ICMA) , pages 543–548, 2023
2023
-
[229]
Baghdadi, Samah Adel Gamel, Man- sourah Aljohani, Fatma M
Hanaa Zaineldin, Nadiah A. Baghdadi, Samah Adel Gamel, Man- sourah Aljohani, Fatma M. Talaat, Amer Malki, Mahmoud Badawy, and Mostafa Elhosseini. Active convolutional neural networks sign language (activecnn-sl) framework: a paradigm shift in deaf-mute communication. Artificia...
2024
-
[230]
Semi- supervised rgb-d hand gesture recognition via mutual learning of self- supervised models
Jian Zhang, Kaihao He, Ting Yu, Jun Yu, and Zhenming Yuan. Semi- supervised rgb-d hand gesture recognition via mutual learning of self- supervised models. ACM Trans. Multimedia Comput. Commun. Appl. , 21(4), March 2025
2025
-
[231]
3d hand pose tracking and estimation using stereo matching
Jiawei Zhang, Jianbo Jiao, Mingliang Chen, Liangqiong Qu, Xiaobin Xu, and Qingxiong Yang. 3d hand pose tracking and estimation using stereo matching. arXiv preprint arXiv:1610.07214 , 2016
2016 arXiv
-
[232]
Attention in convolutional lstm for gesture recognition
Liang Zhang, Guangming Zhu, Lin Mei, Peiyi Shen, Syed Afaq Ali Shah, and Mohammed Bennamoun. Attention in convolutional lstm for gesture recognition. Advances in neural information processing systems, 31, 2018
2018
-
[233]
Handformer2t: A lightweight regression-based model for interacting hands pose estimation from a single rgb image
Pengfei Zhang and Deying Kong. Handformer2t: A lightweight regression-based model for interacting hands pose estimation from a single rgb image. In 2024 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) , pages 6236–6245, 2024
2024
-
[234]
From actemes to action: A strongly-supervised representation for detailed action understanding
Weiyu Zhang, Menglong Zhu, and Konstantinos G Derpanis. From actemes to action: A strongly-supervised representation for detailed action understanding. In Proceedings of the IEEE international conference on computer vision , pages 2248–2255, 2013
2013
-
[235]
Egogesture: A new dataset and benchmark for egocentric hand gesture recognition
Yifan Zhang, Congqi Cao, Jian Cheng, and Hanqing Lu. Egogesture: A new dataset and benchmark for egocentric hand gesture recognition. IEEE Transactions on Multimedia , 20(5):1038–1050, 2018
2018
-
[236]
Human-robot interactive operating system for underwater manipulators based on hand gesture recognition
Yufei Zhang, Zheyu Hu, Dawei Tu, and Xu Zhang. Human-robot interactive operating system for underwater manipulators based on hand gesture recognition. In International Conference on Intelligent Robotics and Applications, pages 575–586. Springer, 2023
2023
-
[237]
Adaptive cross-fusion learning for multi-modal gesture recognition
Benjia Zhou, Jun Wan, Yanyan Liang, and Guodong Guo. Adaptive cross-fusion learning for multi-modal gesture recognition. Virtual Reality & Intelligent Hardware , 3(3):235–247, 2021
2021
-
[238]
A lightweight hand gesture recognition in complex backgrounds
Weina Zhou and Kun Chen. A lightweight hand gesture recognition in complex backgrounds. Displays, 74:102226, 2022
2022
-
[239]
Learning to estimate 3d hand pose from single rgb images
Christian Zimmermann and Thomas Brox. Learning to estimate 3d hand pose from single rgb images. In Proceedings of the IEEE international conference on computer vision , pages 4903–4911, 2017
2017
-
[240]
Freihand: A dataset for markerless capture of hand pose and shape from single rgb images
Christian Zimmermann, Duygu Ceylan, Jimei Yang, Bryan Russell, Max Argus, and Thomas Brox. Freihand: A dataset for markerless capture of hand pose and shape from single rgb images. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2019
2019
-
[241]
Video gesture analysis for autism spectrum disorder detection
Andrea Zunino, Pietro Morerio, Andrea Cavallo, Caterina Ansuini, Jessica Podda, Francesca Battaglia, Edvige Veneselli, Cristina Becchio, and Vittorio Murino. Video gesture analysis for autism spectrum disorder detection. In 2018 24th International Conference on Pattern Recogni...
2018
-
[242]
Bayta¸ s, and Lale Akarun
O ˘gulcan Özdemir, ˙Inci M. Bayta¸ s, and Lale Akarun. Hand graph topology selection for skeleton-based sign language recognition. In 2024 IEEE 18th International Conference on Automatic Face and Gesture Recognition (FG) , pages 1–5, 2024. Manousos Linardakis is a B.Sc. gradua...
2024
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.