{"id":"8f135590-0070-4864-ae8b-2d6f7ae12ccc","arxiv_id":"2501.11992","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A structured survey of recent hand gesture recognition from camera input, with a taxonomy of methods, datasets, metrics, and open problems.","lead":"This paper reviews recent work on recognizing hand gestures from camera images and video, and organizes methods, datasets, metrics, and challenges. It is a survey for researchers and practitioners who need a map of current hand gesture recognition work.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's comprehensiveness claim is undermined by a retrieval pipeline that omits depth-only HGR work and by contradictory paper counts; the selection is not reproducible as reported.","rationale":"Good-faith reading: the survey is a useful organized review, and its taxonomy of tasks, input types, capture methods, and algorithms is reasonable. The overviews of standard algorithms and challenges are competent, and the paper does not claim a formal mathematical result that could be internally inconsistent. The reader's weakest assumption correctly identifies the literature retrieval pipeline as load-bearing. I add a concrete mechanism that makes that concern sharper: the Scopus query design has no depth/RGB-D search term even though depth input is one of the survey's three core input modalities. This is not merely a hypothetical 'some work might have been missed'; it is a systematic blind spot in the query design. The internal count inconsistencies (125 vs. 137, 88 vs. 89) and the disagreement about whether NNMF used abstracts or only titles/keywords further undermine reproducibility. These issues do not warrant rejection: a corrected retrieval and a released paper manifest would likely preserve the survey's organizational value. They do warrant keeping the verdict conditional until the selection is verified and the quantitative claims are recomputed.","tokens_in":46689,"tokens_out":4821,"duration_ms":54139,"concrete_test":"Reproduce the exact four Scopus queries from Table II for 2018-2025 with duplicate removal and NNMF as described, and in parallel run a fifth query: ('depth' OR 'RGB-D') AND ('hand') AND ('gesture' OR 'pose') AND ('recognition' OR 'estimation'), applying the same inclusion/exclusion criteria. Publish both resulting paper manifests and compare unique papers. If the fifth query adds any inclusion-criteria-satisfying depth/RGB-D papers not already in the original 125, or if the reproduced counts do not reconcile to the reported 125, 88, and 137 figures, then the comprehensiveness claim in Section I and the aggregate distributions in Figs. 8, 12, and 16 are not robust and require revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that this survey is a comprehensive synthesis of visual-input hand gesture recognition (Abstract, Section I). That claim rests on the retrieval pipeline in Section III, and the pipeline has a systematic blind spot for the depth/RGB-D modality the survey explicitly claims to cover. In Table II, query 1 requires the term 'RGB', and query 3 permits 'RGB', 'video', 'skeleton', or 'multi modal' but not 'depth' or 'RGB-D'. Consequently, relevant papers on depth-based hand pose estimation or gesture recognition that do not also mention RGB/video/skeleton are invisible to the selection process. Since the survey reports that 19% of selected papers use RGB-D input (Fig. 8a) and devotes Section IV-B to RGB-D methods, this omission directly threatens the representativeness of the descriptive statistics. The pipeline is also internally inconsistent and unreproducible as reported: Table I states 137 papers, while Section III-C states 125; Table II sums to 89 Scopus selections while the text says 88; Section III-A says NNMF was applied to titles, keywords, and abstracts, while Section III-D says it used only titles and keywords; and no NNMF relevance threshold, per-paper topic assignment, or selected-paper manifest is provided. Because the headline claims about trends, method prevalence, and gaps are aggregates over this selected set, a systematic retrieval omission shifts those conclusions. This is a correctness risk in the survey's sampling basis, not a flaw in any mathematical derivation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a structured survey of hand gesture recognition (HGR) and 3D hand pose estimation from visual input (RGB, RGB-D, and video), covering work published between 2018 and 2025. It proposes a taxonomy based on task objective, input type, capture method, and recognition technique; applies non-negative matrix factorization (NNMF) to organize retrieved papers into six topics; tabulates benchmark datasets and state-of-the-art results; and discusses challenges, edge deployment, explainability, bias, and privacy. The stated goal is to provide a comprehensive yet focused alternative to broader recent surveys that include non-visual modalities.","tokens_in":46901,"tokens_out":5238,"duration_ms":56128,"significance":"If its selection process is representative and its counts are correct, the survey would be a useful entry point for researchers: it offers a clear classification scheme, a compact dataset table, an algorithmic overview with common formulations, and welcome sections on deployment and ethical considerations. The use of topic modeling to structure the literature is a constructive addition to typical survey methodology. Its value as a reference, however, depends on reproducible retrieval and internally consistent statistics, both of which currently need repair; those issues affect the headline claims about trends and method prevalence.","major_comments":[{"comment":"The reported paper counts are internally inconsistent: Table I states that the survey reviews 137 papers, Section III-C states 125 studies (37 from Scholar and 88 from Scopus), and Table II's Selected Papers column sums to 89 rather than 88. Since the quantitative claims (venue distribution, input-type percentages, topic timeline) are aggregates over this set, please reconcile the totals and provide a complete list of the included papers, or a public manifest, so the counts can be audited.","section":"Table I, Section III-A, Section III-C"},{"comment":"The Scopus query design systematically excludes depth-only work: query 1 requires the term 'RGB', and query 3 permits 'RGB', 'video', 'skeleton', or 'multi modal' but not 'depth' or 'RGB-D'. Because the survey explicitly claims to cover depth images as input and reports that 19% of its selected papers use RGB-D input (Fig. 8a), any depth-based hand pose estimation or gesture recognition paper that does not mention RGB/video/skeleton is invisible to the selection pipeline. Please add explicit depth/RGB-D queries and rerun the selection, or qualify the comprehensiveness claim accordingly.","section":"Table II, Section IV-B"},{"comment":"The NNMF-based selection step is not reproducible as reported: Section III-A says NNMF was applied to the titles, keywords, and abstracts of the papers, while Section III-D says it was applied only to titles and keywords; no relevance threshold, per-paper topic assignment rule, or list of papers rejected after topic modeling is provided. Please specify the exact text fields used, define the relevance criterion, and release the topic assignments or the selection script so that another group can reproduce the 125/137-paper set.","section":"Section III-A, Section III-D"},{"comment":"The state-of-the-art table requires verification before it can be trusted: the reported best MPJPE of 1.1 on HO3D and 1.18 on FreiHAND are well below typical published results on these benchmarks (which are usually in the millimeter-to-centimeter range depending on the protocol), and the 100% accuracy on ASL is presented without any dataset split or evaluation details. Please state the exact metric unit, evaluation protocol, and data split for each row, or remove values that cannot be substantiated.","section":"Table VII"}],"minor_comments":[{"comment":"The sentence 'a \"true\" vision-based approach reported by in 1993 [161]' is missing the author name and should be reworded.","section":"Section I"},{"comment":"The sentence 'The methodology for extracting the topics from the collection of articles is detailed in Section II-D that follows' should refer to Section III-D, not Section II-D.","section":"Section III-C"},{"comment":"The 'American Sign Language Digits' row cites the same reference [17] as the MUGD row, reports a different sample count, and lists 36 classes for a digit dataset; this appears to be a dataset mislabeling and should be checked.","section":"Table V"},{"comment":"The notation 'WQ.WK.WV' should be written as 'WQ, WK, WV' or an equivalent list, since the periods are ambiguous.","section":"Section V-F"},{"comment":"The citation for Transformer models in the sentence 'Vision Transformers (ViT) [16]' points to a gesture-recognition application paper rather than the original ViT paper (Dosovitskiy et al.); please update the reference.","section":"Section VII-B"}],"recommendation":"major_revision","confidential_remarks":"This is a useful organizational contribution, but the sampling basis and counting inconsistencies are load-bearing for a survey whose central claim is comprehensiveness. The depth-query blind spot and the absence of a reproducible relevance threshold materially affect the reported statistics. I do not see signs of misconduct; the issues look like drafting inconsistencies and verification gaps that a careful revision can resolve."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline: this is a genuinely useful survey of visual-input HGR whose numbers you currently cannot fully trust. The taxonomy is solid; the paper-count inconsistencies and a depth-focused retrieval blind spot need fixing before I'd rely on its quantitative claims.\n\nWhat is new is modest but real. Several 2024 surveys cover adjacent ground (Hashi et al., Shin et al., Tripathi & Verma). This one adds a tighter visual-input scope and an NNMF topic-modeling pass, plus dataset tables grouped by task and input modality. That is a reasonable organizing contribution, not a breakthrough. The circularity burden is essentially nil; the two self-citations are peripheral.\n\nThe soft spots are real and mostly fixable. The paper's own counts are inconsistent: Table I says 137 papers, Section III-C says 125 (37 Scholar + 88 Scopus), and the Scopus query table sums to 89 selected papers while the text says 88. Sloppy, but correctable. More concerning, the retrieval pipeline has a systematic blind spot: the Scopus queries require the terms \"RGB\", \"video\", or \"skeleton\". A depth-only paper that mentions none of these is invisible. The survey explicitly claims to cover depth input and reports that 19% of selected papers use RGB-D, so the sampling basis is skewed and the depth-related trend statistics cannot be trusted. The topic-modeling description also contradicts itself (titles/keywords/abstracts in Section III-A vs titles/keywords in Section III-D). And Table VII's state-of-the-art values are implausible: HO3D MPJPE of 1.1 and FreiHAND of 1.18 are not numbers I can square with the literature in millimeters. My guess is a unit or transcription error. The authors include a disclaimer about cross-study comparison, but that does not justify presenting those values without protocol context.\n\nWho is this for? A researcher entering HGR who wants a quick map of tasks, datasets, and method families. It is one of several recent surveys, not a unique resource, but it is a fair one. It deserves a serious referee: the synthesis is useful and every problem I listed is correctable. I would send it to peer review with a request for revision, specifically to reconcile the paper counts, fix or acknowledge the depth-only retrieval gap, release the selected-paper manifest and NNMF code, and correct or contextualize Table VII. If those changes land, I would cite it.","headline":"A genuinely useful survey of visual-input HGR with a data-quality problem: the taxonomy is solid, but the paper's own counts don't add up and the retrieval pipeline misses depth-only work.","tokens_in":47469,"tokens_out":3513,"would_cite":false,"duration_ms":36484,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey claims that recent hand gesture recognition research from visual input can be organized into six research topics and a five-axis method taxonomy, and that this organization reveals systematic associations—box/filter capture…","keywords":["hand gesture recognition","gesture classification","gesture estimation","sign language recognition","visual input","RGB-D data","benchmark datasets","deep learning"],"falsifier":"Re-run the selection pipeline with an additional broad query such as 'egocentric hand' or 'hand-object interaction' for the same period and same venues, then recompute the reported percentages for input type and recognition method; if the added papers shift those percentages substantially or introduce new frequent topics, the 125-paper corpus is not representative and the associations the survey draws would need re-examination.","tokens_in":46435,"feed_emoji":"👋","tokens_out":6094,"duration_ms":60609,"temperature":0.7,"pith_summary":"This paper tries to show that hand gesture recognition from visual input has matured into a field that can be mapped by a few stable axes: the task (classifying a gesture versus estimating hand pose), the input (RGB, depth, or video, monocular or multi-view), how the hand is captured (skeleton versus bounding-box/filter), and the recognition technique (neural, non-neural, or hybrid). On a corpus of 125 studies from 2018 to 2025, the authors build that map and use it to quantify where the field clusters: video is the most common input, hybrid CNN-based pipelines dominate, and classification is a more frequent goal than estimation. The survey also inventories benchmark datasets, reports state-of-the-art accuracies on key datasets, and lists open challenges such as occlusion, cross-user generalization, and real-time efficiency. A sympathetic reader would take the contribution as an organized, evidence-based snapshot of the field plus a set of associations that help researchers position new work.","feed_headline":"Six themes organize 125 hand-gesture papers","feed_subtitle":"A five-axis taxonomy—task, input, camera setup, capture, method—shows where the field clusters and what remains open.","key_machinery":"The carrying object is the survey's classification framework itself: six topic clusters discovered by non-negative matrix factorization of paper titles, keywords, and abstracts—hand gesture classification, hand gesture estimation, sign language recognition, hand/body reconstruction, multimodal fusion, and real-time recognition—combined with a five-axis methodological table (input type, camera setup, capture method, task goal, recognition method). The framework is used to tabulate all 125 reviewed papers and, through Bayes' theorem, to compute conditional probabilities such as P(classification | box/filter), the quantitative evidence for the paper's associations. The same machinery structures the dataset and challenge sections, making the taxonomy the device that connects selection, analysis, and conclusions.","core_discovery":"The paper's central claim is that a systematic review of visual-input hand gesture recognition, built from top-venue publications and targeted database queries, yields a coherent taxonomy that previous surveys lacked. The taxonomy separates gesture classification from gesture estimation, then cross-cuts those tasks by input modality (RGB, RGB-D, video), camera setup (monocular, multi-view), hand capture representation (skeleton-based versus box/filter-based), and recognition method (neural network, non-neural, or hybrid). Within this corpus the authors report that video input accounts for 53% of studies, hybrid methods for 68%, box/filter capture is about twice as likely to be associated with classification than estimation, and multi-view setups are more strongly associated with estimation. They further provide a dataset inventory showing ASL and HO3D as the most-used benchmarks for classification and estimation respectively and argue that lack of standardized benchmarks is a central limitation of current research.","pith_inferences":["A consequence the authors leave implicit: if the field adopted their taxonomy as a reporting standard, meta-analyses could track shifts in method prevalence over time and test whether the 2024 spike in publications continues.","The Bayes associations are corpus-relative; a broader corpus that included more hand-object interaction and egocentric work would likely raise the estimation share and strengthen the multiview-estimation link.","A testable extension would be to run the same selection pipeline on the 2025-2026 literature and check whether transformer-based methods displace CNN+LSTM hybrids as the dominant approach, a trend the paper identifies as emerging.","The benchmarking gap they identify suggests a concrete next step: a reproducibility study that fixes data splits and metrics, re-evaluating the leading methods on ASL, AUTSL, JESTER, WLASL, HO3D, and FreiHAND under one protocol."],"forward_implications":["Researchers entering the field can use the taxonomy to position a new method against the dominant video-and-hybrid baseline rather than searching across hundreds of papers.","The reported associations give testable expectations: a new box/filter method is more likely aimed at classification, and a multi-view system at hand pose estimation.","Datasets ASL and HO3D function as de facto benchmarks for classification and estimation, so new methods will be compared against the headline numbers the survey compiles.","Because accuracy on some benchmark datasets is already very high (100% on ASL, 98.53% on AUTSL), progress signals will increasingly come from harder, more naturalistic datasets such as WLASL and isoGD, where top accuracies remain below 85%.","The survey's call for unified evaluation frameworks, if heeded, would make its own cross-paper performance table reproducible and comparable."],"supporting_citations":[{"why":"Earlier systematic review of hand gesture recognition techniques, challenges, and applications; used as the baseline showing the need for a visual-input-focused update.","marker":"[225]"},{"why":"Computer-vision-based hand gesture recognition review; provides the prior comparison on techniques, camera types, and recognition rates.","marker":"[149]"},{"why":"HCI-focused review of vision-based HGR methods and databases; supplies the earlier HCI-specific evaluation-metric perspective.","marker":"[176]"},{"why":"Structured review of vision-based hand gesture recognition across camera orientations; defines the prior methodological benchmark this survey extends.","marker":"[5]"},{"why":"Systematic review of hand gesture recognition from 2018 to 2024; used to corroborate the transformer trend and to position this survey's narrower visual-input scope.","marker":"[63]"},{"why":"Decade review of human motion recognition covering vision and wearable sensors; serves as a broader-scope comparison point for this survey's focus.","marker":"[142]"},{"why":"Methodological and structural review across diverse data modalities; the paper contrasts its own visual-input-only scope against this multimodal treatment.","marker":"[183]"},{"why":"Survey on vision-based dynamic hand gesture recognition; provides the prior treatment of dynamic gestures and traditional-versus-deep-learning comparison.","marker":"[200]"},{"why":"Topic modeling of research fields; supplies the topic-modeling methodology the survey applies to group retrieved papers.","marker":"[153]"},{"why":"Historical development of hand gesture recognition; grounds the paper's timeline of vision-based HGR evolution.","marker":"[161]"}],"fun_headline_variants":["Five-axis taxonomy maps 125 hand-gesture papers","Hand gesture survey: classification vs estimation","Survey dissects hand gesture methods and datasets","New taxonomy for visual hand gesture recognition","Hand gesture research: 5 axes, 3 input types"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's claims about trends and associations depend on its literature retrieval pipeline—a crawl of top-venue publication lists plus four database queries, filtered by titles, abstracts, and topic modeling—returning a representative sample of hand gesture recognition research from 2018 to 2025.","fun_headline_variants_meta":{"raw":{"variants":["Five-axis taxonomy maps 125 hand-gesture papers","Hand gesture survey: classification vs estimation","Survey dissects hand gesture methods and datasets","New taxonomy for visual hand gesture recognition","Hand gesture research: 5 axes, 3 input types"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000815,"raw_usage":{"total_tokens":3554,"prompt_tokens":909,"completion_tokens":2645,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":525,"completion_tokens_details":{"reasoning_tokens":2574}},"tokens_in":525,"tokens_out":2645,"duration_ms":22986,"temperature":1.0,"reasoning_tokens":2574,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T17:37:04.918439+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the selection pipeline with an additional broad query such as 'egocentric hand' or 'hand-object interaction' for the same period and same venues, then recompute the reported percentages for input type and recognition method; if the added papers shift those percentages substantially or introduce new frequent topics, the 125-paper corpus is not representative and the associations the survey draws would need re-examination.","supporting_citations":[{"cited_title":"A systematic review on hand gesture recognition techniques, challenges and applications","cited_arxiv_id":null,"evidence_quote":"Earlier systematic review of hand gesture recognition techniques, challenges, and applications; used as the baseline showing the need for a visual-input-focused update."},{"cited_title":"Methods, databases and recent advancement of vision-based hand gesture recognition for hci systems: A review","cited_arxiv_id":null,"evidence_quote":"HCI-focused review of vision-based HGR methods and databases; supplies the earlier HCI-specific evaluation-metric perspective."},{"cited_title":"A methodological and structural review of hand gesture recognition across diverse data modalities","cited_arxiv_id":null,"evidence_quote":"Methodological and structural review across diverse data modalities; the paper contrasts its own visual-input-only scope against this multimodal treatment."},{"cited_title":"Survey on vision-based dynamic hand gesture recognition","cited_arxiv_id":null,"evidence_quote":"Survey on vision-based dynamic hand gesture recognition; provides the prior treatment of dynamic gestures and traditional-versus-deep-learning comparison."}],"review_version":1}