{"id":"d22eb5ac-61e3-42d2-834a-a1afa5046a89","arxiv_id":"2608.10588","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A 160-class, 144,000-image HamNoSys-grounded handshape benchmark is introduced with four baselines; leave-one-subject-out accuracy falls to about 45%.","lead":"This paper describes a new 144,000-image benchmark of 160 handshape classes defined by the HamNoSys sign-language notation chart and reports four baseline models on it. The results show a large performance drop when models are evaluated on signers not seen in training, from about 86% down to 45% top-1 accuracy.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 160-class label inventory rests on a single non-specialist operator's reading of the HamNoSys chart; without a specialist audit the benchmark's central claim is unverified.","rationale":"The reader identified the validity and unambiguity of the 160 labels as the weakest assumption, and my reading converges on the same point. The paper is carefully documented and internally consistent: the arithmetic is correct, the splits are described precisely, the LOSO protocol prevents participant leakage, and the limitations section is honest. I found no mathematical or experimental error that would invalidate the reported numbers conditional on the labels being correct. However, the central claim that this is a 'HamNoSys-grounded' benchmark depends on the 160 classes corresponding to genuinely distinct chart illustrations and on the collected images correctly realising those classes. Neither condition is independently verified. The operator is described as having sign-language experience, not as a HamNoSys annotation specialist, and no agreement statistics are provided. This is not an accusation of carelessness; it is a statement that the load-bearing part of the resource is currently supported only by the authors' internal process. A specialist audit would settle the question. If the audit reveals systematic label ambiguity, the benchmark's utility is substantially weakened, but the appropriate verdict remains conditional rather than reject because the underlying data may still be salvageable through label correction. I therefore recommend no change to the reader's CONDITIONAL verdict. The secondary concern about MediaPipe-based sample selection is real but less central: all four baselines use the same 139,199-image subset, so model comparisons remain fair, and the per-class imbalance in that subset (812-900 images per class) is modest. Still, reporting per-class detection failure rates would strengthen the claim that the common modelling subset is representative.","tokens_in":14546,"tokens_out":4939,"duration_ms":49899,"concrete_test":"Obtain the class-mapping file and a stratified sample of 20 images per class (3,200 images) from the retained test partitions. Have two HamNoSys specialists independently (a) map each referenced chart illustration to its dataset class ID and (b) label the sampled images. Compute Cohen's kappa between the specialists and against the dataset labels. If specialist-chart agreement or specialist-dataset agreement falls below a pre-registered threshold (e.g., kappa < 0.8), the inventory is not reproducible as specified. In addition, compute per-class MediaPipe detection failure rates on the complete 144,000-image set; if failure rates are strongly correlated with particular classes, report the common modelling subset's per-class balance and assess whether the 4,801 excluded images bias the benchmark.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that each of the 160 classes corresponds to a distinct illustrated hand model in the official HamNoSys 4 Handshapes Chart and that the collected images faithfully instantiate those classes. The weakest link is the chain from chart illustration to dataset label. Section 3.1 assigns one class per drawn model without reporting any inter-annotator reliability study, and Section 3.2 states only that 'an operator with sign-language experience' verified participant reproductions against the reference. Sign-language experience is not the same as HamNoSys annotation expertise: no specialist audit, no second annotator, and no agreement metric are reported. The paper's own Limitations section concedes that 'further annotation validation by HamNoSys specialists may strengthen the resource.' This matters because every reported accuracy inherits the label inventory. The five most confused pairs in Figure 9, such as FTI_TFO1 versus FTR_TFO1 with an 11.99% pair error rate, could reflect genuine visual similarity or systematic label ambiguity; the reported numbers cannot distinguish these. If the chart illustrations are ambiguous or the operator systematically misread subtle thumb-opposition distinctions, the subject-dependent and LOSO benchmarks are not measuring the intended handshape categories. The dataset is not yet public, so an independent check of the class-mapping file and the chart-to-class correspondence is currently impossible.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a handshape-recognition dataset of 144,000 RGB images collected from 15 participants, with 160 classes derived from the illustrated HamNoSys 4 Handshapes Chart. Four baseline families (ResNet-18, ViT-B/16, a landmark GCN, and XGBoost) are evaluated under a class-stratified subject-dependent frame-level split and a 15-fold leave-one-subject-out protocol, with additional matched-model experiments on LSWH100 and ASL Fingerspelling Dataset A. The central claims are that the resource is balanced and HamNoSys-grounded, that the evaluation protocols separate seen-signer from unseen-signer performance, and that the baselines provide reproducible reference numbers for the new benchmark.","tokens_in":14754,"tokens_out":4167,"duration_ms":43335,"significance":"If the resource is made publicly available and the label inventory is validated, this would fill a real gap: a broad, phonetically defined handshape benchmark with real images and explicit signer-independent evaluation. The paper's strengths include the systematic chart-to-class mapping procedure, the balanced per-class distribution, the clear separation between subject-dependent and LOSO protocols, and the use of four diverse baselines with matched external datasets. The moderately large gap between subject-dependent and LOSO results (e.g., 86.20% vs. 45.22% for ViT-B/16) is a useful and credible finding about the difficulty of unseen-signer handshape recognition. However, the central resource is not currently publicly inspectable, the label inventory rests on a single non-specialist operator with no reported agreement metric, and the subject-dependent split may be inflated by frame-level leakage from the same recording clips.","major_comments":[{"comment":"The 160-class inventory is defined by treating each drawn model in the HamNoSys chart as a class, and Section 3.2 states that participant reproductions were verified only by 'an operator with sign-language experience.' No second annotator, inter-annotator agreement statistic, or HamNoSys-specialist audit is reported; the Limitations section concedes that 'further annotation validation by HamNoSys specialists may strengthen the resource.' This is load-bearing because every reported accuracy is measured against these labels. A systematic misreading of subtle distinctions, such as the thumb-opposition forms that appear as the most confused pairs in Figure 9, would mean the benchmark is not measuring the intended classes. Please add a specialist audit or a documented inter-annotator reliability study, and release the class-mapping file so the chart-to-class correspondence can be independently checked.","section":"§3.1–3.2 and Limitations"},{"comment":"The subject-dependent protocol is a class-stratified random frame-level split applied after retaining every fifth frame from each 10-second clip. Because frames from the same source clip can appear simultaneously in training, validation, and test partitions, models can exploit clip-level appearance (background, lighting, identity, exact hand pose progression) rather than generalizable handshape information. This likely inflates the Table 9 accuracies and weakens the claim that these numbers serve as a 'seen-participant reference.' Please report an additional clip-disjoint or recording-disjoint subject-dependent split, or at least quantify the leakage effect by comparing against a split in which all frames from any given clip are kept in a single partition.","section":"§3.6.1"},{"comment":"The dataset and the supplementary class-mapping file are only available 'upon reasonable request' from the corresponding author, and no public URL or repository is provided. For a dataset paper whose central contribution is the resource itself, this prevents independent verification of the chart-to-class mapping, label quality, and frame provenance, and it undermines the stated goal of providing a 'reproducible resource.' Please make the data, or at minimum the class-mapping file and a substantial annotated sample with metadata, publicly available under a clear license before acceptance; if full public release is impossible for institutional reasons, state the conditions explicitly and provide a persistent mechanism for access.","section":"Data Availability"}],"minor_comments":[{"comment":"There are several typographical issues: the abstract contains 'pro-vide' with a line break, and the body has missing spaces such as 'overviewThe' and 'areproducibleresource'; please run a proofreading pass.","section":"Abstract and body text"},{"comment":"The Limitations section begins with 'Few limitations should be noted'; this should read 'A few limitations should be noted.'","section":"Limitations"},{"comment":"The landmark preprocessing is described as 'preprocessed as in [43]' (the authors' prior ICPR paper). Please specify exactly which preprocessing steps are reused and whether any modifications were made, so that the baselines can be reproduced independently.","section":"§4.1"},{"comment":"The table reports 'Approx. per class' as 609, 130, and 131, but the actual per-class counts vary because 4,801 images were excluded from the modelling subset; consider reporting the min–max range per partition for clarity.","section":"Table 6"},{"comment":"Figures 4 and 7 are dense and small at the current resolution; consider providing higher-resolution versions or vector graphics so the protocol details are legible.","section":"Figures 4 and 7"}],"recommendation":"major_revision","confidential_remarks":"The main concern is that the paper's central contribution is not currently accessible: the dataset is only 'available upon reasonable request' and the label inventory has not been specialist-validated. Both issues are fixable within the manuscript's scope, but they are load-bearing for the claims of a reproducible and phonetically grounded benchmark. I do not see a circularity problem: the reported accuracies are direct measurements, and the reuse of landmark preprocessing from [43] does not determine the results. If the authors provide public access or a detailed audit package and add a clip-disjoint split, the paper would be suitable for acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look. The concrete new thing is the resource: 160 handshape classes read off the official HamNoSys 4 Handshapes Chart, instantiated in 144,000 real RGB images from 15 signers, with both subject-dependent and leave-one-subject-out baselines on four model families. That combination is genuinely absent from the cited literature — LSWH100 is synthetic, the fingerspelling sets are alphabet-bound, and Deep Hand/PHOENIX labels are auxiliary. The documentation is careful: class counts per chart row, naming conventions, explicit protocols, and honest limitations. The LOSO numbers (roughly 45% top-1 for the best models) show that unseen-signer generalisation on a 160-way fine-grained inventory is hard, while the subject-dependent results (84–86%) provide a sensible seen-signer reference. The matched-model runs on LSWH100 and ASL Dataset A are used as context rather than as direct comparisons, which is the right call. Internal numbers are consistent and I did not find a load-bearing experimental error.\n\nThe soft spots are real but mostly addressable. The label inventory is the load-bearing one: each chart drawing becomes a class, but the mapping was done by a single operator with sign-language experience, not by a HamNoSys specialist, and no inter-annotator reliability is reported. The paper itself concedes as much in Limitations. That matters because every accuracy inherits the mapping; the confusion pairs in Figure 9 could reflect genuine visual similarity or systematic labelling ambiguity. Second, the subject-dependent split is frame-level after every-fifth-frame sampling, so frames from the same 10-second clip can appear in both training and test. That leaks temporal redundancy and inflates the seen-signer reference; a clip-disjoint split would be cleaner. Third, the dataset and code are not actually released — 'upon reasonable request' is not a public benchmark. For a resource paper, that is a major gap. None of these sinks the contribution, but they are exactly what a reviewer should push on.\n\nThe paper is for researchers building handshape recognition models or needing a phonetically grounded, signer-aware evaluation set. It deserves a serious referee. I would send it out, expecting revision to add a specialist label audit, a clip-disjoint split, and a real data release. My own verdict would be conditional accept, not reject.","headline":"Genuinely useful benchmark resource; the chart-to-label mapping is the soft spot, and it isn't publicly released.","tokens_in":15305,"tokens_out":3374,"would_cite":true,"duration_ms":28716,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims a balanced HamNoSys-grounded handshape benchmark is reproducible and that generalising to unseen signers, not within-participant recognition, is the main challenge.","keywords":["sign language","handshape recognition","HamNoSys","dataset","leave-one-subject-out evaluation","fine-grained classification","RGB images"],"falsifier":"Have HamNoSys specialists independently label a stratified sample of around 1,000 images from the modelling subset using the original chart illustrations without seeing the dataset's class codes, and measure agreement with those codes; if per-class agreement is well below the internal verification standard, label noise is large enough to change the reported accuracy gaps.","tokens_in":14358,"feed_emoji":"🖐","tokens_out":9834,"duration_ms":83668,"temperature":0.7,"pith_summary":"The paper claims that a balanced, real-image benchmark for fine-grained handshape recognition can be grounded in the official HamNoSys 4 Handshapes Chart, and that such a benchmark is needed because existing resources either use language-specific fingerspelling alphabets, synthetic images, or weak corpus-derived labels. To establish this, the authors collected 144,000 RGB images from 15 participants performing 160 chart-defined handshape classes and evaluated four baseline model families under both a participant-overlapping split and a leave-one-subject-out protocol. The central result is that seen-participant accuracy reaches about 86 percent at best, but accuracy on unseen participants drops to roughly 45 percent, identifying signer-independent generalisation as the main open problem. The resource is positioned as a reproducible reference for phonology-grounded sign-language transcription and recognition research.","feed_headline":"Unseen signers cut handshape-model accuracy to 45%","feed_subtitle":"A 144,000-image HamNoSys benchmark for 160 handshapes shows signer-independent recognition is the open problem.","key_machinery":"The load-bearing object is the official, non-exhaustive HamNoSys 4 Handshapes Chart: a publicly illustrated reference inventory of hand models used to define a bounded set of 160 classes, with each drawn hand treated as one class and blank or symbol-only cells excluded. Around this, the acquisition pipeline records 10-second clips with a prescribed hand rotation and extracts every fifth frame to yield 60 images per class per participant, giving 144,000 images. The evaluation machinery is the pairing of a class-stratified frame-level split with a 15-fold leave-one-subject-out protocol, which separates seen-participant reference performance from true signer-independent generalisation. A hand-landmark extractor from the literature supplies both the hand crops for the appearance models and the 21-point landmark graphs used by the landmark-based models.","core_discovery":"The paper's central claim is that a balanced, real-image benchmark for fine-grained isolated handshape recognition can be built directly from the official HamNoSys 4 Handshapes Chart, and that this benchmark is needed because existing resources either cover only language-specific fingerspelling alphabets, use synthetic images, or carry weak corpus-derived labels. Treating each distinct illustrated hand model as one class yields 160 classes; a controlled acquisition protocol with operator verification produced 144,000 RGB images from 15 participants. Under a participant-overlapping split, the best model, ViT-B/16, reaches 86.20 percent top-1 accuracy, while under leave-one-subject-out evaluation the same model falls to 45.22 percent, and ResNet-18 falls to 45.38 percent. The conclusion the authors draw is that seen-participant recognition is no longer the limiting factor; unseen-participant generalisation is the principal challenge.","pith_inferences":["An extension the authors leave implicit is that the 160-class inventory could serve as a pretraining task for downstream continuous sign-language recognition, and the per-participant metadata would allow testing whether such pretraining reduces the leave-one-subject-out gap.","The near-balanced class distribution and participant identifiers make this resource usable for fairness and calibration audits across signers, though the paper itself does not perform those analyses.","A testable prediction consistent with the confusion analysis is that synthetic or multi-camera augmentation targeted at the documented close-pair errors would improve unseen-signer accuracy more than generic augmentation.","Because the chart is non-exhaustive, dynamic two-handed transitions are absent; extending the same class-mapping scheme to those forms would be a natural incremental benchmark rather than a redesign."],"forward_implications":["If the central claim is right, the HamNoSys chart can be used as a class inventory without expert transcription, making phonetically defined handshape resources reproducible across labs.","If the central claim is right, participant-overlapping accuracy numbers overstate deployable performance, and leave-one-subject-out evaluation should become the default reporting protocol for signer-independent handshape recognition.","If the central claim is right, the roughly 39-to-41 point drop between the two protocols quantifies how much of the task is signer identity rather than handshape identity.","If the central claim is right, the documented confusion pairs give a concrete error taxonomy for future models to target finger selection, bending, and thumb opposition.","If the central claim is right, matched-model results on LSWH100 and ASL Fingerspelling Dataset A provide external anchors, not claims about which dataset is intrinsically harder."],"supporting_citations":[{"why":"Supplies the official illustrated HamNoSys 4 Handshapes Chart from which all 160 class definitions are derived.","marker":"[19]"},{"why":"Provides the hand-tracking and 21-point landmark extraction used to build the common modelling subset and both landmark-based baselines.","marker":"[41]"},{"why":"Defines the ResNet-18 architecture used as one of the two appearance-based baselines.","marker":"[39]"},{"why":"Defines the ViT-B/16 architecture used as the other appearance-based baseline and the model selected for confusion analysis.","marker":"[40]"},{"why":"Provides the graph convolutional network formulation used to process hand landmarks with anatomical connectivity.","marker":"[42]"},{"why":"Provides the XGBoost gradient-boosted tree baseline operating on landmark-derived feature vectors.","marker":"[44]"},{"why":"Supplies ASL Fingerspelling Dataset A, a real external benchmark used for matched-model comparison.","marker":"[27]"},{"why":"Supplies LSWH100, a synthetic external benchmark used for matched-model comparison.","marker":"[25]"}],"fun_headline_variants":["Signer shift drops handshape AI to 45% accuracy","HamNoSys benchmark: 86% seen, 45% unseen signers","New benchmark exposes signer-independent handshape gap","From 86% to 45%: signer gap in handshape recognition"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that every drawn hand model in the HamNoSys 4 Handshapes Chart is an unambiguous handshape class and that the operator's reference-guided verification ensured participants reproduced those classes, so the reported accuracies stand or fall with that label validity.","fun_headline_variants_meta":{"raw":{"variants":["Signer shift drops handshape AI to 45% accuracy","HamNoSys benchmark: 86% seen, 45% unseen signers","New benchmark exposes signer-independent handshape gap","From 86% to 45%: signer gap in handshape recognition"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000656,"raw_usage":{"total_tokens":3028,"prompt_tokens":991,"completion_tokens":2037,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":607,"completion_tokens_details":{"reasoning_tokens":1961}},"tokens_in":607,"tokens_out":2037,"duration_ms":13501,"temperature":1.0,"reasoning_tokens":1961,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T21:22:56.938871+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have HamNoSys specialists independently label a stratified sample of around 1,000 images from the modelling subset using the original chart illustrations without seeing the dataset's class codes, and measure agreement with those codes; if per-class agreement is well below the internal verification standard, label noise is large enough to change the reported accuracy gaps.","supporting_citations":[{"cited_title":"HamNoSys 4 handshapes chart","cited_arxiv_id":null,"evidence_quote":"Supplies the official illustrated HamNoSys 4 Handshapes Chart from which all 160 class definitions are derived."},{"cited_title":"& Sun, J","cited_arxiv_id":null,"evidence_quote":"Defines the ResNet-18 architecture used as one of the two appearance-based baselines."},{"cited_title":"InInternational Conference on Learning Representations(2021)","cited_arxiv_id":null,"evidence_quote":"Defines the ViT-B/16 architecture used as the other appearance-based baseline and the model selected for confusion analysis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the graph convolutional network formulation used to process hand landmarks with anatomical connectivity."},{"cited_title":"& Guestrin, C","cited_arxiv_id":null,"evidence_quote":"Provides the XGBoost gradient-boosted tree baseline operating on landmark-derived feature vectors."},{"cited_title":"& Bowden, R","cited_arxiv_id":null,"evidence_quote":"Supplies ASL Fingerspelling Dataset A, a real external benchmark used for matched-model comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies LSWH100, a synthetic external benchmark used for matched-model comparison."}],"review_version":1}