{"id":"1eba27f0-7b59-4584-9027-d502648dffb2","arxiv_id":"2509.03467","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Introduces the first continuous Saudi Sign Language dataset (KAU-CSSL, 85 medical sentence classes) and a ResNet-18 + Transformer + BiLSTM classifier reporting 99.02% signer-dependent and 77.71% signer-independent accuracy.","lead":"This preprint introduces KAU-CSSL, the first continuous Saudi Sign Language dataset of medical sentences, plus a transformer-based video classifier reporting 99.02% signer-dependent and 77.71% signer-independent accuracy. The contribution is mainly a new dataset, but the data are not publicly released and the evaluation protocol has several problems.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"KAU-CSSL's 'continuous' status is unsubstantiated: 83% hearing signers mimic a reference video; the reported accuracies may measure imitation, not natural continuous SSL.","rationale":"I agree with the reader's weakest assumption. The recording methodology and signer demographics make the dataset's authenticity the most load-bearing condition. It is not merely about 'consensus' that deaf signers should be involved; the paper's own procedures allow for non-natural production. The phrase 'continuous' is defined operationally as 'no pauses,' which is insufficient for linguistic continuous signing. The imitative setup—copying a reference video—can suppress co-articulation and other natural features. If the test confirms the videos are unnatural, the 'first continuous SSL dataset' claim is misleading, and the model's near-perfect accuracy likely reflects overfitting to a stylized, low-variability task rather than robust continuous SLR. The reader's conditional verdict is appropriate; I recommend maintaining CONDITIONAL (UNCHANGED) because the concern is already identified and the needed evidence (native-signer validation) is absent. My agreement is 'agree.'","tokens_in":20416,"tokens_out":5453,"duration_ms":58376,"concrete_test":"Randomly sample 100 KAU-CSSL videos stratified by signer and class. Have a panel of at least 3 deaf native SSL signers (or certified interpreters with deaf community membership) independently: (1) transcribe the signed sentence; (2) rate naturalness/fluency on a 1-5 scale; (3) judge whether the signing is 'natural continuous signing' or 'staged/imitative/concatenated isolated signs.' Pre-register: if average naturalness < 3 or >20% of transcriptions disagree with the dataset label, the continuous-SSL validity claim fails. Alternatively, compute co-articulation metrics (e.g., inter-sign transition duration, hand velocity continuity) on these videos and compare with a reference set recorded from deaf native signers; a statistically significant reduction in transition smoothness would confirm the concern.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central contribution is the KAU-CSSL dataset, claimed to be 'the first continuous Saudi Sign Language dataset.' The reported accuracies (99.02% signer-dependent, 77.71% signer-independent) and the entire evaluation rest on this dataset being a valid continuous SSL corpus. Section 3 Phase 1 shows 83.33% of signers are hearing; Phase 3 says participants watched an expert translator's reference video 'to ensure they used the exact gestures and followed the same sequence'; Phase 4's acceptance criteria only check correct sign order, no pauses, and no extraneous signs. Nothing verifies that the resulting productions exhibit natural continuous-signing phenomena (co-articulation, movement epenthesis, rhythm, non-manual markers). Because most signers are non-deaf and possibly non-expert, the videos may be imitative, staged, or effectively isolated signs performed in sequence without pauses. If so, the 'continuous' claim is unsupported, and the accuracies are measures of a potentially artificial task, not of continuous SSL recognition. This is the single load-bearing assumption: if false, both headline results are uninterpretable and the dataset's novelty evaporates.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces KAU-CSSL, described as the first continuous Saudi Sign Language (SSL) dataset, containing videos of 85 medical-related sentence classes. It also proposes KAU-SignTransformer, a video-classification model composed of a pretrained ResNet-18 backbone, a Transformer encoder, and a bidirectional LSTM, reporting 99.02% signer-dependent accuracy and 77.71% signer-independent accuracy. The dataset was collected from 24 signers who imitated reference videos recorded by an expert translator; 83.33% of signers are hearing. The authors claim the dataset fills a gap in Arabic SL resources and that the model demonstrates effective continuous SSL recognition.","tokens_in":20773,"tokens_out":4419,"duration_ms":42733,"significance":"If the KAU-CSSL dataset is a valid continuous SSL corpus, it would fill a recognized gap in Arabic sign language resources, particularly for medical communication. The dataset is the paper's main contribution, and the reported accuracies are secondary to that contribution. The ablation study is useful, but it does not compensate for weaknesses in dataset validation and evaluation protocol. The model architecture is standard, so the significance hinges on whether the dataset is a genuine, reproducible, and accessible continuous signing corpus.","major_comments":[{"comment":"Dataset statistics are inconsistent across the manuscript. The abstract and Table 1 report 5,810 videos; §4.1 reports 5,879 videos; §5 reports training/validation/test sizes of 4,085/827/940 (total 5,852), while §4 reports 4,088/870/921 (total 5,879). Additionally, §3 states 24 signers performed each of 85 sentences three times, which would yield 6,120 videos, not 5,810; the per-class max/min/average counts (74/63/68) imply a different total again. Since the dataset is the central contribution, the correct counts and split definitions must be reconciled and precisely stated.","section":"§3, §4.1, §5"},{"comment":"The claim that KAU-CSSL is a 'continuous' SSL dataset is not supported by the collection protocol. Phase 3 says participants watched an expert translator's reference video 'to ensure they used the exact gestures and followed the same sequence.' Phase 4 acceptance criteria check correct sign order, absence of pauses, and absence of extraneous signs, but do not verify natural continuous-signing phenomena such as co-articulation, movement epenthesis, rhythm, or non-manual markers. With 83.33% of signers being hearing, the productions may be imitative and staged rather than natural continuous signing. The paper must provide evidence, e.g., fluency assessment of signers, validation by deaf native signers, or annotation of continuous-signing phenomena, to substantiate the 'continuous' claim.","section":"§3, Phase 3 and Phase 4"},{"comment":"The signer-dependent protocol is likely inflated by near-duplicate same-signer samples. The dataset records each signer performing each sentence three times, but the train/validation/test split is not described in terms of signer overlap. If the same signer's repetitions of the same sentence appear in both training and test, the 99.02% accuracy partly measures memorization of signer-specific idiosyncrasies. Specify whether the split is signer-disjoint, and report a signer-holdout evaluation or per-signer accuracy as a sanity check.","section":"§5.1 and §4"},{"comment":"The model description is internally inconsistent. §4 opens by saying the problem is formulated using 'a pre-trained VideoMAE model,' but §4.2 describes a ResNet-18 backbone with a Transformer encoder and BiLSTM, and the title/abstract call it a 'Vision Transformer.' This is not merely a wording issue; it obscures what architecture was actually evaluated. Correct the description and remove references to VideoMAE if it is not used.","section":"§4 and §4.2"},{"comment":"The signer-independent evaluation lacks essential detail and comparison. The paper does not state how the signer-independent split was constructed, how many signers were held out, or whether any same-signer repetitions leaked into the test set. A single architecture's 77.71% accuracy, without comparison to a standard baseline under the same protocol, is insufficient to support the claim that the model is effective for continuous SSL recognition. Provide the split details and at least one baseline comparison (e.g., I3D, VideoMAE, or CNN-LSTM).","section":"§5.1 and §4.1"}],"minor_comments":[{"comment":"The sentence '85 distinct sentences, which comprise a lexicon of roughly 5,000 words' is implausible given sentences of 3–5 words; likely a numerical error.","section":"§3, Phase 2"},{"comment":"Typo: 'painful_slallowing' should be 'painful_swallowing'.","section":"§4.1"},{"comment":"The text switches between 'validation accuracy of 99.02%' and 'Test Accuracy: 99.02%'; clarify which set the headline number refers to.","section":"§5.1"},{"comment":"References [13] and [16] appear to contain placeholder DOIs; references [5] and [9] seem to be the same article. Please verify and correct.","section":"References"},{"comment":"The dataset is 'available upon request'; for a dataset contribution, consider a public release or a clear access procedure with ethical approval details.","section":"§3, Data Availability"},{"comment":"Numerous typos and grammatical errors (e.g., 'perfromance', 'resourses', 'domnstrate', 'phaze') require proofreading.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The dataset validity is the crux. If the videos are indeed imitative productions by mostly hearing, non-expert signers, the 'continuous' claim is not established and the headlined accuracies are not interpretable. The manuscript would benefit from independent linguistic validation by deaf sign language experts, and the editor may wish to check the reference list for placeholder DOIs."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the KAU-CSSL dataset is a real potential resource for a language that has none, and the paper is worth engaging despite its problems. But the 'continuous' claim and the two headline accuracies are not yet trustworthy: the recording protocol mostly captures hearing non-expert signers imitating a single expert's reference video, and the numbers in the paper do not reconcile.\n\nWhat is new: they are right that no public continuous SSL dataset existed, and a medical-domain sentence-level corpus for Saudi Sign Language is a concrete contribution. The model is a standard ResNet-18 + Transformer + BiLSTM pipeline—no new mechanism, but a reasonable baseline. The ablation study covers the right axes (frames, layers, heads, pretraining, augmentation) and is honestly reported. The attention to signer diversity (gender, attire, niqab) is also a genuine plus.\n\nWhere it gets soft. First, the dataset validity. Section 3 Phase 3 says participants watched an expert's reference video to produce the exact gestures and sequence; Phase 4 acceptance only checks sign order, speed, no pauses, no extraneous signs. With 83% of signers hearing and mostly non-expert, the videos may be accurate imitations of a template rather than natural continuous signing with co-articulation and non-manual markers. That does not kill the dataset as a sentence-level isolated-sequence resource, but it does mean the 'first continuous SSL' claim is overstated. The signer-independent 77.71% is the more honest number; the 99.02% signer-dependent result is likely inflated by near-duplicate videos of the same signer across train/test. The self-citation to their own Efhamni work is fine; that is relevant context, not padding. Second, the reporting is sloppy: the dataset size is 5,810 in the intro/contributions and 5,879 in Section 4.1; the split is 4,088/870/921 in Section 4 and 4,085/827/940 in Section 5; Section 4 also says 'VideoMAE' before describing ResNet-18. These need to be reconciled. Third, 'available upon request' is weak for a dataset paper that is the main contribution.\n\nBottom line: worth a serious referee, not a desk reject. The dataset could be valuable if released and validated with proper splits and a more defensible notion of continuous signing. My recommendation: engage, but send it back for major revision, and ask the authors to publish the data and a clear protocol statement.","headline":"Real dataset potential, but the 'continuous' claim and the headline numbers need serious scrutiny before anyone builds on them.","tokens_in":21195,"tokens_out":3213,"would_cite":false,"duration_ms":34515,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces the first continuous Saudi Sign Language dataset, KAU-CSSL, with 5,810 videos of 85 medical sentences, plus a transformer model reporting 99.02% accuracy on seen signers and 77.71% on unseen signers.","keywords":["Saudi Sign Language","continuous sign language recognition","sentence-level dataset","transformer encoder","ResNet-18","bidirectional LSTM","signer-independent evaluation","medical sign language"],"falsifier":"Record a held-out set of deaf-community signers, not shown the reference video, performing the same 85 sentences in natural conversation; if accuracy on that set falls far below 77.71%, the reported signer-independent result does not measure real-world generalization.","tokens_in":20406,"feed_emoji":"🤟","tokens_out":7429,"duration_ms":73107,"temperature":0.7,"pith_summary":"The paper's goal is to close a resource gap: Arabic sign language research has mostly used isolated signs or a pan-Arab vocabulary, while Saudi Sign Language has had no continuous, sentence-level corpus. The authors introduce KAU-CSSL, a dataset of 5,810 videos in which 24 signers each perform 85 medical sentences three times, with an expert translator's reference video guaranteeing the same sign order and an expert review filtering out flawed takes. On top of this corpus they train KAU-SignTransformer, a model that extracts ResNet-18 frame features, models temporal order with a transformer encoder and bidirectional LSTM, and reports 99.02% accuracy when test signers were seen during training and 77.71% when they were not. If the dataset is accepted as genuine continuous SSL, it supplies a benchmark for sentence-level recognition, a first quantitative target for signer-independent performance, and an ablation showing which model components matter most.","feed_headline":"First continuous Saudi sign language dataset hits 99% accuracy","feed_subtitle":"Sentence-level Saudi signing, 85 medical phrases, and a transformer that also reaches 77.7% on unseen signers.","key_machinery":"The object that carries the argument is the dataset itself: KAU-CSSL is a continuous sign language corpus, defined in contrast to the isolated-word corpora that dominate Arabic sign language research. Its two design controls are (1) a reference-video protocol in which an expert translator pre-records every sentence so that all 24 signers perform the same signs in the same order, and (2) an expert-review gate that discards takes with wrong speed, order, pauses, laughter, or extra signs. Those controls are what make the 85 labels trustworthy enough to train a classifier. The architecture used to demonstrate the corpus is a temporal video classifier: pretrained ResNet-18 per-frame features, lin","core_discovery":"Stated on the paper's own terms: continuous Saudi Sign Language is now a dataset-backed problem. KAU-CSSL contains 5,810 videos of 85 medical-communication sentences, each performed three times by each of 24 signers recruited from deaf, hard-of-hearing, and hearing participants; an expert translator recorded a reference video first, and participants imitated it so that sign order and vocabulary stayed fixed, after which sign-language experts rejected clips that were too slow, wrongly ordered, extraneous, or interrupted by laughter or talk. The recognition model samples 32 frames per video, passes each frame through a pretrained ResNet-18, adds sinusoidal positional encoding, applies a 3-laye","pith_inferences":["A direct next experiment would test accuracy on a separate cohort of deaf signers who did not watch the reference video, since the training signers are 83% hearing and every take imitates a reference clip; that would show whether 77.71% holds outside the studio.","A per-sign evaluation that splits sentences into component signs would reveal whether the 99% figure reflects true sentence recognition or the model latching onto reusable sub-signs such as 'needed', which appear across many of the 85 classes.","The same four-phase protocol could be applied to other under-resourced sign languages or to new SSL domains such as legal and educational settings; adding word-by-word sign labels would turn the corpus from a classification benchmark into a translation benchmark.","A temporal-segmentation or keypoint-augmented version of the model is the obvious stress test for movement epenthesis, since the paper identifies transitions between signs as a core difficulty while the current model pools over uniformly sampled frames."],"forward_implications":["KAU-CSSL provides a common evaluation set for continuous SSL, so future systems can be compared on the same 85 medical sentences and 5,810 videos.","The reported 99.02% signer-dependent accuracy establishes that a transformer-based video classifier can separate these sentence classes almost perfectly when signer identity is not a variable.","The 77.71% signer-independent result gives a first numerical target for generalizing continuous SSL to unseen signers.","The ablation shows pretrained ResNet-18 is the highest-value component, with a 3.47-point accuracy drop when randomly initialized, so transfer learning from generic image features is a sound default for small continuous sign corpora.","Misclassifications concentrate in visually similar and rarer sentences, such as 'oncologist' versus 'pediatrician' and short signs like 'doctor needed', which points to class balancing and finer temporal modeling as the next concrete improvements."],"supporting_citations":[{"why":"Supplies the reference-video recording protocol that KAU-CSSL copies to enforce sign consistency across signers.","marker":"[19]"},{"why":"Establishes that the prior Saudi sign corpus is isolated-sign, defining the gap KAU-CSSL claims to fill.","marker":"[20]"},{"why":"The closest continuous Arabic sign language corpus, with 50 sentences and 6 signers, which KAU-CSSL extends.","marker":"[21]"},{"why":"The published Saudi sign dictionary from which the 85 medical sentences were drawn and validated.","marker":"[30]"},{"why":"Provides the ImageNet normalization statistics and pretrained weights used by the ResNet-18 backbone.","marker":"[31]"},{"why":"Defines the residual network architecture used as the per-frame spatial feature extractor.","marker":"[32]"},{"why":"Defines the transformer encoder architecture used for temporal modeling.","marker":"[33]"}],"fun_headline_variants":["First sentence-level Saudi sign language dataset enables 99% AI accuracy","Saudi Sign Language recognition hits 99% with vision transformer","Transformer reads continuous Saudi sign language at 99% accuracy","New dataset teaches AI continuous Saudi sign language","99% accuracy on Saudi sign language sentences via new dataset"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"Everything depends on the videos being genuine, fluent, correctly labeled Saudi Sign Language sentences rather than imitations of a reference clip, because the model is only learning to match expert-assigned labels.","fun_headline_variants_meta":{"raw":{"variants":["First sentence-level Saudi sign language dataset enables 99% AI accuracy","Saudi Sign Language recognition hits 99% with vision transformer","Transformer reads continuous Saudi sign language at 99% accuracy","New dataset teaches AI continuous Saudi sign language","99% accuracy on Saudi sign language sentences via new dataset"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001514,"raw_usage":{"total_tokens":5946,"prompt_tokens":824,"completion_tokens":5122,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":568,"completion_tokens_details":{"reasoning_tokens":5041}},"tokens_in":568,"tokens_out":5122,"duration_ms":30369,"temperature":1.0,"reasoning_tokens":5041,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T10:52:15.663530+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Record a held-out set of deaf-community signers, not shown the reference video, performing the same 85 sentences in natural conversation; if accuracy on that set falls far below 77.71%, the reported signer-independent result does not measure real-world generalization.","supporting_citations":[{"cited_title":"Duarte, S","cited_arxiv_id":null,"evidence_quote":"Supplies the reference-video recording protocol that KAU-CSSL copies to enforce sign consistency across signers."},{"cited_title":"Al-Hammadi, G","cited_arxiv_id":null,"evidence_quote":"Establishes that the prior Saudi sign corpus is isolated-sign, defining the gap KAU-CSSL claims to fill."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The closest continuous Arabic sign language corpus, with 50 sentences and 6 signers, which KAU-CSSL extends."},{"cited_title":"Guide to Saudi Sign Language","cited_arxiv_id":null,"evidence_quote":"The published Saudi sign dictionary from which the 85 medical sentences were drawn and validated."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the ImageNet normalization statistics and pretrained weights used by the ResNet-18 backbone."},{"cited_title":"Vaswani, N","cited_arxiv_id":null,"evidence_quote":"Defines the transformer encoder architecture used for temporal modeling."}],"review_version":1}