REVIEW 3 major objections 6 minor 1 cited by
Isharah: A Large-Scale Multi-Scene Dataset for Continuous Sign Language Recognition
T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Isharah is the first large-scale continuous sign language dataset recorded with smartphone cameras in unconstrained settings.
desk verdict A genuinely useful new CSLR dataset for Saudi Sign Language, but the paper's own statistics need a careful audit before the benchmarks can be trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the data-construction pipeline. A single SSL expert recorded reference videos for all 2,000 sentences; 18 signers then watched those references and recorded batches of roughly 50 sentences per video with their own smartphones; clips were manually segmented per sentence with LosslessCut; and gloss annotations were produced once for the 2,000 reference sentences, then propagated to the 30,000 clips after verification that signers followed the reference ordering. The design makes annotation feasible but concentrates correctness in the segmentation and propagation steps.
What would settle it
Take a random sample of, say, 300 clips from Isharah-2000, have two independent SSL experts re-segment the batch videos and re-annotate the glosses without seeing the released labels, and compute agreement on clip boundaries and gloss sequences. If agreement is low, or if the propagated labels systematically diverge from the re-annotations, the benchmark numbers and the dataset's value are directly affected.
Extended reading notes
Core claim
The discovery, on the paper's own terms, is a dataset and benchmark suite: Isharah contains 30,000 sentence-level clips spanning 2,000 unique sentences, a gloss vocabulary of 1,132 and an Arabic word vocabulary of 2,700, with every clip gloss-annotated and translated. It is the first CSLR dataset of this scale recorded with signers' smartphone cameras in unconstrained settings, yielding wide variation in resolution, orientation, distance, background, and signing style. The paper argues this variation is exactly what CSLR systems lack from controlled corpora such as Phoenix2014 or CSL-Daily, and benchmarks show the task is hard: best signer-independent WER on Isharah-2000 is 27.4% and best unseen-sentence WER is 38.8%, with gloss-based SLT outperforming gloss-free SLT. The authors also report qualitative error patterns dominated by deletions in signer-independent settings and substitutions in unseen sentences.
Load-bearing premise
The dataset's reliability depends on the manual segmentation of each batch video into sentence clips and on the assumption that signers' performances match the reference videos closely enough that gloss labels copied from the references are correct; the paper asserts these were verified but gives no quantitative measure such as inter-annotator agreement or boundary accuracy.
Editorial extensions
If this is right
- Isharah gives CSLR and SLT researchers a common benchmark on which signer-independent and unseen-sentence generalization can be compared for Saudi Sign Language.
- Because Isharah-500, Isharah-1000, and Isharah-2000 share dev/test sentences in the unseen-sentence setup, results across subsets isolate the effect of training-set size and linguistic diversity on model performance.
- The 300 cross-lingual sentences translated from CSL-Daily, SIGNUM, and Continuous GrSL provide a direct way to study sign language differences and multilingual transfer.
- The reported WER gap between signer-independent and unseen-sentence settings shows that memorization of sentence patterns is a measurable failure mode, and the qualitative errors suggest augmentation with zooming and occlusion handling as concrete next steps.
Reading between the lines
- Because the reference-video propagation design makes annotation quality hinge on verification rather than independent labeling, a natural extension the paper does not report is publishing inter-annotator agreement or boundary statistics alongside the release; adding those numbers would let users weight the benchmark results accordingly.
- The selfie-mode recording direction mentioned in the paper could be tested immediately by acquiring a small smartphone selfie subset and measuring how much WER shifts for models trained on the frontal-camera Isharah clips.
- If temporal boundaries were added to the existing gloss annotations, the same 30,000 clips could support sign spotting and localization tasks without new video collection, which would broaden the dataset's use beyond recognition and translation.
- The observed gap in error rates between deaf signers and interpreters, if it holds up, implies that signer demographics and fluency are confounds that future benchmarks should report explicitly.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Isharah, a continuous sign language recognition (CSLR) dataset for Saudi Sign Language consisting of 30,000 video clips recorded by 18 signers using smartphone cameras in uncontrolled environments. It provides gloss-level annotations and Arabic translations for all clips, and organizes the corpus into Isharah-500, Isharah-1000, and Isharah-2000 subsets. The paper defines signer-independent and unseen-sentence CSLR tasks, as well as gloss-based and gloss-free sign language translation (SLT) tasks, and reports benchmark results for seven CSLR and two SLT models. The central claim is that Isharah is the first large-scale, multi-scene CSLR dataset collected with smartphone cameras and that its annotations are reliable enough to support these benchmarks.
Significance. If the dataset's annotations and split statistics are validated, Isharah would be a timely and valuable resource: it is the first large-scale CSLR corpus for Saudi Sign Language, it is collected in unconstrained settings with substantial variation in backgrounds, cameras, and signers, and it includes bilingual annotations (gloss plus Arabic) that support both CSLR and SLT. The public release pledge, the explicit signer-independent and unseen-sentence evaluation protocols, and the breadth of benchmarked models (VAC, SMKD, TLP, SEN, CorrNet, Swin-MSTP, SlowFastSign, MMTLB, GFSLT-VLP) are concrete strengths. The qualitative error analysis (deletion-dominated errors, signer-speed effects, out-of-frame signs) adds practical value. However, the dataset's utility rests on the correctness of manual segmentation and of gloss labels propagated from 2,000 reference videos; the paper currently does not supply the quantitative evidence needed to establish that reliability.
major comments (3)
- [Table 2] Table 2: the reported video statistics are internally inconsistent. For example, the signer-independent training split of Isharah-1000 contains 10,000 videos, 15.81 hours, and 5,718,897 frames (an average of roughly 572 frames per video, corresponding to about 52.9 hours at 30 fps), whereas Isharah-2000, with 20,000 videos and 28.76 hours, lists only 3,820,660 frames (about 191 frames per video, about 35.4 hours at 30 fps). The 2000-set therefore has fewer training frames than the 1000-set despite twice as many videos, and the frames-per-hour ratios differ by a factor of roughly three. The unseen-sentence rows show the same anomaly (Isharah-1000: 10,000 videos and 6,679,388 frames; Isharah-2000: 13,500 videos and 5,461,292 frames). These numbers must be corrected, or the frame-rate and duration conventions must be stated explicitly, because the split statistics are part of the dataset's central description.
- [Table 3] Table 3: unique-sentence counts contradict the stated inventory and expose unquantified annotation deviation. Isharah-2000 is described in Section 3.1 as built from 2,000 selected sentences, yet the signer-independent training split lists 2,358 unique sentences, the test split lists 2,340, and the unseen-sentence training split lists 2,903; Isharah-1000 lists 1,136 unique training sentences against 1,000 targets. Section 3.5 acknowledges that signers deviated from prescribed sentences and that some videos were individually annotated, so extra strings can arise, but the paper reports no counts of re-recorded versus individually annotated clips, no deviation rate, and no inter-annotator agreement. Without these quantities, the gloss ground truth used in the WER benchmarks is not verifiable, and phrases such as "2,000 unique sentences" in the introduction remain misleading.
- [Sections 3.2 and 3.3] No quantitative validation is reported for the two most labor-intensive annotation steps. The manual segmentation of roughly 50-sentence batches into clips is verified only descriptively ("it was ensured that each sentence video was correctly recorded, segmented, and labeled"), and the two-expert annotation is described as "work[ing] closely" without reporting agreement metrics such as Cohen's kappa or any boundary-error measure. Since the dataset's scale is achieved by propagating annotations from 2,000 reference videos to 30,000 clips, validation of this pipeline is load-bearing. The paper should report at minimum the proportion of clips re-recorded, the proportion individually annotated, the number of annotators per clip, and agreement statistics on a held-out sample.
minor comments (6)
- [Sections 3.1, 3.3, 3.5] Typographical and language issues: Section 3.3 has "processm" for "process"; Section 3.5 has "conversion" where "conversation" is intended; Section 3.1 has "an sign language expert" instead of "a sign language expert".
- [Table 2] In the Isharah-2000 unseen-sentence development row, the frame count "16,3098" appears to be a typo for "163,098"; this should be corrected.
- [Table 3] The table headers contain spacing artifacts ("V ocab. Size", "V ocab Size") and should be reformatted.
- [References] Several reference entries use abbreviated author patterns such as "Junseok et al. Ahn", "Yutong et al. Chen", and "Benjia et al. Zhou"; these should be expanded to full author lists for completeness.
- [References] Reference [20] (ElSabagh et al., "A comprehensive survey on arabic text augmentation") does not appear to be cited anywhere in the body; either cite it or remove it.
- [Figure 6(a)] Figure 6(a) reports an average of 130 frames per video, which is not consistent with Table 2 (for example, Isharah-2000 overall averages about 193 frames per video and Isharah-1000 signer-independent training averages about 572). Please clarify which subset and frame-extraction protocol the figure refers to.
Circularity Check
No significant circularity: Isharah is a dataset construction and benchmarking paper whose reported numbers are empirical measurements, not predictions derived from the authors' own parameters or equations.
full rationale
Isharah is a dataset paper, not a derivation. The paper's central outputs are 30,000 video clips, their gloss and Arabic annotations, split statistics, and baseline benchmark numbers. None of these is a 'prediction' in the sense of a quantity forced by an input parameter fitted by the authors. The annotation pipeline is described as recording reference videos, manual segmentation, verification, and propagation of annotations, with individual annotation for deviating clips; the benchmark WER and BLEU results are empirical measurements of external models against those labels. The apparent inconsistencies in Tables 2 and 3 (for example, unique-sentence counts exceeding the 2,000-sentence inventory, and Isharah-2000 reporting fewer frames than Isharah-1000 for a larger video count) are data-quality or bookkeeping concerns that should be investigated, but they do not constitute circular reasoning under the definition used here, because there is no equation or fitted parameter that is equivalent to the reported result by construction. Self-citations appear in related work and in the Swin-MSTP baseline, but they are not load-bearing premises that force the dataset's existence or the benchmark outcomes. No quoted reduction from an input to an output can be exhibited, so the honest finding is no significant circularity.
Assumptions & free parameters
assumptions (4)
- domain assumption The two SSL experts produced consistent, correct gloss annotations for all 30,000 clips.
- domain assumption Manual segmentation of long batch recordings into sentence clips is accurate enough for sentence-level training.
- domain assumption The 18 signers are fluent and representative enough of Saudi Sign Language for the dataset to support general CSLR/SLT.
- domain assumption Signers' imitation of reference videos preserves the gloss order used for batch annotation.
Cite this review
Pith. "Pith review of Isharah: A Large-Scale Multi-Scene Dataset for Continuous Sign Language Recognition." pith.science (2026). https://pith.science/paper/MBGIDS2R
@misc{pith2026250603615,
author = {Pith},
title = {Pith review of: Isharah: A Large-Scale Multi-Scene Dataset for Continuous Sign Language Recognition},
year = {2026},
howpublished = {\url{https://pith.science/paper/MBGIDS2R}},
note = {Machine review of arXiv:2506.03615}
}
read the original abstract
Current benchmarks for sign language recognition (SLR) focus mainly on isolated SLR, while there are limited datasets for continuous SLR (CSLR), which recognizes sequences of signs in a video. Additionally, existing CSLR datasets are collected in controlled settings, which restricts their effectiveness in building robust real-world CSLR systems. To address these limitations, we present Isharah, a large multi-scene dataset for CSLR. It is the first dataset of its type and size that has been collected in an unconstrained environment using signers' smartphone cameras. This setup resulted in high variations of recording settings, camera distances, angles, and resolutions. This variation helps with developing sign language understanding models capable of handling the variability and complexity of real-world scenarios. The dataset consists of 30,000 video clips performed by 18 deaf and professional signers. Additionally, the dataset is linguistically rich as it provides a gloss-level annotation for all dataset's videos, making it useful for developing CSLR and sign language translation (SLT) systems. This paper also introduces multiple sign language understanding benchmarks, including signer-independent and unseen-sentence CSLR, along with gloss-based and gloss-free SLT. The Isharah dataset is available on https://snalyami.github.io/Isharah_CSLR/.
Figures
Figures from the paper (5 more)
Forward citations
Cited by 1 Pith paper
-
A Signer-Invariant Conformer and Multi-Scale Fusion Transformer for Continuous Sign Language Recognition
The paper reports state-of-the-art WERs of 13.07% (signer-independent) and 47.78% (unseen sentences) on Isharah-1000 using a conformer and a multi-scale fusion transformer.
Reference graph
Works this paper leans on
-
[1]
https://www.who.int/news- room/fact-sheets/detail/deafness-and-hearing-loss
Hearing loss statistics. https://www.who.int/news- room/fact-sheets/detail/deafness-and-hearing-loss. Last visit: Oct. 02, 2024. 1
work page 2024
-
[2]
A Com- prehensive Study on Deep Learning-based Methods for Sign Language Recognition
Nikolas Adaloglou, Theocharis Chatzis, Ilias Papas- tratis, Andreas Stergioulas, and Georgios Th. A Com- prehensive Study on Deep Learning-based Methods for Sign Language Recognition. IEEE Transactions on Multimedia, pages 1–14, 2021. 2, 3, 4, 8
work page 2021
-
[3]
Junseok et al. Ahn. Slowfast network for continuous sign language recognition. In ICASSP, 2024. 3, 6, 9
work page 2024
-
[4]
BSL-1K: Scaling Up Co-articulated Sign Language Recognition Using Mouthing Cues
Samuel Albanie, G ¨ul Varol, Liliane Momeni, Tri- antafyllos Afouras, Joon Son Chung, Neil Fox, and Andrew Zisserman. BSL-1K: Scaling Up Co-articulated Sign Language Recognition Using Mouthing Cues. Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intel- ligence and Lecture Notes in Bioinformatics) , 12356 LNCS:35–53, 2020. 3
work page 2020
-
[5]
Bbc-oxford british sign language dataset
Samuel Albanie, G ¨ul Varol, Liliane Momeni, Hannah Bull, Triantafyllos Afouras, Himel Chowdhury, Neil Fox, Bencie Woll, Rob Cooper, Andrew McParland, et al. Bbc-oxford british sign language dataset. arXiv preprint arXiv:2111.03635, 2021. 3
arXiv 2021
-
[6]
Sarah Alyami and Hamzah Luqman. Swin-mstp: Swin transformer with multi-scale temporal percep- tion for continuous sign language recognition. Neu- rocomputing, 617:129015, 2025. 3, 6, 9
work page 2025
-
[7]
Sarah Alyami, Hamzah Luqman, and Mohammad Hammoudeh. Reviewing 25 years of continuous sign language recognition research: Advances, challenges, and prospects. Information Processing & Manage- ment, 61(5):103774, 2024. 2, 3, 8
work page 2024
-
[8]
PK Athira, CJ Sruthi, and A Lijiya. A signer indepen- dent sign language recognition with co-articulation elimination from live videos: an indian scenario.Jour- nal of King Saud University-Computer and Informa- tion Sciences, 34(3):771–781, 2022. 2
work page 2022
Show all 71 references
-
[9]
Neural Sign Language Translation
Necati Cihan Camgoz, Simon Hadfield, Oscar Koller, Hermann Ney, and Richard Bowden. Neural Sign Language Translation. Proceedings of the IEEE Com- puter Society Conference on Computer Vision and Pat- tern Recognition, pages 7784–7793, 2018. 2, 3, 4
2018
-
[10]
Sign language transformers: Joint end-to-end sign language recognition and trans- lation
Necati Cihan Camg ¨oz, Oscar Koller, Simon Hadfield, and Richard Bowden. Sign language transformers: Joint end-to-end sign language recognition and trans- lation. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recog- nition, pages 10020–1003...
2020
-
[11]
Content4all open research sign language translation datasets
Necati Cihan Camg ¨oz, Ben Saunders, Guillaume Ro- chette, Marco Giovanelli, Giacomo Inches, Robin Nachtrab-Ribback, and Richard Bowden. Content4all open research sign language translation datasets. In 2021 16th IEEE International Conference on Auto- matic Face and Gesture Rec...
2021
-
[12]
Two-stream network for sign language recognition and translation
Yutong Chen, Ronglai Zuo, Fangyun Wei, Yu Wu, Shujie Liu, and Brian Mak. Two-stream network for sign language recognition and translation. Advances in Neural Information Processing Systems, 35:17043– 17056, 2022. 3, 8
2022
-
[13]
Yutong et al. Chen. A simple multi-modality trans- fer learning baseline for sign language translation. In CVPR, 2022. 3, 7, 10
2022
-
[14]
Enhanced elan func- tionality for sign language corpora
Onno Crasborn and Han Sloetjes. Enhanced elan func- tionality for sign language corpora. In 6th Interna- tional Conference on Language Resources and Evalu- ation (LREC 2008)/3rd Workshop on the Representa- tion and Processing of Sign Languages: Construction and Exploitation of...
2008
-
[15]
Spatial–temporal transformer for end-to-end sign language recognition
Zhenchao Cui, Wenbo Zhang, Zhaoxin Li, and Zhaoqi Wang. Spatial–temporal transformer for end-to-end sign language recognition. Complex & Intelligent Sys- tems, pages 1–12, 2023. 3
2023
-
[16]
Enhancing a sign language translation system with vision-based features
Philippe Dreuw, Daniel Stein, and Hermann Ney. Enhancing a sign language translation system with vision-based features. Lecture Notes in Computer Science (including subseries Lecture Notes in Artifi- cial Intelligence and Lecture Notes in Bioinformatics), 5085 LNAI:108–113, 20...
2009
-
[17]
How2sign: a large-scale multimodal dataset for continuous ameri- can sign language
Amanda Duarte, Shruti Palaskar, Lucas Ventura, Deepti Ghadiyaram, Kenneth DeHaan, Florian Metze, Jordi Torres, and Xavier Giro-i Nieto. How2sign: a large-scale multimodal dataset for continuous ameri- can sign language. In Proceedings of the IEEE/CVF conference on computer vis...
2021
-
[18]
A com- prehensive survey and taxonomy of sign language re- search
El-Sayed M El-Alfy and Hamzah Luqman. A com- prehensive survey and taxonomy of sign language re- search. Engineering Applications of Artificial Intelli- gence, 114:105198, 2022. 1
2022
-
[19]
Jumla-qsl-22: A novel qatari sign language continuous dataset
Oussama El Ghoul, Maryam Aziz, and Achraf Oth- man. Jumla-qsl-22: A novel qatari sign language continuous dataset. IEEE Access, 11:112639–112649,
-
[20]
A comprehensive survey on arabic text augmentation: approaches, challenges, and applications
Ahmed Adel ElSabagh, Shahira Shaaban Azab, and Hesham Ahmed Hefny. A comprehensive survey on arabic text augmentation: approaches, challenges, and applications. Neural Computing and Applications , pages 1–34, 2025
2025
-
[21]
Jens Forster, Christoph Schmidt, Thomas Hoyoux, Oscar Koller, Uwe Zelle, Justus Piater, and Hermann Ney. RWTH-PHOENIX-weather: A large vocabulary sign language recognition and translation corpus.Pro- ceedings of the 8th International Conference on Lan- guage Resources and Eval...
2012
-
[22]
Llms are good sign language trans- lators
Jia Gong, Lin Geng Foo, Yixuan He, Hossein Rah- mani, and Jun Liu. Llms are good sign language trans- lators. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 18362–18372, 2024. 3
2024
-
[23]
Distill- ing cross-temporal contexts for continuous sign lan- guage recognition
Leming Guo, Wanli Xue, Qing Guo, Bo Liu, Kaihua Zhang, Tiantian Yuan, and Shengyong Chen. Distill- ing cross-temporal contexts for continuous sign lan- guage recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, pages 10771–10780, 2023. 3
2023
-
[24]
Self- Mutual Distillation Learning for Continuous Sign Language Recognition
Aiming Hao, Yuecong Min, and Xilin Chen. Self- Mutual Distillation Learning for Continuous Sign Language Recognition. Proceedings of the IEEE In- ternational Conference on Computer Vision , pages 11283–11292, 2021. 6, 9
2021
-
[25]
Collaborative Multilingual Continuous Sign Lan- guage Recognition: A Unified Framework
Hezhen Hu, Junfu Pu, Wengang Zhou, and Houqiang Li. Collaborative Multilingual Continuous Sign Lan- guage Recognition: A Unified Framework. IEEE Transactions on Multimedia, 25:7559–7570, 2023. 3
2023
-
[26]
Prior-Aware Cross Modality Aug- mentation Learning for Continuous Sign Language Recognition
Hezhen Hu, Junfu Pu, Wengang Zhou, Hang Fang, and Houqiang Li. Prior-Aware Cross Modality Aug- mentation Learning for Continuous Sign Language Recognition. IEEE Transactions on Multimedia , 26: 593–606, 2024. 3
2024
-
[27]
Temporal lift pooling for continuous sign language recognition
Lianyu Hu, Liqing Gao, Zekang Liu, and Wei Feng. Temporal lift pooling for continuous sign language recognition. In European Conference on Computer Vision, pages 511–527. Springer, 2022. 6, 9
2022
-
[28]
Continuous sign language recognition with correlation network
Lianyu Hu, Liqing Gao, Zekang Liu, and Wei Feng. Continuous sign language recognition with correlation network. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 2529–2539, 2023. 3, 6, 9
2023
-
[29]
Self-emphasizing network for continuous sign lan- guage recognition
Lianyu Hu, Liqing Gao, Zekang Liu, and Wei Feng. Self-emphasizing network for continuous sign lan- guage recognition. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 854–862,
-
[30]
Adabrowse: Adaptive video browser for efficient continuous sign language recognition
Lianyu Hu, Liqing Gao, Zekang Liu, Chi-Man Pun, and Wei Feng. Adabrowse: Adaptive video browser for efficient continuous sign language recognition. In Proceedings of the 31st ACM International Confer- ence on Multimedia, pages 709–718, 2023. 3
2023
-
[31]
Scalable frame resolution for efficient continuous sign language recognition
Lianyu Hu, Liqing Gao, Zekang Liu, and Wei Feng. Scalable frame resolution for efficient continuous sign language recognition. Pattern Recognition , 145: 109903, 2024. 3
2024
-
[32]
Video-based sign language recogni- tion without temporal segmentation
Jie Huang, Wengang Zhou, Qilin Zhang, Houqiang Li, and Weiping Li. Video-based sign language recogni- tion without temporal segmentation. 32nd AAAI Con- ference on Artificial Intelligence, AAAI 2018 , pages 2257–2264, 2018. 2, 3, 4
2018
-
[33]
Publishing DGS corpus data: Different Formats for Different Needs
Elena Jahn, Reiner Konrad, Gabriele Langer, Sven Wagner, and Thomas Hanke. Publishing DGS corpus data: Different Formats for Different Needs. pages 83–90, 2018. 2
2018
-
[34]
Signing outside the studio: Benchmarking background robust- ness for continuous sign language recognition
Youngjoon Jang, Youngtaek Oh, Jae Won Cho, Dong- Jin Kim, Joon Son Chung, and In So Kweon. Signing outside the studio: Benchmarking background robust- ness for continuous sign language recognition. arXiv preprint arXiv:2211.00448, 2022. 3
2022 arXiv
-
[35]
Self-sufficient framework for contin- uous sign language recognition
Youngjoon Jang, Youngtaek Oh, Jae Won Cho, Myungchul Kim, Dong-Jin Kim, In So Kweon, and Joon Son Chung. Self-sufficient framework for contin- uous sign language recognition. InICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) ,...
2023
-
[36]
Cosign: Exploring co-occurrence signals in skeleton-based continuous sign language recognition
Peiqi Jiao, Yuecong Min, Yanan Li, Xiaotao Wang, Lei Lei, and Xilin Chen. Cosign: Exploring co-occurrence signals in skeleton-based continuous sign language recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vi- sion (ICCV), pages 20676–20686, 2023. 3
2023
-
[37]
CoSign: Exploring Co-occurrence Signals in Skeleton-based Continuous Sign Language Recognition
Peiqi Jiao, Yuecong Min, Yanan Li, Xiaotao Wang, Lei Lei, and Xilin Chen. CoSign: Exploring Co-occurrence Signals in Skeleton-based Continuous Sign Language Recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vi- sion, pages 20676–20686, 2023. 3
2023
-
[38]
TheRuSLan: Database of Russian sign language
Ildar Kagirov, Denis Ivanko, Dmitry Ryumin, Alexan- der Axyonov, and Alexey Karpov. TheRuSLan: Database of Russian sign language. LREC 2020 - 12th International Conference on Language Resources and Evaluation, Conference Proceedings , (May):6079– 6085, 2020. 2, 3
2020
-
[39]
Contin- uous sign language recognition: Towards large vocab- ulary statistical recognition systems handling multiple signers
Oscar Koller, Jens Forster, and Hermann Ney. Contin- uous sign language recognition: Towards large vocab- ulary statistical recognition systems handling multiple signers. Computer Vision and Image Understanding, 141:108–125, 2015. 2, 3, 4
2015
-
[40]
Sign boundary and hand articulation feature recognition in sign language videos
Ioannis Koulierakis, Georgios Siolas, Eleni Efthimiou, Stavroula-Evita Fotinea, and Andreas- Georgios Stafylopatis. Sign boundary and hand articulation feature recognition in sign language videos. Machine Translation, 35(3):323–343, 2021. 2
2021
-
[41]
Isolated sign lan- guage recognition based on tree structure skeleton im- ages
David Laines, Miguel Gonzalez-Mendoza, Gilberto Ochoa-Ruiz, and Gissella Bejarano. Isolated sign lan- guage recognition based on tree structure skeleton im- ages. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, pages 276–2...
2023
-
[42]
Word-level deep sign language recognition from video: A new large-scale dataset and methods comparison
Dongxu Li, Cristian Rodriguez, Xin Yu, and Hong- dong Li. Word-level deep sign language recognition from video: A new large-scale dataset and methods comparison. In Proceedings of the IEEE/CVF winter conference on applications of computer vision , pages 1459–1469, 2020. 2
2020
-
[43]
Tsp- net: hierarchical feature learning via temporal seman- tic pyramid for sign language translation
Dongxu Li, Chenchen Xu, Xin Yu, Kaihao Zhang, Ben Swift, Hanna Suominen, and Hongdong Li. Tsp- net: hierarchical feature learning via temporal seman- tic pyramid for sign language translation. In Pro- ceedings of the 34th International Conference on Neu- ral Information Proces...
2020
-
[44]
Multi-view spatial- temporal network for continuous sign language recog- nition
Ronghui Li and Lu Meng. Multi-view spatial- temporal network for continuous sign language recog- nition. arXiv preprint arXiv:2204.08747, 2022. 3
2022 arXiv
-
[45]
Arabsign: A multi-modality dataset and benchmark for continuous arabic sign language recognition
Hamzah Luqman. Arabsign: A multi-modality dataset and benchmark for continuous arabic sign language recognition. In 2023 IEEE 17th International Con- ference on Automatic Face and Gesture Recognition (FG), pages 1–8. IEEE, 2023. 2, 3, 4
2023
-
[46]
Automatic translation of arabic text-to-arabic sign language.Uni- versal Access in the Information Society , 18(4):939– 951, 2019
Hamzah Luqman and Sabri A Mahmoud. Automatic translation of arabic text-to-arabic sign language.Uni- versal Access in the Information Society , 18(4):939– 951, 2019. 1
2019
-
[47]
Visual alignment constraint for continuous sign language recognition
Yuecong Min, Aiming Hao, Xiujuan Chai, and Xilin Chen. Visual alignment constraint for continuous sign language recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vi- sion, pages 11542–11551, 2021. 6, 9
2021
-
[48]
FluentSigners-50: A signer independent benchmark dataset for sign language pro- cessing
Medet Mukushev, Aidyn Ubingazhibov, Aigerim Ky- dyrbekova, Alfarabi Imashev, Vadim Kimmelman, and Anara Sandygulova. FluentSigners-50: A signer independent benchmark dataset for sign language pro- cessing. PLoS ONE, 17(9 September):1–19, 2022. 2, 4
2022
-
[49]
A Hong Kong Sign Language corpus collected from sign-interpreted TV news
Zhe Niu, Ronglai Zuo, Brian Mak, and Fangyun Wei. A Hong Kong Sign Language corpus collected from sign-interpreted TV news. In Proceedings of the 2024 Joint International Conference on Computational Lin- guistics, Language Resources and Evaluation (LREC- COLING 2024), pages 63...
2024
-
[50]
Dilated convolutional network with iterative optimization for continuous sign language recognition
Junfu Pu, Wengang Zhou, and Houqiang Li. Dilated convolutional network with iterative optimization for continuous sign language recognition. IJCAI Inter- national Joint Conference on Artificial Intelligence , 2018-July:885–891, 2018. 3
2018
-
[51]
Itera- tive alignment network for continuous sign language recognition
Junfu Pu, Wengang Zhou, and Houqiang Li. Itera- tive alignment network for continuous sign language recognition. Proceedings of the IEEE Computer So- ciety Conference on Computer Vision and Pattern Recognition, 2019-June:4160–4169, 2019. 3
2019
-
[52]
LSA64 : An Argentinian Sign Language Dataset
Franco Ronchetti, Facundo Quiroga, and Laura Lan- zarini. LSA64 : An Argentinian Sign Language Dataset. Congreso Argentino de Ciencias de la Com- putacion (CACIC), pages 794–803, 2016. 2
2016
-
[53]
Auslan-daily: Australian sign lan- guage translation for daily communication and news
Xin Shen, Shaozu Yuan, Hongwei Sheng, Heming Du, and Xin Yu. Auslan-daily: Australian sign lan- guage translation for daily communication and news. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track ,
-
[54]
Open-domain sign language trans- lation learned from online video
Bowen Shi, Diane Brentari, Greg Shakhnarovich, and Karen Livescu. Open-domain sign language trans- lation learned from online video. arXiv preprint arXiv:2205.12870, 2022. 3
2022 arXiv
-
[55]
Sidig, Hamzah Luqman, Sabri Mah- moud, and Mohamed Mohandes
Ala Addin I. Sidig, Hamzah Luqman, Sabri Mah- moud, and Mohamed Mohandes. KArSL: Arabic Sign Language Database. ACM Transactions on Asian and Low-Resource Language Information Processing , 20 (1):1–19, 2021. 2
2021
-
[56]
AUTSL: A large scale multi-modal Turkish sign lan- guage dataset and baseline methods
Ozge Mercanoglu Sincan and Hacer Yalim Keles. AUTSL: A large scale multi-modal Turkish sign lan- guage dataset and baseline methods. IEEE Access, 8: 181340–181355, 2020. 2
2020
-
[57]
Towards a Video Corpus for Signer-Independent Continuous Sign Language Recognition
Ulrich V on Agris and Karl-Friedrich Kraiss. Towards a Video Corpus for Signer-Independent Continuous Sign Language Recognition. The 7th International Workshop on Gesture in Human-Computer Interaction and Simulation, GW 2007, pages 1–6, 2007. 2, 3, 4
2007
-
[58]
Improving continuous sign language recognition with cross-lingual signs
Fangyun Wei and Yutong Chen. Improving continuous sign language recognition with cross-lingual signs. In Proceedings of the IEEE/CVF International Confer- ence on Computer Vision, pages 23612–23621, 2023. 3
2023
-
[59]
Sign2GPT: Leveraging large language models for gloss-free sign language translation
Ryan Wong, Necati Cihan Camgoz, and Richard Bow- den. Sign2GPT: Leveraging large language models for gloss-free sign language translation. InThe Twelfth In- ternational Conference on Learning Representations ,
-
[60]
Multi-scale local- temporal similarity fusion for continuous sign lan- guage recognition
Pan Xie, Zhi Cui, Yao Du, Mengyi Zhao, Jianwei Cui, Bin Wang, and Xiaohui Hu. Multi-scale local- temporal similarity fusion for continuous sign lan- guage recognition. Pattern Recognition, 136, 2023. 3
2023
-
[61]
SF-Net: Structured Feature Network for Continuous Sign Language Recognition
Zhaoyang Yang, Zhenmei Shi, Xiaoyong Shen, and Yu-Wing Tai. SF-Net: Structured Feature Network for Continuous Sign Language Recognition. 2019. 3
2019
-
[62]
Improving gloss-free sign language translation by reducing representation density
Jinhui Ye, Xing Wang, Wenxiang Jiao, Junwei Liang, and Hui Xiong. Improving gloss-free sign language translation by reducing representation density. Ad- vances in Neural Information Processing Systems, 37: 107379–107402, 2025. 3
2025
-
[63]
Gloss attention for gloss-free sign language translation
Aoxiong Yin, Tianyun Zhong, Li Tang, Weike Jin, Tao Jin, and Zhou Zhao. Gloss attention for gloss-free sign language translation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recogni- tion, pages 2551–2562, 2023. 3
2023
-
[64]
C2ST: Cross-modal Contextualized Se- quence Transduction for Continuous Sign Language Recognition
Huaiwen Zhang, Zihang Guo, Yang Yang, Xin Liu, and De Hu. C2ST: Cross-modal Contextualized Se- quence Transduction for Continuous Sign Language Recognition. 2023 IEEE/CVF International Con- ference on Computer Vision (ICCV) , pages 20996– 21005, 2023. 3
2023
-
[65]
Conditional sentence generation and cross-modal reranking for sign lan- guage translation
Jian Zhao, Weizhen Qi, Wengang Zhou, Nan Duan, Ming Zhou, and Houqiang Li. Conditional sentence generation and cross-modal reranking for sign lan- guage translation. IEEE Transactions on Multimedia, 24:2662–2672, 2022. 3
2022
-
[66]
Cvt-slr: Contrastive visual-textual transformation for sign lan- guage recognition with variational alignment
Jiangbin Zheng, Yile Wang, Cheng Tan, Siyuan Li, Ge Wang, Jun Xia, Yidong Chen, and Stan Z Li. Cvt-slr: Contrastive visual-textual transformation for sign lan- guage recognition with variational alignment. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Patt...
-
[67]
Benjia et al. Zhou. Gloss-free sign language transla- tion: Improving from visual-language pretraining. In CVPR, 2023. 2, 3, 7, 8, 10
2023
-
[68]
H. Zhou, W. Zhou, W. Qi, J. Pu, and H. Li. Improv- ing sign language translation with monolingual data by sign back-translation. In 2021 IEEE/CVF Con- ference on Computer Vision and Pattern Recognition (CVPR), pages 1316–1325, Los Alamitos, CA, USA,
2021
-
[69]
Spatial-temporal multi-cue network for sign lan- guage recognition and translation
Hao Zhou, Wengang Zhou, Yun Zhou, and Houqiang Li. Spatial-temporal multi-cue network for sign lan- guage recognition and translation. IEEE Transactions on Multimedia, 24:768–779, 2021. 3
2021
-
[70]
Improving continuous sign language recognition with consistency constraints and signer removal
Ronglai Zuo and Brian Mak. Improving continuous sign language recognition with consistency constraints and signer removal. ACM Transactions on Multimedia Computing, Communications and Applications, 2024. 3
2024
-
[2021]
2, 3, 4, 6, 8
IEEE Computer Society. 2, 3, 4, 6, 8
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.