Pith. sign in

REVIEW 3 major objections 6 minor 1 cited by

Isharah: A Large-Scale Multi-Scene Dataset for Continuous Sign Language Recognition

T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Isharah is the first large-scale continuous sign language dataset recorded with smartphone cameras in unconstrained settings.

desk verdict A genuinely useful new CSLR dataset for Saudi Sign Language, but the paper's own statistics need a careful audit before the benchmarks can be trusted. read the letter →

arxiv 2506.03615 v1 pith:MBGIDS2R submitted 2025-06-04 cs.CV

classification cs.CV
keywords SaudiSignLanguagecontinuousrecognitiontranslationglossannotationsmartphonevideodatasetmulti-scenesigner-independentevaluationunseen-sentence
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces Isharah, a dataset of 30,000 continuous sign language video clips of Saudi Sign Language, recorded by 18 fluent signers on their own smartphones in homes and offices rather than in a studio. The authors' central claim is that Isharah is the first large-scale CSLR dataset collected in an unconstrained, multi-scene environment, and that its gloss-level annotations for every clip plus Arabic translations make it a resource for both continuous sign recognition and sign language translation. They support this claim with three nested subsets (Isharah-500, Isharah-1000, Isharah-2000), benchmarks over seven CSLR models and two SLT models, and two evaluation setups: signer-independent and unseen-sentence recognition. A sympathetic reader should care because existing CSLR benchmarks are mostly lab-recorded, and a smartphone-captured dataset with real camera variation is closer to the conditions a deployed sign language system would face.

What carries the argument

The load-bearing mechanism is the data-construction pipeline. A single SSL expert recorded reference videos for all 2,000 sentences; 18 signers then watched those references and recorded batches of roughly 50 sentences per video with their own smartphones; clips were manually segmented per sentence with LosslessCut; and gloss annotations were produced once for the 2,000 reference sentences, then propagated to the 30,000 clips after verification that signers followed the reference ordering. The design makes annotation feasible but concentrates correctness in the segmentation and propagation steps.

What would settle it

Take a random sample of, say, 300 clips from Isharah-2000, have two independent SSL experts re-segment the batch videos and re-annotate the glosses without seeing the released labels, and compute agreement on clip boundaries and gloss sequences. If agreement is low, or if the propagated labels systematically diverge from the re-annotations, the benchmark numbers and the dataset's value are directly affected.

Watch

Extended reading notes

Core claim

The discovery, on the paper's own terms, is a dataset and benchmark suite: Isharah contains 30,000 sentence-level clips spanning 2,000 unique sentences, a gloss vocabulary of 1,132 and an Arabic word vocabulary of 2,700, with every clip gloss-annotated and translated. It is the first CSLR dataset of this scale recorded with signers' smartphone cameras in unconstrained settings, yielding wide variation in resolution, orientation, distance, background, and signing style. The paper argues this variation is exactly what CSLR systems lack from controlled corpora such as Phoenix2014 or CSL-Daily, and benchmarks show the task is hard: best signer-independent WER on Isharah-2000 is 27.4% and best unseen-sentence WER is 38.8%, with gloss-based SLT outperforming gloss-free SLT. The authors also report qualitative error patterns dominated by deletions in signer-independent settings and substitutions in unseen sentences.

Load-bearing premise

The dataset's reliability depends on the manual segmentation of each batch video into sentence clips and on the assumption that signers' performances match the reference videos closely enough that gloss labels copied from the references are correct; the paper asserts these were verified but gives no quantitative measure such as inter-annotator agreement or boundary accuracy.

Editorial extensions

If this is right

  • Isharah gives CSLR and SLT researchers a common benchmark on which signer-independent and unseen-sentence generalization can be compared for Saudi Sign Language.
  • Because Isharah-500, Isharah-1000, and Isharah-2000 share dev/test sentences in the unseen-sentence setup, results across subsets isolate the effect of training-set size and linguistic diversity on model performance.
  • The 300 cross-lingual sentences translated from CSL-Daily, SIGNUM, and Continuous GrSL provide a direct way to study sign language differences and multilingual transfer.
  • The reported WER gap between signer-independent and unseen-sentence settings shows that memorization of sentence patterns is a measurable failure mode, and the qualitative errors suggest augmentation with zooming and occlusion handling as concrete next steps.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the reference-video propagation design makes annotation quality hinge on verification rather than independent labeling, a natural extension the paper does not report is publishing inter-annotator agreement or boundary statistics alongside the release; adding those numbers would let users weight the benchmark results accordingly.
  • The selfie-mode recording direction mentioned in the paper could be tested immediately by acquiring a small smartphone selfie subset and measuring how much WER shifts for models trained on the frontal-camera Isharah clips.
  • If temporal boundaries were added to the existing gloss annotations, the same 30,000 clips could support sign spotting and localization tasks without new video collection, which would broaden the dataset's use beyond recognition and translation.
  • The observed gap in error rates between deaf signers and interpreters, if it holds up, implies that signer demographics and fluency are confounds that future benchmarks should report explicitly.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper introduces Isharah, a continuous sign language recognition (CSLR) dataset for Saudi Sign Language consisting of 30,000 video clips recorded by 18 signers using smartphone cameras in uncontrolled environments. It provides gloss-level annotations and Arabic translations for all clips, and organizes the corpus into Isharah-500, Isharah-1000, and Isharah-2000 subsets. The paper defines signer-independent and unseen-sentence CSLR tasks, as well as gloss-based and gloss-free sign language translation (SLT) tasks, and reports benchmark results for seven CSLR and two SLT models. The central claim is that Isharah is the first large-scale, multi-scene CSLR dataset collected with smartphone cameras and that its annotations are reliable enough to support these benchmarks.

Significance. If the dataset's annotations and split statistics are validated, Isharah would be a timely and valuable resource: it is the first large-scale CSLR corpus for Saudi Sign Language, it is collected in unconstrained settings with substantial variation in backgrounds, cameras, and signers, and it includes bilingual annotations (gloss plus Arabic) that support both CSLR and SLT. The public release pledge, the explicit signer-independent and unseen-sentence evaluation protocols, and the breadth of benchmarked models (VAC, SMKD, TLP, SEN, CorrNet, Swin-MSTP, SlowFastSign, MMTLB, GFSLT-VLP) are concrete strengths. The qualitative error analysis (deletion-dominated errors, signer-speed effects, out-of-frame signs) adds practical value. However, the dataset's utility rests on the correctness of manual segmentation and of gloss labels propagated from 2,000 reference videos; the paper currently does not supply the quantitative evidence needed to establish that reliability.

major comments (3)
  1. [Table 2] Table 2: the reported video statistics are internally inconsistent. For example, the signer-independent training split of Isharah-1000 contains 10,000 videos, 15.81 hours, and 5,718,897 frames (an average of roughly 572 frames per video, corresponding to about 52.9 hours at 30 fps), whereas Isharah-2000, with 20,000 videos and 28.76 hours, lists only 3,820,660 frames (about 191 frames per video, about 35.4 hours at 30 fps). The 2000-set therefore has fewer training frames than the 1000-set despite twice as many videos, and the frames-per-hour ratios differ by a factor of roughly three. The unseen-sentence rows show the same anomaly (Isharah-1000: 10,000 videos and 6,679,388 frames; Isharah-2000: 13,500 videos and 5,461,292 frames). These numbers must be corrected, or the frame-rate and duration conventions must be stated explicitly, because the split statistics are part of the dataset's central description.
  2. [Table 3] Table 3: unique-sentence counts contradict the stated inventory and expose unquantified annotation deviation. Isharah-2000 is described in Section 3.1 as built from 2,000 selected sentences, yet the signer-independent training split lists 2,358 unique sentences, the test split lists 2,340, and the unseen-sentence training split lists 2,903; Isharah-1000 lists 1,136 unique training sentences against 1,000 targets. Section 3.5 acknowledges that signers deviated from prescribed sentences and that some videos were individually annotated, so extra strings can arise, but the paper reports no counts of re-recorded versus individually annotated clips, no deviation rate, and no inter-annotator agreement. Without these quantities, the gloss ground truth used in the WER benchmarks is not verifiable, and phrases such as "2,000 unique sentences" in the introduction remain misleading.
  3. [Sections 3.2 and 3.3] No quantitative validation is reported for the two most labor-intensive annotation steps. The manual segmentation of roughly 50-sentence batches into clips is verified only descriptively ("it was ensured that each sentence video was correctly recorded, segmented, and labeled"), and the two-expert annotation is described as "work[ing] closely" without reporting agreement metrics such as Cohen's kappa or any boundary-error measure. Since the dataset's scale is achieved by propagating annotations from 2,000 reference videos to 30,000 clips, validation of this pipeline is load-bearing. The paper should report at minimum the proportion of clips re-recorded, the proportion individually annotated, the number of annotators per clip, and agreement statistics on a held-out sample.
minor comments (6)
  1. [Sections 3.1, 3.3, 3.5] Typographical and language issues: Section 3.3 has "processm" for "process"; Section 3.5 has "conversion" where "conversation" is intended; Section 3.1 has "an sign language expert" instead of "a sign language expert".
  2. [Table 2] In the Isharah-2000 unseen-sentence development row, the frame count "16,3098" appears to be a typo for "163,098"; this should be corrected.
  3. [Table 3] The table headers contain spacing artifacts ("V ocab. Size", "V ocab Size") and should be reformatted.
  4. [References] Several reference entries use abbreviated author patterns such as "Junseok et al. Ahn", "Yutong et al. Chen", and "Benjia et al. Zhou"; these should be expanded to full author lists for completeness.
  5. [References] Reference [20] (ElSabagh et al., "A comprehensive survey on arabic text augmentation") does not appear to be cited anywhere in the body; either cite it or remove it.
  6. [Figure 6(a)] Figure 6(a) reports an average of 130 frames per video, which is not consistent with Table 2 (for example, Isharah-2000 overall averages about 193 frames per video and Isharah-1000 signer-independent training averages about 572). Please clarify which subset and frame-extraction protocol the figure refers to.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: Isharah is a dataset construction and benchmarking paper whose reported numbers are empirical measurements, not predictions derived from the authors' own parameters or equations.

full rationale

Isharah is a dataset paper, not a derivation. The paper's central outputs are 30,000 video clips, their gloss and Arabic annotations, split statistics, and baseline benchmark numbers. None of these is a 'prediction' in the sense of a quantity forced by an input parameter fitted by the authors. The annotation pipeline is described as recording reference videos, manual segmentation, verification, and propagation of annotations, with individual annotation for deviating clips; the benchmark WER and BLEU results are empirical measurements of external models against those labels. The apparent inconsistencies in Tables 2 and 3 (for example, unique-sentence counts exceeding the 2,000-sentence inventory, and Isharah-2000 reporting fewer frames than Isharah-1000 for a larger video count) are data-quality or bookkeeping concerns that should be investigated, but they do not constitute circular reasoning under the definition used here, because there is no equation or fitted parameter that is equivalent to the reported result by construction. Self-citations appear in related work and in the Swin-MSTP baseline, but they are not load-bearing premises that force the dataset's existence or the benchmark outcomes. No quoted reduction from an input to an output can be exhibited, so the honest finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The dataset's scientific value rests on unmeasured assumptions about annotation consistency, segmentation accuracy, and signer fluency. The paper asserts expert verification but does not quantify it, and the recording protocol encouraged imitation of reference videos, which may reduce natural signing variation.

assumptions (4)
  • domain assumption The two SSL experts produced consistent, correct gloss annotations for all 30,000 clips.
    Section 3.3 describes the annotation process but provides no inter-annotator agreement score or error rate.
  • domain assumption Manual segmentation of long batch recordings into sentence clips is accurate enough for sentence-level training.
    Section 3.2 describes segmentation with LosslessCut and verification, but no boundary accuracy measure is reported.
  • domain assumption The 18 signers are fluent and representative enough of Saudi Sign Language for the dataset to support general CSLR/SLT.
    Section 3.1 includes 11 deaf, 3 hard-of-hearing, and 4 translators; fluency and representativeness are asserted, not measured.
  • domain assumption Signers' imitation of reference videos preserves the gloss order used for batch annotation.
    Section 3.2 says samples were aimed to be as close to references as possible, then gloss annotations were propagated from references; deviations were handled individually.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Isharah: A Large-Scale Multi-Scene Dataset for Continuous Sign Language Recognition." pith.science (2026). https://pith.science/paper/MBGIDS2R

@misc{pith2026250603615,
  author       = {Pith},
  title        = {Pith review of: Isharah: A Large-Scale Multi-Scene Dataset for Continuous Sign Language Recognition},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MBGIDS2R}},
  note         = {Machine review of arXiv:2506.03615}
}
read the original abstract

Current benchmarks for sign language recognition (SLR) focus mainly on isolated SLR, while there are limited datasets for continuous SLR (CSLR), which recognizes sequences of signs in a video. Additionally, existing CSLR datasets are collected in controlled settings, which restricts their effectiveness in building robust real-world CSLR systems. To address these limitations, we present Isharah, a large multi-scene dataset for CSLR. It is the first dataset of its type and size that has been collected in an unconstrained environment using signers' smartphone cameras. This setup resulted in high variations of recording settings, camera distances, angles, and resolutions. This variation helps with developing sign language understanding models capable of handling the variability and complexity of real-world scenarios. The dataset consists of 30,000 video clips performed by 18 deaf and professional signers. Additionally, the dataset is linguistically rich as it provides a gloss-level annotation for all dataset's videos, making it useful for developing CSLR and sign language translation (SLT) systems. This paper also introduces multiple sign language understanding benchmarks, including signer-independent and unseen-sentence CSLR, along with gloss-based and gloss-free SLT. The Isharah dataset is available on https://snalyami.github.io/Isharah_CSLR/.

Figures

Figures reproduced from arXiv: 2506.03615 by the authors.

Figure 1
Figure 1. Samples from the Isharah dataset. (a) The dataset was collected using the frontal camera of different smartphones with different [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Overview of the main stages followed in constructing [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Distribution of video resolutions in Isharah, highlighting [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Annotation tool, developed to facilitate the gloss anno [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Video samples from the Isharah dataset with the corresponding Arabic and gloss annotations. [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: Statistics of the Isharah dataset: (a) Distribution of the number of frames across videos under the signer-independent setting (left) [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 7
Figure 7. Figure 7: Qualitative analysis of SlowFastSign and SMKD CSLR [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 8
Figure 8. Figure 8: Qualitative analysis of the output of gloss-based [PITH_FULL_IMAGE:figures/full_fig_p010_8.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Signer-Invariant Conformer and Multi-Scale Fusion Transformer for Continuous Sign Language Recognition

    cs.CV 2025-08 conditional novelty 5.0 of 10

    The paper reports state-of-the-art WERs of 13.07% (signer-independent) and 47.78% (unseen sentences) on Isharah-1000 using a conformer and a multi-scale fusion transformer.

Reference graph

Works this paper leans on

71 extracted references · 67 canonical work pages · cited by 1 Pith paper

  1. [1]

    https://www.who.int/news- room/fact-sheets/detail/deafness-and-hearing-loss

    Hearing loss statistics. https://www.who.int/news- room/fact-sheets/detail/deafness-and-hearing-loss. Last visit: Oct. 02, 2024. 1

  2. [2]

    A Com- prehensive Study on Deep Learning-based Methods for Sign Language Recognition

    Nikolas Adaloglou, Theocharis Chatzis, Ilias Papas- tratis, Andreas Stergioulas, and Georgios Th. A Com- prehensive Study on Deep Learning-based Methods for Sign Language Recognition. IEEE Transactions on Multimedia, pages 1–14, 2021. 2, 3, 4, 8

  3. [3]

    Junseok et al. Ahn. Slowfast network for continuous sign language recognition. In ICASSP, 2024. 3, 6, 9

  4. [4]

    BSL-1K: Scaling Up Co-articulated Sign Language Recognition Using Mouthing Cues

    Samuel Albanie, G ¨ul Varol, Liliane Momeni, Tri- antafyllos Afouras, Joon Son Chung, Neil Fox, and Andrew Zisserman. BSL-1K: Scaling Up Co-articulated Sign Language Recognition Using Mouthing Cues. Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intel- ligence and Lecture Notes in Bioinformatics) , 12356 LNCS:35–53, 2020. 3

  5. [5]

    Bbc-oxford british sign language dataset

    Samuel Albanie, G ¨ul Varol, Liliane Momeni, Hannah Bull, Triantafyllos Afouras, Himel Chowdhury, Neil Fox, Bencie Woll, Rob Cooper, Andrew McParland, et al. Bbc-oxford british sign language dataset. arXiv preprint arXiv:2111.03635, 2021. 3

  6. [6]

    Swin-mstp: Swin transformer with multi-scale temporal percep- tion for continuous sign language recognition

    Sarah Alyami and Hamzah Luqman. Swin-mstp: Swin transformer with multi-scale temporal percep- tion for continuous sign language recognition. Neu- rocomputing, 617:129015, 2025. 3, 6, 9

  7. [7]

    Reviewing 25 years of continuous sign language recognition research: Advances, challenges, and prospects

    Sarah Alyami, Hamzah Luqman, and Mohammad Hammoudeh. Reviewing 25 years of continuous sign language recognition research: Advances, challenges, and prospects. Information Processing & Manage- ment, 61(5):103774, 2024. 2, 3, 8

  8. [8]

    PK Athira, CJ Sruthi, and A Lijiya. A signer indepen- dent sign language recognition with co-articulation elimination from live videos: an indian scenario.Jour- nal of King Saud University-Computer and Informa- tion Sciences, 34(3):771–781, 2022. 2

Show all 71 references
  1. [9]

    Neural Sign Language Translation

    Necati Cihan Camgoz, Simon Hadfield, Oscar Koller, Hermann Ney, and Richard Bowden. Neural Sign Language Translation. Proceedings of the IEEE Com- puter Society Conference on Computer Vision and Pat- tern Recognition, pages 7784–7793, 2018. 2, 3, 4

  2. [10]

    Sign language transformers: Joint end-to-end sign language recognition and trans- lation

    Necati Cihan Camg ¨oz, Oscar Koller, Simon Hadfield, and Richard Bowden. Sign language transformers: Joint end-to-end sign language recognition and trans- lation. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recog- nition, pages 10020–1003...

  3. [11]

    Content4all open research sign language translation datasets

    Necati Cihan Camg ¨oz, Ben Saunders, Guillaume Ro- chette, Marco Giovanelli, Giacomo Inches, Robin Nachtrab-Ribback, and Richard Bowden. Content4all open research sign language translation datasets. In 2021 16th IEEE International Conference on Auto- matic Face and Gesture Rec...

  4. [12]

    Two-stream network for sign language recognition and translation

    Yutong Chen, Ronglai Zuo, Fangyun Wei, Yu Wu, Shujie Liu, and Brian Mak. Two-stream network for sign language recognition and translation. Advances in Neural Information Processing Systems, 35:17043– 17056, 2022. 3, 8

  5. [13]

    Yutong et al. Chen. A simple multi-modality trans- fer learning baseline for sign language translation. In CVPR, 2022. 3, 7, 10

  6. [14]

    Enhanced elan func- tionality for sign language corpora

    Onno Crasborn and Han Sloetjes. Enhanced elan func- tionality for sign language corpora. In 6th Interna- tional Conference on Language Resources and Evalu- ation (LREC 2008)/3rd Workshop on the Representa- tion and Processing of Sign Languages: Construction and Exploitation of...

  7. [15]

    Spatial–temporal transformer for end-to-end sign language recognition

    Zhenchao Cui, Wenbo Zhang, Zhaoxin Li, and Zhaoqi Wang. Spatial–temporal transformer for end-to-end sign language recognition. Complex & Intelligent Sys- tems, pages 1–12, 2023. 3

  8. [16]

    Enhancing a sign language translation system with vision-based features

    Philippe Dreuw, Daniel Stein, and Hermann Ney. Enhancing a sign language translation system with vision-based features. Lecture Notes in Computer Science (including subseries Lecture Notes in Artifi- cial Intelligence and Lecture Notes in Bioinformatics), 5085 LNAI:108–113, 20...

  9. [17]

    How2sign: a large-scale multimodal dataset for continuous ameri- can sign language

    Amanda Duarte, Shruti Palaskar, Lucas Ventura, Deepti Ghadiyaram, Kenneth DeHaan, Florian Metze, Jordi Torres, and Xavier Giro-i Nieto. How2sign: a large-scale multimodal dataset for continuous ameri- can sign language. In Proceedings of the IEEE/CVF conference on computer vis...

  10. [18]

    A com- prehensive survey and taxonomy of sign language re- search

    El-Sayed M El-Alfy and Hamzah Luqman. A com- prehensive survey and taxonomy of sign language re- search. Engineering Applications of Artificial Intelli- gence, 114:105198, 2022. 1

  11. [19]

    Jumla-qsl-22: A novel qatari sign language continuous dataset

    Oussama El Ghoul, Maryam Aziz, and Achraf Oth- man. Jumla-qsl-22: A novel qatari sign language continuous dataset. IEEE Access, 11:112639–112649,

  12. [20]

    A comprehensive survey on arabic text augmentation: approaches, challenges, and applications

    Ahmed Adel ElSabagh, Shahira Shaaban Azab, and Hesham Ahmed Hefny. A comprehensive survey on arabic text augmentation: approaches, challenges, and applications. Neural Computing and Applications , pages 1–34, 2025

  13. [21]

    Jens Forster, Christoph Schmidt, Thomas Hoyoux, Oscar Koller, Uwe Zelle, Justus Piater, and Hermann Ney. RWTH-PHOENIX-weather: A large vocabulary sign language recognition and translation corpus.Pro- ceedings of the 8th International Conference on Lan- guage Resources and Eval...

  14. [22]

    Llms are good sign language trans- lators

    Jia Gong, Lin Geng Foo, Yixuan He, Hossein Rah- mani, and Jun Liu. Llms are good sign language trans- lators. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 18362–18372, 2024. 3

  15. [23]

    Distill- ing cross-temporal contexts for continuous sign lan- guage recognition

    Leming Guo, Wanli Xue, Qing Guo, Bo Liu, Kaihua Zhang, Tiantian Yuan, and Shengyong Chen. Distill- ing cross-temporal contexts for continuous sign lan- guage recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, pages 10771–10780, 2023. 3

  16. [24]

    Self- Mutual Distillation Learning for Continuous Sign Language Recognition

    Aiming Hao, Yuecong Min, and Xilin Chen. Self- Mutual Distillation Learning for Continuous Sign Language Recognition. Proceedings of the IEEE In- ternational Conference on Computer Vision , pages 11283–11292, 2021. 6, 9

  17. [25]

    Collaborative Multilingual Continuous Sign Lan- guage Recognition: A Unified Framework

    Hezhen Hu, Junfu Pu, Wengang Zhou, and Houqiang Li. Collaborative Multilingual Continuous Sign Lan- guage Recognition: A Unified Framework. IEEE Transactions on Multimedia, 25:7559–7570, 2023. 3

  18. [26]

    Prior-Aware Cross Modality Aug- mentation Learning for Continuous Sign Language Recognition

    Hezhen Hu, Junfu Pu, Wengang Zhou, Hang Fang, and Houqiang Li. Prior-Aware Cross Modality Aug- mentation Learning for Continuous Sign Language Recognition. IEEE Transactions on Multimedia , 26: 593–606, 2024. 3

  19. [27]

    Temporal lift pooling for continuous sign language recognition

    Lianyu Hu, Liqing Gao, Zekang Liu, and Wei Feng. Temporal lift pooling for continuous sign language recognition. In European Conference on Computer Vision, pages 511–527. Springer, 2022. 6, 9

  20. [28]

    Continuous sign language recognition with correlation network

    Lianyu Hu, Liqing Gao, Zekang Liu, and Wei Feng. Continuous sign language recognition with correlation network. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 2529–2539, 2023. 3, 6, 9

  21. [29]

    Self-emphasizing network for continuous sign lan- guage recognition

    Lianyu Hu, Liqing Gao, Zekang Liu, and Wei Feng. Self-emphasizing network for continuous sign lan- guage recognition. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 854–862,

  22. [30]

    Adabrowse: Adaptive video browser for efficient continuous sign language recognition

    Lianyu Hu, Liqing Gao, Zekang Liu, Chi-Man Pun, and Wei Feng. Adabrowse: Adaptive video browser for efficient continuous sign language recognition. In Proceedings of the 31st ACM International Confer- ence on Multimedia, pages 709–718, 2023. 3

  23. [31]

    Scalable frame resolution for efficient continuous sign language recognition

    Lianyu Hu, Liqing Gao, Zekang Liu, and Wei Feng. Scalable frame resolution for efficient continuous sign language recognition. Pattern Recognition , 145: 109903, 2024. 3

  24. [32]

    Video-based sign language recogni- tion without temporal segmentation

    Jie Huang, Wengang Zhou, Qilin Zhang, Houqiang Li, and Weiping Li. Video-based sign language recogni- tion without temporal segmentation. 32nd AAAI Con- ference on Artificial Intelligence, AAAI 2018 , pages 2257–2264, 2018. 2, 3, 4

  25. [33]

    Publishing DGS corpus data: Different Formats for Different Needs

    Elena Jahn, Reiner Konrad, Gabriele Langer, Sven Wagner, and Thomas Hanke. Publishing DGS corpus data: Different Formats for Different Needs. pages 83–90, 2018. 2

  26. [34]

    Signing outside the studio: Benchmarking background robust- ness for continuous sign language recognition

    Youngjoon Jang, Youngtaek Oh, Jae Won Cho, Dong- Jin Kim, Joon Son Chung, and In So Kweon. Signing outside the studio: Benchmarking background robust- ness for continuous sign language recognition. arXiv preprint arXiv:2211.00448, 2022. 3

  27. [35]

    Self-sufficient framework for contin- uous sign language recognition

    Youngjoon Jang, Youngtaek Oh, Jae Won Cho, Myungchul Kim, Dong-Jin Kim, In So Kweon, and Joon Son Chung. Self-sufficient framework for contin- uous sign language recognition. InICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) ,...

  28. [36]

    Cosign: Exploring co-occurrence signals in skeleton-based continuous sign language recognition

    Peiqi Jiao, Yuecong Min, Yanan Li, Xiaotao Wang, Lei Lei, and Xilin Chen. Cosign: Exploring co-occurrence signals in skeleton-based continuous sign language recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vi- sion (ICCV), pages 20676–20686, 2023. 3

  29. [37]

    CoSign: Exploring Co-occurrence Signals in Skeleton-based Continuous Sign Language Recognition

    Peiqi Jiao, Yuecong Min, Yanan Li, Xiaotao Wang, Lei Lei, and Xilin Chen. CoSign: Exploring Co-occurrence Signals in Skeleton-based Continuous Sign Language Recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vi- sion, pages 20676–20686, 2023. 3

  30. [38]

    TheRuSLan: Database of Russian sign language

    Ildar Kagirov, Denis Ivanko, Dmitry Ryumin, Alexan- der Axyonov, and Alexey Karpov. TheRuSLan: Database of Russian sign language. LREC 2020 - 12th International Conference on Language Resources and Evaluation, Conference Proceedings , (May):6079– 6085, 2020. 2, 3

  31. [39]

    Contin- uous sign language recognition: Towards large vocab- ulary statistical recognition systems handling multiple signers

    Oscar Koller, Jens Forster, and Hermann Ney. Contin- uous sign language recognition: Towards large vocab- ulary statistical recognition systems handling multiple signers. Computer Vision and Image Understanding, 141:108–125, 2015. 2, 3, 4

  32. [40]

    Sign boundary and hand articulation feature recognition in sign language videos

    Ioannis Koulierakis, Georgios Siolas, Eleni Efthimiou, Stavroula-Evita Fotinea, and Andreas- Georgios Stafylopatis. Sign boundary and hand articulation feature recognition in sign language videos. Machine Translation, 35(3):323–343, 2021. 2

  33. [41]

    Isolated sign lan- guage recognition based on tree structure skeleton im- ages

    David Laines, Miguel Gonzalez-Mendoza, Gilberto Ochoa-Ruiz, and Gissella Bejarano. Isolated sign lan- guage recognition based on tree structure skeleton im- ages. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, pages 276–2...

  34. [42]

    Word-level deep sign language recognition from video: A new large-scale dataset and methods comparison

    Dongxu Li, Cristian Rodriguez, Xin Yu, and Hong- dong Li. Word-level deep sign language recognition from video: A new large-scale dataset and methods comparison. In Proceedings of the IEEE/CVF winter conference on applications of computer vision , pages 1459–1469, 2020. 2

  35. [43]

    Tsp- net: hierarchical feature learning via temporal seman- tic pyramid for sign language translation

    Dongxu Li, Chenchen Xu, Xin Yu, Kaihao Zhang, Ben Swift, Hanna Suominen, and Hongdong Li. Tsp- net: hierarchical feature learning via temporal seman- tic pyramid for sign language translation. In Pro- ceedings of the 34th International Conference on Neu- ral Information Proces...

  36. [44]

    Multi-view spatial- temporal network for continuous sign language recog- nition

    Ronghui Li and Lu Meng. Multi-view spatial- temporal network for continuous sign language recog- nition. arXiv preprint arXiv:2204.08747, 2022. 3

  37. [45]

    Arabsign: A multi-modality dataset and benchmark for continuous arabic sign language recognition

    Hamzah Luqman. Arabsign: A multi-modality dataset and benchmark for continuous arabic sign language recognition. In 2023 IEEE 17th International Con- ference on Automatic Face and Gesture Recognition (FG), pages 1–8. IEEE, 2023. 2, 3, 4

  38. [46]

    Automatic translation of arabic text-to-arabic sign language.Uni- versal Access in the Information Society , 18(4):939– 951, 2019

    Hamzah Luqman and Sabri A Mahmoud. Automatic translation of arabic text-to-arabic sign language.Uni- versal Access in the Information Society , 18(4):939– 951, 2019. 1

  39. [47]

    Visual alignment constraint for continuous sign language recognition

    Yuecong Min, Aiming Hao, Xiujuan Chai, and Xilin Chen. Visual alignment constraint for continuous sign language recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vi- sion, pages 11542–11551, 2021. 6, 9

  40. [48]

    FluentSigners-50: A signer independent benchmark dataset for sign language pro- cessing

    Medet Mukushev, Aidyn Ubingazhibov, Aigerim Ky- dyrbekova, Alfarabi Imashev, Vadim Kimmelman, and Anara Sandygulova. FluentSigners-50: A signer independent benchmark dataset for sign language pro- cessing. PLoS ONE, 17(9 September):1–19, 2022. 2, 4

  41. [49]

    A Hong Kong Sign Language corpus collected from sign-interpreted TV news

    Zhe Niu, Ronglai Zuo, Brian Mak, and Fangyun Wei. A Hong Kong Sign Language corpus collected from sign-interpreted TV news. In Proceedings of the 2024 Joint International Conference on Computational Lin- guistics, Language Resources and Evaluation (LREC- COLING 2024), pages 63...

  42. [50]

    Dilated convolutional network with iterative optimization for continuous sign language recognition

    Junfu Pu, Wengang Zhou, and Houqiang Li. Dilated convolutional network with iterative optimization for continuous sign language recognition. IJCAI Inter- national Joint Conference on Artificial Intelligence , 2018-July:885–891, 2018. 3

  43. [51]

    Itera- tive alignment network for continuous sign language recognition

    Junfu Pu, Wengang Zhou, and Houqiang Li. Itera- tive alignment network for continuous sign language recognition. Proceedings of the IEEE Computer So- ciety Conference on Computer Vision and Pattern Recognition, 2019-June:4160–4169, 2019. 3

  44. [52]

    LSA64 : An Argentinian Sign Language Dataset

    Franco Ronchetti, Facundo Quiroga, and Laura Lan- zarini. LSA64 : An Argentinian Sign Language Dataset. Congreso Argentino de Ciencias de la Com- putacion (CACIC), pages 794–803, 2016. 2

  45. [53]

    Auslan-daily: Australian sign lan- guage translation for daily communication and news

    Xin Shen, Shaozu Yuan, Hongwei Sheng, Heming Du, and Xin Yu. Auslan-daily: Australian sign lan- guage translation for daily communication and news. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track ,

  46. [54]

    Open-domain sign language trans- lation learned from online video

    Bowen Shi, Diane Brentari, Greg Shakhnarovich, and Karen Livescu. Open-domain sign language trans- lation learned from online video. arXiv preprint arXiv:2205.12870, 2022. 3

  47. [55]

    Sidig, Hamzah Luqman, Sabri Mah- moud, and Mohamed Mohandes

    Ala Addin I. Sidig, Hamzah Luqman, Sabri Mah- moud, and Mohamed Mohandes. KArSL: Arabic Sign Language Database. ACM Transactions on Asian and Low-Resource Language Information Processing , 20 (1):1–19, 2021. 2

  48. [56]

    AUTSL: A large scale multi-modal Turkish sign lan- guage dataset and baseline methods

    Ozge Mercanoglu Sincan and Hacer Yalim Keles. AUTSL: A large scale multi-modal Turkish sign lan- guage dataset and baseline methods. IEEE Access, 8: 181340–181355, 2020. 2

  49. [57]

    Towards a Video Corpus for Signer-Independent Continuous Sign Language Recognition

    Ulrich V on Agris and Karl-Friedrich Kraiss. Towards a Video Corpus for Signer-Independent Continuous Sign Language Recognition. The 7th International Workshop on Gesture in Human-Computer Interaction and Simulation, GW 2007, pages 1–6, 2007. 2, 3, 4

  50. [58]

    Improving continuous sign language recognition with cross-lingual signs

    Fangyun Wei and Yutong Chen. Improving continuous sign language recognition with cross-lingual signs. In Proceedings of the IEEE/CVF International Confer- ence on Computer Vision, pages 23612–23621, 2023. 3

  51. [59]

    Sign2GPT: Leveraging large language models for gloss-free sign language translation

    Ryan Wong, Necati Cihan Camgoz, and Richard Bow- den. Sign2GPT: Leveraging large language models for gloss-free sign language translation. InThe Twelfth In- ternational Conference on Learning Representations ,

  52. [60]

    Multi-scale local- temporal similarity fusion for continuous sign lan- guage recognition

    Pan Xie, Zhi Cui, Yao Du, Mengyi Zhao, Jianwei Cui, Bin Wang, and Xiaohui Hu. Multi-scale local- temporal similarity fusion for continuous sign lan- guage recognition. Pattern Recognition, 136, 2023. 3

  53. [61]

    SF-Net: Structured Feature Network for Continuous Sign Language Recognition

    Zhaoyang Yang, Zhenmei Shi, Xiaoyong Shen, and Yu-Wing Tai. SF-Net: Structured Feature Network for Continuous Sign Language Recognition. 2019. 3

  54. [62]

    Improving gloss-free sign language translation by reducing representation density

    Jinhui Ye, Xing Wang, Wenxiang Jiao, Junwei Liang, and Hui Xiong. Improving gloss-free sign language translation by reducing representation density. Ad- vances in Neural Information Processing Systems, 37: 107379–107402, 2025. 3

  55. [63]

    Gloss attention for gloss-free sign language translation

    Aoxiong Yin, Tianyun Zhong, Li Tang, Weike Jin, Tao Jin, and Zhou Zhao. Gloss attention for gloss-free sign language translation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recogni- tion, pages 2551–2562, 2023. 3

  56. [64]

    C2ST: Cross-modal Contextualized Se- quence Transduction for Continuous Sign Language Recognition

    Huaiwen Zhang, Zihang Guo, Yang Yang, Xin Liu, and De Hu. C2ST: Cross-modal Contextualized Se- quence Transduction for Continuous Sign Language Recognition. 2023 IEEE/CVF International Con- ference on Computer Vision (ICCV) , pages 20996– 21005, 2023. 3

  57. [65]

    Conditional sentence generation and cross-modal reranking for sign lan- guage translation

    Jian Zhao, Weizhen Qi, Wengang Zhou, Nan Duan, Ming Zhou, and Houqiang Li. Conditional sentence generation and cross-modal reranking for sign lan- guage translation. IEEE Transactions on Multimedia, 24:2662–2672, 2022. 3

  58. [66]

    Cvt-slr: Contrastive visual-textual transformation for sign lan- guage recognition with variational alignment

    Jiangbin Zheng, Yile Wang, Cheng Tan, Siyuan Li, Ge Wang, Jun Xia, Yidong Chen, and Stan Z Li. Cvt-slr: Contrastive visual-textual transformation for sign lan- guage recognition with variational alignment. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Patt...

  59. [67]

    Benjia et al. Zhou. Gloss-free sign language transla- tion: Improving from visual-language pretraining. In CVPR, 2023. 2, 3, 7, 8, 10

  60. [68]

    H. Zhou, W. Zhou, W. Qi, J. Pu, and H. Li. Improv- ing sign language translation with monolingual data by sign back-translation. In 2021 IEEE/CVF Con- ference on Computer Vision and Pattern Recognition (CVPR), pages 1316–1325, Los Alamitos, CA, USA,

  61. [69]

    Spatial-temporal multi-cue network for sign lan- guage recognition and translation

    Hao Zhou, Wengang Zhou, Yun Zhou, and Houqiang Li. Spatial-temporal multi-cue network for sign lan- guage recognition and translation. IEEE Transactions on Multimedia, 24:768–779, 2021. 3

  62. [70]

    Improving continuous sign language recognition with consistency constraints and signer removal

    Ronglai Zuo and Brian Mak. Improving continuous sign language recognition with consistency constraints and signer removal. ACM Transactions on Multimedia Computing, Communications and Applications, 2024. 3

  63. [2021]

    2, 3, 4, 6, 8

    IEEE Computer Society. 2, 3, 4, 6, 8

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.