Pith. sign in

REVIEW 2 cited by

Towards the extraction of robust sign embeddings for low resource sign language recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.17558 v2 pith:OIYFRDRQ submitted 2023-06-30 cs.CV cs.CL

classification cs.CVcs.CL
keywords signlanguagemodelsembeddingsposekeypoint-basedlearningtransfer
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Isolated Sign Language Recognition (SLR) has mostly been applied on datasets containing signs executed slowly and clearly by a limited group of signers. In real-world scenarios, however, we are met with challenging visual conditions, coarticulated signing, small datasets, and the need for signer independent models. To tackle this difficult problem, we require a robust feature extractor to process the sign language videos. One could expect human pose estimators to be ideal candidates. However, due to a domain mismatch with their training sets and challenging poses in sign language, they lack robustness on sign language data and image-based models often still outperform keypoint-based models. Furthermore, whereas the common practice of transfer learning with image-based models yields even higher accuracy, keypoint-based models are typically trained from scratch on every SLR dataset. These factors limit their usefulness for SLR. From the existing literature, it is also not clear which, if any, pose estimator performs best for SLR. We compare the three most popular pose estimators for SLR: OpenPose, MMPose and MediaPipe. We show that through keypoint normalization, missing keypoint imputation, and learning a pose embedding, we can obtain significantly better results and enable transfer learning. We show that keypoint-based embeddings contain cross-lingual features: they can transfer between sign languages and achieve competitive performance even when fine-tuning only the classifier layer of an SLR model on a target sign language. We furthermore achieve better performance using fine-tuned transferred embeddings than models trained only on the target sign language. The embeddings can also be learned in a multilingual fashion. The application of these embeddings could prove particularly useful for low resource sign languages in the future.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Unsupervised Keypoints for Real-Time Fall Detection: Comparative Analysis Under Real-world Conditions with Predictive Bandwidth Reduction

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Under occlusion, unsupervised keypoints detect falls more reliably than supervised pose keypoints, but supervised keypoints do better when the full body is visible.

  2. TrackStudio: An Integrated Toolkit for Markerless Tracking

    cs.CV 2025-11 conditional novelty 4.0 of 10

    TrackStudio packages MediaPipe and Anipose into a GUI with new validation on 76 participants; it shows stable self-consistency, but its error claims lack ground-truth comparison.

Pith tools