Pith. sign in

REVIEW 2 cited by

Self-Supervised Video Transformers for Isolated Sign Language Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.02450 v1 pith:SF656JWO submitted 2023-09-02 cs.CV

classification cs.CV
keywords languagepre-trainingsignwlasl2000fourislrisolatedlinear
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents an in-depth analysis of various self-supervision methods for isolated sign language recognition (ISLR). We consider four recently introduced transformer-based approaches to self-supervised learning from videos, and four pre-training data regimes, and study all the combinations on the WLASL2000 dataset. Our findings reveal that MaskFeat achieves performance superior to pose-based and supervised video models, with a top-1 accuracy of 79.02% on gloss-based WLASL2000. Furthermore, we analyze these models' ability to produce representations of ASL signs using linear probing on diverse phonological features. This study underscores the value of architecture and pre-training task choices in ISLR. Specifically, our results on WLASL2000 highlight the power of masked reconstruction pre-training, and our linear probing results demonstrate the importance of hierarchical vision transformers for sign language representation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Targeted Linguistic Analysis of Sign Language Models with Minimal Translation Pairs

    cs.CL 2026-04 unverdicted novelty 7.0 of 10

    Introduces ASL-MTP benchmark and shows a state-of-the-art ASL-to-English model relies strongly on manual cues while missing non-manual cues.

  2. Fine-Tuning Video Transformers for Word-Level Bangla Sign Language: A Comparative Analysis for Classification Tasks

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Off-the-shelf video transformers (VideoMAE, ViViT, TimeSformer) fine-tuned on Bangla sign language videos reach 95.5% top-1 accuracy on BdSLW60 and 81.04% on the BdSLW401 front subset.

Pith tools