Pith. sign in

REVIEW 2 cited by

Multi-View Spatial-Temporal Network for Continuous Sign Language Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.08747 v1 pith:5UVDPYDA submitted 2022-04-19 cs.CV

classification cs.CV
keywords languagesignnetworkspatial-temporalcontinuouslearnrecognitionfeatures
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Sign language is a beautiful visual language and is also the primary language used by speaking and hearing-impaired people. However, sign language has many complex expressions, which are difficult for the public to understand and master. Sign language recognition algorithms will significantly facilitate communication between hearing-impaired people and normal people. Traditional continuous sign language recognition often uses a sequence learning method based on Convolutional Neural Network (CNN) and Long Short-Term Memory Network (LSTM). These methods can only learn spatial and temporal features separately, which cannot learn the complex spatial-temporal features of sign language. LSTM is also difficult to learn long-term dependencies. To alleviate these problems, this paper proposes a multi-view spatial-temporal continuous sign language recognition network. The network consists of three parts. The first part is a Multi-View Spatial-Temporal Feature Extractor Network (MSTN), which can directly extract the spatial-temporal features of RGB and skeleton data; the second is a sign language encoder network based on Transformer, which can learn long-term dependencies; the third is a Connectionist Temporal Classification (CTC) decoder network, which is used to predict the whole meaning of the continuous sign language. Our algorithm is tested on two public sign language datasets SLR-100 and PHOENIX-Weather 2014T (RWTH). As a result, our method achieves excellent performance on both datasets. The word error rate on the SLR-100 dataset is 1.9%, and the word error rate on the RWTHPHOENIX-Weather dataset is 22.8%.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Isharah: A Large-Scale Multi-Scene Dataset for Continuous Sign Language Recognition

    cs.CV 2025-06 conditional novelty 7.0 of 10

    Isharah is a new 30,000-clip, multi-scene Saudi Sign Language dataset with gloss and translation annotations, plus signer-independent and unseen-sentence benchmarks for continuous sign language recognition and translation.

  2. A Signer-Invariant Conformer and Multi-Scale Fusion Transformer for Continuous Sign Language Recognition

    cs.CV 2025-08 conditional novelty 5.0 of 10

    The paper reports state-of-the-art WERs of 13.07% (signer-independent) and 47.78% (unseen sentences) on Isharah-1000 using a conformer and a multi-scale fusion transformer.

Pith tools