SiLVERScore, built on the CiCo video-text contrastive model, discriminates correct vs. random sign-video/text pairs with 0.99 ROC AUC and is robust to word reordering and prosody intensity.
Non-Autoregressive Sign Language Production via Knowledge Distillation
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Sign Language Production (SLP) aims to translate expressions in spoken language into corresponding ones in sign language, such as skeleton-based sign poses or videos. Existing SLP models are either AutoRegressive (AR) or Non-Autoregressive (NAR). However, AR-SLP models suffer from regression to the mean and error propagation during decoding. NSLP-G, a NAR-based model, resolves these issues to some extent but engenders other problems. For example, it does not consider target sign lengths and suffers from false decoding initiation. We propose a novel NAR-SLP model via Knowledge Distillation (KD) to address these problems. First, we devise a length regulator to predict the end of the generated sign pose sequence. We then adopt KD, which distills spatial-linguistic features from a pre-trained pose encoder to alleviate false decoding initiation. Extensive experiments show that the proposed approach significantly outperforms existing SLP models in both Frechet Gesture Distance and Back-Translation evaluation.
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
SiLVERScore: Semantically-Aware Embeddings for Sign Language Generation Evaluation
SiLVERScore, built on the CiCo video-text contrastive model, discriminates correct vs. random sign-video/text pairs with 0.99 ROC AUC and is robust to word reordering and prosody intensity.