Pith. sign in

REVIEW 1 cited by

MS2SL: Multimodal Spoken Data-Driven Continuous Sign Language Production

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.12842 v1 pith:AZHKINFP submitted 2024-07-04 cs.CL cs.AI

classification cs.CLcs.AI
keywords signlanguagemodelcontinuousproductiontextaudiospeech
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Sign language understanding has made significant strides; however, there is still no viable solution for generating sign sequences directly from entire spoken content, e.g., text or speech. In this paper, we propose a unified framework for continuous sign language production, easing communication between sign and non-sign language users. In particular, a sequence diffusion model, utilizing embeddings extracted from text or speech, is crafted to generate sign predictions step by step. Moreover, by creating a joint embedding space for text, audio, and sign, we bind these modalities and leverage the semantic consistency among them to provide informative feedback for the model training. This embedding-consistency learning strategy minimizes the reliance on sign triplets and ensures continuous model refinement, even with a missing audio modality. Experiments on How2Sign and PHOENIX14T datasets demonstrate that our model achieves competitive performance in sign language production.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Using the Pepper Robot to Support Sign Language Communication

    cs.RO 2025-09 conditional novelty 4.0 of 10

    Pepper can produce a subset of Italian Sign Language signs that LIS users recognize at the single-sign level, but sentence-level comprehension largely fails (about 8 percent correct).

Pith tools