Pith. sign in

Progressive Transformers for End-to-End Sign Language Production

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

The goal of automatic Sign Language Production (SLP) is to translate spoken language to a continuous stream of sign language video at a level comparable to a human translator. If this was achievable, then it would revolutionise Deaf hearing communications. Previous work on predominantly isolated SLP has shown the need for architectures that are better suited to the continuous domain of full sign sequences. In this paper, we propose Progressive Transformers, a novel architecture that can translate from discrete spoken language sentences to continuous 3D skeleton pose outputs representing sign language. We present two model configurations, an end-to-end network that produces sign direct from text and a stacked network that utilises a gloss intermediary. Our transformer network architecture introduces a counter that enables continuous sequence generation at training and inference. We also provide several data augmentation processes to overcome the problem of drift and improve the performance of SLP models. We propose a back translation evaluation mechanism for SLP, presenting benchmark quantitative results on the challenging RWTH-PHOENIX-Weather-2014T(PHOENIX14T) dataset and setting baselines for future research.

citation-role summary

background 1

citation-polarity summary

fields

cs.HC 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

background 1

representative citing papers

Towards AI-driven Sign Language Generation with Non-manual Markers

cs.HC · 2025-02-08 · conditional · novelty 6.0

The authors combine an LLM, motion matching, and a pose-to-video model to generate ASL videos with non-manual markers, reporting a BLEU-4 of 0.276 for text-to-gloss and a user study where DHH participants rated generated videos lower than human signing.

citing papers explorer

Showing 1 of 1 citing paper.

  • Towards AI-driven Sign Language Generation with Non-manual Markers cs.HC · 2025-02-08 · conditional · none · ref 135 · internal anchor

    The authors combine an LLM, motion matching, and a pose-to-video model to generate ASL videos with non-manual markers, reporting a BLEU-4 of 0.276 for text-to-gloss and a user study where DHH participants rated generated videos lower than human signing.