Pith. sign in

Reconsidering Sentence-Level Sign Language Translation

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Historically, sign language machine translation has been posed as a sentence-level task: datasets consisting of continuous narratives are chopped up and presented to the model as isolated clips. In this work, we explore the limitations of this task framing. First, we survey a number of linguistic phenomena in sign languages that depend on discourse-level context. Then as a case study, we perform the first human baseline for sign language translation that actually substitutes a human into the machine learning task framing, rather than provide the human with the entire document as context. This human baseline -- for ASL to English translation on the How2Sign dataset -- shows that for 33% of sentences in our sample, our fluent Deaf signer annotators were only able to understand key parts of the clip in light of additional discourse-level context. These results underscore the importance of understanding and sanity checking examples when adapting machine learning to new domains.

citation-role summary

background 1

citation-polarity summary

fields

cs.HC 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

background 1

representative citing papers

Towards AI-driven Sign Language Generation with Non-manual Markers

cs.HC · 2025-02-08 · conditional · novelty 6.0

The authors combine an LLM, motion matching, and a pose-to-video model to generate ASL videos with non-manual markers, reporting a BLEU-4 of 0.276 for text-to-gloss and a user study where DHH participants rated generated videos lower than human signing.

citing papers explorer

Showing 1 of 1 citing paper.

  • Towards AI-driven Sign Language Generation with Non-manual Markers cs.HC · 2025-02-08 · conditional · none · ref 147 · internal anchor

    The authors combine an LLM, motion matching, and a pose-to-video model to generate ASL videos with non-manual markers, reporting a BLEU-4 of 0.276 for text-to-gloss and a user study where DHH participants rated generated videos lower than human signing.