Pith. sign in

REVIEW 5 cited by

SLTUNET: A Simple Unified Model for Sign Language Translation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.01778 v1 pith:J6HT2TFJ submitted 2023-05-02 cs.CL cs.CV

classification cs.CLcs.CV
keywords sltunettranslationdatasigncorpusjointlylanguagemodality
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite recent successes with neural models for sign language translation (SLT), translation quality still lags behind spoken languages because of the data scarcity and modality gap between sign video and text. To address both problems, we investigate strategies for cross-modality representation sharing for SLT. We propose SLTUNET, a simple unified neural model designed to support multiple SLTrelated tasks jointly, such as sign-to-gloss, gloss-to-text and sign-to-text translation. Jointly modeling different tasks endows SLTUNET with the capability to explore the cross-task relatedness that could help narrow the modality gap. In addition, this allows us to leverage the knowledge from external resources, such as abundant parallel data used for spoken-language machine translation (MT). We show in experiments that SLTUNET achieves competitive and even state-of-the-art performance on PHOENIX-2014T and CSL-Daily when augmented with MT data and equipped with a set of optimization techniques. We further use the DGS Corpus for end-to-end SLT for the first time. It covers broader domains with a significantly larger vocabulary, which is more challenging and which we consider to allow for a more realistic assessment of the current state of SLT than the former two. Still, SLTUNET obtains improved results on the DGS Corpus. Code is available at https://github.com/bzhangGo/sltunet.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Attention-Steered Vision-Language Models for Sign Language Translation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    AttnSign adds spatial attention supervision and motion-cadence reinforcement learning to a VLM, improving sign language translation accuracy on How2Sign and OpenASL.

  2. Bridging Sign and Spoken Languages: Pseudo Gloss Generation for Sign Language Translation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    LLM-generated pseudo glosses, reordered via weak video supervision, enable sign language translation that rivals gloss-supervised models while needing only 30 gloss examples.

  3. SAGE: Segment-Aware Gloss-Free Encoding for Token-Efficient Sign Language Translation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    SAGE uses a frozen sign-segmentation model to turn sign videos into about half as many visual tokens as prior methods, then aligns those tokens with a language model to reach BLEU-4 of 24.10 on PHOENIX14T.

  4. Sign Spotting Disambiguation using Large Language Models

    cs.CV 2025-07 conditional novelty 5.0 of 10

    LLM-based beam search disambiguation improves dictionary sign spotting WER from 47.2% to 44.4% on an internal BSL dataset.

  5. Exploring Pose-based Sign Language Translation: Ablation Studies and Attention Insights

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Pose normalization based on the signer's signing space substantially improves gloss-free sign language translation with a T5 model, while interpolation and augmentation give smaller, less certain gains.

Pith tools