Pith. sign in

REVIEW 2 cited by

Leveraging Graph-based Cross-modal Information Fusion for Neural Sign Language Translation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.00526 v1 pith:H4MYKXRJ submitted 2022-11-01 cs.CL cs.AI

classification cs.CLcs.AI
keywords neurallanguagegraphsigninformationmodelsmulti-modaltranslation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Sign Language (SL), as the mother tongue of the deaf community, is a special visual language that most hearing people cannot understand. In recent years, neural Sign Language Translation (SLT), as a possible way for bridging communication gap between the deaf and the hearing people, has attracted widespread academic attention. We found that the current mainstream end-to-end neural SLT models, which tries to learning language knowledge in a weakly supervised manner, could not mine enough semantic information under the condition of low data resources. Therefore, we propose to introduce additional word-level semantic knowledge of sign language linguistics to assist in improving current end-to-end neural SLT models. Concretely, we propose a novel neural SLT model with multi-modal feature fusion based on the dynamic graph, in which the cross-modal information, i.e. text and video, is first assembled as a dynamic graph according to their correlation, and then the graph is processed by a multi-modal graph encoder to generate the multi-modal embeddings for further usage in the subsequent neural translation models. To the best of our knowledge, we are the first to introduce graph neural networks, for fusing multi-modal information, into neural sign language translation models. Moreover, we conducted experiments on a publicly available popular SLT dataset RWTH-PHOENIX-Weather-2014T. and the quantitative experiments show that our method can improve the model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DapPep: Domain Adaptive Peptide-agnostic Learning for Universal T-cell Receptor-antigen Binding Affinity Prediction

    q-bio.QM 2024-11 conditional novelty 5.0 of 10

    DapPep, built from ESM-2 with cross-attention and peptide-reconstruction pre-training, reports ROC-AUC 0.816 and PR-AUC 0.836 on unseen peptides, beating PanPep by about 9 to 11 percent.

  2. Pan-protein Design Learning Enables Task-adaptive Generalization for Low-resource Enzyme Design

    q-bio.QM 2024-11 conditional novelty 5.0 of 10

    CrossDesign aligns pretrained protein language models with structure encoders to improve enzyme sequence design and zero-shot mutation fitness prediction.

Pith tools