REVIEW 4 major objections 5 minor 171 references
Towards AI-driven Sign Language Generation with Non-manual Markers
T0 review · 4 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read A modular English-to-ASL pipeline that adds non-manual markers reaches BLEU-4 0.276 and better video quality, while 30 DHH participants found meaning often conveyed but visual quality still lags human signing.
desk verdict A genuine modular SLG pipeline with a well-ablated pose-to-video improvement, but the headline non-manual-marker claim is not supported by the paper's own numbers. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a three-module pipeline whose second stream—non-manual markers—is the element prior systems lack. Module 1 uses GPT-4o with 1,474 in-context English–gloss examples from ASLLRP and a 3,915-entry word-gloss dictionary to produce ASL glosses, and separately uses zero-shot prompting to label each sentence as yes/no question, wh-question, conditional, and/or negated; those labels drive eyebrow position later. Module 2 is Motion Matching: a dictionary of 12,681 skeletal pose clips extracted from ASLLRP isolated-sign videos via a standard pose-estimation pipeline, a greedy selection function that minimizes a weighted Euclidean distance between the end of one sign and the start of the next (economy of motion), and linear blending of clip boundaries, followed by expression blending that raises or lowers the eyebrows according to the Module 1 labels. Module 3 is a U-Net pose-to-image model trained on How2Sign, conditioned on signer identity, with a rasterization function that draws each body part as a shaded convex polygon and encodes hand depth via palm-surface color, trained with whole-frame L1, hand-region L1, and LPIPS losses, and fed only frames that pass optical-flow and landmark-jitter quality checks. The enhanced rasterization and frame filtering are what the paper credits for the video-quality gains; the non-manual stream is what it credits for the grammatical acceptability gains.
What would settle it
Swap only Module 2: take the same English sentences, run Module 1, then render two video sets with the same Module 3—one from Module 2's skeleton output and one from raw ASLLRP skeleton poses of the same sentences—and compare DHH understandability ratings; if the full-pipeline videos are not rated worse than the retargeted ones, the paper's claim that motion synthesis is the current bottleneck, and its full-system evaluation, would be undercut.
Extended reading notes
Core claim
Stated in the paper's own terms, the discovery is that an ASL generation system can be assembled from three independently trained modules—a few-shot LLM that translates English into ASL glosses and classifies the sentence's linguistic features (yes/no question, wh-question, conditional, negation), a Motion Matching synthesizer that selects and blends sign clips from a 12,681-clip dictionary while adjusting eyebrow position according to those features, and a signer-conditioned U-Net that renders skeleton poses into photorealistic video—and that this modular construction yields better translation and video quality than the baselines it compares against. The strongest reported numbers are a BLEU-4 of 0.276 for English-to-ASL gloss translation (compared with 0.191 in the prior work the authors cite), average precision of 0.91 and recall of 0.97 for non-manual information detection, and large improvements on FID/FVD video metrics from the proposed rasterization and frame-quality filtering. The user study adds the qualitative claim: DHH signers rate the meaning as similar or acceptable in 53.8% of full-pipeline videos, rate the full model's translations as more accurate than human-annotated gloss inputs, and prefer the version with non-manual markers over the version without. The paper also establishes a current ceiling: with raw human skeletons the same video model is understandable 60% of the time, versus 21.1% for the full AI pipeline, which the authors interpret as evidence that motion synthesis, not text or rendering, now limits the system.
Load-bearing premise
The entire end-to-end evaluation assumes that a pose-to-video model trained on How2Sign will work on skeleton sequences built from ASLLRP signs, even though those datasets differ in signers, seating, camera angle, and eye contact—discrepancies the paper itself lists in Section 6.2—and if that transfer fails, the user-study ratings cannot be assigned to the translation or motion modules.
Editorial extensions
If this is right
- If the BLEU-4 result transfers beyond ASLLRP, then few-shot LLM prompting with a constrained vocabulary is a workable route for low-resource English-to-ASL gloss translation.
- If non-manual marker detection is as reliable as reported, sign generation systems can obtain facial grammar directly from English text without manual annotation.
- If the video metrics reflect real visual improvement, the proposed rasterization and frame-quality filtering are reusable components for any skeleton-to-video signing system.
- If the user-study ratings generalize, DHH users will accept AI signing for practical use cases such as doctor visits and video conversations, but only after visual quality and naturalness of motion are improved.
- The gap between retargeted skeletons and the full pipeline indicates that the motion-matching module, not the text or rendering modules, currently caps the system's understandability.
Reading between the lines
- The paper does not test this, but a head-to-head run on NCSLGR, rather than comparing across datasets, would determine whether the claimed BLEU-4 advantage over prior work is real or dataset-specific; the authors' own RAG experiment (0.279) suggests prompt-retrieval is an easy further lever.
- A natural next experiment the paper does not report is a forced-choice user study showing the same sentence with and without eyebrow blending, to test whether the non-manual marker actually changes comprehension of a question versus a statement; the current study found no significant rating difference between those two conditions.
- The retargeting result predicts that investing in coarticulation-aware or learned motion matching would raise understandability more than further improving the renderer, because even raw skeletons pass through the same Module 3 and still outperform the full pipeline.
- Resolving the How2Sign/ASLLRP domain shift, for example by collecting a matched skeleton-video corpus, is a prerequisite for any end-to-end evaluation, and the paper itself implies this in Section 6.2.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper describes a three-module pipeline for English-to-ASL video generation: (1) GPT-4o translates English into English-based ASL glosses and classifies four text-derived linguistic features related to non-manual markers (yes/no question, wh-question, conditional, negation); (2) a Motion Matching system converts glosses plus predicted NMM information into skeletal pose sequences using the ASLLRP corpus; (3) a U-Net pose-to-video model trained on How2Sign renders photorealistic signing videos. The authors report a BLEU-4 of 0.276 for text-to-gloss translation, an average NMM-detection precision of 0.91 and recall of 0.97, video-generation improvements over their own baselines across all metrics, and a user study with 30 DHH signers assessing understandability, visual quality, motion naturalness, and translation quality. The paper claims that the prototype makes significant progress on manual and non-manual signing while acknowledging substantial remaining gaps in visual and motion fidelity.
Significance. If the claims were fully established, this would be a valuable contribution to sign-language generation: it is a modular, open-ended system that explicitly models non-manual markers, it reports a strong text-to-gloss result on a real ASL dataset, and its pose-to-video ablations (Table 2) show large, internally consistent gains in rasterization and frame-quality selection. The user study is also a notable strength: 30 screened DHH signers, mixed-effects modeling, and transparent reporting of demographic variables. The paper is unusually candid about dataset inconsistencies and limitations. However, the headline NMM contribution is not yet established. The precision/recall evaluation in Figure 4 is computed against labels that four researchers revised to match the English text, so it does not independently verify that the predicted NMMs are present in the signed video, and the user-study contrast that would support a downstream benefit of NMMs is not statistically significant. Because the title and contributions place NMMs at the center, these issues are load-bearing for the paper's novelty rather than cosmetic.
major comments (4)
- [§3.1 Step 4; Fig. 4] The precision/recall numbers in Figure 4 are computed against labels that four researchers revised 'to better reflect the text content' after inspecting the English text. Since the classifier's input is also the English text, the metric measures agreement with the authors' text-based relabeling, not whether the ASL videos actually displayed the predicted non-manual markers. The corrected labels are a reasonable gold standard for a text-classification task, but the paper's contribution claims 'detecting non-manual information from English text' in a way that implies visual NMM validity. I recommend validating the NMM predictions against video-based NMM annotations (or at least reporting agreement with the original ASLLRP video-derived labels) and/or explicitly reframing Figure 4 as text-based category agreement rather than detection of non-manual information in the signed output.
- [§5.4.2] The text states that incorporating non-manual markers 'resulted in higher acceptance' for the full model compared to the model without expression blending, but the reported LMM contrast between AI (Full) and AI (w/o Expr) is not significant (z=0.907, p=0.364), and the same subsection reports no significant differences among the three AI models for signing quality or facial-expression accuracy. This is an internal contradiction between the prose claim and the statistical results. The claim should be downgraded to a descriptive trend, or the analysis should be augmented, before the user study can be used as evidence for the NMM contribution.
- [§4.1.1, Table 8, §6.1] The headline BLEU-4 of 0.276 is obtained on the ASLLRP test split, whereas the prior scores of 0.002, 0.124, and 0.191 in Table 8 are from different datasets (self-collected and NCSLGR). The paper itself cautions that direct comparison is challenging, but later claims this is 'the highest reported score for such translation task in the literature.' A cross-dataset numeric comparison cannot support a superlative claim. The authors should either run or report a baseline on the same ASLLRP test split, or restrict the claim to 'within our ASLLRP split, the proposed configuration reaches a BLEU-4 of 0.276.'
- [§3.3, §5.1, §6.2] The full-system user study is conducted under a known domain shift: the pose-to-video model is trained on How2Sign, but the skeletons and sentences at inference come from ASLLRP, with the differences in signer, seating, camera angle, and eye contact that the paper lists in §6.2. Because this shift is known to degrade visual and motion quality, the user-study ratings of AI-generated conditions cannot cleanly be attributed to the quality of Modules 1 and 2, nor can they cleanly isolate the effect of expression blending. I recommend either adding a same-domain control condition (for example, rendering the Video Retargeting skeletons with the same pose-to-video model and comparing it to the full pipeline) or explicitly limiting causal claims about translation and NMM contribution to the technical metrics rather than the user-study ratings.
minor comments (5)
- [Throughout] There are several typos, including 'Ground True Correction' (§B.1), 'Represenetations' (heading §4.1), 'Appedix' (§B.4), and 'AI (Annotations))s' (§5.4.2). A careful proofreading pass is needed.
- [Table 1 footnote] The footnote states that 'All BLEU-4 and SacreBLEU scores are identical,' which is confusing: SacreBLEU is a standardized BLEU implementation and need not equal the reported BLEU-4 by construction. Please clarify the relationship between the two columns.
- [§5.3] The evaluator correlation analysis is reported as pairwise Pearson correlations, but for ordinal rating data an intraclass correlation coefficient or weighted kappa would be more appropriate; at minimum, the choice of Pearson correlation should be justified.
- [§3.1 / §B.1] The paper does not state whether the preprocessing scripts, prompt templates, or trained models will be released. Providing even the annotated dataset or the exact prompts would substantially improve reproducibility, especially since the core component is a proprietary API model.
- [§5.2] The study reports 30 participants and 27 sampled sentences, but Section 5.1.2 says each participant viewed 21 videos. Clarifying the exact overlap and randomization protocol would help readers understand the effective sample size per condition.
Circularity Check
NMM precision/recall is partly circular because test labels were re-annotated from the same English text that is the model's only input; the video-side NMM benefit is not independently shown.
-
self definitional
[Section 3.1 (Data Preprocessing), Appendix B.1 'Step 4: Ground True Correction', Section 4.1.2 / Figure 4]
"Four of our researchers conducted a ground truth correction to resolve misalignments between the linguistic labels for the four types of non-manual information and the English text, ensuring the labels more accurately reflected the text content. ... four of our researchers iteratively re-labeled and discussed the test set sentence categories, refining the labels to better reflect the text content. These revised labels were then used as the ground truth, allowing us to calculate precision and recall for each sentence type predictions."
The classifier input is the English sentence itself ('asked the model to predict whether a given English sentence: is (1) a yes-no question, (2) a wh-question, (3) a conditional statement, and/or (4) contains negation'), and the gold labels were revised by humans reading that same English text 'to better reflect the text content.' Therefore the Figure 4 precision/recall measures agreement between GPT-4o's text reading and the annotators' text reading; it is not an independent measurement of whether the signed video actually displayed the corresponding non-manual marker. Both predictor and ground truth are derived from the same evidential source, so the high scores are partly self-consistent by construction and do not validate the NMM generation step.
full rationale
The bulk of the technical pipeline is self-contained and non-circular: the gloss-translation evaluation uses a held-out ASLLRP split with standard metrics; the video-generation ablations (rasterization, frame filtering) are evaluated on held-out How2Sign data against standard image/video metrics; and the user study is an external, out-of-participant check of the full system. The main circularity is confined to the headline non-manual-marker detection claim. The test-set labels for yes/no questions, wh-questions, conditionals, and negation were explicitly re-annotated by four researchers to match the English text, which is the same input the zero-shot LLM classifier receives. Consequently, the reported precision of 0.91 and recall of 0.97 quantify text-to-text label agreement rather than detection of visible non-manual information in signed video. This is a partial, construction-level circularity in the paper's central NMM contribution, but it does not infect the video-generation or gloss-translation results. There are self-citations in the bibliography, but none is load-bearing in the derivation; they are ordinary prior-work references. The score is therefore 6 rather than 8 or 10, because the circular step affects one central advertised result while the rest of the system is independently evaluated.
Assumptions & free parameters
free parameters (4)
- alpha_p body/face/hands weights =
not reported
- clip blend window =
20 frames at 90 Hz
- neutral pose padding =
0.5 seconds
- in-context example count =
1,474 (80% of ASLLRP)
assumptions (4)
- domain assumption English text contains sufficient linguistic cues to classify yes/no questions, wh-questions, conditionals, and negation for ASL non-manual markers.
- domain assumption ASLLRP glosses and English translations are a valid parallel corpus for training and evaluating text-to-gloss translation.
- domain assumption Mediapipe skeletal keypoints are accurate enough to drive pose synthesis and expression blending.
- domain assumption A pose-to-video model trained on How2Sign generalizes to skeleton sequences from ASLLRP.
Cite this review
Pith. "Pith review of Towards AI-driven Sign Language Generation with Non-manual Markers." pith.science (2026). https://pith.science/paper/DSY6HBLH
@misc{pith2026250205661,
author = {Pith},
title = {Pith review of: Towards AI-driven Sign Language Generation with Non-manual Markers},
year = {2026},
howpublished = {\url{https://pith.science/paper/DSY6HBLH}},
note = {Machine review of arXiv:2502.05661}
}
read the original abstract
Sign languages are essential for the Deaf and Hard-of-Hearing (DHH) community. Sign language generation systems have the potential to support communication by translating from written languages, such as English, into signed videos. However, current systems often fail to meet user needs due to poor translation of grammatical structures, the absence of facial cues and body language, and insufficient visual and motion fidelity. We address these challenges by building on recent advances in LLMs and video generation models to translate English sentences into natural-looking AI ASL signers. The text component of our model extracts information for manual and non-manual components of ASL, which are used to synthesize skeletal pose sequences and corresponding video frames. Our findings from a user study with 30 DHH participants and thorough technical evaluations demonstrate significant progress and identify critical areas necessary to meet user needs.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
K Aberman, M Shi, J Liao, D Liscbinski, B Chen, and D Cohen-Or. 2019. Deep Video-Based Performance Cloning. In Computer Graphics Forum, Vol. 38. Wiley- Blackwell Publishing Ltd., 219–233
2019
-
[2]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Floren- cia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)
arXiv 2023
- [3]
-
[4]
Vidia Anindhita and Dessi Puji Lestari. 2016. Designing interaction for deaf youths by using user-centered design approach. In 2016 international conference on advanced informatics: Concepts, theory and application (icaicta) . IEEE, 1–6
2016
-
[5]
Rotem Shalev Arkushin, Amit Moryossef, and Ohad Fried. 2023. Ham2pose: Animating sign language notation into pose sequences. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 21046–21056
2023
-
[6]
Veditz Quote - 1913 (2015)
British Deaf Association. Veditz Quote - 1913 (2015). https://vimeo.com/ 132549587
2015
-
[7]
Charlotte Baker-Shenk. 1985. The facial behavior of deaf signers: Evidence of a complex language. American Annals of the Deaf 130, 4 (1985), 297–304
1985
-
[8]
Charlotte Lee Baker-Shenk. 1983. A microanalysis of the nonmanual components of questions in American Sign Language . University of California, Berkeley
1983
Show all 171 references
-
[9]
Charlotte Lee Baker-Shenk and Dennis Cokely. 1991. American Sign Language: A teacher’s resource text on grammar and culture . Gallaudet University Press
1991
-
[10]
Yogesh Balaji, Martin Renqiang Min, Bing Bai, Rama Chellappa, and Hans Peter Graf. 2019. Conditional GAN with Discriminative Filter Generation for Text-to- Video Synthesis.. In IJCAI, Vol. 1. 2
2019
-
[11]
Vasileios Baltatzis, Rolandos Alexandros Potamias, Evangelos Ververas, Guanx- iong Sun, Jiankang Deng, and Stefanos Zafeiriou. 2024. Neural Sign Actors: A diffusion model for 3D sign language production from text. In Proceedings of the IEEE/CVF Conference on Computer Vision an...
2024
-
[12]
Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, Jade Goldstein...
2005
-
[13]
Patrick Boudreault, Muhammad Abubakar, Andrew Duran, Bridget Lam, Zehui Liu, Christian Vogler, and Raja Kushalnagar. 2024. Closed Sign Language In- terpreting: A Usability Study. In International Conference on Computers Helping People with Special Needs . Springer, 42–49
2024
-
[14]
Danielle Bragg, Naomi Caselli, John W Gallagher, Miriam Goldberg, Courtney J Oka, and William Thies. 2021. ASL sea battle: gamifying sign language data col- lection. In Proceedings of the 2021 CHI conference on human factors in computing systems. 1–13
2021
-
[15]
Danielle Bragg, Naomi Caselli, Julie A Hochgesang, Matt Huenerfauth, Leah Katz-Hernandez, Oscar Koller, Raja Kushalnagar, Christian Vogler, and Richard E Ladner. 2021. The fate landscape of sign language ai datasets: An interdisci- plinary perspective. ACM Transactions on Acce...
2021
-
[16]
Danielle Bragg, Oscar Koller, Mary Bellard, Larwan Berke, Patrick Boudreault, Annelies Braffort, Naomi Caselli, Matt Huenerfauth, Hernisa Kacorri, Tessa Verhoef, Christian Vogler, and Meredith Ringel Morris. 2019. Sign Language Recognition, Generation, and Translation: An Inte...
2019
-
[17]
DIANE Brentari. 1998. A prosodic model of sign language phonology.A Bradford Book (1998)
1998
-
[18]
Diane Brentari and Laurinda Crossley. 2002. Prosody on the hands and face: Evidence from American Sign Language. Sign Language & Linguistics 5, 2 (2002), 105–130
2002
-
[19]
Tim Brooks, Aleksander Holynski, and Alexei A Efros. 2023. Instructpix2pix: Learning to follow image editing instructions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 18392–18402
2023
-
[20]
Tom B Brown. 2020. Language models are few-shot learners. arXiv preprint ArXiv:2005.14165 (2020)
2020 arXiv
-
[21]
Michael Büttner and Simon Clavet. 2015. Motion matching-the road to next gen animation. Proc. of Nucl. ai 1, 2015 (2015), 2
2015
-
[22]
Necati Cihan Camgoz, Simon Hadfield, Oscar Koller, Hermann Ney, and Richard Bowden. 2018. Neural Sign Language Translation. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition . IEEE, Salt Lake City, UT, 7784–7793. https://doi.org/10.1109/CVPR.2018.00812
2018
-
[23]
Brenda Cartwright. 2024. Signing Savvy. https://www.signingsavvy.com/index. php
2024
-
[24]
Naomi K Caselli, Zed Sevcikova Sehyr, Ariel M Cohen-Goldberg, and Karen Emmorey. 2017. ASL-LEX: A lexical database of American Sign Language. Behavior research methods 49 (2017), 784–801
2017
-
[25]
Caroline Chan, Shiry Ginosar, Tinghui Zhou, and Alexei A Efros. 2019. Ev- erybody dance now. In Proceedings of the IEEE/CVF international conference on computer vision. 5933–5942
2019
-
[26]
Yutong Chen, Ronglai Zuo, Fangyun Wei, Yu Wu, Shujie Liu, and Brian Mak. 2022. Two-stream network for sign language recognition and translation. Advances in Neural Information Processing Systems 35 (2022), 17043–17056. Towards AI-driven SLG with Non-manual Markers CHI ’25, Apr...
2022
-
[27]
Simon Clavet et al. 2016. Motion matching and the road to next-gen animation. In Proc. of GDC, Vol. 2. 4
2016
-
[28]
Onno Crasborn and Han Sloetjes. 2008. Enhanced ELAN functionality for sign language corpora. In 6th International Conference on Language Resources and Evaluation (LREC 2008)/3rd Workshop on the Representation and Processing of Sign Languages: Construction and Exploitation of S...
2008
-
[29]
Florinel-Alin Croitoru, Vlad Hondru, Radu Tudor Ionescu, and Mubarak Shah
-
[30]
Anthony Christopher Davison and David Victor Hinkley. 1997. Bootstrap meth- ods and their application . Number 1. Cambridge university press
1997
-
[31]
Aashaka Desai, Lauren Berger, Fyodor Minakov, Nessa Milano, Chinmay Singh, Kriston Pumphrey, Richard Ladner, Hal Daumé III, Alex X Lu, Naomi Caselli, et al. 2024. ASL citizen: a community-sourced dataset for advancing isolated sign language recognition. Advances in Neural Info...
2024
-
[32]
Aashaka Desai, Maartje De Meulder, Julie A Hochgesang, Annemarie Kocab, and Alex X Lu. 2024. Systemic Biases in Sign Language AI Research: A Deaf-Led Call to Reevaluate Research Agendas. arXiv preprint arXiv:2403.02563 (2024)
2024 arXiv
-
[33]
Philippe Dreuw, Carol Neidle, Vassilis Athitsos, Stanley Sclaroff, and Hermann Ney. 2008. Benchmark databases for video-based automatic sign language recognition. In Proceedings of Language Resources and Evaluation Conference (LREC) 2008. EUROPEAN LANGUAGE RESOURCES ASSOC-ELRA
2008
-
[34]
Amanda Duarte, Shruti Palaskar, Lucas Ventura, Deepti Ghadiyaram, Kenneth DeHaan, Florian Metze, Jordi Torres, and Xavier Giro-i Nieto. 2021. How2Sign: A Large-scale Multimodal Dataset for Continuous American Sign Language. In 2021 IEEE/CVF Conference on Computer Vision and Pa...
2021
-
[35]
Sarah Ebling and John Glauert. 2016. Building a Swiss German Sign Language avatar with JASigning and evaluating it among the Deaf community. Universal Access in the Information Society 15 (2016), 577–587
2016
-
[36]
Santiago Egea Gómez, Euan McGill, and Horacio Saggion. 2021. Syntax- aware Transformers for Neural Machine Translation: The Case of Text to Sign Gloss Translation. In Proceedings of the 14th Workshop on Building and Using Comparable Corpora (BUCC 2021) , Reinhard Rapp, Serge S...
2021
-
[37]
Karen Emmorey. 2001. Language, cognition, and the brain: Insights from sign language research. Psychology Press
2001
-
[38]
Michael Erard. 2017. Why Sign-Language Gloves Don’t Help Deaf Peo- ple. https://www.theatlantic.com/technology/archive/2017/11/why-sign- language-gloves-dont-help-deaf-people/545441/
2017
-
[39]
Sen Fang, Chunyu Sui, Xuedong Zhang, and Yapeng Tian. 2023. SignDiff: Learning Diffusion Models for American Sign Language Production. arXiv preprint arXiv:2308.16082 (2023)
2023 arXiv
-
[40]
Sen Fang, Lei Wang, Ce Zheng, Yapeng Tian, and Chen Chen. 2024. Sign- LLM: Sign Languages Production Large Language Models. arXiv preprint arXiv:2405.10718 (2024)
2024 arXiv
-
[41]
Gunnar Farnebäck. 2003. Two-frame motion estimation based on polynomial expansion. In Image Analysis: 13th Scandinavian Conference, SCIA 2003 Halmstad, Sweden, June 29–July 2, 2003 Proceedings 13 . Springer, 363–370
2003
-
[42]
Mengyang Feng, Jinlin Liu, Kai Yu, Yuan Yao, Zheng Hui, Xiefan Guo, Xianhui Lin, Haolan Xue, Chen Shi, Xiaowen Li, et al . 2023. Dreamoving: A human video generation framework based on diffusion models. arXiv e-prints (2023), arXiv–2312
2023
-
[43]
The Academic Center for Excellence. 2023. ASL Grammar Guide. https://germanna.edu/sites/default/files/2023-07/ASL%20Grammar%20Guide% 20%28edit%207-24-23%29.pdf
2023
-
[44]
Jens Forster, Christoph Schmidt, Oscar Koller, Martin Bellgardt, and Hermann Ney. [n. d.]. Extensions of the Sign Language Recognition and Translation Corpus RWTH-PHOENIX-Weather. ([n. d.])
-
[45]
Lynn A Friedman. 1975. Space, time, and person reference in American Sign Language. Language (1975), 940–961
1975
-
[46]
Neil Stephen Glickman. 1993. Deaf identity development: Construction and validation of a theoretical model . University of Massachusetts Amherst
1993
-
[47]
Jan Gugenheimer, Katrin Plaumann, Florian Schaub, Patrizia Di Campli San Vito, Saskia Duck, Melanie Rabus, and Enrico Rukzio. 2017. The impact of assistive technology on communication quality between deaf and hearing individuals. In Proceedings of the 2017 ACM Conference on Co...
2017
-
[48]
Ping Guo, Yubing Ren, Yue Hu, Yunpeng Li, Jiarui Zhang, Xingsheng Zhang, and He-Yan Huang. 2024. Teaching Large Language Models to Translate on Low-resource Languages with Textbook Prompting. In Proceedings of the 2024 Joint International Conference on Computational Linguistic...
2024
-
[49]
Thomas Hanke. 2004. HamNoSys-representing sign language data in language resources and language processing contexts. In LREC, Vol. 4. 1–6
2004
-
[50]
Thomas Hanke, Marc Schulder, Reiner Konrad, and Elena Jahn. 2020. Extending the Public DGS Corpus in size and depth. In Proceedings of the LREC2020 9th workshop on the representation and processing of sign languages: Sign language resources in the service of the language commu...
2020
-
[51]
Vicki L Hanson and Carol A Padden. 2012. Computers and videodisc technology for bilingual ASL/English instruction of deaf children. In Cognition, Education, and Multimedia. Routledge, 49–63
2012
-
[52]
Amr Hendy, Mohamed Abdelrehim, Amr Sharaf, Vikas Raunak, Mohamed Gabr, Hitokazu Matsushita, Young Jin Kim, Mohamed Afify, and Hany Hassan Awadalla. 2023. How good are gpt models at machine translation? a compre- hensive evaluation. arXiv preprint arXiv:2302.09210 (2023)
2023 arXiv
-
[53]
Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. 2022. Prompt-to-prompt image editing with cross attention control. arXiv preprint arXiv:2208.01626 (2022)
2022 arXiv
-
[54]
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems 30 (2017)
2017
-
[55]
Joseph Hill. 2020. Do deaf communities actually want sign language gloves? Nature Electronics 3, 9 (2020), 512–513
2020
-
[56]
JoAndrea Hoegg, Joseph W Alba, and Darren W Dahl. 2010. The good, the bad, and the ugly: Influence of aesthetics on product feature judgments. Journal of Consumer Psychology 20, 4 (2010), 419–430
2010
-
[57]
Annette Hohenberger, Daniela Happ, and Helen Leuninger. 2002. Modality- dependent aspects of sign language production: Evidence from slips of the hands and their repairs in German Sign Language. Modality and structure in signed and spoken languages (2002), 112–142
2002
-
[58]
Daniel Holden, Oussama Kanoun, Maksym Perepichka, and Tiberiu Popa. 2020. Learned motion matching. ACM Transactions on Graphics (TOG) 39, 4 (2020), 53–1
2020
-
[59]
Sture Holm. 1979. A simple sequentially rejective multiple test procedure. Scandinavian journal of statistics (1979), 65–70
1979
-
[60]
Alain Hore and Djemel Ziou. 2010. Image quality metrics: PSNR vs. SSIM. In 2010 20th international conference on pattern recognition . IEEE, 2366–2369
2010
-
[61]
Jeremy Hsu. 2024. AI can turn text into sign language – but it’s often unintelligi- ble. https://www.newscientist.com/article/2436111-ai-can-turn-text-into-sign- language-but-its-often-unintelligible
2024
-
[62]
Li Hu. 2024. Animate anyone: Consistent and controllable image-to-video synthesis for character animation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 8153–8163
2024
-
[63]
Lianghua Huang, Di Chen, Yu Liu, Yujun Shen, Deli Zhao, and Jingren Zhou
-
[64]
Matt Huenerfauth. 2008. Generating American Sign Language animation: over- coming misconceptions and technical challenges. Universal Access in the Infor- mation Society 6 (2008), 419–434
2008
-
[65]
In Proceedings of the 40th International Conference on Machine Learning
Composer: creative and controllable image synthesis with composable con- ditions. In Proceedings of the 40th International Conference on Machine Learning . 13753–13773
-
[66]
Matt Huenerfauth and Vicki Hanson. 2009. Sign language in the interface: access for deaf signers. Universal Access Handbook. NJ: Erlbaum 38 (2009), 14
2009
-
[67]
Matt Huenerfauth. 2009. A linguistically motivated model for speed and pausing in animations of american sign language. ACM Transactions on Accessible Computing (TACCESS) 2, 2 (2009), 1–31
2009
-
[68]
Matt Huenerfauth, Liming Zhao, Erdan Gu, and Jan Allbeck. 2008. Evaluation of American Sign Language Generation by Native ASL Signers. ACM Transactions on Accessible Computing 1, 1 (May 2008), 1–27. https://doi.org/10.1145/1361203. 1361206
2008 doi
-
[69]
Matt Huenerfauth, Liming Zhao, Erdan Gu, and Jan Allbeck. 2007. Evaluating American Sign Language generation through the participation of native ASL signers. In Proceedings of the 9th international ACM SIGACCESS conference on Computers and accessibility. 211–218
2007
-
[70]
Eui Jun Hwang, Huije Lee, and Jong C Park. 2024. A Gloss-Free Sign Lan- guage Production with Discrete Representation. In 2024 IEEE 18th International Conference on Automatic Face and Gesture Recognition (FG) . IEEE, 1–6
2024
-
[71]
Eui Jun Hwang, Sukmin Cho, Huije Lee, Youngwoo Yoon, and Jong C Park. 2024. Universal Gloss-level Representation for Gloss-free Sign Language Translation and Production. arXiv preprint arXiv:2407.02854 (2024)
2024 arXiv
-
[72]
Apple Inc. 2024. Introducing Apple’s On-Device and Server Foundation Models. https://machinelearning.apple.com/research/introducing-apple- foundation-models
2024
-
[73]
Mert Inan, Katherine Atwell, Anthony Sicilia, Lorna Quandt, and Malihe Alikhani. 2024. Generating Signed Language Instructions in Large-Scale Di- alogue Systems. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistic...
2024
-
[74]
Hamid Reza Vaezi Joze and Oscar Koller. 2018. Ms-asl: A large-scale data set and benchmark for understanding american sign language. arXiv preprint arXiv:1812.01053 (2018)
2018 arXiv
-
[75]
Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros. 2017. Image-to- image translation with conditional adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition . 1125–1134. CHI ’25, April 26-May 1, 2025, Yokohama, Japan Z...
2017
-
[76]
Michael Kipp, Alexis Heloir, and Quan Nguyen. 2011. Sign language avatars: Animation and comprehensibility. InIntelligent Virtual Agents: 10th International Conference, IV A 2011, Reykjavik, Iceland, September 15-17, 2011. Proceedings 11 . Springer, 113–126
2011
-
[77]
Jung-Ho Kim, Eui Jun Hwang, Sukmin Cho, Du Hui Lee, and Jong C Park
-
[78]
Edward S Klima and Ursula Bellugi. 1979. The signs of language . Harvard University Press
1979
-
[79]
Paddy Ladd. 2003. Understanding deaf culture: In search of deafhood. Multilingual Matters
2003
-
[80]
Michael Kipp, Quan Nguyen, Alexis Heloir, and Silke Matthes. 2011. Assessing the deaf user perspective on sign language avatars. In The proceedings of the 13th international ACM SIGACCESS conference on Computers and accessibility . 107–114
2011
-
[81]
Joseph Lee Rodgers and W Alan Nicewander. 1988. Thirteen ways to look at the correlation coefficient. The American Statistician 42, 1 (1988), 59–66
1988
-
[82]
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rock- täschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing...
2020
-
[83]
Sooyeon Lee, Abraham Glasser, Becca Dingman, Zhaoyang Xia, Dimitris Metaxas, Carol Neidle, and Matt Huenerfauth. 2021. American sign language video anonymization to support online participation of deaf and hard of hearing users. In Proceedings of the 23rd International ACM SIG...
2021
-
[84]
Scott K Liddell. 2003. Grammar, gesture, and meaning in American Sign Language. Cambridge University Press
2003
-
[85]
Chin-Yew Lin. 2004. ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out. Association for Computational Linguistics, Barcelona, Spain, 74–81. https://aclanthology.org/W04-1013
2004
-
[86]
Dongxu Li, Chenchen Xu, Xin Yu, Kaihao Zhang, Benjamin Swift, Hanna Suomi- nen, and Hongdong Li. 2020. Tspnet: Hierarchical feature learning via temporal semantic pyramid for sign language translation. Advances in Neural Information Processing Systems 33 (2020), 12034–12045
2020
-
[87]
Ceil Lucas. 2001. The sociolinguistics of sign languages . Cambridge University Press
2001
-
[88]
Camillo Lugaresi, Jiuqiang Tang, Hadon Nash, Chris McClanahan, Esha Uboweja, Michael Hays, Fan Zhang, Chuo-Ling Chang, Ming Guang Yong, Juhyun Lee, et al. 2019. Mediapipe: A framework for building perception pipelines. arXiv preprint arXiv:1906.08172 (2019)
2019 arXiv
-
[89]
Lingjie Liu, Weipeng Xu, Michael Zollhoefer, Hyeongwoo Kim, Florian Bernard, Marc Habermann, Wenping Wang, and Christian Theobalt. 2019. Neural ren- dering and reenactment of human actor videos. ACM Transactions on Graphics (TOG) 38, 5 (2019), 1–14
2019
-
[90]
Yongsen Ma, Gang Zhou, Shuangquan Wang, Hongyang Zhao, and Woosub Jung. 2018. SignFi: Sign language recognition using WiFi. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 2, 1 (2018), 1–21
2018
-
[91]
Aleix M Martínez, Ronnie B Wilbur, Robin Shay, and Avinash C Kak. 2002. Purdue RVL-SLLL ASL database for automatic recognition of American Sign Language. In Proceedings. Fourth IEEE International Conference on Multimodal Interfaces. IEEE, 167–172
2002
-
[92]
Xiaohan Ma, Rize Jin, and Tae-Sun Chung. 2024. Multi-Channel Spatio-Temporal Transformer for Sign Language Production. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024). 11699–11712
2024
-
[93]
David C Mohr, Ken R Weingardt, Madhu Reddy, and Stephen M Schueller. 2017. Three problems with current digital mental health research... and three things we can do about them. Psychiatric services 68, 5 (2017), 427–429
2017
-
[94]
Amit Moryossef, Mathias Müller, Anne Göhring, Zifan Jiang, Yoav Goldberg, and Sarah Ebling. 2023. An open-source gloss-based baseline for spoken to signed language translation. In Proceedings of the Second International Workshop on Automatic Translation for Signed and Spoken L...
2023
-
[95]
Ross E Mitchell and Travas A Young. 2023. How many people use sign language? A national health survey-based estimate. Journal of Deaf Studies and Deaf Education 28, 1 (2023), 1–6
2023
-
[96]
Chong Mou, Xintao Wang, Liangbin Xie, Yanze Wu, Jian Zhang, Zhongang Qi, and Ying Shan. 2024. T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 4296–4304
2024
-
[97]
Laura J Muir and Iain EG Richardson. 2005. Perception of sign language and its application to visual communications for deaf people. Journal of Deaf studies and Deaf education 10, 4 (2005), 390–401
2005
-
[98]
Amit Moryossef, Kayo Yin, Graham Neubig, and Yoav Goldberg. 2021. Data Augmentation for Sign Language Gloss Translation. http://arxiv.org/abs/2105. 07476 arXiv:2105.07476 [cs]
2021 arXiv
-
[99]
Jemina Napier. 2002. Sign language interpreting: Linguistic coping strategies . Douglas McLean
2002
-
[100]
Carol Neidle. 2001. SignStream™: A database tool for research on visual-gestural language. Sign language & linguistics 4, 1-2 (2001), 203–214
2001
-
[101]
Mathias Müller, Zifan Jiang, Amit Moryossef, Annette Rios, and Sarah Ebling
-
[102]
In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , Anna Rogers, Jordan Boyd- Graber, and Naoaki Okazaki (Eds.)
Considerations for meaningful sign language machine translation based on glosses. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , Anna Rogers, Jordan Boyd- Graber, and Naoaki Okazaki (Eds.). Association for ...
-
[103]
Carol Neidle. 2017. A User’s guide to SignStream ® 3. (2017)
2017
-
[105]
Carol Neidle. 2002. Signstream annotation: Addendum to conventions used for the american sign language linguistic research project, Report No. 11. (2002)
2002
-
[106]
Carol Neidle. 2007. Signstream annotation: Addendum to conventions used for the american sign language linguistic research project. (2007)
2007
-
[107]
Form Sign Datasets Carol Neidle and Augustine Opoku. [n. d.].Boston University, Boston. Technical Report. MA Report
-
[108]
deafness
Chijioke Obasi. 2008. Seeing the deaf in “deafness”. Journal of Deaf Studies and Deaf Education 13, 4 (2008), 455–465
2008
-
[109]
Carol Neidle, Augustine Opoku, and Dimitris Metaxas. 2022. Asl video corpora & sign bank: Resources available through the american sign language linguistic research project (asllrp). arXiv preprint arXiv:2201.07899 (2022)
2022 arXiv
-
[110]
Carol Neidle and Christian Vogler. 2012. A new web interface to facilitate access to corpora: Development of the ASLLRP data access interface (DAI). In Proc. 5th Workshop on the Representation and Processing of Sign Languages: Interactions between Corpus and Lexicon, LREC , Vo...
2012
-
[111]
Charles E Osgood. 1964. Semantic differential technique in the comparative study of cultures. American anthropologist 66, 3 (1964), 171–200
1964
-
[112]
Carol A Padden and Tom L Humphries. 1988. Deaf in America: Voices from a culture. Harvard University Press
1988
-
[113]
Why Sign-Language Gloves Don’t Help Deaf People
Visual Anthropology of Japan. 2019. "Why Sign-Language Gloves Don’t Help Deaf People" -and- neither does the "’Woman’s hand’ iPhone case to keep you company" -and then- a couple of new products that were made with deaf collaboration. http://visualanthropologyofjapan.blogspot.c...
2019
-
[114]
Charles E Osgood. 1957. The measurement of meaning. Urbana: University of Illinois Press (1957)
1957
-
[115]
José Pinheiro and Douglas Bates. 2006. Mixed-effects models in S and S-PLUS . Springer science & business media
2006
-
[116]
Maja Popović. 2015. chrF: character n-gram F-score for automatic MT evaluation. In Proceedings of the Tenth Workshop on Statistical Machine Translation , Ondřej Bojar, Rajan Chatterjee, Christian Federmann, Barry Haddow, Chris Hokamp, Matthias Huck, Varvara Logacheva, and Pave...
2015
-
[117]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a Method for Automatic Evaluation of Machine Translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics , Pierre Isabelle, Eugene Charniak, and Dekang Lin (Eds...
2002
-
[118]
Keqin Peng, Liang Ding, Qihuang Zhong, Li Shen, Xuebo Liu, Min Zhang, Yuanxin Ouyang, and Dacheng Tao. 2023. Towards Making the Most of Chat- GPT for Machine Translation. In Findings of the Association for Computational Linguistics: EMNLP 2023. 5622–5633
2023
-
[119]
Alfredo Sánchez, and Josefina Guerrero
Soraia Prietch, J. Alfredo Sánchez, and Josefina Guerrero. 2022. A Systematic Review of User Studies as a Basis for the Design of Systems for Automatic Sign Language Processing. ACM Transactions on Accessible Computing 15, 4 (Dec. 2022), 1–33. https://doi.org/10.1145/3563395
2022 doi
-
[120]
Lorna C Quandt, Athena Willis, Melody Schwenk, Kaitlyn Weeks, and Ruthie Ferster. 2022. Attitudes toward signing avatars vary depending on hearing status, age of signed language acquisition, and avatar type. Frontiers in psychology 13 (2022), 730917
2022
-
[121]
Matt Post. 2018. A call for clarity in reporting BLEU scores. arXiv preprint arXiv:1804.08771 (2018)
2018 arXiv
-
[122]
Soraia Prietch, J Alfredo Sánchez, and Josefina Guerrero. 2022. A systematic review of user studies as a basis for the design of systems for automatic sign language processing. ACM Transactions on Accessible Computing 15, 4 (2022), 1–33
2022
-
[123]
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen
-
[124]
Razieh Rastgoo, Kourosh Kiani, Sergio Escalera, and Mohammad Sabokrou. 2021. Sign Language Production: A Review. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) . IEEE, Nashville, TN, USA, 3446–3456. https://doi.org/10.1109/CVPRW53098.2...
2021
-
[125]
David Quinto-Pozos. 2010. Rates of fingerspelling in american sign language. In Poster presented at 10th Theoretical Issues in Sign Language Research conference, Towards AI-driven SLG with Non-manual Markers CHI ’25, April 26-May 1, 2025, Yokohama, Japan West Lafayette, Indian...
2010
-
[126]
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog 1, 8 (2019), 9
2019
-
[127]
Wendy Sandler and Diane Carolyn Lillo-Martin. 2006. Sign language and lin- guistic universals. Cambridge University Press
2006
-
[128]
arXiv preprint arXiv:2204.06125 1, 2 (2022), 3
Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125 1, 2 (2022), 3
2022 arXiv
-
[129]
Ben Saunders, Necati Cihan Camgoz, and Richard Bowden. 2020. Everybody sign now: Translating spoken language to photo realistic sign language video. arXiv preprint arXiv:2011.09846 (2020)
2020 arXiv
-
[130]
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015. U-net: Convolutional networks for biomedical image segmentation. In Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings,...
2015
-
[131]
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. 2022. Photorealistic text-to-image diffusion models with deep language understanding. Advances in neural infor...
2022
-
[132]
Ben Saunders, Necati Cihan Camgoz, and Richard Bowden. 2022. Signing at Scale: Learning to Co-Articulate Signs for Large-Scale Photo-Realistic Sign Language Production. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, New Orleans, LA, USA, ...
2022
-
[133]
Ben Saunders, Necati Cihan Camgoz, and Richard Bowden. 2020. Adver- sarial training for multi-channel sign language production. arXiv preprint arXiv:2008.12405 (2020)
2020 arXiv
-
[134]
Bowen Shi, Diane Brentari, Greg Shakhnarovich, and Karen Livescu. 2022. Open- domain sign language translation learned from online video. arXiv preprint arXiv:2205.12870 (2022)
2022 arXiv
-
[135]
Ben Saunders, Necati Cihan Camgoz, and Richard Bowden. 2020. Progressive Transformers for End-to-End Sign Language Production. http://arxiv.org/abs/ 2004.14874 arXiv:2004.14874 [cs]
2020 arXiv
-
[136]
Ben Saunders, Necati Cihan Camgoz, and Richard Bowden. 2021. Mixed signals: Sign language production via a mixture of motion primitives. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 1919–1929
2021
-
[137]
Stephanie Stoll, Necati Cihan Camgöz, Simon Hadfield, and Richard Bowden
-
[138]
Ben Saunders, Necati Cihan Camgoz, and Richard Bowden. 2022. Signing at scale: Learning to co-articulate signs for large-scale photo-realistic sign language production. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 5141–5151
2022
-
[139]
Valerie Sutton. 1974. SignWriting. Retrieved online at: http://www. signwriting. org/labout/what/what02. html (1974)
1974
-
[140]
Matthew Snover, Bonnie Dorr, Rich Schwartz, Linnea Micciulla, and John Makhoul. 2006. A Study of Translation Edit Rate with Targeted Human An- notation. In Proceedings of the 7th Conference of the Association for Machine Translation in the Americas: Technical Papers. Associati...
2006
-
[141]
William C. Stokoe. 1961. Sign language structure: an outline of the visual communication systems of the American deaf. 1960. Journal of deaf studies and deaf education 10 1 (1961), 3–37. https://api.semanticscholar.org/CorpusID: 5948293
1961
-
[142]
Nina Tran, Richard E Ladner, and Danielle Bragg. 2023. US Deaf Community Perspectives on Automatic Sign Language Translation. InProceedings of the 25th International ACM SIGACCESS Conference on Computers and Accessibility . 1–7
2023
-
[143]
Narek Tumanyan, Michal Geyer, Shai Bagon, and Tali Dekel. 2023. Plug-and-play diffusion features for text-driven image-to-image translation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 1921–1930
2023
-
[144]
Stephanie Stoll, Necati Cihan Camgoz, Simon Hadfield, and Richard Bowden
-
[145]
David Uthus, Garrett Tanzer, and Manfred Georg. 2023. YouTube-ASL: A Large-Scale, Open-Domain American Sign Language-English Parallel Corpus. arXiv:2306.15162 https://arxiv.org/abs/2306.15162
2023 arXiv
-
[146]
Clayton Valli and Ceil Lucas. 2000. Linguistics of American sign language: An introduction. Gallaudet University Press
2000
-
[147]
Garrett Tanzer, Maximus Shengelia, Ken Harrenstien, and David Uthus. 2024. Reconsidering Sentence-Level Sign Language Translation. arXiv preprint arXiv:2406.11049 (2024)
2024 arXiv
-
[148]
Noam Tractinsky, Adi S Katz, and Dror Ikar. 2000. What is beautiful is usable. Interacting with computers 13, 2 (2000), 127–145
2000
-
[149]
Tan Wang, Linjie Li, Kevin Lin, Yuanhao Zhai, Chung-Ching Lin, Zhengyuan Yang, Hanwang Zhang, Zicheng Liu, and Lijuan Wang. 2024. Disco: Disentangled control for realistic human dance generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognit...
2024
-
[150]
Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Guilin Liu, Andrew Tao, Jan Kautz, and Bryan Catanzaro. 2018. Video-to-Video Synthesis. Advances in Neural Information Processing Systems 31 (2018)
2018
-
[151]
Thomas Unterthiner, Sjoerd Van Steenkiste, Karol Kurach, Raphael Marinier, Marcin Michalski, and Sylvain Gelly. 2018. Towards accurate generative models of video: A new metric & challenges. arXiv preprint arXiv:1812.01717 (2018)
2018 arXiv
-
[152]
WFD. 2022. World Federation of the Deaf. https://wfdeaf.org/our-work/
2022
-
[153]
Elizabeth A Winston. 1991. Spatial referencing and cohesion in an American Sign Language text. Sign language studies 73, 1 (1991), 397–410
1991
-
[154]
Adele Vogel and Jessica L Korte. 2024. What Factors Motivate Culturally Deaf People to Want Assistive Technologies?. InExtended Abstracts of the CHI Con- ference on Human Factors in Computing Systems . 1–7
2024
-
[155]
Harry Walsh, Ben Saunders, and Richard Bowden. 2024. Sign Stitching: A Novel Approach to Sign Language Production. http://arxiv.org/abs/2405.07663 arXiv:2405.07663 [cs]
2024 arXiv
-
[156]
Ling Yang, Zhilong Zhang, Yang Song, Shenda Hong, Runsheng Xu, Yue Zhao, Wentao Zhang, Bin Cui, and Ming-Hsuan Yang. 2023. Diffusion models: A comprehensive survey of methods and applications. Comput. Surveys 56, 4 (2023), 1–39
2023
-
[157]
Aoxiong Yin, Tianyun Zhong, Li Tang, Weike Jin, Tao Jin, and Zhou Zhao. 2023. Gloss attention for gloss-free sign language translation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 2551–2562
2023
-
[158]
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. 2004. Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing 13, 4 (2004), 600–612
2004
-
[159]
Albina Zakharenko. 2023. Semantic Differential Scale: Definition, Questions, Examples. https://aidaform.com/blog/semantic-differential-scale-definition- examples.html
2023
-
[160]
Han Zhang, Vedant Das Swain, Leijie Wang, Nan Gao, Yilun Sheng, Xuhai Xu, Flora D Salim, Koustuv Saha, Anind K Dey, and Jennifer Mankoff. 2024. Illuminating the Unseen: A Framework for Designing and Mitigating Context- induced Harms in Behavioral Sensing. arXiv preprint arXiv:...
2024 arXiv
-
[161]
Pan Xie, Qipeng Zhang, Peng Taiying, Hao Tang, Yao Du, and Zexian Li. 2024. G2P-DDM: Generating Sign Pose Sequence from Gloss Sequence with Discrete Diffusion Model. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 6234–6242
2024
-
[162]
Shangqing Xu and Chao Zhang. 2024. Misconfidence-based demonstration selection for llm in-context learning. arXiv preprint arXiv:2401.06301 (2024)
2024 arXiv
-
[163]
Dele Zhu, Vera Czehmann, and Eleftherios Avramidis. 2023. Neural Machine Translation Methods for Translating Text to Sign Language Glosses. In Proceed- ings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Anna Rogers, Jordan...
2023 doi
-
[164]
+” signs), number of signing hands (annotated as “(1h)
Jason E Zinza. 2006. Master ASL. Sign Media Inc (2006). CHI ’25, April 26-May 1, 2025, Yokohama, Japan Zhang et al. A A Review of American Sign Language and Publicly Available Datasets Similar to other sign languages, ASL is also a visual-based natu- ral language, expressed by...
2006
-
[165]
Morteza Zahedi, Daniel Keysers, Thomas Deselaers, and Hermann Ney. 2005. Combination of tangent distance and an image distortion model for appearance- based sign language recognition. In Pattern Recognition: 27th DAGM Symposium, Vienna, Austria, August 31-September 2, 2005. Pr...
2005
-
[168]
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. 2023. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE/CVF Inter- national Conference on Computer Vision . 3836–3847
2023
-
[169]
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang
-
[170]
In Proceedings of the IEEE conference on computer vision and pattern recognition
The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition . 586–595
-
[2018]
In BMVC, Vol
Sign Language Production using Neural Machine Translation and Gener- ative Adversarial Networks.. In BMVC, Vol. 2019. 1–12
2019
-
[2020]
International Journal of Computer Vision 128, 4 (April 2020), 891–908
Text2Sign: Towards Sign Language Production Using Neural Machine Translation and Generative Adversarial Networks. International Journal of Computer Vision 128, 4 (April 2020), 891–908. https://doi.org/10.1007/s11263- 019-01281-2
2020 doi
-
[2022]
In Proceedings of the Thirteenth Language Resources and Evaluation Conference
Sign language production with avatar layering: A critical use case over rare words. In Proceedings of the Thirteenth Language Resources and Evaluation Conference. 1519–1528
-
[2023]
Diffusion models in vision: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 9 (2023), 10850–10869
2023
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.