Pith. sign in

REVIEW 3 major objections 6 minor 29 references

Using the Pepper Robot to Support Sign Language Communication

T0 review · 3 major / 6 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read A standard Pepper robot can produce intelligible Italian Sign Language signs for a limited, co-designed vocabulary, but sentence-level signing remains beyond it.

desk verdict Honest limited empirical step for LIS on Pepper, with a clear selection caveat that caps the generalization but does not sink the modest claim. read the letter →

arxiv 2509.09889 v1 pith:OLNIP4NA submitted 2025-09-11 cs.RO cs.HC

classification cs.ROcs.HC
keywords socialrobotsItalianSignLanguageLISPepperrobothuman-robotinteractionrecognitioninversekinematicsaccessibility
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper investigates whether a commercially available social robot, Pepper, can produce intelligible Italian Sign Language (LIS) signs. Working with a Deaf student and his interpreter, the authors co-designed and implemented 52 signs, then asked 12 proficient LIS users to recognize 15 isolated signs and 2 short signed sentences. The majority of isolated signs were recognized correctly—10 of the 15 significantly above the 25% chance level—while full-sentence recognition was very poor, with only 2 of 24 responses correct. The paper's central claim is that even a mass-produced robot can convey a limited LIS vocabulary intelligibly, but cannot yet support sentence-level signed communication. The authors argue this opens a path toward more inclusive human-robot interaction in public settings, provided multimodal support and participatory design are added.

What carries the argument

The central object is the Pepper humanoid's constrained embodiment—its five fingers move only together, its wrist rotation and elbow mobility are limited, and a chest-mounted tablet blocks torso-close gestures—together with the two sign-authoring pipelines used: a manual keyframe Animation Editor and a MATLAB numerical inverse-kinematics solver that converts 3D hand trajectories into joint values and .qianim animation files. The co-design loop with a Deaf student and interpreter is what filters the 52 signs: each sign is implemented, shown, and iteratively revised until the Deaf collaborators accept it, so the sign set is shaped by both linguistic validity and Pepper's physical affordances.

What would settle it

Run the same recognition questionnaire on a set of 15 signs drawn at random from the 52 (or from the dictionary) without co-designer filtering, and also readminister the sentence test with proper spatial accord added. If random-sign recognition falls to chance, or if accord-full sentences remain near-zero, the paper's 'good intelligibility' and 'sentences fail' claims have to be restated.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that a standard Pepper robot, whose fingers can only open and close as a group, can nonetheless perform a substantial subset of LIS signs with good intelligibility. In a recognition test with 12 Deaf and hearing LIS users, signs such as 'forget,' 'done,' 'shampoo,' and 'university' were recognized perfectly, raising the question of which articulatory features are truly load-bearing for sign recognition. At the same time, combining signs into short sentences dropped comprehension near chance, showing that lexical intelligibility does not by itself scale to utterance-level communication. The authors also show that an inverse-kinematics pipeline can ge

Load-bearing premise

The conclusion rests on the assumption that the 15 signs hand-picked by the Deaf student and interpreter for the test fairly represent the 52 implemented signs and what Pepper can likely sign, and that omitting LIS spatial accord does not change the sentence-difficulty result.

Editorial extensions

If this is right

  • A commercial social robot can act as a partial LIS communicator for isolated words in domains like schools, museums, or information points.
  • Sentence-level signed interaction on this class of hardware is not achievable with arm movements alone; multimodal channels (tablet, subtitles, avatars) will be required.
  • Automated inverse-kinematics generation makes sign authoring fast enough to expand the tested vocabulary, but only for signs whose handshapes are all-fingers-open or all-closed.
  • Recognition of iconic, simplified signs is high while articulation-ambiguous signs fail, so robot sign sets must be selected by Deaf users, not by dictionary alone.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the 15 tested signs were selected with the Deaf co-designers for Pepper-compatibility, the reported 83% isolated-sign accuracy likely overstates how well Pepper would do on a random LIS sample; a random-sample test would separate selection effects from genuine intelligibility.
  • The near-zero sentence performance may be largely attributable to the deliberate omission of spatial accord, which assigns grammatical roles; restoring it would likely lower sentence recognition still further, meaning the robot's communicative ceiling is below what the paper's favorable framing implies.
  • The success of the IK pipeline suggests that motion-capture or avatar-generated trajectories could be retargeted to Pepper at scale, but only for the open/closed finger subset; robots with independently articulable fingers would be the natural next step to test whether sentence-level signing becomes feasible.
  • If a robot is to serve Deaf users in public spaces, the results imply that its signing channel should be considered a fallback or supplement to the screen, not a standalone language channel—a design constraint that argues for co-locating sign output with visual text.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper investigates whether the Pepper humanoid robot can produce intelligible isolated signs and short sentences of Italian Sign Language (LIS). The authors co-designed 52 signs with a Deaf student and his LIS interpreter, using both manual animation and a semi-automated inverse-kinematics pipeline, and then ran an online questionnaire with 12 LIS-proficient participants: 15 four-alternative forced-choice sign recognition items and 2 open-ended sentence interpretation items. They report 119/180 correct sign choices overall, 10/15 signs significant at the nominal 5% level, and much lower sentence-level comprehension. The paper concludes that Pepper can perform a substantial number of LIS signs with good intelligibility, while acknowledging technical limitations and the exploratory nature of the study.

Significance. If the claims are accepted, the paper provides a useful data point for inclusive HRI: a widely available commercial robot can be made to produce a non-trivial subset of signs that LIS users can recognize, with honest reporting of per-sign failures (e.g., Profumo 0%, Insegnare 33%). The participatory co-design and public availability of sign videos are strengths, as is the reported time saving of the IK toolchain. However, the main generalization to 'a substantial number of LIS signs' goes beyond the evaluated 15-sign subset, and the sentence-level result is confounded by the deliberate omission of spatial accord. The contribution is therefore best framed as a feasibility study of a curated subset, not as an established capability for general LIS communication.

major comments (3)
  1. [Section 4.2, Table 1, Section 5] The 15 questionnaire signs were chosen with the help of the LIS interpreter and the Deaf student — the same individuals who co-designed the 52 signs and, per the Acknowledgments, discarded signs that Pepper did not perform correctly. The remaining 37 implemented signs were never evaluated. Thus the aggregate 119/180 and the per-sign p-values support intelligibility of this curated, likely easy-to-animate subset, not the Section 5 claim that Pepper 'is capable of performing a substantial number of LIS signs with a good level of intelligibility.' In addition, the 15 binomial tests are uncorrected; with a Bonferroni threshold (0.05/15 ≈ 0.0033) only 6 of 15 signs remain individually significant. Please either test a representative/random sample of the 52 signs, or substantially qualify the conclusion to the evaluated subset; report multiplicity-corrected results. The limitation paragraph in
  2. [Section 3.2] The feasibility claim for the automated pipeline (38/52 signs) relies on the assertion that the elbow-movement difference between manual and automated signs 'had no impact on the recognition ability of the Deaf student and his sign language interpreter.' This is an anecdotal assessment by the two co-designers, not a measurement from the 12-participant study. Table 1 does not report which evaluated signs were manual vs. automated, so the user study cannot confirm equivalence. Please provide the manual/automated breakdown of the 15 evaluated signs, or temper the claim to anecdotal evidence from the co-design phase.
  3. [Section 4, open-ended questions] The sentence stimuli omit LIS spatial accord by design, and the authors assert this simplification 'so not impact the findings' without supporting evidence. Spatial accord is a grammatical device for marking syntactic roles; removing it changes the linguistic stimulus and may directly lower sentence comprehension. Consequently, the Section 5 conclusion that Pepper has limited capacity to convey temporal and syntactic structure is not supported: the low sentence recognition (2/24 correct) could reflect the missing grammatical device, lexical confusions (e.g., mela/acqua), or timing, rather than Pepper's embodiment per se. A control condition with correctly spatially-marked sentences, or a human-signing baseline, is needed before attributing the deficit to the robot.
minor comments (6)
  1. [Throughout] There are numerous OCR-style typos and formatting artifacts (e.g., 'Communica7on', 'Assis3ve', '0me', 'An.pa.co', 'Fa_o'); please proofread the final manuscript carefully.
  2. [Abstract, Section 5] Grammar: 'a exploratory user study' should be 'an exploratory user study'; 'Results shows' should be 'Results show'.
  3. [Table 1] The column header 'Binomial value p-' is unclear; also the note for Spaventarsi ('None of the above answer, 50% selected it correctly') is ambiguous and should be rewritten.
  4. [Section 3.1] In the description of the Andare sign, 'the leg arm remains stationary' appears to be a typo for 'the left arm remains stationary'.
  5. [Section 3.2] Clarify why the 10 'asymmetric or body-involving' signs could not be reproduced by the pipeline — is this due to the single-arm URDF model? This is relevant for assessing the generality of the toolchain.
  6. [Data availability] The GitHub repository currently contains videos of signs. Consider also releasing the qianim animation files and the IK exporter script to support reproducibility and reuse.

Circularity Check

1 steps flagged · score 3.0 of 10

Test-set selection by the same co-designers who built and filtered the signs makes the 'substantial number of LIS signs' conclusion partly self-confirming.

  1. other [Section 4.2 (Results, Single-Choice Questions) and Acknowledgments]
    "The 15 presented signs were chosen among the implemented signs with the help of the LIS interpreter and the Deaf student. ... We are par3cularly grateful to Flavio and his LIS interpreter Simoneaa for the considerable help in choosing which signs to implement, generate them and discard the ones that Pepper was not performing correctly."

    The evaluation set is not an independent sample of the 52 implemented signs: the same Deaf student and interpreter who co-designed and iteratively refined the signs, and who discarded signs that Pepper did not perform correctly, selected the 15 signs used in the recognition test. The later finding that most of these tested signs were intelligible, and the conclusion that Pepper can perform 'a substantial number of LIS signs with a good level of intelligibility', therefore partly reflects a prior filtering step that removed likely-failing signs before any participant saw them. The unselected 37 signs are never evaluated, so the 'substantial number' generalization is built on a survivorship-biased stimulus set chosen by the people whose design work is being validated.

full rationale

The paper's central evaluation is an independent 12-participant recognition study, so the results are not a pure logical tautology: participants' responses are genuine empirical data, and some signs (e.g., Profumo) scored at chance. No parameter is fitted to the evaluation data, and no load-bearing self-citation or uniqueness theorem is invoked. The main circularity concern is the construction of the test set: Section 4.2 states that the 15 signs were chosen with the help of the same LIS interpreter and Deaf student who, according to the Acknowledgments, also chose which signs to implement and discarded ones Pepper performed incorrectly. This creates a selection channel: the tested signs are a curated subset already judged feasible and satisfactory by the designers, so the confirmation that 'the majority of isolated signs were recognized correctly' is partly an artifact of that curation. Additionally, Section 4 asserts without support that omitting LIS spatial accord 'does not impact the findings'; this is a validity threat rather than a circularity, but it weakens the sentence-level interpretation. Overall, the recognition data are informative about the 15 tested signs, but the generalization to a 'substantial number of LIS signs' is undermined by the non-representative, insider-selected sample, giving a moderate circularity score of 3.

Assumptions & free parameters 2 free parameters · 5 assumptions · 0 invented entities

The paper makes no derivational claim, so there is no fitting hidden as prediction. The central claim rests on engineering choices (IK weights, timing), on the definition of the feasible sign subset (open/closed hands only), and on evaluation-design assumptions: hand-picked signs, neglected spatial accord, video-based rather than live testing. These are stated in the paper, which is to its credit, but they are load-bearing.

free parameters (2)
  • IK solver weights (per sign)
    Passed as user-defined weights to the MATLAB inverseKinematics solver (Section 3.2); values are not reported, and they influence joint solutions and therefore sign shape. They are implementation choices, not fitted to recognition data.
  • Movement timing per sign
    User-defined timing parameters in the Python exporter (Section 3.2); timing is known to affect sign perception but values are not reported and were not fitted to the evaluation results.
assumptions (5)
  • standard math A numerical IK solution exists for each target hand trajectory within Pepper's kinematic constraints; solver convergence implies feasible motion.
    Invoked in Section 3.2 via MATLAB inverseKinematics; 4 signs failed due to kinematic infeasibility or solver instability, so this axiom is partially violated and was treated as an empirical filter.
  • domain assumption LIS can be approximated by torso-arm joint trajectories with all-fingers-open/closed hand states, neglecting finger differentiation and facial non-manual markers.
    Stated in Section 3.1 (fingers not independently controllable; only signs with open/closed hand configurations were implemented). This defines the feasible sign subset.
  • ad hoc to paper The 15 signs chosen with the co-design experts represent the implemented set well enough for the aggregate 'majority recognized' claim.
    Section 4.2: 'The 15 presented signs were chosen among the implemented signs with the help of the LIS interpreter and the Deaf student'; not a random sample.
  • ad hoc to paper Neglecting spatial accord among signs does not distort the sentence-level result.
    Section 4: 'for sake of simplicity, we are neglecting the spatial accord... We believe that this simplification so not impact the findings of this exploratory experiment.' This is a flagged assumption and directly affects the open-ended sentence claims.
  • domain assumption Video-based recognition by 12 LIS-proficient participants is a valid proxy for intelligibility of live robot signing.
    Section 4; all evaluations were on videos in an online questionnaire, not live interaction; contextual cues from the physical robot (size, gaze, tablet) are absent.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Using the Pepper Robot to Support Sign Language Communication." pith.science (2026). https://pith.science/paper/OLNIP4NA

@misc{pith2026250909889,
  author       = {Pith},
  title        = {Pith review of: Using the Pepper Robot to Support Sign Language Communication},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OLNIP4NA}},
  note         = {Machine review of arXiv:2509.09889}
}
read the original abstract

Social robots are increasingly experimented in public and assistive settings, but their accessibility for Deaf users remains quite underexplored. Italian Sign Language (LIS) is a fully-fledged natural language that relies on complex manual and non-manual components. Enabling robots to communicate using LIS could foster more inclusive human robot interaction, especially in social environments such as hospitals, airports, or educational settings. This study investigates whether a commercial social robot, Pepper, can produce intelligible LIS signs and short signed LIS sentences. With the help of a Deaf student and his interpreter, an expert in LIS, we co-designed and implemented 52 LIS signs on Pepper using either manual animation techniques or a MATLAB based inverse kinematics solver. We conducted a exploratory user study involving 12 participants proficient in LIS, both Deaf and hearing. Participants completed a questionnaire featuring 15 single-choice video-based sign recognition tasks and 2 open-ended questions on short signed sentences. Results shows that the majority of isolated signs were recognized correctly, although full sentence recognition was significantly lower due to Pepper's limited articulation and temporal constraints. Our findings demonstrate that even commercially available social robots like Pepper can perform a subset of LIS signs intelligibly, offering some opportunities for a more inclusive interaction design. Future developments should address multi-modal enhancements (e.g., screen-based support or expressive avatars) and involve Deaf users in participatory design to refine robot expressivity and usability.

Figures

Figures reproduced from arXiv: 2509.09889 by the authors.

Figure 1
Figure 1. Fig.1: Star0ng and ending posi0ons for the Italian word amare (to love) in LIS. [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Fig.2: Star0ng and ending posi0ons for the Italian word andare (to go) in LIS. [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Fig.3: A fragment of a qianim [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

29 extracted references · 2 canonical work pages

  1. [2]

    In: 2019 28th IEEE Interna3onal Conference on Robot and Human Interac3ve Communica3on (RO-MAN)

    Axelsson, M., Racca, M., Weir, D., Kyrki, V.: A par3cipatory design process ofa robo3c tutor of assis3ve sign language for children with au3sm. In: 2019 28th IEEE Interna3onal Conference on Robot and Human Interac3ve Communica3on (RO-MAN). IEEE (Oct

  2. [4]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Paaern Recogni3on (CVPR) (2023), haps://arxiv.org/abs/2304.10482

    Biswas, S., Kulkarni, A., et al.: Sgnify: Linguis3cally-aware expressive 3davatar genera3on from sign language videos. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Paaern Recogni3on (CVPR) (2023), haps://arxiv.org/abs/2304.10482

  3. [5]

    Bolkart, T., Johnson, B.D., Habermann, M., et al.: Neural sign actors: A diffusion model for 3d sign language produc3on from text (2023), haps://arxiv.org/abs/2312.02702, arXiv preprint arXiv:2312.02702

  4. [6]

    (eds.): Gramma3ca della lingua dei segni italiana(LIS)

    Branchini, C., Mantovan, L. (eds.): Gramma3ca della lingua dei segni italiana(LIS). Ca’ Foscari, Venezia (2022)

  5. [7]

    Qualita3ve Researchin Psychology 3(2), 77–101 (2006)

    Braun, V., Clarke, V.: Using thema3c analysis in psychology. Qualita3ve Researchin Psychology 3(2), 77–101 (2006). haps://doi.org/10.1191/1478088706qp063oa

  6. [8]

    Springer handbook ofrobo3cs pp

    Breazeal, C., Dautenhahn, K., Kanda, T.: Social robo3cs. Springer handbook ofrobo3cs pp. 1935–1972 (2016)

  7. [9]

    (ed.): Sign Languages

    Brentari, D. (ed.): Sign Languages. Cambridge University Press, Cambridge (2010)

  8. [10]

    Camgoz, N.C., Koller, O., Hadfield, S., Bowden, R.: Mixed signals: Sign language produc3on via a mixture of mo3on primi3ves (2021), haps://arxiv.org/abs/2107.11317, arXiv preprint arXiv:2107.11317

Show all 29 references
  1. [11]

    In: Eghimiou, E., et al

    Cormier, K., Crasborn, O.A., Bank, R.: Digging into signs: Emerging annota3onstandards for sign language corpora. In: Eghimiou, E., et al. (eds.) Proceedings of the 7th Workshop on the Representa3on and Processing of Sign Languages: Corpus Mining. pp. 35–40. ELRA, Portorož, Sl...

  2. [12]

    Fong, T., Nourbakhsh, I., Dautenhahn, K.: A survey of socially interac3ve robots.Robo3cs and autonomous systems 42(3-4), 143–166 (2003)

  3. [13]

    Gago, J.J., Vasco, V., Łukawski, B., Paaacini, U., Tikhanoff, V., Victores, J.G.,Balaguer, C.: Sequence-to-sequence natural language to humanoid robot sign language (2019), haps://doi.org/10.48550/arXiv.1907.04198, preprint

  4. [14]

    Rebaudengo Fossano

    Geraci, C., Mazzei, A.: Last train to “Rebaudengo Fossano”: The case of somenames in avatar transla3on. In: Crasborn, O., Eghimiou, E., Fo3nea, S.E., Hanke, T., Hochgesang, J.A., Kristoffersen, J., Mesch, J. (eds.) Proceedings of the LREC2014 6th Workshop on the Representa3on a...

  5. [15]

    They convey communica0ve intent both verbally and non-verbally, recognize and express emo0ons, and support users, including those with impairments, in everyday social interac0ons

    are designed to engage with humans on a social level. They convey communica0ve intent both verbally and non-verbally, recognize and express emo0ons, and support users, including those with impairments, in everyday social interac0ons. As social robots become more capable and ac...

  6. [16]

    CRC Press(2017)

    Kanda, T., Ishiguro, H.: Human-robot interac3on in social robo3cs. CRC Press(2017)

  7. [17]

    Li, X., Meng, Y ., Wang, X., et al.: Signavatar: A holis3c 3d sign language produc3on framework with a new benchmark (2024), haps://arxiv.org/abs/2405.07974, arXiv preprint arXiv:2405.07974

  8. [18]

    Human–Computer Interac3on 39(1-2), 109–143 (2024)

    Lieto, A., Striani, M., Gena, C., Dolza, E., Marras, A.M., Pozzato, G.L., Damiano, R.: A sensemaking system for grouping and sugges3ng stories from mul3ple affec3ve viewpoints in museums. Human–Computer Interac3on 39(1-2), 109–143 (2024)

  9. [19]

    Interna3onal Journal of Social Robo3cs 14, 1–40 (04 2022)

    Liu, B., Teaeroo, D., Markopoulos, P .: A systema3c review of experimental workon persuasive social robots. Interna3onal Journal of Social Robo3cs 14, 1–40 (04 2022). haps://doi.org/10.1007/s12369-022-00870-5

  10. [20]

    Macis, D., Perilli, S., Gena, C.: Employing socially assis3ve robots in elderly care.In: Adjunct proceedings of the 30th ACM conference on user modeling, adapta3on and personaliza3on. pp. 130–138 (2022)

  11. [21]

    In: CLiC-it 2023 (2023)

    Marchisio, M., Mazzei, A., Sammaruga, D.: Introducing deep learning with dataaugmenta3on and corpus construc3on for lis. In: CLiC-it 2023 (2023)

  12. [23]

    In: Proceedingsof the Seventh Interna3onal Natural Language Genera3on Conference

    Mazzei, A.: Sign language genera3on with expert systems and ccg. In: Proceedingsof the Seventh Interna3onal Natural Language Genera3on Conference. p. 105–109. INLG ’12, Associa3on for Computa3onal Linguis3cs, USA (2012)

  13. [26]

    Deaf around the world: The impact of language pp

    Padden, C.: Sign language geography. Deaf around the world: The impact of language pp. 19–37 (2010)

  14. [27]

    Kappa, Roma (1992), con il contributo del Mason Perkins Deafness Fund e dell’ Associazione Nazionale Logopedis3

    Radutzky, E., Torossi, C.: Dizionario bilingue elementare della lingua italiana deisegni: oltre 2.500 significa3. Kappa, Roma (1992), con il contributo del Mason Perkins Deafness Fund e dell’ Associazione Nazionale Logopedis3

  15. [28]

    Cambridge University Press (2006)

    Sandler, W., Lillo-Mar3n, D.C.: Sign language and linguis3c universals. Cambridge University Press (2006)

  16. [29]

    In: 2021 30th IEEE Interna3onal Conference on Robot & Human Interac3ve Communica3on (RO-MAN)

    Stoeva, D., Frijns, H.A., Gelautz, M., Schürer, O.: Analy3cal solu3on of pepper’sinverse kinema3cs for a pose matching imita3on system. In: 2021 30th IEEE Interna3onal Conference on Robot & Human Interac3ve Communica3on (RO-MAN). pp. 167–174. IEEE (2021)

  17. [30]

    IAESInterna3onal Journal of Robo3cs and Automa3on (IJRA) 12(4), 405–411 (2023)

    Su3kno, T.: An overview of emerging trends in robo3cs and automa3on. IAESInterna3onal Journal of Robo3cs and Automa3on (IJRA) 12(4), 405–411 (2023)

  18. [31]

    Thunberg, S., Ziemke, T.: Are people ready for social robots in public spaces? In: Companion of the 2020 ACM/IEEE Interna3onal Conference on Human-Robot Interac3on. pp. 482–484 (2020)

  19. [32]

    Xu, Z., Chen, M., Ren, X., et al.: Ms2sl: Mul3modal spoken datadriven con3nuous sign language produc3on via sequen3al diffusion (2024), haps://arxiv.org/abs/2407.12842, arXiv preprint arXiv:2407.12842

  20. [2023]

    40:1–40:4 (2023)

    pp. 40:1–40:4 (2023)

  21. [2024]

    Preprint on arXiv:2310.20436

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.