Pith. sign in

REVIEW 1 cited by

EmojiVoice: Towards long-term controllable expressivity in robot speech

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.15085 v2 pith:KAAEWOSP submitted 2025-06-18 cs.RO cs.HC

classification cs.ROcs.HC
keywords speechexpressivityexpressiverobotrobotssocialassistantcase
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Humans vary their expressivity when speaking for extended periods to maintain engagement with their listener. Although social robots tend to be deployed with ``expressive'' joyful voices, they lack this long-term variation found in human speech. Foundation model text-to-speech systems are beginning to mimic the expressivity in human speech, but they are difficult to deploy offline on robots. We present EmojiVoice, a free, customizable text-to-speech (TTS) toolkit that allows social roboticists to build temporally variable, expressive speech on social robots. We introduce emoji-prompting to allow fine-grained control of expressivity on a phase level and use the lightweight Matcha-TTS backbone to generate speech in real-time. We explore three case studies: (1) a scripted conversation with a robot assistant, (2) a storytelling robot, and (3) an autonomous speech-to-speech interactive agent. We found that using varied emoji prompting improved the perception and expressivity of speech over a long period in a storytelling task, but expressive voice was not preferred in the assistant use case.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Prosody of Emojis

    cs.CL 2025-08 conditional novelty 6.0 of 10

    Speakers systematically modify prosody when reading sentences with different emojis, and listeners can recover the intended emoji from prosody alone, with larger semantic differences yielding larger prosodic shifts.

Pith tools