{"id":"68a0d385-a903-48ed-84ca-1d96bf233f22","arxiv_id":"2506.15085","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"EmojiVoice enables phrase-level emoji-prompted expressive TTS for robots, and user studies show varied emoji voices improve perceived expressivity in storytelling but not in assistant dialogues.","lead":"EmojiVoice is an open-source voice toolkit that lets robot researchers control a text-to-speech system phrase by phrase using emojis, generating real-time speech offline. In a study of 24 adults, varying the emoji-selected voice across phrases made a storytelling robot seem more expressive and engaging, though a steady 'pleasant' voice was preferred for a helper robot.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Fixed voice order in Case Study 2 confounds the central storytelling claim: Emoji is always last after Pleasant and Baseline, with the same 24 participants who already rated all voices in Case Study 1, so perceived expressivity gains may be an order effect.","rationale":"The reader's verdict of CONDITIONAL is appropriate. The reader's weakest_assumption already mentioned the fixed presentation order as part of a broader list, and my stress-test singles out that order effect as the most load-bearing. I agree that the toolkit is useful, the system runs in real time, and the qualitative and quantitative data are honestly reported. However, the central claim about storytelling expressivity rests on a within-subjects comparison with a fixed order and a reused participant pool, which is a standard and serious confound in HRI studies. A counterbalanced replication is a concrete, feasible check that would settle the issue. The reader's conditionality (requiring counterbalanced or acoustic validation) already captures this, so no change to the verdict is needed. I do not see a need to reject the paper, since the toolkit contribution stands independently, and the confound specifically threatens the causal interpretation of the storytelling result rather than the existence or practical value of the toolkit.","tokens_in":12915,"tokens_out":3472,"duration_ms":36750,"concrete_test":"Run a replication of Case Study 2 with a fresh participant pool (n ≈ 24) and counterbalanced or randomized order of the three voices (Baseline, Pleasant, Emoji), using the same story, surveys, and robot setup. If the Emoji voice still achieves significantly higher xMOS and first-choice preference than Pleasant after accounting for order (e.g., Latin square design or order as a between-subjects factor), the central claim survives. If the effect diminishes, disappears, or reverses, the fixed-order confound is the likely driver of the original result.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline finding is that 'using varied emoji prompting improved the perception and expressivity of speech over a long period in a storytelling task' (abstract). The evidence is Case Study 2, where the Emoji voice scored significantly higher than the Pleasant voice on expressiveness (xMOS) and was preferred as first choice (Table III). However, Section IV-B.1 states the order of the voices was fixed: Pleasant, Baseline, Emoji. The Emoji condition always immediately followed the other two. Section IV-B.2 further states that Case Study 2 used the same participant pool as Case Study 1, meaning all 24 participants had already heard and rated all three voices in a different task. This creates two confounds: (1) presentation order is perfectly confounded with condition, so the ANOVA/Tukey results cannot separate a genuine effect of emoji prompting from order effects such as fatigue, contrast, or recency; and (2) prior exposure from Case Study 1 may carry over, making the storytelling ratings non-independent and potentially biased by earlier preferences. The pattern of results is consistent with an order-based explanation: in Case Study 1 the order was Baseline, Emoji, Pleasant and Pleasant was preferred; in Case Study 2 the order was Pleasant, Baseline, Emoji and Emoji was preferred. The paper acknowledges it cannot disentangle general expressivity from context appropriateness (Section V-A), but it does not address the fixed-order confound. Without a counterbalanced design, the central causal claim that emoji prompting causes the observed expressivity improvement is not secure.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents EmojiVoice, an open-source text-to-speech toolkit built on Matcha-TTS, and proposes emoji prompting as a method for phrase-level control of expressivity in robot speech. It reports three case studies: a scripted robot assistant conversation (Case Study 1), a storytelling task (Case Study 2), and an autonomous speech-to-speech interactive agent (Case Study 3). The central claim is that phrase-by-phrase emoji variation improves perceived expressivity relative to a constant 'joyful' voice in storytelling, while varied expressivity is not preferred in the assistant use case. The paper also documents the fine-tuning recipe, reports real-time performance, and releases code, checkpoints, and data.","tokens_in":13256,"tokens_out":4594,"duration_ms":48169,"significance":"If the empirical claim were established, the paper would offer a practical and timely contribution to HRI: a lightweight, open-source, controllable TTS system that runs in real time on modest hardware. The release of checkpoints, training scripts, and documentation is a concrete strength, as is the attempt to evaluate the system in multiple robot interaction contexts. However, the central empirical claim is currently undermined by a fixed voice order in Case Study 2 and by reuse of the same participant pool across Case Studies 1 and 2. The toolkit contribution remains plausible, but the validation of the headline storytelling result needs substantial improvement.","major_comments":[{"comment":"The storytelling result that carries the paper's central claim is confounded by presentation order. Section IV-B.1 states 'The order of the voices were: Pleasant, Baseline, Emoji,' so the Emoji condition was always the last of the three within-subject presentations; Section IV-B.2 reports that Case Study 2 'used the same participant pool as Case Study 1,' so all 24 participants had already heard and rated the same voices in a different task. The observed pattern across the two studies is exactly what an order or contrast effect would predict: in Case Study 1, with order Baseline, Emoji, Pleasant, the Pleasant voice was preferred; in Case Study 2, with order Pleasant, Baseline, Emoji, the Emoji voice was preferred. A counterbalanced design, a fresh participant pool, or a direct statistical test for order effects is required before the perceived-expressivity advantage can be attributed to emoji prompting.","section":"IV-B.1 and IV-B.2"},{"comment":"The statistical analyses in Tables II and III appear to treat the data as independent observations even though every participant rated all three voices within each case study, and the same participants were reused across Case Studies 1 and 2. If a between-subjects one-way ANOVA was used, it violates the independence assumption; if a repeated-measures ANOVA was intended, the paper should report the full model, including error degrees of freedom, sphericity corrections, and participant as a random effect. As reported, the p-values and omega-squared values for xMOS may be overstated. A mixed-effects model or a repeated-measures analysis is needed to support the reported differences.","section":"IV-A.3 and IV-B.3"},{"comment":"The paper's interpretation of every comparative result assumes that the 11 emoji voices are distinct, internally consistent, and different from the baseline and Pleasant voices, yet no acoustic verification or perceptual discrimination test is reported. The fine-tuning uses only about 40 sentences per emoji per speaker, so it is possible that some emoji conditions simply captured recording variation or training noise. The claim that emoji prompting produces distinct and controllable vocal styles would be substantially strengthened by an objective analysis of the synthesized voices, such as F0, duration, spectral measures, or a style-classification test.","section":"III-C"}],"minor_comments":[{"comment":"The abstract says 'fine-grained control of expressivity on a phase level'; this should be 'phrase level' to match the terminology used in the rest of the paper.","section":"Abstract"},{"comment":"Case Study 1 also uses a fixed voice order (Baseline, Emoji, Pleasant), so the conclusion that Pleasant is preferred in the assistant task is subject to the same order-confound concern; this should at least be acknowledged in the limitations.","section":"IV-A.1 and IV-B.1"},{"comment":"The paper says in Section III-C that researchers can use 'as little as 3 minutes of speech per emoji' but the data collection section describes 40 training sentences per emoji; the relationship between these two numbers should be clarified.","section":"III-C and IV-B.2"},{"comment":"The Tukey HSD notation in the tables is dense; expanding the caption with a worked example of the 'E-P' notation, or providing a separate pairwise-comparison table, would improve readability.","section":"Tables II and III"}],"recommendation":"major_revision","confidential_remarks":"The toolkit contribution is useful and the open-source release is commendable, but the central empirical claim rests on a confounded within-subject design. Given that the data and checkpoints are public, an additional counterbalanced user study with independent participants is feasible and would materially improve the paper. I would not reject on the basis of the toolkit contribution alone, but the headline storytelling result cannot stand as written."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things up front. EmojiVoice is a practical contribution: open-source code, data, checkpoints, real-time phrase-level control on a small model, and it addresses a real gap for social robots. But the paper's central empirical claim, that varied emoji prompting beats a constant joyful voice in storytelling, is not secure because of the fixed voice order in Case Study 2 (Pleasant, Baseline, Emoji, always) with the same 24 participants who had already rated all three voices in Case Study 1. The stress-test note lands. The order pattern flips between the two studies, so fatigue, contrast, or recency can explain the result as easily as the actual voice variation.\n\nWhat is genuinely new: emoji prompting as a phrase-level control interface, an accessible fine-tuning pipeline, and a real comparison of varied vs single-style expressive voices in long-form narration. The paper is honest about several limitations, including not separating general expressivity from context appropriateness, and the qualitative comments from participants are reported fairly.\n\nThe soft spots, in rough order of size. First, the fixed order and shared participant pool undermine the causal reading of the storytelling result; this is the load-bearing issue. Second, there is no acoustic verification that the 11 emoji styles are distinct and consistent, and with only about 40 training sentences per emoji, style robustness is genuinely untested. Third, the \"first multi-speaker Matcha-TTS\" claim is wrong as written: the base VCTK checkpoint is already multi-speaker. Fourth, Table I's \"No Hallucinations\" row for Matcha-TTS is asserted without evidence in this paper. All of these are fixable with rewording and added acoustic analyses.\n\nOne thing the reader flagged as circularity I would not treat as fatal: training on the same emoji set used in evaluation is by design, and the relevant comparison is Emoji vs Pleasant vs Baseline. That comparison is not circular. It does mean the ratings partly validate the prompt-to-voice mapping rather than the learned styles alone.\n\nBottom line: the toolkit deserves to be cited, and the paper deserves a serious referee. I would send it to review with a request for a counterbalanced replication or clear pilot framing of Case Study 2, acoustic distinctiveness checks, and removal of the overclaimed firsts. As written, treat the storytelling result as suggestive, not demonstrated.","headline":"A genuinely useful open-source expressive TTS toolkit, but the headline storytelling result is confounded by fixed presentation order and needs a counterbalanced replication.","tokens_in":13753,"tokens_out":2589,"would_cite":true,"duration_ms":27212,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that phrase-by-phrase emoji prompting lets a lightweight neural text-to-speech system keep a robot storyteller expressive over long interactions, while a fixed pleasant voice suits short assistant exchanges better.","keywords":["emoji prompting","expressive text-to-speech","robot speech","temporal variability","storytelling robot","human-robot interaction","Matcha-TTS","social robotics"],"falsifier":"Run the same storytelling script with the emoji-varied voice but with emojis randomly reassigned to phrases; if perceived expressiveness drops to the level of the constant pleasant voice, the effect depends on emoji-content matching rather than temporal variation. Alternatively, measure acoustic prosodic features such as pitch range and speaking rate across the 11 emoji conditions; if they do not separate into distinct clusters, the model has not actually learned 11 distinct styles.","tokens_in":12672,"feed_emoji":"🤖","tokens_out":4766,"duration_ms":45805,"temperature":0.7,"pith_summary":"The paper proposes that varying vocal expressivity phrase by phrase, rather than using one fixed expressive voice, is what keeps a robot's speech engaging over long stretches. To make this practical, it introduces emoji prompting: a text-to-speech system fine-tuned to render each emoji as a distinct vocal style, with the emoji appended to each phrase selecting the style at synthesis time. The toolkit, built on the lightweight Matcha-TTS, runs in real time on robot-class hardware and can be customized with about three minutes of speech per style. The central finding is from a storytelling case study: listeners rated the emoji-varied voice as more expressive than a constant joyful voice, while the opposite held for a short assistant conversation.","feed_headline":"Emoji prompts keep a robot's voice expressive through a long story","feed_subtitle":"Switching vocal style phrase by phrase beat a constant joyful voice in storytelling, but not for assistant tasks.","key_machinery":"The central object is emoji prompting: appending one of eleven Unicode emojis to each text phrase and feeding that emoji token both to the text encoder and the flow-prediction network of Matcha-TTS, a non-autoregressive neural TTS trained with optimal-transport conditional flow matching. The fine-tuned checkpoint holds multiple voices in a single 78MB model, treating each emoji as a style label, so the system can switch expressive styles per phrase without reloading the model. This carries the argument by turning expressive control from a global voice setting into a per-phrase local control that costs almost nothing at inference and works naturally with LLM-generated text that already contains emojis.","core_discovery":"On its own terms, the paper's central claim is that phrase-by-phrase emoji prompting of a fine-tuned Matcha-TTS creates long-term temporally variable expressive speech that listeners perceive as more expressive than a single joyful expressive voice in storytelling. In the storytelling case study, the emoji-varied voice scored significantly higher than the baseline on expressiveness, social impression, and suitability, and significantly higher than the pleasant voice on expressiveness, and it was the first-choice voice with high confidence. In the assistant conversation, the pleasant voice was preferred and rated more suitable, showing that the appropriate vocal strategy depends on the task. The claimed mechanism is that variation itself prevents the monotony of a constant expressive style over long text, even though a consistently pleasant voice can be more appropriate for short, task-oriented utterances.","pith_inferences":["If the benefit comes from temporal variation rather than from the specific emoji set, then randomly varying a small set of pleasant styles might reproduce part of the storytelling gain; this is a testable extension the paper does not run.","The one-emoji-per-phrase granularity is a ceiling on control; word-level emoji markup could allow emphasis on individual words and might address the participant report that timing, pitch, and emphasis sometimes felt off.","Because the LLM's emoji choices strongly shaped perceived appropriateness in the autonomous case, future systems should couple emoji selection to conversational content rather than treating emojis as purely style-only tokens.","The fixed presentation order in the first two case studies means the reported preference differences could partly reflect ordering or fatigue effects; a fully randomized within-subject design would test whether the effect persists."],"forward_implications":["Social robots can vary vocal expression per phrase on-board, without a cloud GPU, using a model with about 20.9M parameters and a real-time factor below one.","Long storytelling interactions can be made more engaging with an emoji-varied voice than with a fixed joyful voice, so robot voice design should treat expression as dynamic rather than a single setting.","For assistant tasks, a consistently pleasant voice may be more appropriate than high variability, so the task and interaction context should determine which expressive strategy to deploy.","LLM-based conversational agents can append emojis to route voice styles automatically, making expressive control nearly free for autonomous spoken dialogue systems.","The toolkit lowers the data cost of custom expressive voices to roughly three minutes of speech per style, which could let roboticists build task-specific expressive voices quickly."],"supporting_citations":[{"why":"Supplies the Matcha-TTS architecture and the multi-speaker checkpoint that the toolkit fine-tunes for emoji-specific voices.","marker":"[11]"},{"why":"Provides the MOS-X2 scale used for prosody, intelligibility, and social impression ratings in the case studies.","marker":"[42]"},{"why":"Supplies the xMOS expressiveness metric used to measure perceived expressive intonation.","marker":"[43]"},{"why":"Supplies the voice suitability scale for robot voices used in the surveys.","marker":"[44]"},{"why":"Defines the real-time factor threshold used to claim that the system synthesizes speech in real time.","marker":"[27]"},{"why":"Provides the related storytelling TTS dataset and comparison point for controllable expressive speech.","marker":"[17]"},{"why":"Supplies the spoken-language-interaction recommendations that the toolkit's design and case studies address.","marker":"[45]"}],"fun_headline_variants":["Phrase-level emoji prompts enhance robot storytelling expressivity","Emoji variation in TTS boosts robot speech expressivity long-term","For robots, emoji-prompted voice beats constant joy in stories","Robot voice: emoji-driven variety aids storytelling, not tasks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that fine-tuning Matcha-TTS on roughly 40 sentences per emoji (about three minutes of speech) per speaker yields 11 distinct, consistent, and controllable vocal styles, so the reported perceptual differences come from those learned styles rather than from the specific emoji choices, voice quality, or presentation order.","fun_headline_variants_meta":{"raw":{"variants":["Phrase-level emoji prompts enhance robot storytelling expressivity","Emoji variation in TTS boosts robot speech expressivity long-term","For robots, emoji-prompted voice beats constant joy in stories","Robot voice: emoji-driven variety aids storytelling, not tasks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00022,"raw_usage":{"total_tokens":1421,"prompt_tokens":893,"completion_tokens":528,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":509,"completion_tokens_details":{"reasoning_tokens":455}},"tokens_in":509,"tokens_out":528,"duration_ms":5366,"temperature":1.0,"reasoning_tokens":455,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:43:18.118747+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same storytelling script with the emoji-varied voice but with emojis randomly reassigned to phrases; if perceived expressiveness drops to the level of the constant pleasant voice, the effect depends on emoji-content matching rather than temporal variation. Alternatively, measure acoustic prosodic features such as pitch range and speaking rate across the 11 emoji conditions; if they do not separate into distinct clusters, the model has not actually learned 11 distinct styles.","supporting_citations":[{"cited_title":"Matcha- TTS: A fast TTS architecture with conditional flow matching,","cited_arxiv_id":null,"evidence_quote":"Supplies the Matcha-TTS architecture and the multi-speaker checkpoint that the toolkit fine-tunes for emoji-specific voices."},{"cited_title":"Investigating mos-x ratings of synthetic and human voices,","cited_arxiv_id":null,"evidence_quote":"Provides the MOS-X2 scale used for prosody, intelligibility, and social impression ratings in the case studies."},{"cited_title":"Word-level text markup for prosody control in speech synthesis,","cited_arxiv_id":null,"evidence_quote":"Supplies the xMOS expressiveness metric used to measure perceived expressive intonation."},{"cited_title":"Giving robots a voice: Human-in-the-loop voice creation and open- ended labeling,","cited_arxiv_id":null,"evidence_quote":"Supplies the voice suitability scale for robot voices used in the surveys."},{"cited_title":"Scaling up online speech recognition using convnets,","cited_arxiv_id":null,"evidence_quote":"Defines the real-time factor threshold used to claim that the system synthesizes speech in real time."},{"cited_title":"Storytts: A highly expressive text-to-speech dataset with rich textual expressiveness annotations,","cited_arxiv_id":null,"evidence_quote":"Provides the related storytelling TTS dataset and comparison point for controllable expressive speech."},{"cited_title":"Spoken language interaction with robots: Recom- mendations for future research,","cited_arxiv_id":null,"evidence_quote":"Supplies the spoken-language-interaction recommendations that the toolkit's design and case studies address."}],"review_version":2}