Pith. sign in

REVIEW 4 major objections 5 minor 46 references

The Prosody of Emojis

T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Emojis systematically change how people say the same sentence, and listeners can hear the intended emoji in the audio.

desk verdict New dataset of human emoji-enriched speech is a real contribution, but the paper's advertised semantic-distance-to-prosody hierarchy does not survive the direct statistical test. read the letter →

arxiv 2508.00537 v2 pith:EZD6W2TC submitted 2025-08-01 cs.CL

classification cs.CL
keywords emojiprosodyspeechproductionperceptionexpressivetext-to-speechmixed-effectsmodellingcomputer-mediatedcommunicationprosodicembeddings
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that emojis are not just decorative marks in text but systematic cues that change how the same sentence is spoken aloud. Using a controlled production task in which different speakers read identical sentences with different emoji prompts, the researchers show that prosody—pitch, intensity, timing, and speech rate—shifts with the emoji, and that listeners can identify the intended emoji from the audio alone at rates well above chance. The paper's central result is a hierarchy: the greater the semantic difference between two emojis, the larger the measured prosodic distance between the two spoken versions, an effect that shows up across several acoustic features and in learned prosodic embeddings. If correct, these findings give a quantitative bridge between emoji semantics and speech production, with direct relevance to expressive text-to-speech and speech-understanding systems.

What carries the argument

The load-bearing design is a paired-recordings paradigm: each speaker reads the same sentence under different emoji conditions (and a no-emoji condition), and the analysis compares prosodic distances between those pairs. Prosodic features are extracted as pitch contours (FCNF0++), intensity (Praat/Parselmouth), speech rate, and a temporal stretch measure derived from Dynamic Time Warping alignment paths, plus 256-dimensional frame-wise prosodic embeddings from a self-supervised Masked Prosody Model. Distances between pairs are computed with (multivariate) DTW and fed into linear mixed-effects models with maximal random effects (by-speaker random slopes and by-utterance intercepts). The semantic side of the hierarchy is built from participant-generated intent and interpretation words, embedded and clustered with HDBSCAN into semantic categories; it is these clusters, not Unicode groups, that define same versus different category conditions.

What would settle it

Record a set of speakers reading each sentence twice with no emoji in either reading, measure the same DTW-based prosodic distances between those no-emoji pairs, and compare with distances between different-category emoji pairs of the same sentences; if the no-emoji baseline equals or exceeds the emoji-condition distances, the claim that emoji semantics drive prosodic divergence is falsified.

Watch

Extended reading notes

Core claim

The paper reports that emoji cues systematically modulate prosodic realisation when lexical content is held constant, that listeners recover the intended emoji from prosodic variation alone significantly above chance, and that prosodic divergence increases with semantic dissimilarity between emojis. Evidence includes listener judgments in which the most expressive speakers yield 64.68% correct emoji identification versus 24.74% incorrect, and inter-speaker pairs judged prosodically convergent produce 66.67% full agreement. Mixed-effects models on acoustic distances show that pairs of emojis from different semantic categories produce larger prosodic distances than adding an emoji to an emoji-free sentence (significant for learned prosodic embeddings, intensity, and speech rate), with the same positive direction, though not significant, for same-category versus different-category contrasts across pitch, intensity, and temporal stretch. The semantic categories come from unsupervised clustering of participant-supplied interpretations, rather than Unicode groupings, which correlate poorly with actual usage.

Load-bearing premise

The results stand or fall on the assumption that the measured acoustic differences between two readings of the same sentence are caused by the emoji cue itself, not by ordinary within-speaker reading variability—yet the design collected no control pairs in which the same speaker reads the same sentence twice with no emoji.

Editorial extensions

If this is right

  • Emoji labels can serve as a cheap, human-validated prosody annotation for expressive text-to-speech: a sentence plus an emoji carries enough prosodic information for a listener to recover the intended symbol.
  • Speech-understanding systems that ignore prosody are throwing away a recoverable channel of symbolic intent; emoji-conditioned prosodic models could predict the emoji (or a class of pragmatic meaning) from audio alone.
  • The semantic-distance-to-prosodic-distance mapping gives a way to embed emojis in an acoustic space, allowing quantitative comparisons of emoji meaning across the boundary between text and speech.
  • Speaker variability is large but structured: some speakers are consistent prosodic encoders, and their recordings are the ones listeners decode best, suggesting personalised prosody models could target or adapt to high-expressiveness strategies.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the hierarchy is robust, a practical extension would be to build an emoji-to-prosodic-vector dictionary from crowd-sourced readings, then use it to condition neural TTS without manual prosody labels.
  • The consistently positive but non-significant Diff-vs-Same coefficients suggest the effect may be driven mainly by categorical distinction (any category difference) rather than graded semantic distance; a study with a continuous semantic distance measure (e.g., cosine distance in the emoji embedding space) would test this directly.
  • The correlation between low emoji ambiguity and fewer incorrect listener interpretations hints that prosodic convergence across speakers could serve as a behavioural measure of emoji semantic stability, complementing text-only ambiguity scores.
  • A direct test of the causal attribution would need the missing control condition: pairs of no-emoji readings of the same sentence by the same speaker, to bound within-speaker variability.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper studies whether emojis affect the prosody of read speech and whether listeners can recover emoji meaning from prosodic variation alone. The authors collect a new dataset through four online experiments: annotators attach emojis to YouTube comments (Experiment 1), speakers read the emoji-enriched and original sentences aloud (Experiment 2), listeners judge whether pairs of same-speaker recordings differ prosodically and match them to emoji variants (Experiment 3), and listeners judge inter-speaker pairs and try to identify the shared emoji (Experiment 4). Prosodic distances are computed between paired recordings using DTW on pitch, intensity, and prosodic embeddings, and mixed-effects models test whether distance depends on whether the paired emojis belong to the same semantic cluster, different clusters, or one is absent. The abstract and introduction claim a 'clear hierarchy' where greater semantic difference between emojis yields larger prosodic divergence, but the statistical evidence for this ordering is weak, and the production design lacks a repeated no-emoji baseline.

Significance. If the claims were fully supported, this would be a novel empirical contribution linking emoji semantics to speech prosody, with implications for expressive speech synthesis and multimodal NLP. The dataset collection is substantial, the experiments are described in enough detail to be replicable, and the authors commit to releasing the data. The perception results (Experiments 3-4) are interesting and may survive a more cautious framing. However, the headline hierarchy result is not supported by the reported statistics, and the absence of a no-emoji repeated-reading control undermines the causal attribution of prosodic shifts to emoji cues. The manuscript's contribution is therefore presently overstated relative to the evidence.

major comments (4)
  1. [§5.3, Table 2] The abstract and Section 1 claim a 'clear hierarchy' where greater semantic differences between emojis correspond to increased prosodic divergence. The direct test of this hierarchy is the Diff vs Same comparison in Table 2, and none of the five features reaches p < 0.05 (embedding p=0.215, pitch p=0.156, intensity p=0.144, speech rate p=0.543, stretch p=0.124). The significant Added vs Diff results show only that pairs with two different emojis differ more than pairs with one emoji versus none; they do not establish that the magnitude of prosodic divergence increases monotonically with semantic dissimilarity. The authors should either provide stronger evidence for the hierarchy (e.g., Bayesian posterior probabilities or effect sizes with confidence intervals) or explicitly demote this claim from a confirmed result to a directional trend.
  2. [§3.2, Experiment 2] The production experiment records each sentence once per emoji variant and once without emoji, but never records the same speaker reading the same no-emoji sentence twice. Without this repeated no-emoji baseline, the measured prosodic differences between emoji conditions cannot be unambiguously attributed to the emoji cue rather than ordinary within-speaker reading variability. This affects the interpretation of RQ1 and the distance comparisons in Section 5.3. Adding such a baseline (or otherwise demonstrating that within-speaker variability is negligible compared to the observed differences) is necessary to support the causal language used throughout the paper.
  3. [§5.2.2; abstract] The abstract promises 'Bayesian multilevel modelling,' but Section 5.2.2 describes a frequentist linear mixed-effects model with p-values from the fitted model. No priors, posterior distributions, or Bayes factors are reported, and no Bayesian model is specified anywhere. This is a mismatch between the stated and actual methodology. In addition, five prosodic features and three pairwise contrasts are tested in Table 2 without any multiple-comparison correction; the reported p-values, especially the borderline ones, should be interpreted with that caveat or adjusted.
  4. [§5.1, §3.2] The semantic categories used as the independent variable in RQ4 are derived from the one-word interpretations collected from the same participants who read the utterances aloud in Experiment 2. This creates a possible coupling between the semantic clustering and the speakers' own production: a participant's interpretation of an emoji may be influenced by the way they pronounced it, or the clustering may reflect idiosyncratic associations that also shape their prosody. The manuscript should address this potential non-independence or at least justify why it does not inflate the Diff vs Same contrast.
minor comments (5)
  1. [§6.3] The sentence 'Listeners, in both cases, favoured the emoji' appears to be missing the intended emoji glyph after 'the'; this makes the qualitative example difficult to follow.
  2. [§4.3, Table 1] The 'Overall' column in Table 1 reports 34.74% and 65.26% for Same and Different, but the 'Correct' column entry for Same (66.67%) seems to refer to a different base than the row percentage; clarify whether these are conditional rates and what the denominators are.
  3. [§5.2.1] The choice of the 7th layer of the Masked Prosody Model is mentioned without justification; a brief explanation of the layer selection or a reference would be helpful.
  4. [Appendix C] The PCA dimensionality is set to 49 components to preserve ~90% variance and HDBSCAN min_cluster_size is set to 3; these choices are not varied or justified in the main text, and the sensitivity of the Diff/Same classification to these parameters is not assessed.
  5. [§4.1, Figure 2] The histogram bins in Figure 2 are labelled by upper bound (e.g., '40' represents 31–40%), but the x-axis in the figure is not visible in the text; ensure the axis label is clear in the final version.

Circularity Check

0 steps flagged · score 0.0 of 10

No definitional circularity: semantic categories and prosodic distances are measured in independent spaces; the abstract's overstatement is an inference issue, not a circularity.

full rationale

The paper's RQ4 test separates the semantic variable from the acoustic outcome. Semantic categories are derived from one-word participant descriptions (Section 5.1) and clustered with HDBSCAN; prosodic divergence is computed from pitch, intensity, speech rate, stretch, and prosodic embeddings (Section 5.2.1). The hypothesis is not true by construction: a pair of utterances can be in the 'Same' category yet have large acoustic distance, and the Diff-vs-Same coefficients in Table 2 are mostly non-significant, which would be impossible if the comparison were definitionally forced. Self-citations (Zhou et al. 2024a, 2024b) appear only in background sections and do not carry the derivation. No parameter is fitted and renamed as a prediction, no uniqueness theorem is imported from the authors' prior work, and no ansatz is smuggled in by citation. The main caveat—semantic labels and speech come from the same participant pool—is a potential endogeneity or validity concern, not a reduction of the predicted quantity to the input quantity.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

The central claim rests on several unexamined modeling choices: the semantic categories come from a clustering of participant descriptions with hand-set parameters, the acoustic features come from pretrained models taken on faith, and the causal interpretation assumes away ordinary rereading variability. The study is empirical, not a derivation, so no free parameter is fitted to the target result, but the hand-set thresholds and pretrained tools do shape the outcome.

free parameters (5)
  • PCA dimensionality = 49
    Selected to retain ~90% variance before HDBSCAN clustering; affects the semantic cluster assignments that define Same/Diff conditions in RQ4.
  • HDBSCAN min_cluster_size = 3
    Chosen to allow fine-grained clusters; directly determines which emojis are considered semantically similar, shaping the main comparison in RQ4.
  • WER exclusion threshold = 20%
    Recordings with >20% word error rate were excluded, affecting the final 3153-recording dataset.
  • Low prosodic variation exclusion cutoff = 30% similar judgments
    Speakers whose recordings were judged prosodically similar in more than 30% of cases were excluded from Experiment 4, which can inflate convergence results.
  • Prosody model layer = 7th layer
    The paper uses the 7th layer of the Masked Prosody Model for embedding extraction; no ablation shows this is optimal.
assumptions (4)
  • domain assumption Pretrained prosody and transcription models (PENN, Parselmouth, Masked Prosody Model, Whisper) provide valid estimates of pitch, intensity, prosody, and word alignment in this dataset.
    Invoked throughout Section 5.2.1; if these tools are inaccurate on British English spontaneous read speech, all distances are biased.
  • domain assumption One-word participant descriptions in Experiments 1 and 2 capture the semantic properties of emojis well enough to form reliable embeddings and clusters.
    The semantic categories in Section 5.1 are built from these descriptions; the entire RQ4 test depends on the validity of these clusters.
  • domain assumption Dynamic Time Warping distance and the other acoustic distance metrics reflect perceptually relevant prosodic differences.
    The distance measures are used as dependent variables in the mixed-effects models; no validation that they track human prosodic similarity.
  • ad hoc to paper The observed prosodic differences between paired readings are attributable to the emoji condition rather than to ordinary within-speaker reading variability.
    No no-emoji baseline pair was recorded; Experiment 2 design (Section 3.2) reads each sentence once per emoji, so the causal interpretation is an unverified assumption.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Prosody of Emojis." pith.science (2026). https://pith.science/paper/EZD6W2TC

@misc{pith2026250800537,
  author       = {Pith},
  title        = {Pith review of: The Prosody of Emojis},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/EZD6W2TC}},
  note         = {Machine review of arXiv:2508.00537}
}
read the original abstract

Prosodic features such as pitch, timing, and intonation are central to spoken communication, conveying emotion, intent, and discourse structure. In text-based settings, where these cues are absent, emojis act as visual surrogates that add affective and pragmatic nuance. This study examines how emojis influence prosodic realisation in speech and how listeners interpret prosodic cues to recover emoji meanings. Unlike previous work, we directly link prosody and emojis by analysing human speech data collected through a controlled elicited production task. Using Bayesian multilevel modelling, we show that speakers systematically adapt their prosody based on emoji cues, and that listeners can recover intended meanings significantly above chance. Furthermore, our results reveal a clear hierarchy in prosodic shifts: greater semantic differences between emojis correspond to increased prosodic divergence. These findings suggest that emojis are meaningful carriers of prosodic intent that bridge the gap between digital text and spoken production.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

46 extracted references · 37 canonical work pages

  1. [1]

    Alhaj, and Asad Ali

    Ehab Saleh Alnuzaili, Muhammad Waqar Amin, Sami Saad Alghamdi, Nazir Ahmed Malik, Abdulbasit A. Alhaj, and Asad Ali. 2024. Emojis as graphic equivalents of prosodic features in natural speech: evidence from computer-mediated discourse of whatsapp and facebook. Cogent Arts & Humanities, 11(1):2391646

  2. [2]

    Dane Archer and Robin M Akert. 1977. Words and everything else: Verbal and nonverbal cues in social interpretation. Journal of personality and social psychology, 35(6):443

  3. [3]

    Dale J Barr, Roger Levy, Christoph Scheepers, and Harry J Tily. 2013. Random effects structure for confirmatory hypothesis testing: Keep it maximal. Journal of memory and language, 68(3):255--278

  4. [4]

    Ricardo JGB Campello, Davoud Moulavi, and J \"o rg Sander. 2013. Density-based clustering based on hierarchical density estimates. In Pacific-Asia conference on knowledge discovery and data mining, pages 160--172. Springer

  5. [5]

    Jie Chi, Maureen de Seyssel, and Natalie Schluter. 2025. https://doi.org/10.18653/v1/2025.findings-naacl.471 The role of prosody in spoken question answering . In Findings of the Association for Computational Linguistics: NAACL 2025, pages 8468--8479, Albuquerque, New Mexico. Association for Computational Linguistics

  6. [6]

    Jennifer Cole. 2015. Prosody in context: A review. Language, Cognition and Neuroscience, 30(1-2):1--31

  7. [7]

    David Crystal. 2001. Language and the Internet. Cambridge University Press

  8. [8]

    Justyna Czkestochowska, Kristina Gligori \'c , Maxime Peyrard, Yann Mentha, Micha Bie \'n , Andrea Gr \"u tter, Anita Auer, Aris Xanthos, and Robert West. 2022. On the context-free ambiguity of emoji. In Proceedings of the International AAAI Conference on Web and Social Media, volume 16, pages 1388--1392

Show all 46 references
  1. [9]

    Daantje Derks, Arjan ER Bos, and Jasper Von Grumbkow. 2008. Emoticons in computer-mediated communication: Social motives and social context. Cyberpsychology & behavior, 11(1):99--101

  2. [10]

    LouAnn Gerken and Karla McGregor. 1998. An overview of prosody and its role in normal and disordered child language. American Journal of Speech-Language Pathology, 7(2):38--48

  3. [11]

    u ge T G \

    T \"u ge T G \"u l s en. 2016. You tell me in emojis. In Computational and cognitive approaches to narratology, pages 354--375. IGI Global

  4. [12]

    Nele Hellbernd and Daniela Sammler. 2016. Prosody conveys speaker’s intentions: Acoustic cues for speech act perception. Journal of Memory and Language, 88:70--86

  5. [13]

    Julia Hirschberg. 2002. Communication and prosody: Functional aspects of prosody. Speech Communication, 36(1-2):31--43

  6. [14]

    Thomas Holtgraves and Caleb Robinson. 2020. Emoji can facilitate recognition of conveyed indirect meaning. PloS one, 15(4):e0232361

  7. [15]

    Matthew Honnibal, Ines Montani, Sofie Van Landeghem, Adriane Boyd, et al. 2020. spacy: Industrial-strength natural language processing in python

  8. [16]

    Jiaxiong Hu, Qianyao Xu, Limin Paul Fu, and Yingqing Xu. 2019. Emojilization: An automated method for speech to emoji-labeled text. In Extended Abstracts of the 2019 CHI Conference on Human Factors in Computing Systems, pages 1--6

  9. [17]

    Lawrence Hubert and Phipps Arabie. 1985. Comparing partitions. Journal of classification, 2:193--218

  10. [18]

    Yannick Jadoul, Bill Thompson, and Bart de Boer. 2018. https://doi.org/https://doi.org/10.1016/j.wocn.2018.07.001 Introducing P arselmouth: A P ython interface to P raat . Journal of Phonetics, 71:1--15

  11. [19]

    Allan James. 2017. Prosody and paralanguage in speech and the social media: The vocal and graphic realisation of affective meaning. Linguistica, 57(1):137--149

  12. [20]

    Elsi Kaiser. 2021. Focus marking with emoji: On the relation between information structure and expressive meaning. In Colloque de syntaxe et s \'e mantique \'a Paris (CSSP), Paris, France

  13. [21]

    Joseph Kane, Michael N Johnstone, and Patryk Szewczyk. 2024. Voice synthesis improvement by machine learning of natural prosody. Sensors, 24(5):1624

  14. [22]

    Ilse Lehiste and Norman J Lass. 1976. Suprasegmental features of speech. Contemporary issues in experimental phonetics, 225:239

  15. [23]

    Li Li and Yue Yang. 2018. Pragmatic functions of emoji in internet-based communication---a corpus-based study. Asian-Pacific Journal of Second and Foreign Language Education, 3:1--12

  16. [24]

    Leland McInnes, John Healy, Nathaniel Saul, and Lukas Grossberger. 2018. Umap: Uniform manifold approximation and projection. The Journal of Open Source Software, 3(29):861

  17. [25]

    Hannah Miller, Daniel Kluver, Jacob Thebault-Spieker, Loren Terveen, and Brent Hecht. 2017. Understanding emoji ambiguity in context: The role of text in emoji-related miscommunication. In Eleventh international AAAI conference on web and social media

  18. [26]

    Max Morrison, Caedon Hsieh, Nathan Pruyne, and Bryan Pardo. 2023. Cross-domain neural pitch and periodicity estimation. In arXiv preprint arXiv:2301.12258

  19. [27]

    Noa Na ' aman, Hannah Provenza, and Orion Montoya. 2017. https://aclanthology.org/P17-3022/ Varying linguistic purposes of emoji in ( T witter) context . In Proceedings of ACL 2017, Student Research Workshop , pages 136--141, Vancouver, Canada. Association for Computational Li...

  20. [28]

    Louise AG Neel, Jacqui G McKechnie, Christopher M Robus, and Christopher J Hand. 2023. Emoji alter the perception of emotion in affectively neutral text messages. Journal of Nonverbal Behavior, 47(1):83--97

  21. [29]

    Valeria A Pfeifer, Emma L Armstrong, and Vicky Tzuyin Lai. 2022. Do all facial emojis communicate emotion? the impact of facial emojis on perceived sender emotion and text processing. Computers in Human Behavior, 126:107016

  22. [30]

    Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023. Robust speech recognition via large-scale weak supervision. In Proceedings of the 40th International Conference on Machine Learning, ICML'23. JMLR.org

  23. [31]

    Nils Reimers and Iryna Gurevych. 2019. https://doi.org/10.18653/v1/D19-1410 Sentence- BERT : Sentence embeddings using S iamese BERT -networks . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference...

  24. [32]

    Monica A Riordan. 2017. Emojis as tools for emotion work: Communicating affect in text messages. Journal of Language and Social Psychology, 36(5):549--567

  25. [33]

    Seamless Communication et al . 2023. Seamless: Multilingual expressive and streaming speech translation. arXiv preprint arXiv:2312.05187

  26. [34]

    RJ Skerry-Ryan, Eric Battenberg, Ying Xiao, Yuxuan Wang, Daisy Stanton, Joel Shor, Ron Weiss, Rob Clark, and Rif A Saurous. 2018. Towards end-to-end prosody transfer for expressive speech synthesis with tacotron. In international conference on machine learning, pages 4693--4702. PMLR

  27. [35]

    Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. 2020. Mpnet: Masked and permuted pre-training for language understanding. Advances in neural information processing systems, 33:16857--16867

  28. [36]

    Ioannis Tsiamas, Matthias Sperber, Andrew Finch, and Sarthak Garg. 2024. https://doi.org/10.18653/v1/2024.wmt-1.119 Speech is more than words: Do speech-to-text translation systems leverage prosody? In Proceedings of the Ninth Conference on Machine Translation, pages 1235--125...

  29. [37]

    Paige Tutt \"o s \' , Shivam Mehta, Zachary Syvenky, Bermet Burkanova, Gustav Eje Henter, and Angelica Lim. 2025. Emojivoice: Towards long-term controllable expressivity in robot speech. arXiv preprint arXiv:2506.15085

  30. [38]

    Nguyen Xuan Vinh, Julien Epps, and James Bailey. 2010. http://jmlr.org/papers/v11/vinh10a.html Information theoretic measures for clusterings comparison: Variants, properties, normalization and correction for chance . Journal of Machine Learning Research, 11(95):2837--2854

  31. [39]

    Michael Wagner. 2016. How to be kind with prosody. in Speech prosody, 1:250--1253

  32. [40]

    Sarenne Wallbridge, Christoph Minixhofer, Catherine Lai, and Peter Bell. 2025. Prosodic structure beyond lexical content: a study in self-supervised learning. In proceedings of Interspeech 2025

  33. [41]

    Nigel Ward and Gina-Anne Levow. 2021. https://doi.org/10.18653/v1/2021.acl-tutorials.5 Prosody: Models, methods, and applications . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural...

  34. [42]

    Benjamin Weissman and Darren Tanner. 2018. A strong wink between verbal and emoji-based irony: How the brain processes ironic emojis during language comprehension. PloS one, 13(8):e0201727

  35. [43]

    Lukas Wolf, Tiago Pimentel, Evelina Fedorenko, Ryan Cotterell, Alex Warstadt, Ethan Wilcox, and Tamar Regev. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.606 Quantifying the redundancy between prosody and text . In Proceedings of the 2023 Conference on Empirical Methods i...

  36. [44]

    Giulio Zhou, Sydelle De Souza, Ella Markham, Oghenetekevwe Kwakpovwe, and Sumin Zhao. 2024 a . https://doi.org/10.18653/v1/2024.emnlp-main.1041 Semantics and sentiment: Cross-lingual variations in emoji use . In Proceedings of the 2024 Conference on Empirical Methods in Natura...

  37. [45]

    Giulio Zhou, Tsz Kin Lam, Alexandra Birch, and Barry Haddow. 2024 b . https://aclanthology.org/2024.findings-eacl.46/ Prosody in cascade and direct speech-to-text translation: a case study on K orean wh-phrases . In Findings of the Association for Computational Linguistics: EA...

  38. [46]

    Juan Zuluaga-Gomez, Sara Ahmed, Danielius Visockas, and Cem Subakan. 2023. https://arxiv.org/abs/2305.18283 Commonaccent: Exploring large acoustic pretrained models for accent classification based on common voice . Interspeech 2023

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.