REVIEW 4 major objections 6 minor 1 cited by
Irony in Emojis: A Comparative Study of Human and LLM Interpretation
T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read This paper claims that GPT-4o systematically overestimates the likelihood that emojis are used ironically, relative to human usage patterns measured in a Chinese social media corpus, and that model and human scores agree only weakly.
desk verdict Useful extension of prior emoji-LLM work, but the central overestimation claim compares inverse conditionals and does not hold as stated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The transfer rule of Equation (1) is the hinge: it assigns to each emoji e the average irony rating R(p) of all posts containing e, so human emoji-level irony is never directly annotated but inferred from post-level ratings. The paper compares these derived human scores with GPT-4o's ratings obtained from an 11-point likelihood prompt, with candidates presented as images, and then rescales the model ratings to the 1–5 human scale. The statistical machinery consists of the Wilcoxon signed-rank test for median difference and the Spearman rank correlation for agreement, with prompt variations that insert gender and age labels.
What would settle it
Collect a set of posts, have human annotators rate the irony of each emoji directly in context rather than the whole post, and compare emoji-level human scores to GPT-4o's likelihood ratings on the same emojis; if the model's overestimation disappears or reverses under direct emoji annotation, the paper's central gap is an artifact of the post-to-emoji transfer.
Extended reading notes
Core claim
On the paper's own terms, the central finding is a measurable gap: GPT-4o, prompted to act as a social media user, assigns significantly higher likelihood-of-irony ratings to emojis than the irony scores humans produce through natural usage in the Ciron corpus, and the agreement between model and human scores is positive but weak (ρ = 0.28). This overestimation appears across the emoji set and is accompanied by larger variance in the model's ratings. The paper attributes the gap to possible training-data skew toward ironic emoji usage, overgeneralization of irony patterns, and the English-centric orientation of GPT-4o relative to the Chinese-language dataset, and it treats the demographic prompt results as evidence that the model tends to assign lower irony scores when prompted with older ages.
Load-bearing premise
The whole comparison rests on treating the irony rating of a post as the irony level of every emoji inside it, so that one post-level number stands in for each emoji's human score; if that transfer is wrong, the human benchmark is not actually measuring emoji irony.
Editorial extensions
If this is right
- Applications that rely on GPT-4o to gauge emoji sentiment can expect it to flag irony more often than users intend, risking over-detection in chatbot and sentiment pipelines.
- The weak correlation means the model's ranking of which emojis are most ironic is not a reliable proxy for human rankings.
- If the age effect is real and stable, LLM-based social-media personas for older demographics will produce more literal emoji interpretations, which may matter for behavioral simulation.
- The larger variance in model scores suggests its interpretations are less anchored than human usage patterns, so single-shot ratings should not be treated as calibrated.
- The use of a Chinese-language dataset with an English-trained model means the observed overestimation may partly reflect cross-linguistic transfer rather than a general property of the model.
Reading between the lines
- A cleaner test would annotate emojis directly in context; if such annotation lowered human emoji-level irony scores, the GPT-4o overestimation gap would widen, and if it raised them, the gap could narrow or disappear. This is my inference, not the paper's.
- The rescaling of the 11-point model scale to a 1–5 scale is monotone but arbitrary; a different mapping could change the magnitude of the reported median difference, although it would not affect Spearman correlation.
- The study's binary gender and five age bins likely compress real demographic variation; a prompt study with non-binary genders and finer age gradations could reveal effects the current design cannot see.
- Because the correlation is computed on only 82 emojis with many ties, a permutation or bootstrap test would give a more robust sense of whether ρ = 0.28 is stable.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper compares GPT-4o's irony ratings for emojis with "human-perceived" irony scores obtained from the Ciron corpus. The human score S(e) in Eq. (1) is the average post-level irony rating over posts containing emoji e. GPT-4o is prompted to rate, on an 11-point scale, how likely it would be to use a given emoji to express irony; these ratings are rescaled to a 5-point scale and compared with S(e) using a Wilcoxon signed-rank test and a Spearman correlation. The paper reports that GPT-4o assigns significantly higher irony scores than humans (W = 918.5, p < .001) and that the two sets of scores are weakly but significantly correlated (rho = 0.28, p < .05). It also explores prompts with demographic information and reports an age-related trend in GPT-4o's scores.
Significance. If the comparison were valid, the finding that GPT-4o systematically overestimates the ironic potential of emojis relative to human usage would be practically useful for emoji-aware sentiment analysis and human-behavior simulation. The paper has several strengths: it uses an external human-annotated corpus rather than fitting model parameters to the human data, the model is a frozen system queried with a transparent prompt, and the nonparametric rank tests are appropriate for the discrete, skewed scores involved. The limitation paragraph honestly acknowledges the single-model and Chinese-corpus constraints. However, the central comparison is undermined by a conceptual mismatch between what the human score measures and what the prompt elicits, so the headline overestimation claim is not currently supported.
major comments (4)
- [Human Perception of Emoji Irony, Eq. (1)] The sentence "We use the irony rating of a post to represent the irony level of the emojis found within it" is an untested and load-bearing assumption. A post-level irony rating is assigned to the entire post, not to each emoji in it; an ironic post may contain emojis that are not themselves ironic, and a non-ironic post may contain an emoji used ironically. Because S(e) is the entire human benchmark used in the Wilcoxon and Spearman comparisons, this assumption must be validated with emoji-level annotations or at least with a per-emoji agreement analysis. Without such validation, the comparison is not a clean measure of emoji-level human perception.
- [GPT-4o's Classification of Emoji Irony and Results] The prompt asks GPT-4o to "rate your likelihood of using this emoji if your intention is to express irony," which elicits P(use e | ironic intent), whereas Eq. (1) computes S(e) = E[irony rating | e appears], an estimate of P(ironic intent | e). These are inverse conditionals related by Bayes' rule, and they can diverge substantially when emoji base rates are skewed, as they are here (82 emojis across about 3,000 posts). Consequently, the Wilcoxon result W = 918.5 does not establish that GPT-4o "systematically overestimates" human perception; it may simply reflect the difference between asking "how likely would you use this emoji ironically?" and "how often does this emoji appear in ironic posts?" The paper's own smirk-emoji example in Results illustrates exactly this distinction: GPT-4o's high rating is justified by possible ironic contexts, while the Ciron posts containing that emoji are mostly non-ironic. A valid comparison would require eliciting human likelihood ratings under the same prompt, or annotating emoji irony directly in context.
- [Results, rescaling paragraph] The paper states that the model's ratings are "rescaled from a range of 1–11 to align with the 1–5 scale" but does not give the transformation. This matters because the Wilcoxon signed-rank test operates on the signs of paired differences, and an affine transformation of the model scores with a nonzero intercept can change the sign of some differences. For example, the natural mapping x' = 0.4x + 0.6 adds a positive constant to all model scores, which can flip small negative differences to positive ones and alter W and the reported p < .001. The authors should specify the exact rescaling, justify it, and show that the statistical conclusions are robust to reasonable alternatives, or use a comparison that does not depend on the arbitrary intercept.
- [Prompts with Demographic Information, Figure 2] The demographic section states that "no significant differences in irony scores are observed between prompts specifying female or male gender" and that scores "tend to decrease on average as the specified age in the prompt increases," but no test statistic, p-value, effect size, or confidence interval is reported for any of these comparisons. If these demographic findings are part of the paper's contribution, they need inferential support; as written, the claims rest on visual inspection of Figure 2. The paper should also explain how the five age groups and the male/female conditions were compared (e.g., repeated-measures tests across the 82 emojis) and whether any multiple-comparison correction was applied.
minor comments (6)
- [Results, Wilcoxon reporting] The text says the "median irony score assigned by GPT-4o is significantly higher" than the human score, but the Wilcoxon signed-rank test tests the median of paired differences (pseudomedian), not the difference of the two marginal medians. The wording should be adjusted to avoid a technically incorrect interpretation.
- [Figure 1] The Spearman correlation is reported only as p < .05; the exact p-value and the sample size (N = 82) should be stated, and the caption should note that the correlation uses tied discrete scores.
- [Human Perception of Emoji Irony] The paper does not report how the 82 unique emojis were extracted or their frequency distribution across the approximately 3,000 emoji-containing posts. Many emojis likely appear only once or twice, making their S(e) estimates very noisy; reporting the counts, or at least their minimum, median, and maximum, would help readers assess the reliability of Eq. (1).
- [Experiment Setting] The paper averages three GPT-4o queries per emoji at temperature 0.5 but does not report the variance or agreement across those queries. Reporting the standard deviation or the range of the three ratings would strengthen the claim that the model scores are stable.
- [Discussion and Conclusions] The speculation that GPT-4o "may have been trained on data with a disproportionate representation of ironic emoji usage" is presented without evidence; it is a plausible hypothesis, but it should be labeled as a hypothesis rather than a conclusion drawn from the current results.
- [Throughout] Several emoji glyphs do not render in the plain-text version of the manuscript (e.g., the kiss emoji in Eq. (2) and the smirk emoji in Results). The final camera-ready version must ensure the glyphs display correctly, since the emojis are the experimental stimuli.
Circularity Check
No significant circularity; the human–LLM comparison rests on an external benchmark and frozen-model outputs.
full rationale
The claimed derivation chain is not circular. The human benchmark S(e) is defined by Eq. (1) directly from post-level irony ratings in the externally compiled Ciron dataset (Xiang et al. 2020); this is an operational definition, not an output of the model. GPT-4o's irony ratings are obtained by prompting the frozen model with an independent question about likelihood of use for ironic intent, and no GPT-4o parameter or prompt output is fitted to the Ciron ratings or to the computed S(e). The only transformation applied to model ratings is a rescaling from the 1–11 prompt scale to the 1–5 human scale; because the reported tests are rank-based (Wilcoxon signed-rank and Spearman), this monotone rescaling cannot manufacture the median difference or the correlation. The paper's self-citations (Lyu et al. 2024a, 2024b) are used as background motivation and as precedent for the image-format prompt, not as the evidence establishing the comparison. The inverse-conditional concern raised in the skeptic note—that Eq. (1) estimates E[irony | emoji appears] while the prompt elicits P(use emoji | ironic intent)—is a substantive construct-validity threat to the interpretation of the difference, but it is not a circularity: the two quantities are independently measured and the reported statistics would still be computed from those independent measurements. Similarly, the assumption that post-level irony transfers to contained emojis is an external validity assumption, not a step that defines the model output in terms of the human output. Accordingly, no step in the derivation reduces by construction to its own inputs.
Assumptions & free parameters
free parameters (1)
- Rescaling method from 11-point to 5-point scale =
Not specified
assumptions (3)
- domain assumption Post-level irony ratings transfer to all emojis within a post.
- domain assumption Ciron annotations are a valid ground truth for human irony perception.
- domain assumption GPT-4o's self-reported likelihood ratings are comparable to human usage-derived scores.
Cite this review
Pith. "Pith review of Irony in Emojis: A Comparative Study of Human and LLM Interpretation." pith.science (2026). https://pith.science/paper/XT6YW7NO
@misc{pith2026250111241,
author = {Pith},
title = {Pith review of: Irony in Emojis: A Comparative Study of Human and LLM Interpretation},
year = {2026},
howpublished = {\url{https://pith.science/paper/XT6YW7NO}},
note = {Machine review of arXiv:2501.11241}
}
read the original abstract
Emojis have become a universal language in online communication, often carrying nuanced and context-dependent meanings. Among these, irony poses a significant challenge for Large Language Models (LLMs) due to its inherent incongruity between appearance and intent. This study examines the ability of GPT-4o to interpret irony in emojis. By prompting GPT-4o to evaluate the likelihood of specific emojis being used to express irony on social media and comparing its interpretations with human perceptions, we aim to bridge the gap between machine and human understanding. Our findings reveal nuanced insights into GPT-4o's interpretive capabilities, highlighting areas of alignment with and divergence from human behavior. Additionally, this research underscores the importance of demographic factors, such as age and gender, in shaping emoji interpretation and evaluates how these factors influence GPT-4o's performance.
Figures
Forward citations
Cited by 1 Pith paper
-
When Smiley Turns Hostile: Interpreting How Emojis Trigger LLMs' Toxicity
Emojis in harmful prompts bypass LLM safety more effectively than plain text, across 7 models and 5 languages, through a heterogeneous tokenization channel.
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Ai, W.; Lu, X.; Liu, X.; Wang, N.; Huang, G.; and Mei, Q. 2017. Untangling emoji popularity through semantic embeddings. In Proceedings of the international AAAI conference on web and social media, volume 11, 2--11
work page 2017
-
[4]
Chen, Y.; Yang, X.; Howman, H.; and Filik, R. 2024. Individual differences in emoji comprehension: Gender, age, and culture. Plos one, 19(2): e0297379
work page 2024
-
[5]
Chen, Z.; Lu, X.; Ai, W.; Li, H.; Mei, Q.; and Liu, X. 2018. Through a gender lens: Learning usage patterns of emojis from large-scale android users. In Proceedings of the 2018 world wide web conference, 763--772
work page 2018
-
[6]
Częstochowska, J.; Gligorić, K.; Peyrard, M.; Mentha, Y.; Bień, M.; Grütter, A.; Auer, A.; Xanthos, A.; and West, R. 2022. On the Context-Free Ambiguity of Emoji. Proceedings of the International AAAI Conference on Web and Social Media, 16(1): 1388--1392
work page 2022
-
[7]
Garcia, C.; Țurcan, A.; Howman, H.; and Filik, R. 2022. Emoji as a tool to aid the comprehension of written sarcasm: Evidence from younger and older adults. Computers in Human Behavior, 126: 106971
work page 2022
-
[8]
C.; Li, M.; Tay, L.; and Ungar, L
Guntuku, S. C.; Li, M.; Tay, L.; and Ungar, L. H. 2019. Studying cultural differences in emoji usage across the east and the west. In Proceedings of the international AAAI conference on web and social media, volume 13, 226--235
work page 2019
Show all 24 references
-
[9]
Nice picture comment!
Herring, S.; and Dainas, A. 2017. “Nice picture comment!” Graphicons in Facebook comment threads
2017
-
[10]
Hu, T.; Guo, H.; Sun, H.; Nguyen, T.-v.; and Luo, J. 2017. Spice up your chat: the intentions and sentiment effects of using emojis. In Proceedings of the International AAAI Conference on Web and Social Media, volume 11, 102--111
2017
-
[11]
Kotek, H.; Dockum, R.; and Sun, D. 2023. Gender bias and stereotypes in large language models. In Proceedings of the ACM collective intelligence conference, 12--24
2023
-
[12]
Li, W.; Chen, Y.; Hu, T.; and Luo, J. 2018. Mining the relationship between emoji usage patterns and personality. In Proceedings of the international AAAI conference on web and social media, volume 12
2018
-
[13]
Lyu, H.; Huang, J.; Zhang, D.; Yu, Y.; Mou, X.; Pan, J.; Yang, Z.; Wei, Z.; and Luo, J. 2024 a . GPT-4V(ision) as A Social Media Analysis Engine. ACM Trans. Intell. Syst. Technol. Just Accepted
2024
-
[14]
Lyu, H.; Qi, W.; Wei, Z.; and Luo, J. 2024 b . Human vs. LMMs: Exploring the Discrepancy in Emoji Interpretation and Usage in Digital Communication. In Proceedings of the International AAAI Conference on Web and Social Media, volume 18, 2104--2110
2024
-
[15]
S.; O'Brien, J.; Cai, C
Park, J. S.; O'Brien, J.; Cai, C. J.; Morris, M. R.; Liang, P.; and Bernstein, M. S. 2023. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th annual acm symposium on user interface software and technology, 1--22
2023
-
[16]
Qiu, Z.; Qiu, K.; Lyu, H.; Xiong, W.; and Luo, J. 2024. Semantics Preserving Emoji Recommendation with Large Language Models. arXiv preprint arXiv:2409.10760
2024 arXiv
-
[17]
Van Hee, C.; Lefever, E.; and Hoste, V. 2018. S em E val-2018 Task 3: Irony Detection in E nglish Tweets. In Apidianaki, M.; Mohammad, S. M.; May, J.; Shutova, E.; Bethard, S.; and Carpuat, M., eds., Proceedings of the 12th International Workshop on Semantic Evaluation, 39--50...
2018
-
[18]
kelly is a warm person, joseph is a role model
Wan, Y.; Pu, G.; Sun, J.; Garimella, A.; Chang, K.-W.; and Peng, N. 2023. " kelly is a warm person, joseph is a role model": Gender biases in llm-generated reference letters. arXiv preprint arXiv:2310.09219
2023 arXiv
-
[19]
Wang, S. 2022. Sarcastic meaning of the slightly smiling face emoji from Chinese Twitter users: When a smiling face does not show friendliness. International Journal of Languages, Literature and Linguistics, 8(2): 65--73
2022
-
[20]
Wankhade, M.; Rao, A. C. S.; and Kulkarni, C. 2022. A survey on sentiment analysis methods, applications, and challenges. Artificial Intelligence Review, 55(7): 5731--5780
2022
-
[21]
Weissman, B.; and Tanner, D. 2018. A strong wink between verbal and emoji-based irony: How the brain processes ironic emojis during language comprehension. PloS one, 13(8): e0201727
2018
-
[22]
Xiang, R.; Gao, X.; Long, Y.; Li, A.; Chersoni, E.; Lu, Q.; Huang, C.-R.; et al. 2020. Ciron: a new benchmark dataset for Chinese irony detection
2020
-
[23]
Zhang, S.; Zhang, X.; Chan, J.; and Rosso, P. 2019. Irony detection via sentiment-based transfer learning. Information Processing & Management, 56(5): 1633--1644
2019
-
[24]
Zhou, Y.; Xu, P.; Wang, X.; Lu, X.; Gao, G.; and Ai, W. 2024. Emojis decoded: Leveraging chatgpt for enhanced understanding in social media communications. arXiv preprint arXiv:2402.01681
2024 arXiv
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.