Pith. sign in

REVIEW 4 major objections 6 minor 1 cited by

Irony in Emojis: A Comparative Study of Human and LLM Interpretation

T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This paper claims that GPT-4o systematically overestimates the likelihood that emojis are used ironically, relative to human usage patterns measured in a Chinese social media corpus, and that model and human scores agree only weakly.

desk verdict Useful extension of prior emoji-LLM work, but the central overestimation claim compares inverse conditionals and does not hold as stated. read the letter →

arxiv 2501.11241 v1 pith:XT6YW7NO submitted 2025-01-20 cs.CL cs.CVcs.SI

classification cs.CLcs.CVcs.SI
keywords emojiironyGPT-4oLLMinterpretationhumanperceptiondetectionCirondatasetdemographicfactorssocialmedia
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether GPT-4o interprets the ironic potential of emojis the way human social media users do. Using the Ciron dataset of Chinese Weibo posts, the authors derive a human irony score for each of 82 emojis by averaging the irony ratings of posts that contain it, then prompt GPT-4o to rate how likely it would use each emoji to express irony. They report that GPT-4o's median irony score is significantly higher than the human-derived score (Wilcoxon W = 918.5, p < .001) and that the two sets of scores correlate only weakly (Spearman ρ = 0.28, p < .05). In other words, the model overestimates how ironic emojis look, and its ordering of emojis by irony aligns poorly with human usage. The paper also finds that when prompts specify a demographic, simulated age shifts GPT-4o's scores downward while simulated gender has little effect.

What carries the argument

The transfer rule of Equation (1) is the hinge: it assigns to each emoji e the average irony rating R(p) of all posts containing e, so human emoji-level irony is never directly annotated but inferred from post-level ratings. The paper compares these derived human scores with GPT-4o's ratings obtained from an 11-point likelihood prompt, with candidates presented as images, and then rescales the model ratings to the 1–5 human scale. The statistical machinery consists of the Wilcoxon signed-rank test for median difference and the Spearman rank correlation for agreement, with prompt variations that insert gender and age labels.

What would settle it

Collect a set of posts, have human annotators rate the irony of each emoji directly in context rather than the whole post, and compare emoji-level human scores to GPT-4o's likelihood ratings on the same emojis; if the model's overestimation disappears or reverses under direct emoji annotation, the paper's central gap is an artifact of the post-to-emoji transfer.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central finding is a measurable gap: GPT-4o, prompted to act as a social media user, assigns significantly higher likelihood-of-irony ratings to emojis than the irony scores humans produce through natural usage in the Ciron corpus, and the agreement between model and human scores is positive but weak (ρ = 0.28). This overestimation appears across the emoji set and is accompanied by larger variance in the model's ratings. The paper attributes the gap to possible training-data skew toward ironic emoji usage, overgeneralization of irony patterns, and the English-centric orientation of GPT-4o relative to the Chinese-language dataset, and it treats the demographic prompt results as evidence that the model tends to assign lower irony scores when prompted with older ages.

Load-bearing premise

The whole comparison rests on treating the irony rating of a post as the irony level of every emoji inside it, so that one post-level number stands in for each emoji's human score; if that transfer is wrong, the human benchmark is not actually measuring emoji irony.

Editorial extensions

If this is right

  • Applications that rely on GPT-4o to gauge emoji sentiment can expect it to flag irony more often than users intend, risking over-detection in chatbot and sentiment pipelines.
  • The weak correlation means the model's ranking of which emojis are most ironic is not a reliable proxy for human rankings.
  • If the age effect is real and stable, LLM-based social-media personas for older demographics will produce more literal emoji interpretations, which may matter for behavioral simulation.
  • The larger variance in model scores suggests its interpretations are less anchored than human usage patterns, so single-shot ratings should not be treated as calibrated.
  • The use of a Chinese-language dataset with an English-trained model means the observed overestimation may partly reflect cross-linguistic transfer rather than a general property of the model.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A cleaner test would annotate emojis directly in context; if such annotation lowered human emoji-level irony scores, the GPT-4o overestimation gap would widen, and if it raised them, the gap could narrow or disappear. This is my inference, not the paper's.
  • The rescaling of the 11-point model scale to a 1–5 scale is monotone but arbitrary; a different mapping could change the magnitude of the reported median difference, although it would not affect Spearman correlation.
  • The study's binary gender and five age bins likely compress real demographic variation; a prompt study with non-binary genders and finer age gradations could reveal effects the current design cannot see.
  • Because the correlation is computed on only 82 emojis with many ties, a permutation or bootstrap test would give a more robust sense of whether ρ = 0.28 is stable.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This paper compares GPT-4o's irony ratings for emojis with "human-perceived" irony scores obtained from the Ciron corpus. The human score S(e) in Eq. (1) is the average post-level irony rating over posts containing emoji e. GPT-4o is prompted to rate, on an 11-point scale, how likely it would be to use a given emoji to express irony; these ratings are rescaled to a 5-point scale and compared with S(e) using a Wilcoxon signed-rank test and a Spearman correlation. The paper reports that GPT-4o assigns significantly higher irony scores than humans (W = 918.5, p < .001) and that the two sets of scores are weakly but significantly correlated (rho = 0.28, p < .05). It also explores prompts with demographic information and reports an age-related trend in GPT-4o's scores.

Significance. If the comparison were valid, the finding that GPT-4o systematically overestimates the ironic potential of emojis relative to human usage would be practically useful for emoji-aware sentiment analysis and human-behavior simulation. The paper has several strengths: it uses an external human-annotated corpus rather than fitting model parameters to the human data, the model is a frozen system queried with a transparent prompt, and the nonparametric rank tests are appropriate for the discrete, skewed scores involved. The limitation paragraph honestly acknowledges the single-model and Chinese-corpus constraints. However, the central comparison is undermined by a conceptual mismatch between what the human score measures and what the prompt elicits, so the headline overestimation claim is not currently supported.

major comments (4)
  1. [Human Perception of Emoji Irony, Eq. (1)] The sentence "We use the irony rating of a post to represent the irony level of the emojis found within it" is an untested and load-bearing assumption. A post-level irony rating is assigned to the entire post, not to each emoji in it; an ironic post may contain emojis that are not themselves ironic, and a non-ironic post may contain an emoji used ironically. Because S(e) is the entire human benchmark used in the Wilcoxon and Spearman comparisons, this assumption must be validated with emoji-level annotations or at least with a per-emoji agreement analysis. Without such validation, the comparison is not a clean measure of emoji-level human perception.
  2. [GPT-4o's Classification of Emoji Irony and Results] The prompt asks GPT-4o to "rate your likelihood of using this emoji if your intention is to express irony," which elicits P(use e | ironic intent), whereas Eq. (1) computes S(e) = E[irony rating | e appears], an estimate of P(ironic intent | e). These are inverse conditionals related by Bayes' rule, and they can diverge substantially when emoji base rates are skewed, as they are here (82 emojis across about 3,000 posts). Consequently, the Wilcoxon result W = 918.5 does not establish that GPT-4o "systematically overestimates" human perception; it may simply reflect the difference between asking "how likely would you use this emoji ironically?" and "how often does this emoji appear in ironic posts?" The paper's own smirk-emoji example in Results illustrates exactly this distinction: GPT-4o's high rating is justified by possible ironic contexts, while the Ciron posts containing that emoji are mostly non-ironic. A valid comparison would require eliciting human likelihood ratings under the same prompt, or annotating emoji irony directly in context.
  3. [Results, rescaling paragraph] The paper states that the model's ratings are "rescaled from a range of 1–11 to align with the 1–5 scale" but does not give the transformation. This matters because the Wilcoxon signed-rank test operates on the signs of paired differences, and an affine transformation of the model scores with a nonzero intercept can change the sign of some differences. For example, the natural mapping x' = 0.4x + 0.6 adds a positive constant to all model scores, which can flip small negative differences to positive ones and alter W and the reported p < .001. The authors should specify the exact rescaling, justify it, and show that the statistical conclusions are robust to reasonable alternatives, or use a comparison that does not depend on the arbitrary intercept.
  4. [Prompts with Demographic Information, Figure 2] The demographic section states that "no significant differences in irony scores are observed between prompts specifying female or male gender" and that scores "tend to decrease on average as the specified age in the prompt increases," but no test statistic, p-value, effect size, or confidence interval is reported for any of these comparisons. If these demographic findings are part of the paper's contribution, they need inferential support; as written, the claims rest on visual inspection of Figure 2. The paper should also explain how the five age groups and the male/female conditions were compared (e.g., repeated-measures tests across the 82 emojis) and whether any multiple-comparison correction was applied.
minor comments (6)
  1. [Results, Wilcoxon reporting] The text says the "median irony score assigned by GPT-4o is significantly higher" than the human score, but the Wilcoxon signed-rank test tests the median of paired differences (pseudomedian), not the difference of the two marginal medians. The wording should be adjusted to avoid a technically incorrect interpretation.
  2. [Figure 1] The Spearman correlation is reported only as p < .05; the exact p-value and the sample size (N = 82) should be stated, and the caption should note that the correlation uses tied discrete scores.
  3. [Human Perception of Emoji Irony] The paper does not report how the 82 unique emojis were extracted or their frequency distribution across the approximately 3,000 emoji-containing posts. Many emojis likely appear only once or twice, making their S(e) estimates very noisy; reporting the counts, or at least their minimum, median, and maximum, would help readers assess the reliability of Eq. (1).
  4. [Experiment Setting] The paper averages three GPT-4o queries per emoji at temperature 0.5 but does not report the variance or agreement across those queries. Reporting the standard deviation or the range of the three ratings would strengthen the claim that the model scores are stable.
  5. [Discussion and Conclusions] The speculation that GPT-4o "may have been trained on data with a disproportionate representation of ironic emoji usage" is presented without evidence; it is a plausible hypothesis, but it should be labeled as a hypothesis rather than a conclusion drawn from the current results.
  6. [Throughout] Several emoji glyphs do not render in the plain-text version of the manuscript (e.g., the kiss emoji in Eq. (2) and the smirk emoji in Results). The final camera-ready version must ensure the glyphs display correctly, since the emojis are the experimental stimuli.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the human–LLM comparison rests on an external benchmark and frozen-model outputs.

full rationale

The claimed derivation chain is not circular. The human benchmark S(e) is defined by Eq. (1) directly from post-level irony ratings in the externally compiled Ciron dataset (Xiang et al. 2020); this is an operational definition, not an output of the model. GPT-4o's irony ratings are obtained by prompting the frozen model with an independent question about likelihood of use for ironic intent, and no GPT-4o parameter or prompt output is fitted to the Ciron ratings or to the computed S(e). The only transformation applied to model ratings is a rescaling from the 1–11 prompt scale to the 1–5 human scale; because the reported tests are rank-based (Wilcoxon signed-rank and Spearman), this monotone rescaling cannot manufacture the median difference or the correlation. The paper's self-citations (Lyu et al. 2024a, 2024b) are used as background motivation and as precedent for the image-format prompt, not as the evidence establishing the comparison. The inverse-conditional concern raised in the skeptic note—that Eq. (1) estimates E[irony | emoji appears] while the prompt elicits P(use emoji | ironic intent)—is a substantive construct-validity threat to the interpretation of the difference, but it is not a circularity: the two quantities are independently measured and the reported statistics would still be computed from those independent measurements. Similarly, the assumption that post-level irony transfers to contained emojis is an external validity assumption, not a step that defines the model output in terms of the human output. Accordingly, no step in the derivation reduces by construction to its own inputs.

Assumptions & free parameters 1 free parameters · 3 assumptions · 0 invented entities

The central claim is an empirical comparison and relies on no fitted parameters. It does, however, depend on two problematic domain assumptions: the transfer of post-level irony to individual emojis, and the comparability of hypothetical model likelihood ratings to usage-derived human scores. The rescaling of the model's scale is an unspecified methodological choice. No new entities are introduced.

free parameters (1)
  • Rescaling method from 11-point to 5-point scale = Not specified
    The model ratings are 'rescaled from a range of 1-11 to align with the 1-5 scale' (Results), but the exact transformation (e.g., linear (x-1)/10*4+1) is not given. The Wilcoxon and Spearman tests are rank-based and thus invariant to monotone rescaling, so the central significance claims are robust, but the reported means and the descriptive comparison may depend on the transformation.
assumptions (3)
  • domain assumption Post-level irony ratings transfer to all emojis within a post.
    The paper computes emoji irony score S(e) as the average of post irony ratings R(p) for posts containing e (Equation 1), explicitly assuming the post's irony label applies to each emoji. This is unvalidated and could be false when irony is carried by text alone.
  • domain assumption Ciron annotations are a valid ground truth for human irony perception.
    The study treats the five-annotator ratings from Xiang et al. (2020) as the human reference, without reporting inter-annotator agreement or re-validating the labels. If the annotations are noisy, the human scores are noisy.
  • domain assumption GPT-4o's self-reported likelihood ratings are comparable to human usage-derived scores.
    The model gives a hypothetical likelihood ('rate your likelihood of using this emoji if your intention is to express irony'), while the human score is a frequency-weighted average of post ratings. These are different constructs, and the paper compares them on a rescaled axis without establishing measurement invariance.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Irony in Emojis: A Comparative Study of Human and LLM Interpretation." pith.science (2026). https://pith.science/paper/XT6YW7NO

@misc{pith2026250111241,
  author       = {Pith},
  title        = {Pith review of: Irony in Emojis: A Comparative Study of Human and LLM Interpretation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XT6YW7NO}},
  note         = {Machine review of arXiv:2501.11241}
}
read the original abstract

Emojis have become a universal language in online communication, often carrying nuanced and context-dependent meanings. Among these, irony poses a significant challenge for Large Language Models (LLMs) due to its inherent incongruity between appearance and intent. This study examines the ability of GPT-4o to interpret irony in emojis. By prompting GPT-4o to evaluate the likelihood of specific emojis being used to express irony on social media and comparing its interpretations with human perceptions, we aim to bridge the gap between machine and human understanding. Our findings reveal nuanced insights into GPT-4o's interpretive capabilities, highlighting areas of alignment with and divergence from human behavior. Additionally, this research underscores the importance of demographic factors, such as age and gender, in shaping emoji interpretation and evaluates how these factors influence GPT-4o's performance.

Figures

Figures reproduced from arXiv: 2501.11241 by the authors.

Figure 2
Figure 2. When the prompt includes demographic informa [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 1
Figure 1. While GPT-4o generally rates the same emoji as more likely to be used for expressing irony compared to hu￾man perception, the irony scores assigned by GPT-4o and those perceived by humans show a significant correlation. Emojis positioned closer to the dashed line indicate greater alignment between GPT-4o’s classification of their use for expressing irony and human perception. Moreover, as shown in [PITH_FULL_IMAGE:… view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When Smiley Turns Hostile: Interpreting How Emojis Trigger LLMs' Toxicity

    cs.CL 2025-09 conditional novelty 6.0 of 10

    Emojis in harmful prompts bypass LLM safety more effectively than plain text, across 7 models and 5 languages, through a heterogeneous tokenization channel.

Reference graph

Works this paper leans on

24 extracted references · 18 canonical work pages · cited by 1 Pith paper

  1. [1]

    , " * write output.state after.block = add.period write newline

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Ai, W.; Lu, X.; Liu, X.; Wang, N.; Huang, G.; and Mei, Q. 2017. Untangling emoji popularity through semantic embeddings. In Proceedings of the international AAAI conference on web and social media, volume 11, 2--11

  4. [4]

    Chen, Y.; Yang, X.; Howman, H.; and Filik, R. 2024. Individual differences in emoji comprehension: Gender, age, and culture. Plos one, 19(2): e0297379

  5. [5]

    Chen, Z.; Lu, X.; Ai, W.; Li, H.; Mei, Q.; and Liu, X. 2018. Through a gender lens: Learning usage patterns of emojis from large-scale android users. In Proceedings of the 2018 world wide web conference, 763--772

  6. [6]

    Częstochowska, J.; Gligorić, K.; Peyrard, M.; Mentha, Y.; Bień, M.; Grütter, A.; Auer, A.; Xanthos, A.; and West, R. 2022. On the Context-Free Ambiguity of Emoji. Proceedings of the International AAAI Conference on Web and Social Media, 16(1): 1388--1392

  7. [7]

    Garcia, C.; Țurcan, A.; Howman, H.; and Filik, R. 2022. Emoji as a tool to aid the comprehension of written sarcasm: Evidence from younger and older adults. Computers in Human Behavior, 126: 106971

  8. [8]

    C.; Li, M.; Tay, L.; and Ungar, L

    Guntuku, S. C.; Li, M.; Tay, L.; and Ungar, L. H. 2019. Studying cultural differences in emoji usage across the east and the west. In Proceedings of the international AAAI conference on web and social media, volume 13, 226--235

Show all 24 references
  1. [9]

    Nice picture comment!

    Herring, S.; and Dainas, A. 2017. “Nice picture comment!” Graphicons in Facebook comment threads

  2. [10]

    Hu, T.; Guo, H.; Sun, H.; Nguyen, T.-v.; and Luo, J. 2017. Spice up your chat: the intentions and sentiment effects of using emojis. In Proceedings of the International AAAI Conference on Web and Social Media, volume 11, 102--111

  3. [11]

    Kotek, H.; Dockum, R.; and Sun, D. 2023. Gender bias and stereotypes in large language models. In Proceedings of the ACM collective intelligence conference, 12--24

  4. [12]

    Li, W.; Chen, Y.; Hu, T.; and Luo, J. 2018. Mining the relationship between emoji usage patterns and personality. In Proceedings of the international AAAI conference on web and social media, volume 12

  5. [13]

    Lyu, H.; Huang, J.; Zhang, D.; Yu, Y.; Mou, X.; Pan, J.; Yang, Z.; Wei, Z.; and Luo, J. 2024 a . GPT-4V(ision) as A Social Media Analysis Engine. ACM Trans. Intell. Syst. Technol. Just Accepted

  6. [14]

    Lyu, H.; Qi, W.; Wei, Z.; and Luo, J. 2024 b . Human vs. LMMs: Exploring the Discrepancy in Emoji Interpretation and Usage in Digital Communication. In Proceedings of the International AAAI Conference on Web and Social Media, volume 18, 2104--2110

  7. [15]

    S.; O'Brien, J.; Cai, C

    Park, J. S.; O'Brien, J.; Cai, C. J.; Morris, M. R.; Liang, P.; and Bernstein, M. S. 2023. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th annual acm symposium on user interface software and technology, 1--22

  8. [16]

    Qiu, Z.; Qiu, K.; Lyu, H.; Xiong, W.; and Luo, J. 2024. Semantics Preserving Emoji Recommendation with Large Language Models. arXiv preprint arXiv:2409.10760

  9. [17]

    Van Hee, C.; Lefever, E.; and Hoste, V. 2018. S em E val-2018 Task 3: Irony Detection in E nglish Tweets. In Apidianaki, M.; Mohammad, S. M.; May, J.; Shutova, E.; Bethard, S.; and Carpuat, M., eds., Proceedings of the 12th International Workshop on Semantic Evaluation, 39--50...

  10. [18]

    kelly is a warm person, joseph is a role model

    Wan, Y.; Pu, G.; Sun, J.; Garimella, A.; Chang, K.-W.; and Peng, N. 2023. " kelly is a warm person, joseph is a role model": Gender biases in llm-generated reference letters. arXiv preprint arXiv:2310.09219

  11. [19]

    Wang, S. 2022. Sarcastic meaning of the slightly smiling face emoji from Chinese Twitter users: When a smiling face does not show friendliness. International Journal of Languages, Literature and Linguistics, 8(2): 65--73

  12. [20]

    Wankhade, M.; Rao, A. C. S.; and Kulkarni, C. 2022. A survey on sentiment analysis methods, applications, and challenges. Artificial Intelligence Review, 55(7): 5731--5780

  13. [21]

    Weissman, B.; and Tanner, D. 2018. A strong wink between verbal and emoji-based irony: How the brain processes ironic emojis during language comprehension. PloS one, 13(8): e0201727

  14. [22]

    Xiang, R.; Gao, X.; Long, Y.; Li, A.; Chersoni, E.; Lu, Q.; Huang, C.-R.; et al. 2020. Ciron: a new benchmark dataset for Chinese irony detection

  15. [23]

    Zhang, S.; Zhang, X.; Chan, J.; and Rosso, P. 2019. Irony detection via sentiment-based transfer learning. Information Processing & Management, 56(5): 1633--1644

  16. [24]

    Zhou, Y.; Xu, P.; Wang, X.; Lu, X.; Gao, G.; and Ai, W. 2024. Emojis decoded: Leveraging chatgpt for enhanced understanding in social media communications. arXiv preprint arXiv:2402.01681

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.