REVIEW 4 major objections 6 minor 1 cited by
Optimism, Expectation, or Sarcasm? Multi-Class Hope Speech Detection in Spanish and English
T0 review · 4 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read PolyHope V2, a corpus of 30,000 English and Spanish tweets labeled for four hope subtypes, shows that fine-tuned transformers outperform large language models at detecting hope and sarcasm, and that the hardest remaining errors lie among…
desk verdict The hope-subtype benchmark is probably fine, but the paper's main novelty—the sarcasm layer—is built from the authors' own classifiers and machine translation, so the headline claims about sarcasm need independent human annotation before this is a reliable resource. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the annotated dataset itself: roughly 30,000 tweets with a four-way hope taxonomy (Generalized, Realistic, Unrealistic, Sarcastic), built from the earlier English PolyHope corpus, a Spanish counterpart collected with translated hope trigger words, and a sarcasm layer mined from two existing sarcasm datasets plus GPT-4o-translated Spanish sarcasm. The paper's quantitative findings all flow from pairing this annotation schema with supervised fine-tuning of four transformer architectures under stratified 5-fold cross-validation; the fine-tuning procedure, token-level training, and fixed label vocabulary are what the paper credits for beating prompt-based LLMs.
What would settle it
Take a random sample of 500 English and 500 Spanish tweets labeled Sarcasm, have independent human annotators label them for sarcasm without seeing the paper's labels, and measure agreement; if agreement is at chance level, or if models retrained on the human labels score below the reported 94 percent sarcasm recall, the reported advantage and benchmark comparisons rest on unreliable ground truth.
Extended reading notes
Core claim
The paper's discovery is that explicitly labeling sarcasm within hope speech changes the task: sarcastic tweets behave like not-hope in binary settings, and keeping sarcasm as a separate fifth class lets models learn the pragmatic cues that distinguish ironic hope from sincere hope. Across English and Spanish, fine-tuned transformers (RoBERTa for English, RoBERTa or Albert for Spanish) deliver the highest and most class-balanced scores, around 86 percent macro F1 in binary settings and 72 to 76 percent macro F1 in multiclass settings, while GPT-4 and Llama 3 in zero- and few-shot modes fall 8 to 9 weighted-F1 points behind even with demonstrations and collapse to 36 to 40 percent macro F1 in zero-shot multiclass. The confusion matrices show the remaining errors are not between hope and sarcasm but among the three sincere hope subtypes, especially realistic versus generalized hope.
Load-bearing premise
The sarcasm labels are treated as ground truth even though they were generated by the authors' own classifiers (English) and by GPT-4o translation with only light expert verification (Spanish), so any bias in those generators would inflate the reported sarcasm performance and the comparisons built on it.
Editorial extensions
If this is right
- Binary hope detection improves when sarcastic tweets are merged with not-hope, sharpening the decision boundary the classifier must learn.
- Multiclass accuracy drops by about ten to twelve points relative to binary, and the hardest errors are among the three sincere hope subtypes, not between hope and sarcasm.
- Sarcasm is detected at over 94 percent recall by the best fine-tuned transformer in both languages, suggesting the pragmatic cues are learnable from the annotation layer.
- Prompt-based LLMs lag by 8 to 9 weighted-F1 points even with ten demonstrations and by far more in zero-shot multiclass, so fine-tuning remains the stronger route for this task.
Reading between the lines
- If the labels hold up, sarcasm-aware hope detection could transfer to other languages by translating the taxonomy rather than retraining from scratch, a direction the paper does not develop.
- Adding temporal or evidential features (futurity markers, probability adverbs, plausibility knowledge) could push realistic-versus-generalized accuracy above the confusion levels the paper reports.
- The four-way distinction maps naturally onto psychological hope constructs, which could make the classifier useful beyond social media, for example in mental-health screening where unrealistic hope may signal a different state than grounded optimism.
- A single prompt template may understate LLM ability; prompt-optimized runs could close part of the gap, a question the paper leaves open.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper introduces PolyHope V2, a bilingual (English/Spanish) tweet dataset for fine-grained hope speech detection with four classes: Generalized, Realistic, Unrealistic, and Sarcastic hope. The English part derives from the existing PolyHope corpus, extended with sarcasm instances selected from iSarcasmEval and Sarcasm Corpus V2 using the authors' own hope and sarcasm classifiers; the Spanish part is based on the MindHope corpus, with sarcasm instances produced by GPT-4o translation of English sarcasm tweets. The authors benchmark fine-tuned transformers (RoBERTa, ALBERT, ELECTRA, DistilBERT) against GPT-4 and Llama-3 in zero-shot and few-shot settings, reporting that fine-tuned transformers consistently outperform prompt-based LLMs, especially for sarcasm. The paper also includes qualitative error analysis and confusion matrices for the best models.
Significance. If the dataset and benchmark are valid, PolyHope V2 would fill a gap by adding a sarcasm dimension to hope speech detection in two languages and by providing a systematic comparison of fine-tuned vs. prompt-based models on this task. The evaluation is standard (5-fold cross-validation, multiple metric families) and the paper is generally transparent about construction choices. However, the central novel component—the sarcasm labels—is generated by the authors' own classifiers (English) and by machine translation (Spanish), with limited human verification. Until the reliability of those labels is established, the reported superiority of fine-tuned transformers on sarcasm, and the dataset itself as a benchmark, are not fully supported. The internal inconsistency in the Table 3 class counts further weakens confidence in the resource.
major comments (4)
- [Table 3] The dataset statistics in Table 3 are internally inconsistent. For English, the binary Hope count is 4,434, whereas the sum of the three hope subtypes (Generalized 2,335 + Realistic 982 + Unrealistic 858) is 4,175, and the binary Not-Hope count is 5,081, whereas the sum of multiclass Not-Hope and Sarcasm (4,081 + 1,259) is 5,340. The same 259-instance offset appears in Spanish (binary Hope 9,654 vs. subtype sum 9,395; binary Not-Hope 10,788 vs. multiclass Not-Hope + Sarcasm 11,047). Since Section 3.3 states that sarcasm is merged into Not-Hope in the binary setting, these numbers cannot both be correct. This inconsistency undermines the dataset description and must be resolved before the corpus statistics can be used.
- [Section 3.3] The English Sarcasm class is constructed by filtering iSarcasmEval and Sarcasm Corpus V2 with the authors' own Hope and Sarcasm classifiers from [9], and the paper reports no human verification of these labels. The same transformer architectures are then fine-tuned and evaluated on this data (Tables 6 and 8), creating a circularity: the 94%+ sarcasm recall reported in Section 6 may reflect the models' ability to reproduce the decision boundary of the label generator rather than to recognize human-annotated sarcasm. Since sarcasm is the paper's main novel contribution, the claim that the dataset contains reliable sarcasm labels is not supported. Please provide evidence of label quality, such as human agreement on a sample, or temper the claims accordingly.
- [Section 3.3] The Spanish Sarcasm instances are translations of the English instances produced by GPT-4o, with only two annotators verifying 'tone.' This does not establish that the Spanish items are natural expressions of sarcasm in Spanish social media, and the protocol for adjudicating tone is not described. The paper's claim of a 'multilingual' sarcasm resource is therefore overstated; the Spanish sarcasm subset is a machine-translated artifact. Please report the verification protocol and inter-annotator agreement, and consider whether the multilingual claim can be sustained.
- [Section 4.2] The few-shot evaluation is underspecified. Section 4.2.1 says '5 balanced examples per class,' but Section 4.2.2 says 'a sample of 10 per label' and Section 5 refers to 'ten in-context demonstrations.' It is also not stated whether the demonstration examples are drawn from the training folds or from the full dataset. If test instances are included among the demonstrations, the reported few-shot scores would be inflated. Please clarify the exact sampling procedure and the number of demonstrations, and ensure no test leakage.
minor comments (6)
- [Section 3.2] The paragraph reports both 'roughly 33,300 unique' tweets after preprocessing and 'only 35,000 tweets were left'; these numbers are inconsistent and should be reconciled.
- [Section 4.2] The few-shot example count is given as '5 balanced examples per class' in Section 4.2.1, '10 per label' in Section 4.2.2, and 'ten in-context demonstrations' in Section 5; please unify the description.
- [Section 5] There is a typo 'RoBER T a' in the paragraph after Table 8; it should read 'RoBERTa'.
- [References] Several reference entries contain placeholder '???' (e.g., [6], [13]), which should be replaced with complete publication details.
- [Abstract / Section 1] The abstract and introduction state 'over 30,000 annotated tweets,' but the sum of the English and Spanish totals (9,515 + 20,442 = 29,957) is just under 30,000; please adjust the wording.
- [Figures 1 and 2] The confusion matrices referenced as Figures 1 and 2 are not visible in the submitted text; please ensure they are included and legible in the final version.
Circularity Check
The novel Sarcasm labels are generated by the authors' own transformers (English) and GPT-4o translation (Spanish), so the reported transformer advantage on sarcasm partly measures recovery of machine-generated labels.
-
fitted input called prediction
[Section 3.3; results in Section 6 and Abstract]
"We combined these datasets and used our Hope classifier (the best transformer model described in [9] fine-tuned on the PolyHope dataset), filtered all instances of sarcasm that can be detected as Hope in any category. Similarly, we developed a deep learning classifier (the best transformer model described in [9] fine-tuned on the Sarcasm dataset) and filtered all instances of potential sarcasm in other classes of Hope, i.e., unrealistic hope."
The novel Sarcasm gold labels are not independent human annotations; they are outputs of the authors' own transformer classifiers from [9]. The paper then reports that the best transformer 'separates... Sarcasm... in more than 94% of cases' (Section 6) and claims fine-tuned transformers 'especially' outperform LLMs in sarcasm (Abstract). Because the evaluated transformers belong to the same model family as the label generator, high sarcasm recall measures agreement with a transformer-generated decision boundary, not with human sarcasm, while the prompt-based LLMs are scored against labels they did not generate. The comparison for the paper's main novel class is therefore biased by construction.
-
other
[Section 3.3, Spanish Sarcasm generation]
"The Spanish instances of Sarcasm were synthetically generated using GPT-4o, by translating the English instances and keeping the sarcastic tone in Spanish. The new labels were then verified using two expert annotators."
The Spanish Sarcasm instances are translations of the English instances, so they inherit the machine-generated English labels rather than providing an independent Spanish sarcasm annotation. The two-expert verification of tone is not a fresh sarcasm labeling. Consequently, the Spanish sarcasm results (also reported as >94% recall) evaluate recovery of GPT-4o-translated, transformer-generated labels, making the bilingual sarcasm claim circular with respect to the paper's own label-generation pipeline.
full rationale
The circularity is localized to the Sarcasm class, which is the paper's stated novel contribution. The English sarcasm gold labels are produced by the authors' own transformer classifiers from [9], and the Spanish labels are GPT-4o translations of those instances; the reported sarcasm recall and the transformer-vs-LLM comparison for that class therefore reduce to agreement with machine-generated labels. The hope-subtype component, by contrast, is built on prior human-annotated PolyHope corpora [9, 40], so the binary and multiclass hope comparisons retain independent content. This warrants a moderate score of 4 rather than a higher score: the central hope-detection benchmark is externally grounded, and the label-generation procedure is disclosed in the paper, but the headline claim about sarcasm is not yet supported as a claim about human sarcasm.
Assumptions & free parameters
assumptions (3)
- domain assumption Sarcastic posts rarely convey genuine hope and can be merged with 'not hope' in the binary setting.
- domain assumption The classifiers used to filter English sarcasm instances (the authors' own Hope and Sarcasm models from [9]) are accurate enough to produce valid ground-truth labels.
- domain assumption GPT-4o-translated sarcastic tweets preserve the sarcastic tone and intent in Spanish.
Cite this review
Pith. "Pith review of Optimism, Expectation, or Sarcasm? Multi-Class Hope Speech Detection in Spanish and English." pith.science (2026). https://pith.science/paper/G6UMURP5
@misc{pith2026250417974,
author = {Pith},
title = {Pith review of: Optimism, Expectation, or Sarcasm? Multi-Class Hope Speech Detection in Spanish and English},
year = {2026},
howpublished = {\url{https://pith.science/paper/G6UMURP5}},
note = {Machine review of arXiv:2504.17974}
}
read the original abstract
Hope is a complex and underexplored emotional state that plays a significant role in education, mental health, and social interaction. Unlike basic emotions, hope manifests in nuanced forms ranging from grounded optimism to exaggerated wishfulness or sarcasm, making it difficult for Natural Language Processing systems to detect accurately. This study introduces PolyHope V2, a multilingual, fine-grained hope speech dataset comprising over 30,000 annotated tweets in English and Spanish. This resource distinguishes between four hope subtypes Generalized, Realistic, Unrealistic, and Sarcastic and enhances existing datasets by explicitly labeling sarcastic instances. We benchmark multiple pretrained transformer models and compare them with large language models (LLMs) such as GPT 4 and Llama 3 under zero-shot and few-shot regimes. Our findings show that fine-tuned transformers outperform prompt-based LLMs, especially in distinguishing nuanced hope categories and sarcasm. Through qualitative analysis and confusion matrices, we highlight systematic challenges in separating closely related hope subtypes. The dataset and results provide a robust foundation for future emotion recognition tasks that demand greater semantic and contextual sensitivity across languages.
Forward citations
Cited by 1 Pith paper
-
Towards High-Level Semantic Intelligence
A survey proposing that AI's next stage should be understood as High-Level Semantic Intelligence: mastering humor, sarcasm, metaphor, empathy, persuasion, and narrative across modalities.
Reference graph
Works this paper leans on
-
[9]
Expert Systems with Applications 225, 120078 (2023)
Balouchzahi, F., Sidorov, G., Gelbukh, A.: Polyhope: Two-level hope speech detection from tweets. Expert Systems with Applications 225, 120078 (2023)
work page 2023
-
[1]
arXiv preprint arXiv:1610.08815 (2016)
Poria, S., Cambria, E., Hazarika, D., Vij, P.: A deeper look into sarcastic tweets using deep convolutional neural networks. arXiv preprint arXiv:1610.08815 (2016)
arXiv 2016
-
[2]
IEEE Intelligent systems 28(2), 15–21 (2013)
Cambria, E., Schuller, B., Xia, Y., Havasi, C.: New avenues in opinion mining and sentiment analysis. IEEE Intelligent systems 28(2), 15–21 (2013)
work page 2013
-
[3]
Bustos-L´ opez, M., Cruz-Ram´ ırez, N., Guerra-Hern´ andez, A., S´ anchez-Morales, L.N., Alor-Hern´ andez, G.: Emotion detection from text in learning environments: 17 a review. New Perspectives on Enterprise Decision-Making Applying Artificial Intelligence Techniques, 483–508 (2021)
work page 2021
-
[4]
In: Proceedings of the 2023 6th International Conference on Information Science and Systems, pp
Gomez, L.R., Watt, T., Babaagba, K.O., Chrysoulas, C., Homay, A., Rangarajan, R., Liu, X.: Emotion recognition on social media using natural language process- ing (nlp) techniques. In: Proceedings of the 2023 6th International Conference on Information Science and Systems, pp. 113–118 (2023)
work page 2023
-
[5]
Mathematical Modelling of Engineering Problems12(2) (2025)
Mishra, A.R., Rai, A., Nandan, D., Kshirsagar, U., Singh, M.K.: Unveiling emo- tions: Nlp-based mood classification and well-being tracking for enhanced mental health awareness. Mathematical Modelling of Engineering Problems12(2) (2025)
work page 2025
-
[6]
Mohammad, S.M.: Sentiment analysis: Automatically detecting valence, emo- tions, and other affectual states from text. In: Emotion Measurement, pp. 323–379. Elsevier, ??? (2021)
work page 2021
-
[7]
arXiv preprint arXiv:2205.01996 (2022)
Buechel, S., Hahn, U.: Emobank: Studying the impact of annotation perspec- tive and representation format on dimensional emotion analysis. arXiv preprint arXiv:2205.01996 (2022)
arXiv 2022
Show all 46 references
-
[8]
In: Proceedings of the 13th International Workshop on Semantic Evaluation, pp
Chatterjee, A., Narahari, K.N., Joshi, M., Agrawal, P.: Semeval-2019 task 3: Emocontext contextual emotion detection in text. In: Proceedings of the 13th International Workshop on Semantic Evaluation, pp. 39–48 (2019)
2019
-
[10]
Computers & Education 140, 103599 (2019)
Xie, H., Chu, H.-C., Hwang, G.-J., Wang, C.-C.: Trends and development in technology-enhanced adaptive/personalized learning: A systematic review of journal publications from 2007 to 2017. Computers & Education 140, 103599 (2019)
2019
-
[11]
In: Proceedings of the 2nd Workshop on Computational Lin- guistics and Clinical Psychology: from Linguistic Signal to Clinical Reality, pp
Resnik, P., Armstrong, W., Claudino, L., Nguyen, T., Nguyen, V.-A., Boyd- Graber, J.: Beyond lda: exploring supervised topic modeling for depression-related language in twitter. In: Proceedings of the 2nd Workshop on Computational Lin- guistics and Clinical Psychology: from Li...
2015
-
[12]
Available at SSRN 4930517
Balouchzahi, F., Butt, S., Sarker, A., MA, A.-G., Sidorov, G., Gelbukh, A.: Anal- ysis of expressions of hope and regret associated with nonmedical prescription drug use in x chatter. Available at SSRN 4930517
-
[13]
Oxford University Press, ??? (1991)
Lazarus, R.S.: Emotion and Adaptation. Oxford University Press, ??? (1991)
1991
-
[14]
Cognition & emotion6(3-4), 169–200 (1992) 18
Ekman, P.: An argument for basic emotions. Cognition & emotion6(3-4), 169–200 (1992) 18
1992
-
[15]
In: Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pp
Riloff, E., Qadir, A., Surve, P., De Silva, L., Gilbert, N., Huang, R.: Sarcasm as contrast between a positive sentiment and negative situation. In: Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pp. 704–714 (2013)
2013
-
[16]
IEEE Intelligent Systems 32(6), 74–80 (2017)
Cambria, E., Poria, S., Gelbukh, A., Thelwall, M.: Sentiment analysis is a big suitcase. IEEE Intelligent Systems 32(6), 74–80 (2017)
2017
-
[17]
Ullah, F., Zamir, M.T., Ahmad, M., Sidorov, G., Gelbukh, A.: Hope: A mul- tilingual approach to identifying positive communication in social media. In: Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2024), Co- located with the 40th Conference of the Spanish Soc...
2024
-
[18]
In: Pro- ceedings of the 2008 ACM Symposium on Applied Computing, pp
Strapparava, C., Mihalcea, R.: Learning to identify emotions in text. In: Pro- ceedings of the 2008 ACM Symposium on Applied Computing, pp. 1556–1560 (2008)
2008
-
[19]
In: ECAI 2020, pp
Palakodety, S., KhudaBukhsh, A.R., Carbonell, J.G.: Hope speech detection: A computational analysis of the voice of peace. In: ECAI 2020, pp. 1881–1889. IOS Press, ??? (2020)
2020
-
[20]
Applied Sciences 13(6), 3983 (2023)
Sidorov, G., Balouchzahi, F., Butt, S., Gelbukh, A.: Regret and hope on trans- formers: An analysis of transformers on regret and hope speech detection datasets. Applied Sciences 13(6), 3983 (2023)
2023
-
[21]
Annual Review of Linguistics 2(1), 325–347 (2016)
Taboada, M.: Sentiment analysis: An overview from linguistics. Annual Review of Linguistics 2(1), 325–347 (2016)
2016
-
[22]
In: Proceedings of the Third Workshop on Com- putational Modeling of People’s Opinions, Personality, and Emotion’s in Social Media, pp
Chakravarthi, B.R.: Hopeedi: A multilingual hope speech detection dataset for equality, diversity, and inclusion. In: Proceedings of the Third Workshop on Com- putational Modeling of People’s Opinions, Personality, and Emotion’s in Social Media, pp. 41–53 (2020)
2020
-
[23]
: Overview of the shared task on hope speech detection for equality, diversity, and inclusion
Chakravarthi, B.R., Muralidaran, V., Priyadharshini, R., Cn, S., McCrae, J.P., Garc´ ıa, M.´A., Jim´ enez-Zafra, S.M., Valencia-Garc´ ıa, R., Kumaresan, P., Pon- nusamy, R., et al. : Overview of the shared task on hope speech detection for equality, diversity, and inclusion. I...
2022
-
[24]
In: Proceedings of the First Workshop on Language Technology for Equality, Diversity and Inclusion, pp
Arunima, S., Ramakrishnan, A., Balaji, A., Thenmozhi, D., et al.: ssn dibertsity@ lt-edi-eacl2021: hope speech detection on multilingual youtube comments via transformer based approach. In: Proceedings of the First Workshop on Language Technology for Equality, Diversity and In...
2021
-
[25]
Knowledge-Based Systems308, 112746 19 (2025)
Balouchzahi, F., Butt, S., Amjad, M., Sidorov, G., Gelbukh, A.: Urduhope: Analy- sis of hope and hopelessness in urdu texts. Knowledge-Based Systems308, 112746 19 (2025)
2025
-
[26]
Ponnusamy, K.K., Vegupatti, M., Kumaresan, P.K., Priyadharshini, R., Buite- laar, P., Chakravarthi, B.R.: Vel@ iberlef 2024: Hope speech detection in spanish social media comments using bert pre-trained model. In: Proceedings of the Iberian Languages Evaluation Forum (IberLEF ...
2024
-
[27]
International Journal of Data Science and Analytics 14(4), 389–406 (2022)
Chakravarthi, B.R.: Multilingual hope speech detection in english and dravidian languages. International Journal of Data Science and Analytics 14(4), 389–406 (2022)
2022
-
[28]
Journal of King Saud University-Computer and Information Sciences 35(8), 101736 (2023)
Malik, M.S.I., Nazarova, A., Jamjoom, M.M., Ignatov, D.I.: Multilingual hope speech detection: A robust framework using transfer learning of fine-tuning roberta model. Journal of King Saud University-Computer and Information Sciences 35(8), 101736 (2023)
2023
-
[29]
arXiv preprint arXiv:2409.10965 (2024)
Thangaraj, H., Chenat, A., Walia, J.S., Marivate, V.: Cross-lingual trans- fer of multilingual models on low resource african languages. arXiv preprint arXiv:2409.10965 (2024)
2024 arXiv
-
[30]
Language Resources and Evaluation 57(4), 1487–1514 (2023)
Garc´ ıa-Baena, D., Garc´ ıa-Cumbreras, M.´A., Jim´ enez-Zafra, S.M., Garc´ ıa-D´ ıaz, J.A., Valencia-Garc´ ıa, R.: Hope speech detection in spanish: The lgbt case. Language Resources and Evaluation 57(4), 1487–1514 (2023)
2023
-
[31]
Garc´ ıa-Baena, D., Balouchzahi, F., Butt, S., Garc´ ıa-Cumbreras, M. ´A., Tonja, A.L., Garc´ ıa-D´ ıaz, J.A., Bozkurt, S., Chakravarthi, B.R., Ceballos, H.G., Valencia-Garc´ ıa, R.,et al.: Overview of hope at iberlef 2024: Approaching hope speech detection in social media fro...
2024
-
[32]
Language Resources and Evaluation (2023)
Nath, T., Singh, V.K., Gupta, V.: Bonghope: An annotated corpus for bengali hope speech detection. Language Resources and Evaluation (2023)
2023
-
[33]
Procesamiento del lenguaje natural 71, 371–381 (2023)
Jim´ enez-Zafra, S.M., Garcia-Cumbreras, M.´A., Garc´ ıa-Baena, D., Garcia-D´ ıaz, J.A., Chakravarthi, B.R., Valencia-Garc´ ıa, R., Ure˜ na-L´ opez, L.A.: Overview of hope at iberlef 2023: Multilingual hope speech detection. Procesamiento del lenguaje natural 71, 371–381 (2023)
2023
-
[34]
In: Proceedings of the 12th Annual Meeting of the Forum for Information Retrieval Evaluation, pp
Mandl, T., Modha, S., Kumar M, A., Chakravarthi, B.R.: Overview of the hasoc track at fire 2020: Hate speech and offensive language identification in tamil, malayalam, hindi, english and german. In: Proceedings of the 12th Annual Meeting of the Forum for Information Retrieval ...
2020
-
[35]
In: Proceedings of the EACL Hackashop on News Media Content Analysis and Automated Report Generation, pp
Pelicon, A., Shekhar, R., Martinc, M., ˇSkrlj, B., Purver, M., Pollak, S.,et al.: Zero- shot cross-lingual content filtering: Offensive language and hate speech detection. In: Proceedings of the EACL Hackashop on News Media Content Analysis and Automated Report Generation, pp....
2021
-
[36]
In: Proceedings of the Second Workshop on Figurative Language Processing, pp
Srivastava, H., Varshney, V., Kumari, S., Srivastava, S.: A novel hierarchical bert architecture for sarcasm detection. In: Proceedings of the Second Workshop on Figurative Language Processing, pp. 93–97 (2020)
2020
-
[37]
In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp
Chauhan, D.S., Dhanush, S., Ekbal, A., Bhattacharyya, P.: Sentiment and emo- tion help sarcasm? a multi-task learning framework for multi-modal sarcasm, sentiment and emotion analysis. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics,...
2020
-
[38]
In: Proceedings of the AAAI Conference on Artificial Intelligence, vol
Jia, M., Xie, C., Jing, L.: Debiasing multimodal sarcasm detection with contrastive learning. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, pp. 18354–18362 (2024)
2024
-
[39]
In: Proceedings of the Second Workshop on Figurative Language Processing, pp
Dong, X., Li, C., Choi, J.D.: Transformer-based context-aware sarcasm detec- tion in conversation threads from social media. In: Proceedings of the Second Workshop on Figurative Language Processing, pp. 276–280 (2020)
2020
-
[40]
Sidorov, G., Balouchzahi, F., Ramos, L., G´ omez-Adorno, H., Gelbukh, A.: Mind- hope: Multilingual identification of nuanced dimensions of hope (2024)
2024
-
[41]
In: The 16th International Workshop on Semantic Evaluation 2022, pp
Farha, I.A., Oprea, S., Wilson, S., Magdy, W.: Semeval-2022 task 6: isarcasmeval, intended sarcasm detection in english and arabic. In: The 16th International Workshop on Semantic Evaluation 2022, pp. 802–814 (2022). Association for Computational Linguistics
2022
-
[42]
In: Proceedings of the 17th Annual Meeting of the Special Interest Group on Discourse and Dialogue, pp
Oraby, S., Harrison, V., Reed, L., Hernandez, E., Riloff, E., Walker, M.: Creating and characterizing a diverse corpus of sarcasm in dialogue. In: Proceedings of the 17th Annual Meeting of the Special Interest Group on Discourse and Dialogue, pp. 31–41 (2016)
2016
-
[43]
arXiv preprint arXiv:2303.08774 (2023)
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F.L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al.: Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)
2023 arXiv
-
[44]
Bubeck, S., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y.T., Li, Y., Lundberg, S., et al.: Sparks of artificial general intelligence: Early experiments with gpt-4 (2023)
2023
-
[45]
arXiv preprint arXiv:2407.21783 (2024) 21
Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Vaughan, A., et al.: The llama 3 herd of models. arXiv preprint arXiv:2407.21783 (2024) 21
2024 arXiv
-
[46]
In: The Eleventh International Conference on Learning Representations (2022) 22
Zhou, Y., Muresanu, A.I., Han, Z., Paster, K., Pitis, S., Chan, H., Ba, J.: Large language models are human-level prompt engineers. In: The Eleventh International Conference on Learning Representations (2022) 22
2022
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.