REVIEW 3 major objections 5 minor 45 references
RCLMuFN: Relational Context Learning and Multiplex Fusion Network for Multimodal Sarcasm Detection
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Sarcasm detector hits 93.09% accuracy with relational context fusion.
desk verdict The architecture is a plausible new combination of known blocks, but the SOTA claim is shaky because the three fusion weights are selected on the test sets (or at least that's not ruled out) and there are no error bars or code. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a multiplex fusion of two feature streams: H_deep, produced by the relational context learning module (shallow cross-attention between BERT and ResNet features, per-modality co-attention, then cross-modal co-attention), and H_CLIP, produced by the CLIP-view feature fusion module (two co-attention blocks over the raw CLIP text and image embeddings). These streams are combined with scalar weights alpha, beta, and gamma that interpolate between directional co-attention paths, letting the model balance visual and textual evidence while preserving the original CLIP semantic alignment.
What would settle it
Re-run the same training and evaluation with alpha, beta, and gamma selected only on the validation split, then test once on the held-out test sets; if accuracy and F1 fall back to near the best prior baselines (about 90.5% on MMSD and 86.5% on MMSD 2.0), the claimed advantage is an artifact of test-set peeking.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that explicitly modeling the relational context between text and image—through shallow cross-attention, stacked self-attention, and deep co-attention, supplemented by a second fusion stream from CLIP—yields state-of-the-art multimodal sarcasm detection. The method reports improvements over the strongest prior model of 2.49 accuracy points on MMSD and 5.03 accuracy points on MMSD 2.0. The authors interpret this as evidence that dynamically learning contextual relations, instead of relying on static graph construction from external knowledge, better captures the way sarcastic meaning shifts as context evolves.
Load-bearing premise
The reported state-of-the-art results rest on fusion weights alpha, beta, and gamma that were chosen by trying different values on the test sets, so if that selection is not a valid model-selection procedure, the accuracy and F1 gains could shrink or vanish on new data.
Editorial extensions
If this is right
- On both public benchmarks, the model reports the highest accuracy and F1 among the compared methods, so a practitioner could adopt it as a strong baseline without building graph structures.
- The ablation study shows that removing the CLIP-view fusion stream causes the largest accuracy drop, indicating that preserving the raw CLIP alignment is essential to the model's performance.
- The dynamic relational context module is claimed to improve cross-dataset generalization, since the same architecture trained on the cleaner MMSD 2.0 data still surpasses prior methods.
- The method currently targets English content; extending the four encoders to multilingual models would be the natural next step.
- The multiplex fusion idea could be reused for other multimodal tasks where image-text incongruity matters, such as humor or offensive-language detection.
Reading between the lines
- The reported alpha, beta, and gamma values were chosen in Section 6.8 by evaluating accuracy and F1 directly on the test sets; if that counts as test-set peeking, the headline gains are likely optimistic and a validation-set selection would be needed to confirm generalization.
- The ablation pattern suggests that the CLIP-view fusion module, not the relational context learning module, contributes the largest share of accuracy; the paper's framing emphasizes relational context, but the numbers point to the CLIP stream as the primary driver.
- A natural testable extension is to apply the same relational-context-plus-multiplex-fusion architecture to other multimodal incongruity tasks, where the dynamic-context hypothesis can be checked independently of sarcasm datasets.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes RCLMuFN, a multimodal sarcasm detection model that extracts text and image features using CLIP, BERT, and ResNet-50, then applies a shallow feature interaction module, a relational context learning module with co-attention, a CLIP-view feature fusion module, and a multiplex feature fusion module. The method is evaluated on the MMSD and MMSD 2.0 datasets, reporting 93.09% accuracy / 91.52% F1 on MMSD and 91.57% / 90.25% on MMSD 2.0, which the authors claim is state-of-the-art. The paper also includes ablations, visualizations of attention and feature distributions, error analysis, and a parameter analysis of three fusion weights α, β, and γ.
Significance. If the empirical results are sound, the paper presents a competitive architecture that avoids the graph-construction overhead of several recent baselines while achieving large reported gains, especially on MMSD 2.0. The proposed relational context learning module and multiplex fusion are plausible and the ablations show that removing the CLIP-view fusion causes a dramatic drop, suggesting the components matter. However, the contribution is primarily empirical and incremental: there is no theoretical analysis, no released code, no error bars over seeds, and the parameter-selection protocol is ambiguous. These gaps currently prevent the reader from verifying whether the reported state-of-the-art margins reflect genuine generalization or optimization on the test split; the claimed significance is therefore conditional on resolving the evaluation-protocol issues.
major comments (3)
- The parameter analysis sweeps α, β, and γ over 0.1 to 0.9 and selects values with the best Acc/F1, but it does not state whether the sweep is performed on the validation split or the test split. Since Table 1 defines a validation set and the final Table 2 numbers are the ones used for the SOTA claim, this distinction is load-bearing. If the sweep used the test sets, the reported results are the best among many configurations selected by peeking at test labels, and the claimed margins over prior work are inflated. The authors must clarify the split used for parameter selection, and ideally report the validation-curve results, fix the selected weights before touching the test set, and state this explicitly. If test-set selection was in fact used, the experiments should be redone with validation-based selection or the claims should be reframed as upper bounds rather than generalization results.
- All reported results are single-run point estimates with no error bars, no multiple seeds, and no significance tests. The test sets are small (2,409 samples each), and the claimed improvements over the strongest baselines (e.g., +2.49 Acc on MMSD, +5.03 Acc on MMSD 2.0) could plausibly arise from run-to-run variance. The authors should report means and standard deviations over at least three to five random seeds for the main comparison and the ablations, and where feasible include statistical significance tests against the strongest baselines. Without such information, the SOTA claim is not empirically substantiated.
- The paper promises to release code only after acceptance and does not provide sufficient implementation details for reproduction: the number of attention heads, projection and MLP hidden sizes, dropout rates, and any learning-rate scheduling are omitted. The architectural description is clear at a high level, but the exact model sizes and training details are needed to verify the results. The authors should release the code (or a detailed pseudo-code with all hyperparameters) as part of the submission, or at minimum provide a complete hyperparameter table and commit to a public repository.
minor comments (5)
- [Section 1, abstract intro] The introduction states that the method produces detection accuracy 3.91% higher than SOTA on MMSD 2.0, but Table 2 shows the accuracy gain is 5.03% and the F1 gain is 3.91%; the abstract does not give numbers, so this discrepancy should be corrected for consistency.
- [Section 6.8 text, Figure 8] In the parameter analysis, the text uses the symbol β when discussing the optimal values of α ("when β is taken as 0.9 and 0.6") and γ ("when β is taken as 0.3 and 0.5"); these should be α and γ, respectively, to avoid confusion about which parameter is being swept.
- [References] References [19] and [20] are duplicates (InCrossMGs), as are [21] and [22] (CMGCN) and [35] and [36] (Tang et al.); the bibliography should be consolidated.
- [Section 6.6, Figure 6] The claim that MuFFM makes sarcasm and non-sarcasm representations 'more dispersed' and therefore 'better generalized' is based on visual inspection of a t-SNE plot of 200 samples; this claim would be stronger if accompanied by a quantitative metric such as average inter-class vs. intra-class distances.
- [Equation (19)] Equation (19) contains an extra opening parenthesis: 'F*_fuse = γ · Sigmoid(F_fuse) · F_fuse) + (1−γ) · H_CLIP' should be cleaned up for readability.
Circularity Check
No derivation-chain circularity; SOTA claim rests on external benchmarks, but the §6.8 selection of fusion weights by test-set Acc/F1 introduces a mild evaluation circularity.
-
fitted input called prediction
[Section 6.8 'Parameter Analysis'; Eqs. (16), (17), (19); Table 2]
"To assess the impact on model performance when features from different paths are given different weights. We analyzed three weighting coefficients α, β, and γ in depth on two sarcasm detection datasets. The performance of the model is evaluated mainly with the Acc and F1 scores as the reference metrics."
The reported SOTA numbers in Table 2 are obtained using the α, β, γ values that Section 6.8 chooses by optimizing Acc/F1 on the two evaluation datasets. The paper does not state that this sweep is restricted to the validation split; it says only that the weights were analyzed 'on two sarcasm detection datasets' and that 'model performance is optimal' at the chosen values. As described, the reported Acc/F1 figures are therefore not independent predictions of generalization but the selected optima of the same metric over the swept weights. The architecture itself is still externally benchmarked, so this is a partial evaluation circularity rather than a fully forced result.
full rationale
The paper's central claim is an empirical accuracy/F1 result on two public benchmarks, not a derived law. The architecture equations (2)-(21) compose standard pretrained encoders, attention modules, and weighted sums; no output quantity is defined in terms of the reported result, and no prior result by these authors is invoked to force the design. The only place where a reported number could reduce to a fitting choice is Section 6.8: the fusion weights α, β, γ are selected by evaluating Acc/F1 on the two datasets, and the same metric values are then reported as the SOTA result in Table 2. Because the paper does not state that the sweep is confined to the validation split, the reported numbers are at least partly selection outcomes rather than independent predictions. I treat this as a mild evaluation circularity (score 2), not a forced derivation, since the architecture itself is benchmarked against external datasets and the ablation results show substantial performance from the components independent of these three scalar weights.
Assumptions & free parameters
free parameters (3)
- alpha (fusion weight in RCLM, Eq. 16) =
0.9 on MMSD, 0.6 on MMSD 2.0
- beta (fusion weight in CLIP-VFFM, Eq. 17) =
0.5 on MMSD, 0.7 on MMSD 2.0
- gamma (fusion weight in MuFFM, Eq. 19) =
0.3 on MMSD, 0.5 on MMSD 2.0
assumptions (4)
- standard math Cross-attention and self-attention operations are mathematically well-defined and suitable for feature interaction.
- domain assumption Pretrained CLIP, BERT, and ResNet encoders provide useful features for sarcasm detection.
- domain assumption The MMSD and MMSD 2.0 datasets are reliable, and their test labels correctly reflect sarcasm.
- ad hoc to paper Selecting the fusion weights alpha, beta, gamma based on test-set accuracy and F1 does not invalidate the generalization claims.
Cite this review
Pith. "Pith review of RCLMuFN: Relational Context Learning and Multiplex Fusion Network for Multimodal Sarcasm Detection." pith.science (2026). https://pith.science/paper/5SE3LVUO
@misc{pith2026241213008,
author = {Pith},
title = {Pith review of: RCLMuFN: Relational Context Learning and Multiplex Fusion Network for Multimodal Sarcasm Detection},
year = {2026},
howpublished = {\url{https://pith.science/paper/5SE3LVUO}},
note = {Machine review of arXiv:2412.13008}
}
read the original abstract
Sarcasm typically conveys emotions of contempt or criticism by expressing a meaning that is contrary to the speaker's true intent. Accurate detection of sarcasm aids in identifying and filtering undesirable information on the Internet, thereby reducing malicious defamation and rumor-mongering. Nonetheless, the task of automatic sarcasm detection remains highly challenging for machines, as it critically depends on intricate factors such as relational context. Most existing multimodal sarcasm detection methods focus on introducing graph structures to establish entity relationships between text and images while neglecting to learn the relational context between text and images, which is crucial evidence for understanding the meaning of sarcasm. In addition, the meaning of sarcasm changes with the evolution of different contexts, but existing methods may not be accurate in modeling such dynamic changes, limiting the generalization ability of the models. To address the above issues, we propose a relational context learning and multiplex fusion network (RCLMuFN) for multimodal sarcasm detection. Firstly, we employ four feature extractors to comprehensively extract features from raw text and images, aiming to excavate potential features that may have been previously overlooked. Secondly, we utilize the relational context learning module to learn the contextual information of text and images and capture the dynamic properties through shallow and deep interactions. Finally, we employ a multiplex feature fusion module to enhance the generalization of the model by penetratingly integrating multimodal features derived from various interaction contexts. Extensive experiments on two multimodal sarcasm detection datasets show that our proposed method achieves state-of-the-art performance.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Wallace, Hao Lyu, Paula Carvalho, and Mário J
Silvio Amir, Byron C. Wallace, Hao Lyu, Paula Carvalho, and Mário J. Silva. 2016. Modelling Context with User Embeddings for Sarcasm Detection in Social Media. In Proceedings of the 20th SIGNLL Conference on Computational Natural Language Learning, CoNLL 2016, Berlin, Germany, August 11-12, 2016 . ACL, 167–177
work page 2016
-
[2]
Christos Baziotis, Athanasiou Nikolaos, Pinelopi Papalampidi, Athanasia Kolovou, Georgios Paraskevopoulos, Nikolaos Ellinas, and Alexandros Potamianos. 2018. NTUA-SLP at SemEval-2018 Task 3: Tracking Ironic Tweets using Ensembles of Word and Character Level Attentive RNNs. InProceedings of The 12th International Workshop on Semantic Evaluation, SemEval@NA...
work page 2018
-
[3]
All Your Products Are Incredibly Amazing!!!
Mondher Bouazizi and Tomoaki Ohtsuki. 2015. Sarcasm Detection in Twitter: "All Your Products Are Incredibly Amazing!!!" - Are They Really?. In2015 IEEE Global Communications Conference, GLOBECOM 2015, San Diego, CA, USA, December 6-10, 2015. IEEE, 1–6
work page 2015
-
[4]
Yitao Cai, Huiyu Cai, and Xiaojun Wan. 2019. Multi-Modal Sarcasm Detection in Twitter with Hierarchical Fusion Model. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers . Association for Computational Linguistics, 2506–2515
work page 2019
-
[5]
Santiago Castro, Devamanyu Hazarika, Verónica Pérez-Rosas, Roger Zimmer- mann, Rada Mihalcea, and Soujanya Poria. 2019. Towards Multimodal Sarcasm Detection (An _Obviously_ Perfect Paper). In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers . Asso...
work page 2019
- [6]
-
[7]
Zixin Chen, Hongzhan Lin, Ziyang Luo, Mingfei Cheng, Jing Ma, and Guang Chen. 2024. CofiPara: A Coarse-to-fine Paradigm for Multimodal Sarcasm Target Identification with Large Multimodal Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Association for Computational Linguistics, Ba...
work page 2024
-
[8]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Associa- tion for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, ...
work page 2019
Show all 45 references
-
[9]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recogn...
2021
-
[10]
Raymond W Gibbs. 1986. On the psycholinguistics of sarcasm. Journal of experimental psychology: general 115, 1 (1986), 3
1986
-
[11]
González-Ibáñez, Smaranda Muresan, and Nina Wacholder
Roberto I. González-Ibáñez, Smaranda Muresan, and Nina Wacholder. 2011. Iden- tifying Sarcasm in Twitter: A Closer Look. In The 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, Pro- ceedings of the Conference, 19-24 June, 2011,...
2011
-
[12]
Alex Graves and Jürgen Schmidhuber. 2005. Framewise phoneme classification with bidirectional LSTM and other neural network architectures.Neural Networks 18, 5-6 (2005), 602–610
2005
-
[13]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep Residual Learning for Image Recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016 . IEEE Computer Society, 770–778
2016
-
[14]
Guimin Hu, Ting-En Lin, Yi Zhao, Guangming Lu, Yuchuan Wu, and Yongbin Li
-
[15]
Mengzhao Jia, Can Xie, and Liqiang Jing. 2024. Debiasing Multimodal Sarcasm Detection with Contrastive Learning. In Thirty-Eighth AAAI Conference on Artifi- cial Intelligence, AAAI 2024, Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence, IAAI 2024, ...
2024
-
[16]
Aditya Joshi, Pushpak Bhattacharyya, and Mark James Carman. 2017. Automatic Sarcasm Detection: A Survey. ACM Comput. Surv. 50, 5 (2017), 73:1–73:22
2017
-
[17]
Yoon Kim. 2014. Convolutional Neural Networks for Sentence Classification. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing, EMNLP 2014, October 25-29, 2014, Doha, Qatar, A meeting of SIGDAT, a Special Interest Group of the ACL . ACL, 1746–1751
2014
-
[18]
Lingyao Li, Lizhou Fan, Shubham Atreja, and Libby Hemphill. 2024. “HOT” ChatGPT: The Promise of ChatGPT in Detecting and Discriminating Hateful, Offensive, and Toxic Comments on Social Media. ACM Trans. Web 18, 2, Article 30 (March 2024), 36 pages
2024
-
[20]
Bin Liang, Chenwei Lou, Xiang Li, Lin Gui, Min Yang, and Ruifeng Xu. 2021. Multi-Modal Sarcasm Detection with Interactive In-Modal and Cross-Modal Graphs. In MM ’21: ACM Multimedia Conference, Virtual Event, China, October 20 - 24, 2021. ACM, 4707–4715
2021
-
[22]
Bin Liang, Chenwei Lou, Xiang Li, Min Yang, Lin Gui, Yulan He, Wenjie Pei, and Ruifeng Xu. 2022. Multi-Modal Sarcasm Detection via Cross-Modal Graph Con- volutional Network. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: L...
2022
-
[23]
Hui Liu, Wenya Wang, and Haoliang Li. 2022. Towards Multi-Modal Sarcasm Detection via Hierarchical Congruity Modeling with Knowledge Enhancement. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022, Abu Dhabi, United Arab Emira...
2022
-
[24]
Lisong Ou and Zhixin Li. 2024. Modeling Multi-Task Joint Training of Aggre- gate Networks for Multi-Modal Sarcasm Detection. In Proceedings of the 2024 International Conference on Multimedia Retrieval (Phuket, Thailand) (ICMR ’24). Association for Computing Machinery, New York...
2024
-
[26]
Hongliang Pan, Zheng Lin, Peng Fu, Yatao Qi, and Weiping Wang. 2020. Modeling Intra and Inter-modality Incongruity for Multi-Modal Sarcasm Detection. In Findings of the Association for Computational Linguistics: EMNLP 2020, Online Event, 16-20 November 2020 (Findings of ACL, V...
2020
-
[27]
Soujanya Poria, Erik Cambria, Devamanyu Hazarika, and Prateek Vij. 2016. A Deeper Look into Sarcastic Tweets Using Deep Convolutional Neural Networks. In COLING 2016, 26th International Conference on Computational Linguistics, Proceedings of the Conference: Technical Papers, D...
2016
-
[28]
Yang Qiao, Liqiang Jing, Xuemeng Song, Xiaolin Chen, Lei Zhu, and Liqiang Nie
-
[29]
Libo Qin, Shijue Huang, Qiguang Chen, Chenran Cai, Yudi Zhang, Bin Liang, Wanxiang Che, and Ruifeng Xu. 2023. MMSD2.0: Towards a Reliable Multi-modal Sarcasm Detection System. In Findings of the Association for Computational Lin- guistics: ACL 2023, Toronto, Canada, July 9-14,...
2023
-
[30]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. In Proceedings...
2021
-
[31]
Reimann, Florian A
Merle M. Reimann, Florian A. Kunneman, Catharine Oertel, and Koen V. Hindriks
-
[32]
Ellen Riloff, Ashequl Qadir, Prafulla Surve, Lalindra De Silva, Nathan Gilbert, and Ruihong Huang. 2013. Sarcasm as Contrast between a Positive Sentiment and Negative Situation. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, EMNLP 20...
2013
-
[33]
Tetreault, and Liangliang Cao
Rossano Schifanella, Paloma de Juan, Joel R. Tetreault, and Liangliang Cao. 2016. Detecting Sarcasm in Multimodal Social Platforms. InProceedings of the 2016 ACM Conference on Multimedia Conference, MM 2016, Amsterdam, The Netherlands, October 15-19, 2016. ACM, 1136–1145. RCLM...
2016
-
[34]
Guixin Su, Mingmin Wu, Zhongqiang Huang, Yongcheng Zhang, Tongguan Wang, Yuxue Hu, and Ying Sha. 2024. Refine, Align, and Aggregate: Multi- view Linguistic Features Enhancement for Aspect Sentiment Triplet Extraction. In Findings of the Association for Computational Linguistic...
2024
-
[36]
Binghao Tang, Boda Lin, Haolong Yan, and Si Li. 2024. Leveraging Generative Large Language Models with Visual Instruction and Demonstration Retrieval for Multimodal Sarcasm Detection. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Co...
2024
-
[37]
Yuan Tian, Nan Xu, Ruike Zhang, and Wenji Mao. 2023. Dynamic Routing Transformer Network for Multimodal Sarcasm Detection. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2...
2023
-
[38]
Oren Tsur, Dmitry Davidov, and Ari Rappoport. 2010. ICWSM—a great catchy name: Semi-supervised recognition of sarcastic sentences in online product reviews. In Proceedings of the International AAAI Conference on Web and Social Media, Vol. 4. 162–169
2010
-
[39]
Yiwei Wei, Shaozu Yuan, Hengyang Zhou, Longbiao Wang, Zhiling Yan, Ruosong Yang, and Meng Chen. 2024. G2SAM: Graph-Based Global Semantic Awareness Method for Multimodal Sarcasm Detection. In Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, February 20-27, 2...
2024
-
[40]
Changsong Wen, Guoli Jia, and Jufeng Yang. 2023. DIP: Dual Incongruity Per- ceiving Network for Sarcasm Detection. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24,
2023
-
[41]
Chuhan Wu, Fangzhao Wu, Sixing Wu, Junxin Liu, Zhigang Yuan, and Yongfeng Huang. 2018. THU_NGN at SemEval-2018 Task 3: Tweet Irony Detection with Densely connected LSTM and Multi-task Learning. In Proceedings of The 12th International Workshop on Semantic Evaluation, SemEval@N...
2018
-
[42]
Tao Xiong, Peiran Zhang, Hongbo Zhu, and Yihui Yang. 2019. Sarcasm Detection with Self-matching Networks and Low-rank Bilinear Pooling. In The World Wide Web Conference, WWW 2019, San Francisco, CA, USA, May 13-17, 2019 . ACM, 2115–2124
2019
-
[43]
Nan Xu, Zhixiong Zeng, and Wenji Mao. 2020. Reasoning with Multimodal Sarcastic Tweets via Modeling Cross-Modality Contrast and Semantic Association. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020 ....
2020
-
[44]
Meishan Zhang, Yue Zhang, and Guohong Fu. 2016. Tweet Sarcasm Detection Using Deep Neural Network. In COLING 2016, 26th International Conference on Computational Linguistics, Proceedings of the Conference: Technical Papers, December 11-16, 2016, Osaka, Japan . ACL, 2449–2460
2016
-
[45]
Zhihong Zhu, Xuxin Cheng, Guimin Hu, Yaowei Li, Zhiqi Huang, and Yuexian Zou. 2024. Towards Multi-modal Sarcasm Detection via Disentangled Multi- grained Multi-modal Distilling. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Re...
2024
-
[46]
Zhihong Zhu, Xianwei Zhuang, Yunyan Zhang, Derong Xu, Guimin Hu, Xian Wu, and Yefeng Zheng. 2024. Tfcd: Towards multi-modal sarcasm detection via training-free counterfactual debiasing. In Proc. of IJCAI
2024
-
[2022]
In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing
UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, Abu Dhabi, United Arab Emirates, 7837–7851
2022
-
[2023]
Mutual-Enhanced Incongruity Learning Network for Multi-Modal Sarcasm Detection. In Thirty-Seventh AAAI Conference on Artificial Intelligence, AAAI 2023, Thirty-Fifth Conference on Innovative Applications of Artificial Intelligence, IAAI 2023, Thirteenth Symposium on Educationa...
2023
-
[2024]
ACM Trans
A Survey on Dialogue Management in Human-robot Interaction. ACM Trans. Hum. Robot Interact. 13, 2 (2024), 22:1–22:22
2024
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.