REVIEW 3 major objections 2 minor 68 references
A Multimodal-Multitask Framework with Cross-modal Relation and Hierarchical Interactive Attention for Semantic Comprehension
T0 review · 3 major / 2 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read This paper's abstract announces a multimodal-multitask fusion framework (MM-ORIENT), but the full text is a different manuscript on survival trees for length-biased data; the abstract's claims appear nowhere in the body.
desk verdict The submission is a survival-tree paper wearing a multimodal abstract; no MM-ORIENT content exists in the body, so the claimed results are unsupported in this document. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Three components carry the argument. (1) The conditional inference tree (CIT) skeleton: recursive partitioning that selects splitting variables by permutation tests of independence rather than impurity minimization, avoiding the selection bias of greedy impurity search; the paper keeps this skeleton and swaps in length-bias components. (2) The LBRC score function derived from the full likelihood: δ + log Ŝ(Z) minus a population constant that cancels on standardization, so splitting reduces to testing association with the censored event and the log-survival estimate. (3) Two nonparametric estimators of the unbiased survival function: MFLE, an EM-computed estimator maximizing the full likeliho
What would settle it
A decisive check: simulate a prevalent cohort under the paper's own covariate-dependent truncation scenario at a mild departure strength, then compare the LBRC forest against its left-truncated counterpart on integrated prediction error; if the advantage vanishes or reverses, the efficiency claim depends too tightly on the stationarity assumption. For the abstract's claim, the check is simpler: nothing in the manuscript describes MM-ORIENT's architecture, datasets, or results, so there is no artifact to point to.
Extended reading notes
Core claim
The full-text manuscript claims survival trees and forests should be purpose-built for length-biased right-censored (LBRC) data rather than inherited from the general left-truncated toolbox. It proposes LBRC-CIT and LBRC-CIF, conditional inference trees and forests that split by a permutation test on a full-likelihood score (effectively δ + log Ŝ(Z)) and predict with either a full-likelihood nonparametric maximum-likelihood estimator (MFLE) or a closed-form composite conditional-likelihood estimator (MCLE). The claim: exploiting the uniform truncation-time distribution implied by a stationary onset process yields efficiency gains in tree recovery and prediction. Simulations over four hazard
Load-bearing premise
The load-bearing premise is that disease onset follows a stationary Poisson process, so the time from diagnosis to study enrollment is uniform; if a real cohort violates that, the length-biased estimators everything is built on are misspecified, and the paper's own sensitivity analysis shows covariate-dependent violations shift split selection toward the covariate that drives the truncation.
Editorial extensions
If this is right
- Prevalent-cohort studies currently running left-truncated conditional inference trees or forests can switch to the LBRC variants with the same pipeline and expect better tree recovery and lower prediction error, with the largest gains under heavy censoring.
- The closed-form MCLE variant nearly matches the full-likelihood MFLE variant in tree recovery while running in about a third of the time, so the computationally cheap option does not obviously cost accuracy.
- Efficiency gains appear across tree, linear, nonlinear, and interaction data structures, not only the tree structure the method assumes, though gains concentrate when the true partition is tree-structured.
- Under mild violation of the stationarity assumption, variable selection stays approximately unbiased and recovery remains better than the left-truncated baseline; under severe covariate-dependent violation, split selection shifts toward the covariate driving the truncation.
- In the real-data lung-cancer application, the LBRC forest variants give the lowest cross-validated integrated Brier scores among tree methods, and the LBRC Cox model is also competitive with them.
Reading between the lines
- Editorial: because the splitting score's distinguishing term cancels on standardization, the LBRC split-selection advantage flows entirely through the choice of survival estimator Ŝ; a clean ablation that changes only the estimator while holding the score fixed would isolate where the gain lives.
- Editorial: the plug-in design means future length-biased estimators — for restricted mean survival time or quantile residual lifetime — could be dropped into the same tree/forest shell, making the pairing a template rather than a single method.
- Editorial: the paper's own sensitivity results imply a practical warning it never states outright: analysts should test the stationarity assumption before deploying these trees, since covariate-dependent onset rates silently convert valid split selection into selection of the truncation-driving covariate.
- Editorial: as for the abstract, the MM-ORIENT claim about reducing latent noise in multimodal fusion cannot be tested from this manuscript, which contains no architecture, datasets, results, or code for it; the multimodal paper, if it exists, must be evaluated from its own text.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The submission consists of an abstract proposing a multimodal-multitask framework, MM-ORIENT, which uses cross-modal relation graphs and Hierarchical Interactive Monomodal Attention (HIMA) to reconstruct monomodal features and fuse them in a way that allegedly reduces latent-stage noise; the abstract further claims that extensive experiments on three datasets demonstrate the framework's effectiveness. However, the supplied full text is an entirely different manuscript: "Tree-based methods for length-biased survival data" (arXiv:2508.16312v4 [stat.ME]), which presents survival trees and forests for length-biased right-censored data, including simulations and a lung-cancer application. None of the MM-ORIENT components, equations, or the claimed three-dataset evaluation appear in the body of the submission.
Significance. If the MM-ORIENT method existed as described, it could be a relevant contribution to multimodal fusion and multitask learning, especially the idea of suppressing cross-modal noise by reconstructing monomodal features from neighborhoods determined by another modality and then applying per-modality attention before late fusion. The abstract articulates a testable thesis about noise reduction at the latent stage. However, the submitted manuscript does not contain the method, its formalisms, or the reported experiments. The actual full text is a survival-analysis methodology paper with its own scope, simulations, and real-data application; although that separate work may have independent merit, it provides no evidence for the abstract's claims about MM-ORIENT. The central claim of this submission is therefore unsupported in the provided document.
major comments (3)
- [Title page and Sections 2-4 of the full text] The submitted full text is not the paper described in the abstract. The title page identifies the work as "Tree-based methods for length-biased survival data" (arXiv:2508.16312v4 [stat.ME]), and the body develops survival trees and forests for length-biased right-censored data. The terms MM-ORIENT, cross-modal relation graph, and HIMA appear only in the abstract; no equation or algorithm in the full text defines them. Consequently, the abstract's assertion of "extensive experimental evaluation on three datasets" has no supporting evidence in this submission.
- [Abstract, second sentence and fourth sentence] The abstract claims the approach acquires multimodal representations "without explicit interaction between different modalities," yet it also states that features are reconstructed based on node neighborhoods "decided by the features of a different modality." Using one modality's features to define neighborhoods for reconstructing the other is itself a cross-modal interaction. Without a formal specification of the graph construction and reconstruction objective, the claimed distinction is not established. Since the noise-reduction property is the central motivation, this is a load-bearing gap.
- [Abstract, final sentence] No datasets, baselines, metrics, ablations, or error bars are reported anywhere in the manuscript for the claimed multimodal multitask framework. The full text describes an unrelated survival-analysis evaluation on a lung-cancer registry, not the three multimodal datasets promised in the abstract. The central empirical claim is therefore unfalsifiable in this submission.
minor comments (2)
- [Abstract, component name] The name "Hierarchical Interactive Monomadal Attention" appears to contain a typo; it should probably read "Monomodal Attention."
- [General] The arXiv identifier in the full text (2508.16312v4) differs from the submission identifier (2508.16300), and the titles, abstracts, and subject areas are entirely inconsistent. This should be resolved before any further review, since it prevents the assigned paper from being evaluated as a coherent submission.
Circularity Check
MM-ORIENT abstract is self-contradictory: its cross-modal relation graph is itself an explicit cross-modal interaction, and the cited three-dataset evaluation is absent from the supplied survival-tree manuscript.
-
self definitional
[Abstract (arXiv:2508.16300), MM-ORIENT paragraph]
"The proposed approach acquires multimodal representations cross-modally without explicit interaction between different modalities... we propose cross-modal relation graphs that reconstruct monomodal features... where the neighborhood is decided by the features of a different modality."
The defining mechanism is itself a cross-modal interaction: one modality's features determine the neighborhood used to reconstruct the other modality's features. Therefore the property 'without explicit interaction' is contradicted by the construction that is supposed to realize it. The claimed noise-reduction mechanism—avoiding explicit interaction—has no independent content because the proposed cross-modal relation graph is, by the paper's own account, an explicit cross-modal interaction. The framework name 'crOss-modal Relation' and 'hIErarchical iNteractive aTtention' reinforce that the interaction-free premise is definitionally excluded.
full rationale
The supplied full text is not the MM-ORIENT paper: it is arXiv:2508.16312v4 [stat.ME], 'Tree-based methods for length-biased survival data' by different authors. That body contains no MM-ORIENT equations, no three datasets, no baselines, and no experimental results, so the abstract's assertion 'extensive experimental evaluation on three datasets demonstrates...' is an unsupported claim within this submission. That is a missing-support/integrity problem rather than a circular derivation, and I have not counted it as a circular step itself. The one concrete circularity-adjacent defect is definitional: the abstract's central premise, that multimodal representations are acquired 'without explicit interaction between different modalities,' is negated by the proposed cross-modal relation graph, which uses one modality's features to decide the neighborhood for reconstructing the other. Because the supposed interaction-free property is contradicted by the method's own defining components, the central claimed advantage cannot stand as an independent derivation. No self-citation chain or fitted-input-called-prediction pattern is present, so the score reflects this single self-definitional contradiction plus the complete absence of the claimed supporting experiments.
Assumptions & free parameters
assumptions (3)
- domain assumption Noise in individual modalities degrades multimodal representations obtained through explicit cross-modal interactions.
- ad hoc to paper Reconstructing monomodal features using a graph whose node neighborhood is defined by the other modality reduces noise and preserves discriminative information.
- domain assumption Evaluation on three datasets is sufficient to demonstrate effectiveness across multiple tasks.
invented entities (2)
-
Cross-modal relation graphs
-
Hierarchical Interactive Monomodal Attention (HIMA)
Cite this review
Pith. "Pith review of A Multimodal-Multitask Framework with Cross-modal Relation and Hierarchical Interactive Attention for Semantic Comprehension." pith.science (2026). https://pith.science/paper/JOWFSWPD
@misc{pith2026250816300,
author = {Pith},
title = {Pith review of: A Multimodal-Multitask Framework with Cross-modal Relation and Hierarchical Interactive Attention for Semantic Comprehension},
year = {2026},
howpublished = {\url{https://pith.science/paper/JOWFSWPD}},
note = {Machine review of arXiv:2508.16300}
}
read the original abstract
A major challenge in multimodal learning is the presence of noise within individual modalities. This noise inherently affects the resulting multimodal representations, especially when these representations are obtained through explicit interactions between different modalities. Moreover, the multimodal fusion techniques while aiming to achieve a strong joint representation, can neglect valuable discriminative information within the individual modalities. To this end, we propose a Multimodal-Multitask framework with crOss-modal Relation and hIErarchical iNteractive aTtention (MM-ORIENT) that is effective for multiple tasks. The proposed approach acquires multimodal representations cross-modally without explicit interaction between different modalities, reducing the noise effect at the latent stage. To achieve this, we propose cross-modal relation graphs that reconstruct monomodal features to acquire multimodal representations. The features are reconstructed based on the node neighborhood, where the neighborhood is decided by the features of a different modality. We also propose Hierarchical Interactive Monomadal Attention (HIMA) to focus on pertinent information within a modality. While cross-modal relation graphs help comprehend high-order relationships between two modalities, HIMA helps in multitasking by learning discriminative features of individual modalities before late-fusing them. Finally, extensive experimental evaluation on three datasets demonstrates that the proposed approach effectively comprehends multimodal content for multiple tasks.
Reference graph
Works this paper leans on
-
[1]
Hierarchical interactive multimodal transformer for aspect-based multimodal sentiment analysis
Jianfei Yu, Kai Chen, and Rui Xia. Hierarchical interactive multimodal transformer for aspect-based multimodal sentiment analysis. IEEE Transactions on Affective Computing, 14 0 (3): 0 1966--1978, 2023. doi:10.1109/TAFFC.2022.3171091
-
[2]
N. Majumder, D. Hazarika, A. Gelbukh, E. Cambria, and S. Poria. Multimodal sentiment analysis using hierarchical fusion with context modeling. Knowledge-Based Systems, 161: 0 124--133, 2018. ISSN 0950-7051. doi:https://doi.org/10.1016/j.knosys.2018.07.041
-
[3]
Xian Sun, Fanglong Yao, and Chibiao Ding. Modeling high-order relationships: Brain-inspired hypergraph-induced multimodal-multitask framework for semantic comprehension. IEEE Transactions on Neural Networks and Learning Systems, pages 1--15, 2023. doi:10.1109/TNNLS.2023.3252359
-
[4]
Peng Xu, Xiatian Zhu, and David A. Clifton. Multimodal learning with transformers: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45 0 (10): 0 12113--12132, 2023. doi:10.1109/TPAMI.2023.3275156
arXiv 2023
-
[5]
Mohammad Zia Ur Rehman, Anukriti Bhatnagar, Omkar Kabde, Shubhi Bansal, and Dr. Nagendra Kumar. I mpli H ate V id: A benchmark dataset and two-stage contrastive learning framework for implicit hate speech detection in videos. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 17209--17...
-
[6]
Ankita Gandhi, Kinjal Adhvaryu, Soujanya Poria, Erik Cambria, and Amir Hussain. Multimodal sentiment analysis: A systematic review of history, datasets, multimodal fusion methods, applications, challenges and future directions. Information Fusion, 91: 0 424--444, 2023 a
work page 2023
-
[7]
Krishanu Maity, Prince Jha, Sriparna Saha, and Pushpak Bhattacharyya. A multitask framework for sentiment, emotion and sarcasm aware cyberbullying detection from multi-modal code-mixed memes. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, page 1739–1749, 2022. ISBN 9781450387323. doi:10....
arXiv 2022
-
[8]
Anusha Chhabra and Dinesh Kumar Vishwakarma. Multimodal hate speech detection via multi-scale visual kernels and knowledge distillation architecture. Engineering Applications of Artificial Intelligence, 126: 0 106991, 2023. ISSN 0952-1976. doi:https://doi.org/10.1016/j.engappai.2023.106991
arXiv 2023
Show all 68 references
-
[9]
Aditya Joshi, Pushpak Bhattacharyya, and Mark J. Carman. Automatic sarcasm detection: A survey. ACM Comput. Surv., 50 0 (5): 0 1–22, 2017. doi:10.1145/3124420
2017 doi
-
[10]
o nig, Eva-Maria Me ner, Alan Cowen, Erik Cambria, and Bj\
Shahin Amiriparian, Lukas Christ, Andreas K\" o nig, Eva-Maria Me ner, Alan Cowen, Erik Cambria, and Bj\" o rn W. Schuller. Muse 2022 challenge: Multimodal humour, emotional reactions, and stress. In Proceedings of the 30th ACM International Conference on Multimedia, page 7389...
2022
-
[11]
Shamim Hossain
Yazhou Zhang, Prayag Tiwari, Qian Zheng, Abdulmotaleb El Saddik, and M. Shamim Hossain. A multimodal coupled graph attention network for joint traffic event detection and sentiment classification. IEEE Transactions on Intelligent Transportation Systems, 24 0 (8): 0 8542--8554,...
2023
-
[12]
Mahaemosen: Towards emotion-aware multimodal marathi sentiment analysis
Prasad Chaudhari, Pankaj Nandeshwar, Shubhi Bansal, and Nagendra Kumar. Mahaemosen: Towards emotion-aware multimodal marathi sentiment analysis. ACM Trans. Asian Low-Resour. Lang. Inf. Process., 22 0 (9): 0 24, 1-24 2023. ISSN 2375-4699. doi:10.1145/3618057
2023 doi
-
[13]
Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov
Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J. Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov. Multimodal transformer for unaligned multimodal language sequences. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, p...
2019 doi
-
[14]
A hybrid deep neural network for multimodal personalized hashtag recommendation
Shubhi Bansal, Kushaan Gowda, and Nagendra Kumar. A hybrid deep neural network for multimodal personalized hashtag recommendation. IEEE Transactions on Computational Social Systems, 10 0 (5): 0 2439--2459, 2023. doi:10.1109/TCSS.2022.3184307
2023
-
[15]
All-but-the-top: Simple and effective post-processing for word representations
Jiaqi Mu and Pramod Viswanath. All-but-the-top: Simple and effective post-processing for word representations. In 6th International Conference on Learning Representations, ICLR 2018, 2018
2018
-
[16]
Multimodal information bottleneck: Learning minimal sufficient unimodal and multimodal representations
Sijie Mai, Ying Zeng, and Haifeng Hu. Multimodal information bottleneck: Learning minimal sufficient unimodal and multimodal representations. IEEE Transactions on Multimedia, 25: 0 4121--4134, 2022. doi:10.1109/TMM.2022.3171679
2022
-
[17]
Divide, conquer and combine: Hierarchical feature fusion network with local and global perspectives for multimodal affective computing
Sijie Mai, Haifeng Hu, and Songlong Xing. Divide, conquer and combine: Hierarchical feature fusion network with local and global perspectives for multimodal affective computing. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4...
2019 doi
-
[18]
User-aware multilingual abusive content detection in social media
Mohammad Zia Ur Rehman , Somya Mehta, Kuldeep Singh, Kunal Kaushik, and Nagendra Kumar. User-aware multilingual abusive content detection in social media. Information Processing and Management, 60 0 (5): 0 103450, 2023. ISSN 0306-4573. doi:https://doi.org/10.1016/j.ipm.2023.103450
2023
-
[19]
Hierarchical attention networks for document classification
Zichao Yang, Diyi Yang, Chris Dyer, Xiaodong He, Alex Smola, and Eduard Hovy. Hierarchical attention networks for document classification. In Proceedings of the 2016 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Tech...
2016 doi
-
[20]
Bi-stream graph learning based multimodal fusion for emotion recognition in conversation
Nannan Lu, Zhiyuan Han, Min Han, and Jiansheng Qian. Bi-stream graph learning based multimodal fusion for emotion recognition in conversation. Information Fusion, 106: 0 102272, 2024
2024
-
[21]
Integrating gin-based multimodal feature transformation and multi-feature combination voting for irony-aware cyberbullying detection
Tingting Li, Ziming Zeng, Qingqing Li, and Shouqiang Sun. Integrating gin-based multimodal feature transformation and multi-feature combination voting for irony-aware cyberbullying detection. Information Processing & Management, 61 0 (3): 0 103651, 2024
2024
-
[22]
Gcnet: Graph completion network for incomplete multimodal learning in conversation
Zheng Lian, Lan Chen, Licai Sun, Bin Liu, and Jianhua Tao. Gcnet: Graph completion network for incomplete multimodal learning in conversation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45 0 (7): 0 8419--8432, 2023. doi:10.1109/TPAMI.2023.3234553
2023
-
[23]
Graphcfc: A directed graph based cross-modal feature complementation approach for multimodal conversational emotion recognition
Jiang Li, Xiaoping Wang, Guoqing Lv, and Zhigang Zeng. Graphcfc: A directed graph based cross-modal feature complementation approach for multimodal conversational emotion recognition. IEEE Transactions on Multimedia, pages 1--13, 2023. doi:10.1109/TMM.2023.3260635
2023
-
[24]
Graph neural network meets sparse representation: Graph sparse neural networks via exclusive group lasso
Bo Jiang, Beibei Wang, Si Chen, Jin Tang, and Bin Luo. Graph neural network meets sparse representation: Graph sparse neural networks via exclusive group lasso. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45 0 (10): 0 12692--12698, 2023. doi:10.1109/TPAMI.2...
2023
-
[25]
Heterogeneous graph contrastive learning network for personalized micro-video recommendation
Desheng Cai, Shengsheng Qian, Quan Fang, Jun Hu, Wenkui Ding, and Changsheng Xu. Heterogeneous graph contrastive learning network for personalized micro-video recommendation. IEEE Transactions on Multimedia, 25: 0 2761--2773, 2023. doi:10.1109/TMM.2022.3151026
2023
-
[26]
Hgber: Heterogeneous graph neural network with bidirectional encoding representation
Yanbei Liu, Lianxi Fan, Xiao Wang, Zhitao Xiao, Shuai Ma, Yanwei Pang, and Jerry Chun-Wei Lin. Hgber: Heterogeneous graph neural network with bidirectional encoding representation. IEEE Transactions on Neural Networks and Learning Systems, pages 1--12, 2023. doi:10.1109/TNNLS....
2023
-
[27]
Align before attend: Aligning visual and textual features for multimodal hateful content detection
Eftekhar Hossain, Omar Sharif, Mohammed Moshiul Hoque, and Sarah Masud Preum. Align before attend: Aligning visual and textual features for multimodal hateful content detection. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational ...
2024
-
[28]
Hatefusion: Harnessing attention-based techniques for enhanced filtering and detection of implicit hate speech
Ashok Yadav and Vrijendra Singh. Hatefusion: Harnessing attention-based techniques for enhanced filtering and detection of implicit hate speech. IEEE Transactions on Computational Social Systems, 2024
2024
-
[29]
Mimicking the brain’s cognition of sarcasm from multidisciplines for twitter sarcasm detection
Fanglong Yao, Xian Sun, Hongfeng Yu, Wenkai Zhang, Wei Liang, and Kun Fu. Mimicking the brain’s cognition of sarcasm from multidisciplines for twitter sarcasm detection. IEEE Transactions on Neural Networks and Learning Systems, 34 0 (1): 0 228--242, 2023. doi:10.1109/TNNLS.20...
2023
-
[30]
A multitask learning model for multimodal sarcasm, sentiment and emotion recognition in conversations
Yazhou Zhang, Jinglin Wang, Yaochen Liu, Lu Rong, Qian Zheng, Dawei Song, Prayag Tiwari, and Jing Qin. A multitask learning model for multimodal sarcasm, sentiment and emotion recognition in conversations. Information Fusion, 93: 0 282--301, 2023 b . ISSN 1566-2535. doi:https:...
2023 doi
-
[31]
COMMA - DEER : CO mmon-sense aware multimodal multitask approach for detection of emotion and emotional reasoning in conversations
Soumitra Ghosh, Gopendra Vikram Singh, Asif Ekbal, and Pushpak Bhattacharyya. COMMA - DEER : CO mmon-sense aware multimodal multitask approach for detection of emotion and emotional reasoning in conversations. In Proceedings of the 29th International Conference on Computationa...
2022
-
[32]
Progressive modality reinforcement for human multimodal emotion recognition from unaligned multimodal sequences
Fengmao Lv, Xiang Chen, Yanyong Huang, Lixin Duan, and Guosheng Lin. Progressive modality reinforcement for human multimodal emotion recognition from unaligned multimodal sequences. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pa...
2021
-
[33]
Shah, Surendrabikram Thapa, Usman Naseem, and Mehwish Nasim
Aashish Bhandari, Siddhant B. Shah, Surendrabikram Thapa, Usman Naseem, and Mehwish Nasim. Crisishatemm: Multimodal analysis of directed and undirected hate speech in text-embedded images from russia-ukraine conflict. In Proceedings of the IEEE/CVF Conference on Computer Visio...
1993
-
[34]
Universal multimodal representation for language understanding
Zhuosheng Zhang, Kehai Chen, Rui Wang, Masao Utiyama, Eiichiro Sumita, Zuchao Li, and Hai Zhao. Universal multimodal representation for language understanding. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45 0 (7): 0 9169--9185, 2023 c . doi:10.1109/TPAMI.20...
2023
-
[35]
EDA : Easy data augmentation techniques for boosting performance on text classification tasks
Jason Wei and Kai Zou. EDA : Easy data augmentation techniques for boosting performance on text classification tasks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Proces...
2019 doi
-
[36]
Robust training under linguistic adversity
Yitong Li, Trevor Cohn, and Timothy Baldwin. Robust training under linguistic adversity. In Proceedings of the 15th Conference of the E uropean Chapter of the Association for Computational Linguistics: Volume 2, Short Papers , pages 21--27, 2017
2017
-
[37]
T iny BERT : Distilling BERT for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. T iny BERT : Distilling BERT for natural language understanding. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4163--4174, 2020. doi:10.18653/v1/20...
2020 doi
-
[38]
Transformer-based multimodal emotional perception for dynamic facial expression recognition in the wild
Xiaoqin Zhang, Min Li, Sheng Lin, Hang Xu, and Guobao Xiao. Transformer-based multimodal emotional perception for dynamic facial expression recognition in the wild. IEEE Transactions on Circuits and Systems for Video Technology, pages 1--1, 2023 d . doi:10.1109/TCSVT.2023.3312858
2023
-
[39]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770--778, 2016
2016
-
[40]
Random erasing data augmentation
Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang. Random erasing data augmentation. Proceedings of the AAAI Conference on Artificial Intelligence, 34 0 (07): 0 13001--13008, 2020. doi:10.1609/aaai.v34i07.7000
2020 doi
-
[41]
Free-form image inpainting with gated convolution
Jiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen, Xin Lu, and Thomas S Huang. Free-form image inpainting with gated convolution. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4471--4480, 2019
2019
-
[42]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pa...
2021
-
[43]
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Langu...
2019
-
[44]
Detectron2
Yuxin Wu, Alexander Kirillov, Francisco Massa, Wan-Yen Lo, and Ross Girshick. Detectron2. https://github.com/facebookresearch/detectron2, 2019
2019
-
[45]
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Doll \'a r, and Ross Girshick. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 2961--2969, 2017
2017
-
[46]
The stanford corenlp natural language processing toolkit
Christopher D Manning, Mihai Surdeanu, John Bauer, Jenny Rose Finkel, Steven Bethard, and David McClosky. The stanford corenlp natural language processing toolkit. In Proceedings of 52nd annual meeting of the association for computational linguistics: system demonstrations, pa...
2014
-
[47]
Hierarchical attention-enhanced contextual capsulenet for multilingual hope speech detection
Mohammad Zia Ur Rehman, Devraj Raghuvanshi, Harshit Pachar, Chandravardhan Singh Raghaw, and Nagendra Kumar. Hierarchical attention-enhanced contextual capsulenet for multilingual hope speech detection. Expert Systems with Applications, 268: 0 126285, 2025 b . doi:https://doi....
2025
-
[48]
A context-aware attention and graph neural network-based multimodal framework for misogyny detection
Mohammad Zia Ur Rehman, Sufyaan Zahoor, Areeb Manzoor, Musharaf Maqbool, and Nagendra Kumar. A context-aware attention and graph neural network-based multimodal framework for misogyny detection. Information Processing & Management, 62 0 (1): 0 103895, 2025 c . doi:https://doi....
2025
-
[49]
Inductive representation learning on large graphs
Will Hamilton, Zhitao Ying, and Jure Leskovec. Inductive representation learning on large graphs. Advances in neural information processing systems, 30, 2017
2017
-
[50]
o rn Gamb \
Chhavi Sharma, Deepesh Bhageria, William Scott, Srinivas Pykl, Amitava Das, Tanmoy Chakraborty, Viswanath Pulabaigari, and Bj \"o rn Gamb \"a ck. Semeval-2020 task 8: Memotion analysis-the visuo-lingual metaphor! In Proceedings of the Fourteenth Workshop on Semantic Evaluation...
2020
-
[51]
Exploring hate speech detection in multimodal publications
Raul Gomez, Jaume Gibert, Lluis Gomez, and Dimosthenis Karatzas. Exploring hate speech detection in multimodal publications. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2020
2020
-
[52]
Detecting harmful memes and their targets
Shraman Pramanick, Dimitar Dimitrov, Rituparna Mukherjee, Shivam Sharma, Md Shad Akhtar, Preslav Nakov, and Tanmoy Chakraborty. Detecting harmful memes and their targets. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 2783--2796, 2021 a
2021
-
[53]
Multimodal sentiment analysis: A systematic review of history, datasets, multimodal fusion methods, applications, challenges and future directions
Ankita Gandhi, Kinjal Adhvaryu, Soujanya Poria, Erik Cambria, and Amir Hussain. Multimodal sentiment analysis: A systematic review of history, datasets, multimodal fusion methods, applications, challenges and future directions. Information Fusion, 91: 0 424--444, 2023 b
2023
-
[54]
Multi-modality cross attention network for image and sentence matching
Xi Wei, Tianzhu Zhang, Yan Li, Yongdong Zhang, and Feng Wu. Multi-modality cross attention network for image and sentence matching. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10941--10950, 2020
2020
-
[55]
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017
2017
-
[56]
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. Visualbert: A simple and performant baseline for vision and language. arXiv preprint arXiv:1908.03557, 2019
1908 arXiv
-
[57]
Vilt: Vision-and-language transformer without convolution or region supervision
Wonjae Kim, Bokyung Son, and Ildoo Kim. Vilt: Vision-and-language transformer without convolution or region supervision. In International Conference on Machine Learning, pages 5583--5594. PMLR, 2021
2021
-
[58]
Disentangling hate in online memes
Roy Ka-Wei Lee, Rui Cao, Ziqing Fan, Jing Jiang, and Wen-Haw Chong. Disentangling hate in online memes. In Proceedings of the 29th ACM International Conference on Multimedia, MM '21, page 5138–5147, 2021. doi:10.1145/3474085.3475625
2021
-
[59]
Momenta: A multimodal framework for detecting harmful memes and their targets
Shraman Pramanick, Shivam Sharma, Dimitar Dimitrov, Md Shad Akhtar, Preslav Nakov, and Tanmoy Chakraborty. Momenta: A multimodal framework for detecting harmful memes and their targets. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 4439--4455, 2021 b
2021
-
[60]
Beneath the surface: Unveiling harmful memes with multimodal reasoning distilled from large language models
Hongzhan Lin, Ziyang Luo, Jing Ma, and Long Chen. Beneath the surface: Unveiling harmful memes with multimodal reasoning distilled from large language models. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 9114--9128, 2023
2023
-
[61]
Multi-interactive memory network for aspect based multimodal sentiment analysis
Nan Xu, Wenji Mao, and Guandan Chen. Multi-interactive memory network for aspect based multimodal sentiment analysis. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 371--378, 2019
2019
-
[62]
Modeling intra and inter-modality incongruity for multi-modal sarcasm detection
Hongliang Pan, Zheng Lin, Peng Fu, Yatao Qi, and Weiping Wang. Modeling intra and inter-modality incongruity for multi-modal sarcasm detection. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1383--1392, 2020. doi:10.18653/v1/2020.findings-emnlp.124
2020 doi
-
[63]
Ensemble pretrained models for multimodal sentiment analysis using textual and video data fusion
Zhicheng Liu, Ali Braytee, Ali Anaissi, Guifu Zhang, Lingyun Qin, and Junaid Akram. Ensemble pretrained models for multimodal sentiment analysis using textual and video data fusion. In Companion Proceedings of the ACM Web Conference 2024, pages 1841--1848, 2024
2024
-
[64]
O zlem \
Eniafe Festus Ayetiran and \"O zlem \"O zg \"o bek. An inter-modal attention-based deep learning framework using unified modality for multimodal fake news, hate speech and offensive language detection. Information Systems, 123: 0 102378, 2024
2024
-
[65]
Mmffhs: Multi-modal feature fusion for hate speech detection on social media
Pradeep Kumar Roy. Mmffhs: Multi-modal feature fusion for hate speech detection on social media. IEEE Transactions on Big Data, 11 0 (03): 0 1247--1258, 2025
2025
-
[66]
Emotion-aware multimodal fusion for meme emotion detection
Shivam Sharma, S Ramaneswaran, Md Shad Akhtar, and Tanmoy Chakraborty. Emotion-aware multimodal fusion for meme emotion detection. IEEE Transactions on Affective Computing, 15 0 (3): 0 1800--1811, 2024
2024
-
[67]
C hat G P T 4o
OpenAI. C hat G P T 4o. https://platform.openai.com/docs/models/gpt-4o. [Accessed 17-05-2024]
2024
-
[68]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.