Pith. sign in

REVIEW 2 major objections 5 minor 156 references

Multimodal fusion lifts document-classification accuracy by about five points; multiview gains are smaller but steadier, and success hinges on matching method to task, not on complexity.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-13 09:19 UTC pith:VZHEQZXL

load-bearing objection First usable meta-analysis of fusion gains in document classification, plus a clean formal map and a blunt diagnosis of how poorly the primary literature validates itself. the 2 major comments →

arxiv 2605.23910 v2 pith:VZHEQZXL submitted 2026-04-07 cs.CL cs.AI

Document Classification Pattern Recognition via Information Fusion: A Systematic Review of Multimodal and Multiview Representation Approaches

classification cs.CL cs.AI
keywords information fusiondocument classificationmultimodal learningmultiview learningsystematic reviewmeta-analysisrepresentation learning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This review of 139 studies asks whether combining multiple data sources (multimodal) or multiple representations of the same text (multiview) actually improves document classification, and by how much. It supplies a formal framework that treats every pipeline stage as a representation, a pattern, and a recognition rule, then uses that framework to organise the literature and to run the first random-effects meta-analysis focused on this task. Multimodal fusion yields a mean accuracy gain of +5.28 percentage points; multiview fusion yields consistent but smaller gains on accuracy, F1 and recall. The review also shows that only a small fraction of papers use statistical tests or share code, so many reported wins rest on weak evidence. The practical takeaway is that fusion works when the chosen method is aligned with the task and the views are complementary, not when the algorithm is simply more elaborate.

Core claim

A random-effects meta-analysis of eligible primary studies shows that multimodal fusion improves accuracy by a mean of +5.28 percentage points (p=0.0016), while multiview fusion produces consistent modest gains for accuracy (+4.67 %), F1-score (+3.08 %) and recall (all p<0.05). Successful fusion depends on strategic alignment of method with task context rather than on algorithmic complexity, and most published claims lack the statistical tests needed to support them.

What carries the argument

The representation–pattern–model triple R=(F,E,M), P=(S,C,T), M=(P,R,RR): a uniform notation that maps every document-classification pipeline onto classical information-fusion operators and makes early, late and hybrid fusion comparable.

Load-bearing premise

The studies that both report absolute scores for a fusion model and a comparable non-fusion baseline on the same split, and for which missing standard deviations can be filled by the pooled average, form an unbiased sample of the true effect of fusion.

What would settle it

A new set of head-to-head experiments on the same datasets and splits that report full means, standard deviations and statistical tests, and that show either zero or negative accuracy gains for the fusion models that previously claimed large positive effects.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This systematic review analyses 139 primary studies on multimodal and multiview information fusion for document classification. It introduces a formal framework (representations R=(F,E,M), patterns P=(S,C,T), models M=(P,R,RR)) that maps modern architectures onto classical fusion concepts, provides a data-driven qualitative taxonomy of trends/challenges, and conducts the first random-effects meta-analysis focused on this task. The meta-analysis reports a significant multimodal accuracy gain of +5.28 percentage points (p=0.0016; F1 directionally positive but non-significant in the primary model) and consistent modest multiview gains for accuracy (+4.67%), F1 (+3.08%) and recall (all p<0.05). Qualitative synthesis highlights severe methodological gaps (only 11.8% multimodal / 23.3% multiview studies use statistical tests; low reproducibility). The authors conclude that fusion success depends on strategic alignment of method with task context rather than algorithmic complexity, and supply practitioner guidelines plus open data/code.

Significance. If the results hold, the paper supplies the first quantitative evidence base for fusion efficacy in document classification, a reusable formal bridge between ad-hoc ML architectures and classical information-fusion operators (Bayesian pools, Dempster–Shafer, etc.), and actionable guidelines. Strengths that raise confidence include the PRISMA-adapted protocol, two-stage analysis (nonparametric robustness checks + inverse-variance random-effects meta-analysis with heterogeneity statistics, funnel plots and Egger tests), explicit reporting of high I² and model-dependent F1 fragility, and full Zenodo release of data, R scripts and technical reports. These elements make the work a solid benchmark and roadmap for a field previously lacking synthesis.

major comments (2)
  1. §7.1 eligibility criteria and SD-imputation procedure: Restricting the meta-analysis to studies that report absolute scores for both a fusion model and a same-split non-fusion baseline, then imputing missing SDs from the pooled average, is a reasonable practical choice but remains the load-bearing assumption for the numerical claims. While the authors already run permutation/Wilcoxon checks, report high I² for accuracy, and note funnel-plot results, a sensitivity analysis that excludes all imputed-SD studies (or uses a range of imputation priors) should be added to confirm that the +5.28 pp multimodal accuracy result and the multiview consistency findings are not artefacts of the better-reported subset.
  2. §5.2 and §8.3: The claim that successful fusion depends on strategic alignment rather than complexity is well-supported by the qualitative high-impact case studies, yet the mapping of concrete architectures (attention, ensembles, contrastive heads) onto classical operators (linear/log opinion pools, mass functions, Kalman-style sequential updates) remains largely illustrative. Strengthening this mapping with 2–3 fully worked examples that show how a published model’s RR component realises a specific fusion operator would make the theoretical contribution more operational and directly testable.
minor comments (5)
  1. Figure 4 and Figure 5: Axis labels and legend fonts are small; increasing size and adding explicit n= values next to each pie slice would improve readability.
  2. Table 12 and Table 14: The column headers mix Est.∆, d_Cohen, d_Cliff and r without a short legend explaining the exact effect-size definitions used; a one-line footnote would help non-meta-analysis readers.
  3. §3.2 / Table 2: The exact Scopus query string is relegated to the Supplement; placing a shortened version (or a DOI-linked permanent link) in the main text would aid immediate reproducibility.
  4. Throughout: Occasional minor inconsistencies in hyphenation (multi-view vs multiview, multi-modal vs multimodal) and a few long sentences in §6.3 that could be split for clarity.
  5. Appendix B: The generative-AI disclosure is appropriately placed, but a brief statement that all statistical results and code were human-verified would further reassure readers.

Circularity Check

0 steps flagged

No circularity: meta-analytic effect sizes are aggregated from independent primary studies under explicit eligibility rules; the formal R/P/M framework is definitional/organizational rather than a derived prediction that collapses to its inputs.

full rationale

The paper is a PRISMA-style systematic review plus random-effects meta-analysis of 139 primary studies. Its central quantitative claims (multimodal accuracy +5.28 pp, multiview accuracy/F1/recall gains) are computed from absolute performance differences reported by those independent studies that meet the Section 7.1 criteria (fusion vs. non-fusion baseline on identical splits, with SD imputation only when necessary). No parameter is fitted to a subset and then re-presented as a prediction of a related quantity; no equation reduces a claimed first-principles result to a definitional identity. The R=(F,E,M), P=(S,C,T), M=(P,R,RR) framework is introduced as a unifying notation for organising existing work, not as a theorem derived from the meta-data it later summarises. The single self-citation of the author’s prior multiview paper appears merely as one of the 139 primary studies and is not invoked as a uniqueness theorem or load-bearing premise for the effect-size estimates or the strategic-alignment conclusion. The derivation chain is therefore self-contained against external benchmarks and exhibits none of the six circularity patterns.

Axiom & Free-Parameter Ledger

2 free parameters · 3 axioms · 2 invented entities

As a systematic review the paper rests mainly on standard statistical assumptions of random-effects meta-analysis and on the operational definitions of its formal framework. Free parameters are the usual meta-analytic quantities (imputed SDs, study weights). Invented entities are the three-component formal objects introduced to organise the literature; they are definitional and carry no independent empirical claim.

free parameters (2)
  • imputed study-level SDs
    When primary studies omitted variance, SDs were replaced by the pooled average of studies that reported them; the central effect-size estimates depend on this imputation.
  • random-effects τ²
    Estimated from the observed study-level differences; high values (I²≈80–90 % for accuracy) drive the width of the confidence intervals.
axioms (3)
  • domain assumption Random-effects model is appropriate for synthesising heterogeneous machine-learning experiments
    Invoked throughout Section 7; standard in meta-analysis but not proven for the particular distribution of document-classification results.
  • ad hoc to paper A study is eligible for quantitative synthesis only if it reports absolute scores for both a fusion model and a non-fusion baseline on identical data splits
    Section 7.1 eligibility criteria; determines which of the 139 papers enter the meta-analysis.
  • ad hoc to paper Representation R = (F, E, M) and pattern P = (S, C, T) capture the essential structure of document-classification pipelines
    Sections 4–5; definitional scaffolding used to organise the qualitative taxonomy.
invented entities (2)
  • R = (F, E, M) representation triple no independent evidence
    purpose: Uniform notation for format, encoding and meaning of any document view or modality
    Introduced in Section 4.1; purely definitional, no external falsifiable prediction.
  • P = (S, C, T) pattern triple no independent evidence
    purpose: Uniform notation for structure, constraints and transformations of classification patterns
    Introduced in Section 4.2; definitional scaffolding.

pith-pipeline@v1.1.0-grok45 · 50136 in / 2588 out tokens · 26546 ms · 2026-07-13T09:19:41.752037+00:00 · methodology

0 comments
read the original abstract

Information fusion is used widely to improve document classification by the integration of multiple data sources (multimodal) or representations (multiview). However, the field lacks a unified framework, a quantitative synthesis of its effectiveness, and clear guidance for practitioners. This systematic review addresses these gaps by analysing 139 primary studies. It introduces a formal framework to structure the field, presents the results of a qualitative analysis to identify key trends, and performs a random-effects meta-analysis (to our knowledge, the first focused on document classification) to quantify performance gains. Our meta-analysis reveals that multimodal fusion improves accuracy (mean gain of +5.28 percentage points, $p=0.0016$) significantly -- the F1-score effect is directionally positive but statistically non-significant in our primary model. Multiview fusion provides consistent but modest gains for accuracy (+4.67\%), F1-score (+3.08\%), and recall (all $p<0.05$). Critically, our qualitative synthesis uncovers challenges in reproducibility in methodological rigour: only 11.8\% (multimodal) and 23.3\% (multiview) of the studies use statistical tests to validate their findings, which undermines the reliability of many of their results. This review's primary contributions are a unifying framework, the first quantitative evidence base, and data-driven guidelines. This review concludes that successful information fusion depends not on algorithmic complexity, but on the strategic alignment of the fusion method with the task context and a commitment to more rigorous validation.

Figures

Figures reproduced from arXiv: 2605.23910 by Marcin Micha{\l} Miro\'nczuk.

Figure 1
Figure 1. Figure 1: The evolution of information fusion research through the prism of review articles. [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: A distribution of the studies on document classification reviews. [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: A graphical representation of the review method for the studies. [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: The left plot presents the number of publications per year from 2001 to 2024; the right plot presents the distribution [PITH_FULL_IMAGE:figures/full_fig_p011_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: The left plot presents the number of publications on multimodal, multiview, and both learning approach; the right [PITH_FULL_IMAGE:figures/full_fig_p012_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: The progression from raw data, through features and representations, to pattern recognition in document classification [PITH_FULL_IMAGE:figures/full_fig_p014_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: A multimodal learning framework that illustrates the integration of different data modalities (text, image, audio) [PITH_FULL_IMAGE:figures/full_fig_p022_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: A multiview learning framework that illustrates how different representations of the same data type are processed and [PITH_FULL_IMAGE:figures/full_fig_p023_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Fusion strategies connected to the R = (F, E, M) framework: (a) early fusion combines encodings at the E level; (b) late fusion combines classifier outputs at the RR level; (c) hybrid fusion integrates features at both E and RR levels. 23 [PITH_FULL_IMAGE:figures/full_fig_p023_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: An overview of evaluation practices in multimodal experiments: The top-left plot presented the frequency distribution [PITH_FULL_IMAGE:figures/full_fig_p038_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: An overview of the validation practices in multimodal experiments. The plots depict use of: a held-out test set or [PITH_FULL_IMAGE:figures/full_fig_p038_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: The distribution of performance gains for multimodal over unimodal approaches. The box plots visualise the [PITH_FULL_IMAGE:figures/full_fig_p039_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: An overview of the evaluation practices in multiview experiments: The top-left plot depicts the frequency distribution [PITH_FULL_IMAGE:figures/full_fig_p041_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: An overview of the validation practices in multiview experiments. The plots depict use of: a held-out test set or [PITH_FULL_IMAGE:figures/full_fig_p042_14.png] view at source ↗
Figure 15
Figure 15. Figure 15: The distribution of performance gains for multiview over singleview approaches. The box plots visualise the [PITH_FULL_IMAGE:figures/full_fig_p043_15.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

156 extracted references · 124 canonical work pages · 1 internal anchor

  1. [1]

    A. A. Goshtasby, Three-dimensional model construction from multiview range images: Survey with new results, Pattern Recognition 31 (1998) 1705–1714.doi:10.1016/S0031-3203(98)00047-8. URLhttps://linkinghub.elsevier.com/retrieve/pii/S0031320398000478

  2. [2]

    J. Du, W. Li, K. Lu, B. Xiao, An overview of multi-modal medical image fusion, Neurocomputing 215 (2016) 3–20.doi:10.1016/j.neucom.2015.07.160

  3. [3]

    X. Zhao, Q. Luo, B. Han, Survey on robot multi-sensor information fusion technology, in: 2008 7th World Congress on Intelligent Control and Automation, IEEE, 2008, pp. 5019–5023.doi:10.1109/wcica.2008.45937 42

  4. [4]

    H. Qian, M. Wang, M. Zhu, H. Wang, A review of multi-sensor fusion in autonomous driving, Sensors 25 (19) (2025) 6033.doi:10.3390/s25196033

  5. [5]

    Li, F.-X

    Y. Li, F.-X. Wu, A. Ngom, A review on machine learning principles for multi-view biological data integration, Briefings in Bioinformatics 19 (2018) 325–340.doi:10.1093/bib/bbw113

  6. [6]

    Y. Yang, H. Wang, Multi-view clustering: A survey, Big Data Mining and Analytics 1 (2018) 83–107.doi: 10.26599/BDMA.2018.9020003

  7. [7]

    Ho, K.-J

    Y.-S. Ho, K.-J. Oh, Overview of multi-view video coding, in: 2007 IWSSIP and EC-SIPMCS - Proc. 2007 14th Int. Workshop on Systems, Signals and Image Processing, and 6th EURASIP Conf. Focused on Speech and Image Processing, Multimedia Communications and Services, 2007, pp. 5–12.doi:10.1109/IWSSIP.2007.4381085

  8. [8]

    Y. Li, L. Pan, Y. Peng, X. Li, X. Wang, L. Qu, Q. Song, Q. Liang, S. Peng, Application of deep learning-based multimodal fusion technology in cancer diagnosis: A survey, Engineering Applications of Artificial Intelligence 143 (2025).doi:10.1016/j.engappai.2024.109972

  9. [9]

    Krones, U

    F. Krones, U. Marikkar, G. Parsons, A. Szmul, A. Mahdi, Review of multimodal machine learning approaches in healthcare, Information Fusion 114 (2025).doi:10.1016/j.inffus.2024.102690

  10. [10]

    Soleymani, D

    M. Soleymani, D. Garcia, B. Jou, B. Schuller, S.-F. Chang, M. Pantic, A survey of multimodal sentiment analysis, Image and Vision Computing 65 (2017) 3–14.doi:10.1016/j.imavis.2017.08.003

  11. [11]

    Chandrasekaran, T

    G. Chandrasekaran, T. Nguyen, J. H. D., Multimodal sentimental analysis for social media applications: A comprehensive review, Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery 11 (2021). doi:10.1002/widm.1415

  12. [12]

    Zhang, Z

    J. Zhang, Z. Yin, P. Chen, S. Nichele, Emotion recognition using multi-modal data and machine learning techniques: A tutorial and review, Information Fusion 59 (2020) 103–126.doi:10.1016/j.inffus.2020.01.011

  13. [13]

    R. Kaur, S. Kautish, Multimodal sentiment analysis: A survey and comparison, International Journal of Service Science, Management, Engineering, and Technology 10 (2019) 38–58.doi:10.4018/IJSSMET.2019040103

  14. [14]

    S. Wang, A. Shibghatullah, T. Iqbal, K. Keoy, A review of multimodal-based emotion recognition techniques for cyberbullying detection in online social media platforms, Neural Computing and Applications 36 (2024) 21923–21956.doi:10.1007/s00521-024-10371-3

  15. [15]

    Y. Liu, B. Pang, X. Wang, Opinion spam detection by incorporating multimodal embedded representation into a probabilistic review graph, Neurocomputing 366 (2019) 276–283.doi:10.1016/j.neucom.2019.08.013

  16. [16]

    Koromilas, T

    P. Koromilas, T. Giannakopoulos, Deep multimodal emotion recognition on human speech: A review, Applied Sciences (Switzerland) 11 (2021).doi:10.3390/app11177962

  17. [17]

    Jiang, S

    F. Jiang, S. Yang, M. Jones, L. Zhang, From attributes to natural language: A survey and foresight on text-based person re-identification, Information Fusion 118 (2025).doi:10.1016/j.inffus.2024.102879

  18. [18]

    Balazs, J

    J. Balazs, J. Velásquez, Opinion mining and information fusion: A survey, Information Fusion 27 (2016) 95–110. doi:10.1016/j.inffus.2015.06.002

  19. [19]

    Y. Zhu, H. Liu, Y. Du, Z. Wu, Ifspard: An information fusion-based framework for spam review detection, in: The Web Conference 2021 - Proceedings of the World Wide Web Conference, WWW 2021, 2021, pp. 507–517. doi:10.1145/3442381.3449920. 50

  20. [20]

    Y. Li, M. Yang, Z. Zhang, A survey of multi-view representation learning, IEEE Transactions on Knowledge and Data Engineering 31 (2019) 1863–1883.doi:10.1109/TKDE.2018.2872063

  21. [21]

    J. Tang, Q. Yi, S. Fu, Y. Tian, Incomplete multi-view learning: Review, analysis, and prospects, Applied Soft Computing 153 (2024).doi:10.1016/j.asoc.2024.111278

  22. [22]

    W. Wei, J. Liang, Information fusion in rough set theory : An overview, Information Fusion 48 (2019) 107–118. doi:10.1016/j.inffus.2018.08.007

  23. [23]

    X. Mi, H. Liao, X. Wu, Z. Xu, Probabilistic linguistic information fusion: A survey on aggregation operators in terms of principles, definitions, classifications, applications, and challenges, International Journal of Intelligent Systems 35 (2020) 529–556.doi:10.1002/int.22216

  24. [24]

    S. Song, X. Li, S. Li, S. Zhao, J. Yu, J. Ma, X. Mao, W. Zhang, M. Wang, How to bridge the gap between modalities: Survey on multimodal large language model, IEEE Transactions on Knowledge and Data Engineering (2025).doi:10.1109/TKDE.2025.3527978

  25. [25]

    Y. Qian, Y. Wang, J. Liu, Q. Zou, Y. Ding, X. Guo, W. Ding, A survey on multi-view fusion for predicting links in biomedical bipartite networks: Methods and applications, Information Fusion 117 (2025). doi: 10.1016/j.inffus.2024.102894

  26. [26]

    H. Xiao, F. Zhou, X. Liu, T. Liu, Z. Li, X. Liu, X. Huang, A comprehensive survey of large language models and multimodal large language models in medicine, Information Fusion 117 (2025).doi:10.1016/j.inffus.2 024.102888

  27. [27]

    Akpatsa, X

    S. Akpatsa, X. Li, H. Lei, A Survey and Future Perspectives of Hybrid Deep Learning Models for Text Classification, Vol. 12736 LNCS, 2021.doi:10.1007/978-3-030-78609-0_31

  28. [28]

    Singh, W

    S. Singh, W. Singh, Ai-based personality prediction for human well-being from text data: a systematic review, Multimedia Tools and Applications 83 (2024) 46325–46368.doi:10.1007/s11042-023-17282-w

  29. [29]

    Vinodhini, R

    G. Vinodhini, R. Chandrasekaran, A comparative performance evaluation of neural network based approach for sentiment classification of online reviews, Journal of King Saud University - Computer and Information Sciences 28 (2016) 2–12.doi:10.1016/j.jksuci.2014.03.024

  30. [30]

    Abdar, F

    M. Abdar, F. Pourpanah, S. Hussain, D. Rezazadegan, L. Liu, M. Ghavamzadeh, P. Fieguth, X. Cao, A. Khosravi, U. Acharya, V. Makarenkov, S. Nahavandi, A review of uncertainty quantification in deep learning: Techniques, applications and challenges, Information Fusion 76 (2021) 243–297.doi:10.1016/j.inffus.2021.05.008

  31. [31]

    Anggrainingsih, G

    R. Anggrainingsih, G. Hassan, A. Datta, Transformer-based models for combating rumours on microblogging platforms: a review, Artificial Intelligence Review 57 (2024).doi:10.1007/s10462-024-10837-9

  32. [32]

    Mamani-Coaquira, E

    Y. Mamani-Coaquira, E. Villanueva, A review on text sentiment analysis with machine learning and deep learning techniques, IEEE Access 12 (2024) 193115–193130.doi:10.1109/ACCESS.2024.3513321

  33. [33]

    Zaheer, M

    H. Zaheer, M. Bashir, Detecting fake news for covid-19 using deep learning: a review, Multimedia Tools and Applications 83 (2024) 74469–74502.doi:10.1007/s11042-024-18564-7

  34. [34]

    Fields, K

    J. Fields, K. Chovanec, P. Madiraju, A survey of text classification with transformers: How wide? how large? how long? how accurate? how expensive? how safe?, IEEE Access 12 (2024) 6518–6531. doi: 10.1109/ACCESS.2024.3349952

  35. [35]

    Biradar, S

    S. Biradar, S. Saumya, A. Chauhan, Combating the infodemic: Covid-19 induced fake news recognition in social media networks, Complex and Intelligent Systems 9 (2023) 2879–2891.doi:10.1007/s40747-022-00672-2

  36. [36]

    Snidaro, J

    L. Snidaro, J. García, J. Llinas, Context-based information fusion: A survey and discussion, Information Fusion 25 (2015) 16–31.doi:10.1016/J.INFFUS.2015.01.002

  37. [37]

    Sun, A survey of multi-view machine learning, Neural Computing and Applications 23 (7–8) (2013) 2031–2038

    S. Sun, A survey of multi-view machine learning, Neural Computing and Applications 23 (7–8) (2013) 2031–2038. doi:10.1007/s00521-013-1362-6

  38. [38]

    J. Zhao, X. Xie, X. Xu, S. Sun, Multi-view learning overview: Recent progress and new challenges, Information Fusion 38 (2017) 43–54.doi:10.1016/j.inffus.2017.02.007

  39. [39]

    Baltrusaitis, C

    T. Baltrusaitis, C. Ahuja, L.-P. Morency, Multimodal machine learning: A survey and taxonomy, IEEE Transactions on Pattern Analysis and Machine Intelligence 41 (2) (2019) 423–443.doi:10.1109/tpami.2018.2 798607

  40. [40]

    W. Guo, J. Wang, S. Wang, Deep multimodal representation learning: A survey, IEEE Access 7 (2019) 63373–63394.doi:10.1109/access.2019.2916887

  41. [41]

    Y. Qin, X. Zhang, S. Yu, G. Feng, A survey on representation learning for multi-view data, Neural Networks 181 (2025) 106842.doi:10.1016/J.NEUNET.2024.106842

  42. [42]

    Moher, A

    D. Moher, A. Liberati, J. Tetzlaff, D. Altman, Preferred reporting items for systematic reviews and meta-analyses: The prisma statement, International Journal of Surgery 8 (2010) 336–341.doi:10.1016/j.ijsu.2010.02.007. 51

  43. [43]

    M. J. Page, J. E. McKenzie, P. M. Bossuyt, I. Boutron, T. C. Hoffmann, C. D. Mulrow, L. Shamseer, J. M. Tetzlaff, E. A. Akl, S. E. Brennan, R. Chou, J. Glanville, J. M. Grimshaw, A. Hróbjartsson, M. M. Lalu, T. Li, E. W. Loder, E. Mayo-Wilson, S. McDonald, L. A. McGuinness, L. A. Stewart, J. Thomas, A. C. Tricco, V. A. Welch, P. Whiting, D. Moher, The pri...

  44. [44]

    H. Zou, M. Shen, C. Chen, Y. Hu, D. Rajan, E. S. Chng, Unis-mmc: Multimodal classification via unimodality- supervised multimodal contrastive learning, in: Findings of the Association for Computational Linguistics: ACL 2023, Association for Computational Linguistics (ACL), 2023, pp. 659–672.doi:10.18653/v1/2023.finding s-acl.41

  45. [45]

    S. Garg, H. SS, S. Kumar, On-device document classification using multimodal features, in: Proceedings of the 3rd ACM India Joint International Conference on Data Science & Management of Data (8th ACM IKDD CODS & 26th COMAD), CODS COMAD 2021, ACM, 2021, pp. 203–207.doi:10.1145/3430984.3431030

  46. [46]

    L. Chen, H. W. Chou, Utilizing cross-modal contrastive learning to improve item categorization bert model, in: Proceedings of The Fifth Workshop on e-Commerce and NLP (ECNLP 5), Association for Computational Linguistics (ACL), 2022, pp. 217–223.doi:10.18653/v1/2022.ecnlp-1.25

  47. [47]

    Hessel, L

    J. Hessel, L. Lee, Does my multimodal model learn cross-modal interactions? it’s harder to tell than you might think!, in: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Association for Computational Linguistics (ACL), 2020, pp. 861–877.doi:10.18653/v1/2020.emnlp-main.62

  48. [48]

    Zhang, Image Information Prompt: Tips for Learning Large Language Models, IOS Press, 2024, pp

    Y. Zhang, Image Information Prompt: Tips for Learning Large Language Models, IOS Press, 2024, pp. undefined–undefined.doi:10.3233/atde231197. URLhttps://www.mendeley.com/catalogue/cb65323d-510b-31b5-a865-34cc24fff44c/

  49. [49]

    Guélorget, G

    P. Guélorget, G. Gadek, T. Zaharia, B. Grilheres, Active learning to measure opinion and violence in french newspapers, Procedia Computer Science 192 (2021) 202–211.doi:10.1016/j.procs.2021.08.021

  50. [50]

    N. A. Andriyanov, Combining text and image analysis methods for solving multimodal classification problems, Pattern Recognition and Image Analysis 32 (3) (2022) 489–494.doi:10.1134/s1054661822030026

  51. [51]

    Gupta, I

    D. Gupta, I. Sen, N. Sachdeva, P. Kumaraguru, A. B. Buduru, Empowering first responders through automated multimodal content moderation, in: 2018 IEEE International Conference on Cognitive Computing (ICCC), IEEE, 2018, pp. 1–8.doi:10.1109/iccc.2018.00008

  52. [52]

    E. Setiawan, Multiview sentiment analysis with image-text-concept features of indonesian social media posts, International Journal of Intelligent Engineering and Systems 14 (2) (2021) 521–535.doi:10.22266/ijies2021 .0430.47

  53. [53]

    T. A. G. da Costa, R. I. Meneguette, J. Ueyama, Providing a greater precision of situational awareness of urban floods through multimodal fusion, Expert Systems with Applications 188 (2022) 115923. doi: 10.1016/j.eswa.2021.115923. URLhttps://linkinghub.elsevier.com/retrieve/pii/S095741742101277X

  54. [54]

    Jiang, J

    S. Jiang, J. Hu, C. L. Magee, J. Luo, Deep learning for technical document classification, IEEE Transactions on Engineering Management 71 (2024) 1163–1179.doi:10.1109/tem.2022.3152216

  55. [55]

    T. Liu, Y. Hu, J. Gao, J. Wang, Y. Sun, B. Yin, Multi-modal long document classification based on hierarchical prompt and multi-modal transformer, Neural Networks 176 (2024) 106322.doi:10.1016/j.neunet.2024.1063 22

  56. [56]

    L. Braz, V. Teixeira, H. Pedrini, Z. Dias, ImTeNet: Image-Text Classification Network for Abnormality Detection and Automatic Reporting on Musculoskeletal Radiographs, Vol. 12558 LNBI, Springer International Publishing, 2020, pp. 150–161.doi:10.1007/978-3-030-65775-8_14

  57. [57]

    P. H. Luz de Araujo, A. P. G. S. de Almeida, F. Ataides Braz, N. Correia da Silva, F. de Barros Vidal, T. E. de Campos, Sequence-aware multimodal page classification of brazilian legal documents, International Journal on Document Analysis and Recognition (IJDAR) 26 (1) (2022) 33–49.doi:10.1007/s10032-022-00406-7

  58. [58]

    Alqaraleh, H

    S. Alqaraleh, H. Sirin, Multimodal Classifier for Disaster Response, Vol. 1983 CCIS, Springer Nature Switzerland, 2023, pp. 1–13.doi:10.1007/978-3-031-50920-9_1

  59. [59]

    Martinc, B

    M. Martinc, B. Škrlj, S. Pollak, Multilingual gender classification with multi-view deep learning notebook for pan at clef 2018, Vol. 2125, CEUR-WS, 2018

  60. [60]

    Liparas, Y

    D. Liparas, Y. HaCohen-Kerner, A. Moumtzidou, S. Vrochidis, I. Kompatsiaris, News Articles Classification Using Random Forests and Weighted Multimodal Features, Vol. 8849, Springer International Publishing, 2014, pp. 63–75.doi:10.1007/978-3-319-12979-2_6

  61. [61]

    Rasheed, A

    A. Rasheed, A. I. Umar, S. H. Shirazi, Z. Khan, M. Shahzad, Cover-based multiple book genre recognition using an improved multimodal network, International Journal on Document Analysis and Recognition (IJDAR) 52 26 (1) (2022) 65–88.doi:10.1007/s10032-022-00413-8

  62. [62]

    Ortiz-Perez, P

    D. Ortiz-Perez, P. Ruiz-Ponce, D. Tomás, J. Garcia-Rodriguez, M. F. Vizcaya-Moreno, M. Leo, A deep learning-based multimodal architecture to predict signs of dementia, Neurocomputing 548 (2023) 126413. doi:10.1016/j.neucom.2023.126413

  63. [63]

    Akhtiamov, V

    O. Akhtiamov, V. Palkov, Gaze, Prosody and Semantics: Relevance of Various Multimodal Signals to Addressee Detection in Human-Human-Computer Conversations, Vol. 11096 LNAI, Springer International Publishing, 2018, pp. 1–10.doi:10.1007/978-3-319-99579-3_1

  64. [64]

    Ravikiran, K

    M. Ravikiran, K. Madgula, Fusing deep quick response code representations improves malware text classification, in: Proceedings of the ACM Workshop on Crossmodal Learning and Application, ICMR ’19, Association for Computing Machinery (ACM), 2019, pp. 11–18.doi:10.1145/3326459.3329166

  65. [65]

    Y. Li, X. Zheng, M. Zhu, J. Mei, Z. Chen, Y. Tao, Compact bilinear pooling and multi-loss network for social media multimodal classification, Signal, Image and Video Processing 18 (11) (2024) 8403–8412.doi: 10.1007/s11760-024-03482-w

  66. [66]

    T. Liu, Y. Hu, J. Gao, Y. Sun, B. Yin, Cross-modal multiple granularity interactive fusion network for long document classification, ACM Transactions on Knowledge Discovery from Data 18 (4) (2024) 1–24. doi:10.1145/3631711

  67. [67]

    C. Ma, A. Shen, H. Yoshikawa, T. Iwakura, D. Beck, T. Baldwin, On the (in)effectiveness of images for text classification, in: Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, Association for Computational Linguistics, 2021, pp. 42–48. doi:10.18653/v1/2021.eacl-main.4

  68. [68]

    Bakkali, Z

    S. Bakkali, Z. Ming, M. Coustaty, M. Rusiñol, O. R. Terrades, Vlcdoc: Vision-language contrastive pre-training model for cross-modal document classification, Pattern Recognition 139 (2023) 109419.doi:10.1016/j.patcog .2023.109419

  69. [69]

    M. A. Álvarez Carmona, E. Villatoro Tello, M. Montes y Gómez, L. Villaseñor Pineda, Author profiling in social media with multimodal information, Computación y Sistemas 24 (3) (2020) 1289–1304.doi:10.13053/c ys-24-3-3488

  70. [70]

    D. Guo, J. Zhang, B. Yang, Y. Lin, A comparative study of speaker role identification in air traffic communication using deep learning approaches, ACM Transactions on Asian and Low-Resource Language Information Processing 22 (4) (2023) 1–17.doi:10.1145/3572792

  71. [71]

    M. A. Wajid, A. Zafar, M. S. Wajid, A deep learning approach for image and text classification using neutrosophy, International Journal of Information Technology 16 (2) (2023) 853–859.doi:10.1007/s41870-023-01529-8

  72. [72]

    B. Liu, L. He, Y. Xie, Y. Xiang, L. Zhu, W. Ding, Minjot: Multimodal infusion joint training for noise learning in text and multimodal classification problems, Information Fusion 102 (2024) 102071.doi:10.1016/j.inffus .2023.102071

  73. [73]

    J. A. de Bruijn, H. de Moel, A. H. Weerts, M. C. de Ruiter, E. Basar, D. Eilander, J. C. J. H. Aerts, Improving the classification of flood tweets with contextual hydrological information in a multimodal neural network, Computers & Geosciences 140 (2020) 104485.doi:10.1016/j.cageo.2020.104485

  74. [74]

    Fujinuma, S

    Y. Fujinuma, S. Varia, N. Sankaran, S. Appalaraju, B. Min, Y. Vyas, A multi-modal multilingual benchmark for document image classification, in: Findings of the Association for Computational Linguistics: EMNLP 2023, Association for Computing Machinery (ACM), 2023, pp. 14361–14376.doi:10.18653/v1/2023.findings-emn lp.958

  75. [75]

    T. Ange, N. Roger, D. Aude, F. Claude, Semi-supervised multimodal deep learning model for polarity detection in arguments, in: 2018 International Joint Conference on Neural Networks (IJCNN), Vol. 2018-July, IEEE, 2018, pp. 1–8.doi:10.1109/ijcnn.2018.8489342

  76. [76]

    Lincker, C

    E. Lincker, C. Guinaudeau, O. Pons, J. Dupire, C. Hudelot, V. Mousseau, I. Barbet, C. Huron, Noisy and unbalanced multimodal document classification: Textbook exercises as a use case, in: 20th International Conference on Content-based Multimedia Indexing, CBMI 2023, Association for Computing Machinery (ACM), 2023, pp. 71–78.doi:10.1145/3617233.3617239

  77. [77]

    Sapena, E

    O. Sapena, E. Onaindia, Multimodal classification of teaching activities from university lecture recordings, Applied Sciences 12 (9) (2022) 4785.doi:10.3390/app12094785

  78. [78]

    R. Jain, C. Wigington, Multimodal document image classification, in: 2019 International Conference on Document Analysis and Recognition (ICDAR), IEEE, 2019, pp. 71–77.doi:10.1109/icdar.2019.00021

  79. [79]

    Z. Chen, S. Diao, B. Wang, G. Li, X. Wan, Towards unifying medical vision-and-language pre-training via soft prompts, in: 2023 IEEE/CVF International Conference on Computer Vision (ICCV), IEEE, 2023, pp. 23346–23356.doi:10.1109/iccv51070.2023.02139. 53

  80. [80]

    L. Rei, D. Mladenic, M. Dorozynski, F. Rottensteiner, T. Schleider, R. Troncy, J. S. Lozano, M. G. Salvatella, Multimodal metadata assignment for cultural heritage artifacts, Multimedia Systems 29 (2) (2022) 847–869. doi:10.1007/s00530-022-01025-2

Showing first 80 references.