Pith. sign in

REVIEW 9 cited by

Knowledge Graphs Meet Multi-Modal Learning: A Comprehensive Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.05391 v4 pith:W6LD32XS submitted 2024-02-08 cs.AI cs.CVcs.IRcs.LG

classification cs.AIcs.CVcs.IRcs.LG
keywords multi-modalresearchknowledgelearningtasksmmkgsurveycomprehensive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Knowledge Graphs (KGs) play a pivotal role in advancing various AI applications, with the semantic web community's exploration into multi-modal dimensions unlocking new avenues for innovation. In this survey, we carefully review over 300 articles, focusing on KG-aware research in two principal aspects: KG-driven Multi-Modal (KG4MM) learning, where KGs support multi-modal tasks, and Multi-Modal Knowledge Graph (MM4KG), which extends KG studies into the MMKG realm. We begin by defining KGs and MMKGs, then explore their construction progress. Our review includes two primary task categories: KG-aware multi-modal learning tasks, such as Image Classification and Visual Question Answering, and intrinsic MMKG tasks like Multi-modal Knowledge Graph Completion and Entity Alignment, highlighting specific research trajectories. For most of these tasks, we provide definitions, evaluation benchmarks, and additionally outline essential insights for conducting relevant research. Finally, we discuss current challenges and identify emerging trends, such as progress in Large Language Modeling and Multi-modal Pre-training strategies. This survey aims to serve as a comprehensive reference for researchers already involved in or considering delving into KG and multi-modal learning research, offering insights into the evolving landscape of MMKG research and supporting future work.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. RADD: Retrieval-Augmented Discrete Diffusion for Multi-Modal Knowledge Graph Completion

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    RADD decouples retrieval and reranking in multi-modal KGC via a relation-aware KGE retriever and conditional discrete denoiser, reporting state-of-the-art results on three benchmarks.

  2. When Modalities Remember: Continual Learning for Multimodal Knowledge Graphs

    cs.CL 2026-04 unverdicted novelty 6.0 of 10

    MRCKG combines a multimodal-structural curriculum, cross-modal preservation, and contrastive replay to let multimodal knowledge graphs learn new entities and relations over time without catastrophic forgetting.

  3. Structured and Abstractive Reasoning on Multi-modal Relational Knowledge Images

    cs.CV 2025-10 unverdicted novelty 6.0 of 10

    Authors build a synthetic data generator and two-stage training pipeline for structured abstractive reasoning on multi-modal relational knowledge images, releasing STAR-64K and showing 3B/7B models outperforming GPT-4o.

  4. NLKI: A lightweight Natural Language Knowledge Integration Framework for Improving Small VLMs in Commonsense VQA Tasks

    cs.CL 2025-08 conditional novelty 6.0 of 10

    NLKI combines fine-tuned dense retrieval, LLM-generated explanations, and noise-robust losses to improve small VLMs on commonsense VQA.

  5. RADD: Retrieval-Augmented Discrete Diffusion for Multi-Modal Knowledge Graph Completion

    cs.AI 2026-04 conditional novelty 5.0 of 10

    A retrieve-then-rerank framework using a KGE shortlist and a discrete diffusion reranker reports SOTA MMKGC scores, but the diffusion mechanism is underspecified and not isolated from a generic reranker.

  6. Multi-Perspective Evidence Synthesis and Reasoning for Unsupervised Multimodal Entity Linking

    cs.CL 2026-04 unverdicted novelty 5.0 of 10

    MSR-MEL synthesizes instance-centric, group-level, lexical, and statistical evidence with LLMs and asymmetric teacher-student GNNs to outperform prior unsupervised methods on multimodal entity linking benchmarks.

  7. I2CR: Intra- and Inter-modal Collaborative Reflections for Multimodal Entity Linking

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A text-first, multi-round visual feedback framework reports state-of-the-art top-1 accuracy on WikiMEL, WikiDiverse, and RichMEL.

  8. Evidence-Based Intelligent Diagnostic and Therapeutic Visualization System with Large Language Models: Multi-Turn Interaction and Multimodal Treatment Plan Generation

    cs.AI 2026-06 unverdicted novelty 4.0 of 10

    The system integrates a Neo4j knowledge graph, four-stage symptom matching with LLM verification, genetic-algorithm-optimized proactive questioning, and multimodal evidence-based visualizations to improve diagnostic t...

  9. The Master-Slave Encoder Model for Improving Patent Text Summarization: A New Approach to Combining Specifications and Claims

    cs.CL 2024-11 unverdicted novelty 4.0 of 10

    MSEA uses a master-slave encoder architecture on patent specifications and claims, enhanced with pointer networks and repetition suppression, to generate better summaries as measured by small ROUGE score gains.

Pith tools