Pith. sign in

REVIEW 4 cited by

A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.14056 v1 pith:OZ6VG7YB submitted 2024-12-18 cs.CV cs.AIcs.CLcs.LGcs.MM

classification cs.CVcs.AIcs.CLcs.LGcs.MM
keywords mxaireviewexplainablemodelsmultimodalartificialchallengesdatasets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Artificial intelligence (AI) has rapidly developed through advancements in computational power and the growth of massive datasets. However, this progress has also heightened challenges in interpreting the "black-box" nature of AI models. To address these concerns, eXplainable AI (XAI) has emerged with a focus on transparency and interpretability to enhance human understanding and trust in AI decision-making processes. In the context of multimodal data fusion and complex reasoning scenarios, the proposal of Multimodal eXplainable AI (MXAI) integrates multiple modalities for prediction and explanation tasks. Meanwhile, the advent of Large Language Models (LLMs) has led to remarkable breakthroughs in natural language processing, yet their complexity has further exacerbated the issue of MXAI. To gain key insights into the development of MXAI methods and provide crucial guidance for building more transparent, fair, and trustworthy AI systems, we review the MXAI methods from a historical perspective and categorize them across four eras: traditional machine learning, deep learning, discriminative foundation models, and generative LLMs. We also review evaluation metrics and datasets used in MXAI research, concluding with a discussion of future challenges and directions. A project related to this review has been created at https://github.com/ShilinSun/mxai_review.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Feature-level Interaction Explanations in Multimodal Transformers

    cs.LG 2026-03 conditional novelty 6.0 of 10

    FL-I2MoE separates unique, synergistic, and redundant cross-modal evidence at the token/patch level and uses SII and redundancy-gap scores to rank pairs whose removal degrades performance more than random masking.

  2. Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment

    cs.AI 2026-07 conditional novelty 5.0 of 10

    For document visual grounding at 4B scale, reasoning-free GRPO training outperforms a reasoning-enabled variant, and the reasoning variant compresses its traces during training.

  3. Regulatory Graphs and GenAI for Real-Time Transaction Monitoring and Compliance Explanation in Banking

    cs.AI 2025-06 reject novelty 3.0 of 10

    A GNN plus retrieval-augmented LLM pipeline for transaction monitoring is described, but the 98.2% F1 claim rests on an incomparable baseline setup and unreleased synthetic data.

  4. Empowering Multimodal LLMs with External Tools: A Comprehensive Survey

    cs.CV 2025-08 unverdicted novelty 2.0 of 10

    A survey paper maps how external tools are used to augment multimodal large language models across data, tasks, evaluation, and future directions.

Pith tools