Multimodal Explanations: Justifying Decisions and Pointing to the Evidence

Anna Rohrbach; Bernt Schiele; Dong Huk Park; Lisa Anne Hendricks; Marcus Rohrbach; Trevor Darrell; Zeynep Akata

arxiv: 1802.08129 · v1 · pith:XSVD7IYZnew · submitted 2018-02-15 · 💻 cs.AI · cs.CL· cs.CV

Multimodal Explanations: Justifying Decisions and Pointing to the Evidence

Dong Huk Park , Lisa Anne Hendricks , Zeynep Akata , Anna Rohrbach , Bernt Schiele , Trevor Darrell , Marcus Rohrbach This is my paper

classification 💻 cs.AI cs.CLcs.CV

keywords textualexplanationmodelsmultimodalvisualattentionbetterdatasets

0 comments

read the original abstract

Deep models that are both effective and explainable are desirable in many settings; prior explainable models have been unimodal, offering either image-based visualization of attention weights or text-based generation of post-hoc justifications. We propose a multimodal approach to explanation, and argue that the two modalities provide complementary explanatory strengths. We collect two new datasets to define and evaluate this task, and propose a novel model which can provide joint textual rationale generation and attention visualization. Our datasets define visual and textual justifications of a classification decision for activity recognition tasks (ACT-X) and for visual question answering tasks (VQA-X). We quantitatively show that training with the textual explanations not only yields better textual justification models, but also better localizes the evidence that supports the decision. We also qualitatively show cases where visual explanation is more insightful than textual explanation, and vice versa, supporting our thesis that multimodal explanation models offer significant benefits over unimodal approaches.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

Extraction and Analysis of Multimodal Concepts in Vision Language Models through Sparse Autoencoders
cs.CV 2026-06 unverdicted novelty 6.0

A new SAE-based framework extracts visual, textual, and multimodal concepts from VLMs and reports up to 45% better visual concept quality on a VQA dataset while identifying multimodal concepts.
A Single Rewrite Suffices: Empirical Lessons from Production Skill Description Optimization
cs.CL 2026-06 unverdicted novelty 5.0

A single LLM rewrite of skill descriptions using false positive and negative cases matches manual optimization performance in production, with most other pipeline components adding little value.