Pith. sign in

REVIEW 2 cited by

A First Look: Towards Explainable TextVQA Models via Visual and Textual Explanations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2105.02626 v1 pith:3T44ORLS submitted 2021-04-29 cs.CV cs.AIcs.CLcs.LG

classification cs.CVcs.AIcs.CLcs.LG
keywords explanationsmultimodalmodelstextualexplainablehelptrainingunimodal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Explainable deep learning models are advantageous in many situations. Prior work mostly provide unimodal explanations through post-hoc approaches not part of the original system design. Explanation mechanisms also ignore useful textual information present in images. In this paper, we propose MTXNet, an end-to-end trainable multimodal architecture to generate multimodal explanations, which focuses on the text in the image. We curate a novel dataset TextVQA-X, containing ground truth visual and multi-reference textual explanations that can be leveraged during both training and evaluation. We then quantitatively show that training with multimodal explanations complements model performance and surpasses unimodal baselines by up to 7% in CIDEr scores and 2% in IoU. More importantly, we demonstrate that the multimodal explanations are consistent with human interpretations, help justify the models' decision, and provide useful insights to help diagnose an incorrect prediction. Finally, we describe a real-world e-commerce application for using the generated multimodal explanations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MEGL: Multimodal Explanation-Guided Learning

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A multimodal explanation-guided learning framework that jointly uses visual saliency maps and textual rationales to train image classifiers, improving accuracy, visual explanation overlap, and text explanation scores ...

  2. Beyond Logit Lens: Contextual Embeddings for Robust Hallucination Detection & Grounding in VLMs

    cs.CL 2024-11 conditional novelty 5.0 of 10

    Middle-layer contextual embeddings, not logit-lens readings, improve hallucination detection in VLMs and enable bounding-box grounding for visual question answering.

Pith tools