Pith. sign in

REVIEW 4 cited by

Look Before You Decide: Prompting Active Deduction of MLLMs for Assumptive Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.12966 v5 pith:R6S5TYPO submitted 2024-04-19 cs.CV cs.AI

classification cs.CVcs.AI
keywords mllmsreasoningtextbfdeductionmethodabilitiesactiveassumptive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, Multimodal Large Language Models (MLLMs) have achieved significant success across multiple disciplines due to their exceptional instruction-following capabilities and extensive world knowledge. However, whether these MLLMs possess human-like compositional reasoning abilities remains an open problem. To unveil their reasoning behaviors, we first curate a \textbf{M}ultimodal \textbf{A}ssumptive \textbf{R}ea\textbf{s}oning Benchmark (MARS-Bench) in this paper. Interestingly, we find that most prevalent MLLMs can be easily fooled by the introduction of a presupposition into the question, whereas such presuppositions appear naive to human reasoning. Besides, we also propose a simple yet effective method, Active Deduction (AD), a novel reinforcement learning paradigm to encourage the model to actively perform composite deduction before reaching a final decision. Equipped with the proposed AD method, a MLLM demonstrates significant improvements in assumptive reasoning abilities without compromising its general-purpose question-answering performance. We also provide extensive evaluations of both open-source and private MLLMs on MARS-Bench, along with experimental analyses of the AD method.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Visual Credit Audit for Multimodal Spatial Reasoning

    cs.CV 2026-07 conditional novelty 6.0 of 10

    VCA finds 12.73–26.25% of spatial decisions are correct yet uncredited by the image, and separates marginal image support from relation-specific visual response.

  2. Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes?

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A new 1,430-item multimodal benchmark shows that leading multimodal LLMs rarely notice small visual traps needed for commonsense safety reasoning, with top scores near 0.46 on a scale whose maximum is about 0.97.

  3. A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future

    cs.CV 2024-12 conditional novelty 4.0 of 10

    A historical review that organizes multimodal explainability methods into four chronological eras and three explainability types, extending coverage to generative LLMs.

  4. Retrieval Augmented Recipe Generation

    cs.CV 2024-11 conditional novelty 4.0 of 10

    A retrieval-augmented LMM with stochastic retrieval sampling and self-consistency voting improves recipe generation from food images on Recipe1M.

Pith tools