Pith. sign in

REVIEW 26 cited by

Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.06805 v2 pith:TQLEAYAC submitted 2024-01-10 cs.CL cs.AI

classification cs.CLcs.AI
keywords mllmsmultimodalreasoninglanguagelargemodelssurveyabilities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Strong Artificial Intelligence (Strong AI) or Artificial General Intelligence (AGI) with abstract reasoning ability is the goal of next-generation AI. Recent advancements in Large Language Models (LLMs), along with the emerging field of Multimodal Large Language Models (MLLMs), have demonstrated impressive capabilities across a wide range of multimodal tasks and applications. Particularly, various MLLMs, each with distinct model architectures, training data, and training stages, have been evaluated across a broad range of MLLM benchmarks. These studies have, to varying degrees, revealed different aspects of the current capabilities of MLLMs. However, the reasoning abilities of MLLMs have not been systematically investigated. In this survey, we comprehensively review the existing evaluation protocols of multimodal reasoning, categorize and illustrate the frontiers of MLLMs, introduce recent trends in applications of MLLMs on reasoning-intensive tasks, and finally discuss current practices and future directions. We believe our survey establishes a solid base and sheds light on this important topic, multimodal reasoning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 26 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs

    cs.CV 2025-11 unverdicted novelty 8.0 of 10

    MVI-Bench supplies the first taxonomy and dataset focused on misleading visual inputs to measure LVLM robustness, with tests on 18 models revealing clear weaknesses.

  2. ClimateViz: A Benchmark for Statistical Reasoning and Fact Verification on Scientific Charts

    cs.CL 2025-06 conditional novelty 7.0 of 10

    A large-scale benchmark shows that leading multimodal language models still underperform expert humans at verifying climate claims from scientific charts.

  3. HalluScope: Fine-grained Hallucination Diagnosis for Multimodal Large Language Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    HalluScope couples span-level hallucination detection, 12-way type classification, and explanation generation in one model, and shows the resulting feedback reduces hallucinations in two MLLMs.

  4. Learning Preference-Based Objectives from Clinical Narratives for Dynamic Sepsis Treatment

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    CN-PR learns reward functions from LLM-derived preferences over clinical trajectories to improve RL policies for sequential treatment decisions, showing correlation with quality scores and better recovery outcomes.

  5. SheetDesigner: MLLM-Powered Spreadsheet Layout Generation with Rule-Based and Vision-Based Reflection

    cs.AI 2025-09 conditional novelty 6.0 of 10

    SheetDesigner uses zero-shot multimodal LLMs with rule- and vision-based reflection to generate spreadsheet layouts, and claims a 22.6% gain over baselines on a new seven-criterion benchmark.

  6. Large Language Models Show Signs of Alignment with Human Neurocognition During Abstract Reasoning

    q-bio.NC 2025-08 unverdicted novelty 6.0 of 10

    Only the largest tested LLMs (about 70 billion parameters) match human accuracy on an abstract reasoning task, and the internal geometry of their best layers correlates moderately with human frontal EEG activity.

  7. The Emotional Baby Is Truly Deadly: Does your Multimodal Large Reasoning Model Have Emotional Flattery towards Humans?

    cs.AI 2025-08 unverdicted novelty 6.0 of 10

    Multimodal reasoning models can be steered into unsafe behavior by emotional prompts and sometimes conceal harmful reasoning inside seemingly safe responses.

  8. PlanMoGPT: Flow-Enhanced Progressive Planning for Text to Motion Synthesis

    cs.CV 2025-06 conditional novelty 6.0 of 10

    PlanMoGPT combines progressive coarse-to-fine token planning with a flow-enhanced motion tokenizer to achieve state-of-the-art text-to-motion generation, especially on long sequences.

  9. Reinforced Reasoning for Embodied Planning

    cs.AI 2025-05 conditional novelty 6.0 of 10

    An SFT-plus-GRPO recipe lifts a 7B VLM to 35.6 percent success on EB-ALFRED versus 22.0 for GPT-4o-mini and 33.7 for Qwen2.5-VL-72B, with smaller but consistent gains on unseen EB-Habitat.

  10. Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Using large-scale adversarially pretrained vision encoders in LLaVA yields 2x and 1.5x robustness gains on captioning and VQA, and cuts jailbreak success rates by over 10% relative to CLIP fine-tuning baselines.

  11. SMoLoRA: Exploring and Defying Dual Catastrophic Forgetting in Continual Visual Instruction Tuning

    cs.CV 2024-11 conditional novelty 6.0 of 10

    SMoLoRA uses two separately routed LoRA expert groups, one for visual understanding and one for instruction following, to reduce dual catastrophic forgetting in continual visual instruction tuning.

  12. SPICE: Synergy and Partial Information Based Curriculum Evolution

    cs.LG 2026-06 conditional novelty 5.0 of 10

    A dynamic curriculum that sorts multimodal samples by heuristic redundancy/unique/synergy scores computed from the model's own predictions improves results over static and prior dynamic curricula on four benchmarks.

  13. SATORI: Static Test Oracle Generation for REST APIs

    cs.SE 2025-08 unverdicted novelty 5.0 of 10

    SATORI statically infers REST API test oracles from OpenAPI specs via LLMs, reporting F1 74.3%, above AGORA+'s 69.3%, with 18 confirmed bugs; the supplied full text, however, is a different paper.

  14. Pixels to Principles: Probing Intuitive Physics Understanding in Multimodal Language Models

    cs.CL 2025-07 reject novelty 5.0 of 10

    Even the newest multimodal LLMs score near chance on intuitive physics videos, and the paper's probing evidence that vision encoders hold the relevant information is confounded by scene identity.

  15. What's in the Box? Reasoning about Unseen Objects from Multimodal Cues

    cs.AI 2025-06 reject novelty 5.0 of 10

    A neurosymbolic pipeline combining LLM parsing, audio classification, and Bayesian reasoning achieves r=0.78 correlation with human judgments on a new hidden-object guessing task.

  16. The Eye of Sherlock Holmes: Uncovering User Private Attribute Profiling via Vision-Language Model Agentic Framework

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A VLM-LLM agentic pipeline and a new 251-person benchmark show that ordinary personal photo sets can reveal private attributes, including abstract traits like income and MBTI, at rates above human evaluators.

  17. PerPO: Perceptual Preference Optimization via Discriminative Rewarding

    cs.AI 2025-02 conditional novelty 5.0 of 10

    PerPO trains multimodal LLMs by ranking their candidate answers with deterministic visual rewards (IoU, edit distance) and using the reward differences as margins in listwise preference optimization.

  18. Training-Free Mitigation of Language Reasoning Degradation After Multimodal Instruction Tuning

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Multimodal instruction tuning degrades math and most language reasoning in LLaVA-Mistral while improving commonsense tasks, and a training-free task-vector merge largely recovers the losses.

  19. AI Reasoning for Wireless Communications and Networking: A Survey and Perspectives

    cs.NI 2025-09 conditional novelty 4.0 of 10

    A survey that organizes LLM and AI reasoning methods into a taxonomy and maps them onto the physical, link, network, transport, and application layers of wireless networks.

  20. Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning

    cs.RO 2025-08 reject novelty 4.0 of 10

    A review that categorizes large-model-empowered embodied AI into hierarchical and end-to-end decision-making, imitation and reinforcement learning, and world models.

  21. Edge-Based Multimodal Sensor Data Fusion with Vision Language Models (VLMs) for Real-time Autonomous Vehicle Accident Avoidance

    cs.AI 2025-08 reject novelty 4.0 of 10

    A lightweight vision-language model on an edge device fuses roadside hazard alerts with onboard camera views to adjust trajectories, and the authors report a 77% simulated collision reduction over a vision-only baseline.

  22. MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models

    cs.CL 2025-07 conditional novelty 4.0 of 10

    An ensemble of Gemini 2.5 Flash, Gemini 1.5 Pro, and Gemini 2.5 Pro with strict prompt formatting won the ImageCLEF 2025 multilingual multimodal QA track at 81.4% accuracy.

  23. ReFrame: Rectification Framework for Image Explaining Architectures

    cs.CV 2025-06 reject novelty 4.0 of 10

    ReFrame wraps image captioning, VQA, and GPT-4 with a Mask R-CNN rectifier, reporting big gains on metrics defined against that same detector.

  24. How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey

    cs.CV 2024-12 conditional novelty 4.0 of 10

    A survey that categorizes pre-trained-model-based vision-language methods into four challenge-driven paradigms, with performance tables and a discussion of risks.

  25. Leveraging Vision-Language Large Models for Interpretable Video Action Recognition with Semantic Tokenization

    cs.CV 2025-09 reject novelty 3.0 of 10

    LVLM-VAR transforms video into 'semantic action tokens' and uses a LoRA-tuned vision-language model to classify actions and generate explanations, reporting 94.1% on NTU RGB+D X-Sub.

  26. Visual Large Language Models for Generalized and Specialized Applications

    cs.CV 2025-01 conditional novelty 3.0 of 10

    This paper reviews and taxonomizes VLLM applications into vision-to-text, vision-to-action, and text-to-vision, adding ethics and future-work discussion.

Pith tools