REVIEW 26 cited by
Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Strong Artificial Intelligence (Strong AI) or Artificial General Intelligence (AGI) with abstract reasoning ability is the goal of next-generation AI. Recent advancements in Large Language Models (LLMs), along with the emerging field of Multimodal Large Language Models (MLLMs), have demonstrated impressive capabilities across a wide range of multimodal tasks and applications. Particularly, various MLLMs, each with distinct model architectures, training data, and training stages, have been evaluated across a broad range of MLLM benchmarks. These studies have, to varying degrees, revealed different aspects of the current capabilities of MLLMs. However, the reasoning abilities of MLLMs have not been systematically investigated. In this survey, we comprehensively review the existing evaluation protocols of multimodal reasoning, categorize and illustrate the frontiers of MLLMs, introduce recent trends in applications of MLLMs on reasoning-intensive tasks, and finally discuss current practices and future directions. We believe our survey establishes a solid base and sheds light on this important topic, multimodal reasoning.
Forward citations
Cited by 26 Pith papers
-
MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs
MVI-Bench supplies the first taxonomy and dataset focused on misleading visual inputs to measure LVLM robustness, with tests on 18 models revealing clear weaknesses.
-
ClimateViz: A Benchmark for Statistical Reasoning and Fact Verification on Scientific Charts
A large-scale benchmark shows that leading multimodal language models still underperform expert humans at verifying climate claims from scientific charts.
-
HalluScope: Fine-grained Hallucination Diagnosis for Multimodal Large Language Models
HalluScope couples span-level hallucination detection, 12-way type classification, and explanation generation in one model, and shows the resulting feedback reduces hallucinations in two MLLMs.
-
Learning Preference-Based Objectives from Clinical Narratives for Dynamic Sepsis Treatment
CN-PR learns reward functions from LLM-derived preferences over clinical trajectories to improve RL policies for sequential treatment decisions, showing correlation with quality scores and better recovery outcomes.
-
SheetDesigner: MLLM-Powered Spreadsheet Layout Generation with Rule-Based and Vision-Based Reflection
SheetDesigner uses zero-shot multimodal LLMs with rule- and vision-based reflection to generate spreadsheet layouts, and claims a 22.6% gain over baselines on a new seven-criterion benchmark.
-
Large Language Models Show Signs of Alignment with Human Neurocognition During Abstract Reasoning
Only the largest tested LLMs (about 70 billion parameters) match human accuracy on an abstract reasoning task, and the internal geometry of their best layers correlates moderately with human frontal EEG activity.
-
The Emotional Baby Is Truly Deadly: Does your Multimodal Large Reasoning Model Have Emotional Flattery towards Humans?
Multimodal reasoning models can be steered into unsafe behavior by emotional prompts and sometimes conceal harmful reasoning inside seemingly safe responses.
-
PlanMoGPT: Flow-Enhanced Progressive Planning for Text to Motion Synthesis
PlanMoGPT combines progressive coarse-to-fine token planning with a flow-enhanced motion tokenizer to achieve state-of-the-art text-to-motion generation, especially on long sequences.
-
Reinforced Reasoning for Embodied Planning
An SFT-plus-GRPO recipe lifts a 7B VLM to 35.6 percent success on EB-ALFRED versus 22.0 for GPT-4o-mini and 33.7 for Qwen2.5-VL-72B, with smaller but consistent gains on unseen EB-Habitat.
-
Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models
Using large-scale adversarially pretrained vision encoders in LLaVA yields 2x and 1.5x robustness gains on captioning and VQA, and cuts jailbreak success rates by over 10% relative to CLIP fine-tuning baselines.
-
SMoLoRA: Exploring and Defying Dual Catastrophic Forgetting in Continual Visual Instruction Tuning
SMoLoRA uses two separately routed LoRA expert groups, one for visual understanding and one for instruction following, to reduce dual catastrophic forgetting in continual visual instruction tuning.
-
SPICE: Synergy and Partial Information Based Curriculum Evolution
A dynamic curriculum that sorts multimodal samples by heuristic redundancy/unique/synergy scores computed from the model's own predictions improves results over static and prior dynamic curricula on four benchmarks.
-
SATORI: Static Test Oracle Generation for REST APIs
SATORI statically infers REST API test oracles from OpenAPI specs via LLMs, reporting F1 74.3%, above AGORA+'s 69.3%, with 18 confirmed bugs; the supplied full text, however, is a different paper.
-
Pixels to Principles: Probing Intuitive Physics Understanding in Multimodal Language Models
Even the newest multimodal LLMs score near chance on intuitive physics videos, and the paper's probing evidence that vision encoders hold the relevant information is confounded by scene identity.
-
What's in the Box? Reasoning about Unseen Objects from Multimodal Cues
A neurosymbolic pipeline combining LLM parsing, audio classification, and Bayesian reasoning achieves r=0.78 correlation with human judgments on a new hidden-object guessing task.
-
The Eye of Sherlock Holmes: Uncovering User Private Attribute Profiling via Vision-Language Model Agentic Framework
A VLM-LLM agentic pipeline and a new 251-person benchmark show that ordinary personal photo sets can reveal private attributes, including abstract traits like income and MBTI, at rates above human evaluators.
-
PerPO: Perceptual Preference Optimization via Discriminative Rewarding
PerPO trains multimodal LLMs by ranking their candidate answers with deterministic visual rewards (IoU, edit distance) and using the reward differences as margins in listwise preference optimization.
-
Training-Free Mitigation of Language Reasoning Degradation After Multimodal Instruction Tuning
Multimodal instruction tuning degrades math and most language reasoning in LLaVA-Mistral while improving commonsense tasks, and a training-free task-vector merge largely recovers the losses.
-
AI Reasoning for Wireless Communications and Networking: A Survey and Perspectives
A survey that organizes LLM and AI reasoning methods into a taxonomy and maps them onto the physical, link, network, transport, and application layers of wireless networks.
-
Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
A review that categorizes large-model-empowered embodied AI into hierarchical and end-to-end decision-making, imitation and reinforcement learning, and world models.
-
Edge-Based Multimodal Sensor Data Fusion with Vision Language Models (VLMs) for Real-time Autonomous Vehicle Accident Avoidance
A lightweight vision-language model on an edge device fuses roadside hazard alerts with onboard camera views to adjust trajectories, and the authors report a 77% simulated collision reduction over a vision-only baseline.
-
MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models
An ensemble of Gemini 2.5 Flash, Gemini 1.5 Pro, and Gemini 2.5 Pro with strict prompt formatting won the ImageCLEF 2025 multilingual multimodal QA track at 81.4% accuracy.
-
ReFrame: Rectification Framework for Image Explaining Architectures
ReFrame wraps image captioning, VQA, and GPT-4 with a Mask R-CNN rectifier, reporting big gains on metrics defined against that same detector.
-
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey
A survey that categorizes pre-trained-model-based vision-language methods into four challenge-driven paradigms, with performance tables and a discussion of risks.
-
Leveraging Vision-Language Large Models for Interpretable Video Action Recognition with Semantic Tokenization
LVLM-VAR transforms video into 'semantic action tokens' and uses a LoRA-tuned vision-language model to classify actions and generate explanations, reporting 94.1% on NTU RGB+D X-Sub.
-
Visual Large Language Models for Generalized and Specialized Applications
This paper reviews and taxonomizes VLLM applications into vision-to-text, vision-to-action, and text-to-vision, adding ethics and future-work discussion.
Discussion (0). Continue with ORCID to comment.