MMSearch-R1 uses reinforcement learning to train multimodal models for on-demand multi-turn internet search with image and text tools, outperforming same-size RAG baselines and matching larger ones while cutting search calls by over 30%.
Murag: Multimodal retrieval-augmented generator for open question answering over images and text
5 Pith papers cite this work. Polarity classification is still indexing.
verdicts
UNVERDICTED 5representative citing papers
SAR-RAG augments an MLLM baseline with semantic retrieval of similar known SAR target images, yielding measurable gains in classification accuracy and dimension regression.
MoRE enables MLLMs to dynamically coordinate heterogeneous retrieval experts via Step-GRPO training, yielding over 7% average gains on open-domain QA benchmarks.
mPLUG-Owl3 introduces hyper attention blocks to integrate vision and language for long image-sequence understanding and reports SOTA results on single-image, multi-image, and video benchmarks.
A retrieval-augmented multi-agent system for traceable fault diagnosis in battery energy storage systems, with BESS-specific routing, schema-constrained DB access, hybrid retrieval, and preliminary internal evaluation.
citing papers explorer
-
MMSearch-R1: Incentivizing LMMs to Search
MMSearch-R1 uses reinforcement learning to train multimodal models for on-demand multi-turn internet search with image and text tools, outperforming same-size RAG baselines and matching larger ones while cutting search calls by over 30%.
-
SAR-RAG: ATR Visual Question Answering by Semantic Search, Retrieval, and MLLM Generation
SAR-RAG augments an MLLM baseline with semantic retrieval of similar known SAR target images, yielding measurable gains in classification accuracy and dimension regression.
-
Mixture-of-Retrieval Experts for Reasoning-Guided Multimodal Knowledge Exploitation
MoRE enables MLLMs to dynamically coordinate heterogeneous retrieval experts via Step-GRPO training, yielding over 7% average gains on open-domain QA benchmarks.
-
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
mPLUG-Owl3 introduces hyper attention blocks to integrate vision and language for long image-sequence understanding and reports SOTA results on single-image, multi-image, and video benchmarks.
-
Traceable Fault Diagnosis for Battery Energy Storage Systems via Retrieval-Augmented Multi-Agent O&M Assistant
A retrieval-augmented multi-agent system for traceable fault diagnosis in battery energy storage systems, with BESS-specific routing, schema-constrained DB access, hybrid retrieval, and preliminary internal evaluation.