CGC improves fine-grained multi-image understanding in MLLMs by constructing contrastive training instances from existing single-image annotations and adding a rule-based spatial reward, achieving SOTA on MIG-Bench and VLM2-Bench with transfer gains to other multimodal tasks.
Mico: Multi-image contrast for reinforcement visual reasoning.arXiv preprint arXiv:2506.22434, 2025
4 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 4roles
background 2polarities
background 2representative citing papers
A 9B multimodal model learns to tailor raw video/GUI streams into schema-aligned training data, matching a proprietary annotator on downstream tasks; the abstract's capacity-scaling claims are not supported by the body.
RSICCLLM introduces a post-training framework with RSICI dataset, difference-aware supervised fine-tuning, and dual-negative preference optimization that claims to outperform much larger models on remote sensing image change captioning.
The survey formalizes MLLM perception as a unified vision-language capability and traces its evolution via a new five-stage taxonomy while outlining future challenges.
citing papers explorer
-
CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding
CGC improves fine-grained multi-image understanding in MLLMs by constructing contrastive training instances from existing single-image annotations and adding a rule-based spatial reward, achieving SOTA on MIG-Bench and VLM2-Bench with transfer gains to other multimodal tasks.
-
DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams
A 9B multimodal model learns to tailor raw video/GUI streams into schema-aligned training data, matching a proprietary annotator on downstream tasks; the abstract's capacity-scaling claims are not supported by the body.
-
RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning
RSICCLLM introduces a post-training framework with RSICI dataset, difference-aware supervised fine-tuning, and dual-negative preference optimization that claims to outperform much larger models on remote sensing image change captioning.
-
From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models
The survey formalizes MLLM perception as a unified vision-language capability and traces its evolution via a new five-stage taxonomy while outlining future challenges.