Pith. sign in

Ouro: A self-bootstrapped frame- work for enhancing multimodal scene understanding

3 Pith papers cite this work. Polarity classification is still indexing.

3 Pith papers citing it
abstract

Recent advances in multimodal language models (MLLMs) have made thinking with images a dominant paradigm for multimodal reasoning. However, existing methods still fail to ensure evidence-answer consistency, where correct answers must be supported by correct visual evidence. To address this issue, we propose DeFacto, a counterfactual reasoning framework that explicitly aligns visual evidence with final answers. Our approach integrates three complementary training paradigms: positive, counterfactual, and random-masking. We further develop a language-guided evidence construction pipeline that automatically localizes question-relevant regions and generates counterfactual variants, resulting in DeFacto-100K. Building on this dataset, we train MLLMs with GRPO-based reinforcement learning and design three complementary rewards to promote correct answering, structured reasoning, and consistent evidence selection. Moreover, we introduce DeFacto-1.5K, a human-annotated benchmark for systematically evaluating evidence-grounded consistency beyond answer accuracy. Experiments on diverse benchmarks demonstrate that DeFacto substantially improves both answer accuracy and evidence-answer consistency over strong baselines.

fields

cs.CV 2 cs.CL 1

years

2026 3

verdicts

UNVERDICTED 3

representative citing papers

Semantic-Enriched Latent Visual Reasoning

cs.CV · 2026-05-19 · unverdicted · novelty 5.0 · 2 refs

SLVR is a two-stage method that enriches region-centric latent representations with fine-grained attribute semantics and aligns them via M-GRPO across multiple queries on the same region, supported by new SLV-Set dataset and SV-QA benchmark.

citing papers explorer

Showing 3 of 3 citing papers.

  • PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning cs.CL · 2026-05-13 · unverdicted · none · ref 37 · internal anchor

    PDCR improves vision-language reasoning by computing separate normalized confidence advantages for perception steps and reasoning steps after unsupervised decomposition.

  • Staying VIGILant: Mitigating Visual Laziness via Counterfactual Visual Alignment in MLLMs cs.CV · 2026-06-24 · unverdicted · none · ref 48 · internal anchor

    VIGIL is a counterfactual RL alignment method that reduces visual hallucinations in MLLMs by enforcing visual grounding via masked attention penalties, outperforming baselines with 25% of the data and showing emergent spatial capabilities.

  • Semantic-Enriched Latent Visual Reasoning cs.CV · 2026-05-19 · unverdicted · none · ref 17 · 2 links · internal anchor

    SLVR is a two-stage method that enriches region-centric latent representations with fine-grained attribute semantics and aligns them via M-GRPO across multiple queries on the same region, supported by new SLV-Set dataset and SV-QA benchmark.