ROSE benchmark shows MLLMs drop up to 44.5 percentage points from counting tasks to region-conditioned action on identical scenes, with the gap persisting even when counts are correct.
Contextual: Evaluating context- sensitive text-rich visual reasoning in large multimodal models
3 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CV 3representative citing papers
FinCriticalED benchmark reveals that OCR and MLLM systems frequently fail to preserve critical financial facts such as numbers and monetary units even when lexical accuracy is high.
Iterative SFT-RL cycles enable a 7B LVLM to develop sophisticated visual chain-of-thought reasoning and improve performance on math and general reasoning benchmarks.
citing papers explorer
-
ROSE: Benchmarking the Perception-to-Action Gap in Multimodal Models
ROSE benchmark shows MLLMs drop up to 44.5 percentage points from counting tasks to region-conditioned action on identical scenes, with the gap persisting even when counts are correct.
-
FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR
FinCriticalED benchmark reveals that OCR and MLLM systems frequently fail to preserve critical financial facts such as numbers and monetary units even when lexical accuracy is high.
-
OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles
Iterative SFT-RL cycles enable a 7B LVLM to develop sophisticated visual chain-of-thought reasoning and improve performance on math and general reasoning benchmarks.