A multimodal LLM trained with deliberate, format-constrained reasoning but evaluated with free-form reasoning outperforms models that keep the constraints at test time.
#with Qwen2.5-VL(b) Compare D2I!
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Learning Deliberately, Acting Intuitively: Unlocking Test-Time Reasoning in Multimodal LLMs
A multimodal LLM trained with deliberate, format-constrained reasoning but evaluated with free-form reasoning outperforms models that keep the constraints at test time.