A new benchmark, RUMS, shows multimodal LLMs frequently fail to flag missing, ambiguous, contradictory, or infeasible instructions, but simple prompts to ask clarifying questions recover much of the gap.
If the object *is actually present*, FAIL
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Hidden in Plain Sight: Reasoning in Underspecified and Misspecified Scenarios for Multimodal LLMs
A new benchmark, RUMS, shows multimodal LLMs frequently fail to flag missing, ambiguous, contradictory, or infeasible instructions, but simple prompts to ask clarifying questions recover much of the gap.