A new benchmark, TempVS, shows that state-of-the-art multimodal LLMs largely fail at multi-event temporal reasoning across image sequences, despite being able to ground individual events.
Title resolution pending
1 Pith paper cite this work, alongside 9 external citations. Polarity classification is still indexing.
1
Pith paper citing it
9
external citations · OpenAlex
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences?
A new benchmark, TempVS, shows that state-of-the-art multimodal LLMs largely fail at multi-event temporal reasoning across image sequences, despite being able to ground individual events.