PunchBench, a 54,000-question benchmark spanning cartoons, posts, comments, and memes, shows that multimodal LLMs lag humans in punchline comprehension, and its Simple-to-Complex Chain-of-Question prompt yields small accuracy gains.
Lawrence Zitnick, and Devi Parikh
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
PunchBench, a 54,000-question benchmark spanning cartoons, posts, comments, and memes, shows that multimodal LLMs lag humans in punchline comprehension, and its Simple-to-Complex Chain-of-Question prompt yields small accuracy gains.