REVIEW 3 major objections 6 minor 66 references
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read PunchBench: AI scores 53% on image-caption jokes where humans score 91%
desk verdict PunchBench is a serious benchmark, but its shortcut-elimination claim rests on unvalidated caption transformations. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the generation of synonymous and antonymous captions, carried out by gpt-3.5-turbo-0125 through word substitution and inversion, with context-consistency adaptation for captions that contain internal contradictions. By forcing the model to face rephrased and semantically flipped versions of the same caption, the benchmark removes the biased words and text-only inconsistencies that allow shortcuts. The second mechanism is the Simple-to-Complex Chain-of-Question prompting strategy, which orders questions from easy formats (Yes/No) to hard ones (Multi-option or Generation) so that earlier responses feed into the final answer.
What would settle it
Ask fresh human annotators to label a random sample of the generated antonymous captions paired with their images; if a substantial share of antonymous pairs are still judged to contain punchlines, the antonym flip does not hold and the benchmark's shortcut-free claim loses its basis.
Extended reading notes
Core claim
The paper's central claim is that PunchBench gives an accurate and comprehensive measure of multimodal punchline comprehension, and that evaluations with it show current MLLMs fall far short of human ability. The benchmark holds 6,000 image-caption pairs, each expanded into original, synonymous, and antonymous captions, and 54,000 question-answer pairs covering punchline perception and reasoning, four question formats, and domains from cartoons to memes. The labels for the rewritten captions follow a rule: synonymous captions keep the original pair's label, antonymous captions receive the opposite. Across ten MLLMs, the strongest, GPT-4o, reaches 80.7 percent on Yes/No perception but only 53.1 percent on multi-option QA, against human scores of 98.3 and 90.7 percent respectively, which the paper reads as evidence that these models lean on shallow caption cues instead of understanding the image-caption interplay.
Load-bearing premise
The shortcut-free evaluation depends on the generated synonymous and antonymous captions actually preserving or inverting the punchline's meaning, since their labels are assigned by rule rather than verified by human judgment.
Editorial extensions
If this is right
- A single question format overstates MLLM punchline ability, since rankings and raw scores shift substantially across the Yes/No, Matching, Multi-option, and Generation formats.
- The performance drop when original captions are replaced with synonymous or antonymous ones indicates that current MLLMs rely on surface wording rather than the image-caption incongruity that creates the punchline.
- Explaining why a pair is funny is consistently harder for MLLMs than detecting that it is funny, so evaluation should separate perception from reasoning.
- SC-CoQ improves accuracy on all question formats and caption versions, and beats both 3-shot in-context learning and chain-of-thought prompting.
- PunchBench's 54,000 question-answer pairs provide a reusable testbed for developing and auditing multimodal humor systems.
Reading between the lines
- The synonymous/antonymous caption perturbation could be adapted to other caption-dependent vision-language benchmarks as a cheap probe for how much of a model's score comes from text priors rather than image understanding.
- The large drop on antonymous captions suggests a testable hypothesis beyond humor: current MLLMs may not reliably track semantic negation and contrast in multimodal settings.
- Because the antonymous labels are assigned by rule, a natural next step is human verification of whether each generated caption actually flips the punchline, or reformulating the task as predicting whether a rewrite preserves meaning.
- The near-ceiling human scores imply that the bottleneck is joint visual-textual reasoning rather than joke difficulty, so extending the design to video punchlines would be a demanding follow-up, as the paper itself notes.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces PunchBench, a benchmark for multimodal punchline comprehension in image-caption pairs. The dataset contains 6,000 image-caption pairs collected from prior datasets and social media, annotated by crowd voting and human-written reasoning sentences. The authors generate synonymous and antonymous captions with GPT-3.5 to mitigate text-only shortcuts, construct perception and reasoning tasks with four question formats, and evaluate ten MLLMs, reporting a large performance gap between MLLMs and humans. They also propose SC-CoQ, a simple-to-complex chain-of-question prompting strategy that improves performance over in-context learning and chain-of-thought. The paper claims that PunchBench provides accurate and comprehensive evaluation and that state-of-the-art MLLMs substantially lag human-level punchline comprehension.
Significance. If the benchmark construction is valid, PunchBench is a meaningful resource: it broadens domain coverage to cartoons, posts, comments, and memes; introduces multiple question formats for both perception and reasoning; and attempts to reduce shortcuts by creating synonymous and antonymous captions. The dataset and code are publicly released, and the human annotation process adds credibility. The consistent gains from SC-CoQ across ten models are also of interest. However, the central claim of accurate, shortcut-free evaluation rests on the validity of the GPT-generated caption transformations and on a measured human baseline. Both currently require strengthening before the headline human-MLLM gap and the shortcut-free claim can be accepted.
major comments (3)
- [Sections 3.2-3.4, Figure 15] The labels for the GPT-3.5-generated synonymous and antonymous captions are assigned by fiat rather than verified by human judgment: the synonymous caption P_s inherits the original label and the antonymous caption P_a receives the opposite label. The quality check in Section 3.4 samples only 100 instructions per format and tests whether the instructions are answerable; it does not test whether synonymous captions preserve the punchline or antonymous captions eliminate it. The indirect evidence in Section 5.4, namely that model accuracy drops on the transformed captions, cannot distinguish successful shortcut removal from noise introduced by semantically drifted rewrites. Please add a human validation study on a representative sample of transformed caption-image pairs, report per-caption-type accuracy and inter-annotator agreement, and include a caption-only baseline on the transformed subsets to confirm that text-only shortcuts are actually removed.
- [Section 5.1, Human Baseline; Table 1] The Generation QA human baseline is set to 100% by construction because the manually annotated reasoning sentences serve as the reference answer for the automatic evaluation. This is not a measured human performance; it trivially defines the ceiling and inflates the reported human-MLLM gap in punchline reasoning. Please provide a measured human baseline for generation, for example by having annotators write free-form reasoning sentences and scoring those with the same evaluation protocol used for MLLM outputs. In addition, the 100-instruction human baselines for the other question formats should be reported separately for original, synonymous, and antonymous captions, so the reader can see whether the transformed captions were validated by the human annotators.
- [Table 1 caption, Section 5.3] The claim that 'the P-value between SC-CoQ performance and other prompting method results is consistently less than 0.01' is not supported by a described statistical test, by error bars, or by any multiple-comparison correction across the twelve models and six question formats. Please specify the test used (e.g., paired bootstrap or Wilcoxon signed-rank), report the number of replicates and the test statistic or exact P-values, and provide confidence intervals or standard errors. Without this, the central claim that SC-CoQ outperforms in-context learning and chain-of-thought cannot be properly evaluated.
minor comments (6)
- [Table 4] The benchmark name is misspelled as 'PunchBech' in the table header, and the HUB row contains 'Matching,Ranking and ExplanationSingle' with missing spaces.
- [Figure 21] The prompt text contains the typo 'sunmmarize'; it should be 'summarize'.
- [Section 5.3] The word 'corss' should be 'cross' in the sentence describing SC-CoQ performance across question formats.
- [Section 5.1] The description of the human baseline is ambiguous: it states that 100 instructions were selected for punchline perception, but Table 1 also reports human numbers for punchline reasoning. Please clarify that the same procedure was applied to both tasks.
- [Section 5.2, footnote 2] The reference for gpt-3.5-turbo-0125 is given as 'https://chatgpt.com/', which points to the ChatGPT web interface rather than the model API documentation; please cite the appropriate API reference.
- [Section 3.1] The text contains the typo 'multimeda platforms' in the sentence about Appendix F; it should be 'multimedia platforms'.
Circularity Check
One reported human score is definitional: the Generation QA human baseline is set to the gold answer itself, making part of the MLLM-vs-human gap an input-as-output artifact.
-
self definitional
[Section 5.1 (Human Baseline), Table 1(b) Generation QA]
"Notably, the manually annotated reasoning sentences serve as the performance of human baseline for the Generation QA."
The human baseline for Generation QA is not measured from independent human subjects; it is defined as the gold reference answer itself, i.e., the manually annotated reasoning sentence. Therefore the reported Human score of 100.0 for Generation QA in Table 1(b) is correct by construction. The headline gap in that row (e.g., GPT-4o SC-CoQ 50.1 vs Human 100.0) is partly an artifact of setting the human output equal to the evaluation target rather than an empirical measurement of human generation ability. The perception tasks and other reasoning formats have independently measured human scores, so the overall gap claim does not collapse entirely, but this specific comparison is an input-as-output result.
full rationale
The central benchmark labels are anchored to human crowd voting and human-written reasoning annotations, so the MLLM accuracy numbers are genuinely measured against an external standard rather than fitted. There is no load-bearing self-citation chain, no imported uniqueness theorem, and no ansatz smuggled in via citation; the self-citations present are contextual only. The one clear reduction-by-construction is the Generation QA human baseline, where the 'human' score is defined as the gold reference sentence, inflating the human side of the gap by definition. The synonymous and antonymous caption labels are assigned by fiat (P_s inherits the original label and P_a gets the opposite), and the 100-instruction-per-format quality check does not directly verify semantic preservation or inversion; this is a real construct-validity threat to the 'shortcut-free' claim, but it is not itself a circular derivation because the models' accuracy values are not determined by that label assignment. Similarly, using GPT-4o/GPT-3.5 to generate distractors and options raises contamination concerns without making the measured scores equivalent to the inputs. Overall, the paper's central claim of a large MLLM-human gap has substantial independent support, but the Generation QA human baseline makes part of the reported gap definitional rather than empirical.
Assumptions & free parameters
free parameters (1)
- Crowd voting agreement threshold =
80%
assumptions (3)
- ad hoc to paper GPT-3.5-generated synonymous and antonymous captions preserve or invert the punchline semantics as labeled
- domain assumption Crowd voting with greater than 80% agreement yields correct punchline labels
- domain assumption GPT-3.5 automatic judge for Generation QA aligns with human judgment
Cite this review
Pith. "Pith review of PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension." pith.science (2026). https://pith.science/paper/DVX7O6MH
@misc{pith2026241211906,
author = {Pith},
title = {Pith review of: PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension},
year = {2026},
howpublished = {\url{https://pith.science/paper/DVX7O6MH}},
note = {Machine review of arXiv:2412.11906}
}
read the original abstract
Multimodal punchlines, which involve humor or sarcasm conveyed in image-caption pairs, are a popular way of communication on online multimedia platforms. With the rapid development of multimodal large language models (MLLMs), it is essential to assess their ability to effectively comprehend these punchlines. However, existing benchmarks on punchline comprehension suffer from three major limitations: 1) language shortcuts that allow models to solely rely on text, 2) lack of question diversity, and 3) narrow focus on a specific domain of multimodal content (e.g., cartoon). To address these limitations, we introduce a multimodal \textbf{Punch}line comprehension \textbf{Bench}mark, named \textbf{PunchBench}, which is tailored for accurate and comprehensive evaluation of punchline comprehension. To enhance the evaluation accuracy, we generate synonymous and antonymous captions by modifying original captions, which mitigates the impact of shortcuts in the captions. To provide a comprehensive evaluation, PunchBench incorporates diverse question formats and image-captions from various domains. On this basis, we conduct extensive evaluations and reveal a significant gap between state-of-the-art MLLMs and humans in punchline comprehension. To improve punchline comprehension, we propose Simple-to-Complex Chain-of-Question (SC-CoQ) strategy, enabling the models to incrementally address complicated questions by first mastering simple ones. SC-CoQ effectively enhances the performance of various MLLMs on PunchBench, surpassing in-context learning and chain-of-thought.
Figures
Figures from the paper (20 more)
Reference graph
Works this paper leans on
-
[1]
URL: " 'urlintro :=
ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRINGS urlintro eprinturl eprintpr...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
AI@Meta. 2024. https://github.com/meta-llama/llama3/blob/main/MODEL_CARD.md Llama 3 model card
2024
-
[4]
Lawrence Zitnick, and Devi Parikh
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015. VQA: visual question answering. In ICCV , pages 2425--2433. IEEE Computer Society
work page 2015
-
[5]
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023 a . https://doi.org/10.48550/ARXIV.2308.12966 Qwen-vl: A frontier large vision-language model with versatile abilities . CoRR, abs/2308.12966
-
[6]
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023 b . Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond. arXiv preprint arXiv:2308.12966
arXiv 2023
-
[7]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert - Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litw...
2020
-
[8]
Yitao Cai, Huiyu Cai, and Xiaojun Wan. 2019. https://doi.org/10.18653/V1/P19-1239 Multi-modal sarcasm detection in twitter with hierarchical fusion model . In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers , pages 2506--2515. Association for Com...
Show all 66 references
-
[9]
Zheng Cai, Maosong Cao, Haojiong Chen, Kai Chen, Keyu Chen, Xin Chen, Xun Chen, Zehui Chen, Zhi Chen, Pei Chu, et al. 2024. Internlm2 technical report. arXiv preprint arXiv:2403.17297
2024 arXiv
-
[10]
Santiago Castro, Devamanyu Hazarika, Ver \'o nica P \'e rez-Rosas, Roger Zimmermann, Rada Mihalcea, and Soujanya Poria. 2019. https://doi.org/10.18653/v1/P19-1455 Towards multimodal sarcasm detection (an \_ O bviously \_ perfect paper) . In Proceedings of the 57th Annual Meeti...
2019 doi
-
[11]
Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, et al. 2024 a . Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling. arXiv preprint arXiv:2412.05271
2024 arXiv
-
[12]
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. 2024 b . Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In Proceedings of the IEEE/CVF conference on com...
2024
-
[13]
Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, and Lidong Bing. 2024. https://arxiv.org/abs/2406.07476 Videollama 2: Advancing spatial-temporal modeling and audio understanding in video-llms . arXiv preprint arXiv...
2024 arXiv
-
[14]
Shad Akhtar
Poorav Desai, Tanmoy Chakraborty, and Md. Shad Akhtar. 2022. Nice perfume. how long did you marinate in it? multimodal sarcasm explanation. In Thirty-Sixth AAAI Conference on Artificial Intelligence, AAAI , pages 10563--10571. AAAI
2022
-
[15]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. https://openreview.net/forum?id=YicbFdNTTy An image is worth 1...
2021
-
[16]
Team GLM, Aohan Zeng, Bin Xu, Bowen Wang, Chenhui Zhang, Da Yin, Diego Rojas, Guanyu Feng, Hanlin Zhao, Hanyu Lai, Hao Yu, Hongning Wang, Jiadai Sun, Jiajie Zhang, Jiale Cheng, Jiayi Gui, Jie Tang, Jing Zhang, Juanzi Li, Lei Zhao, Lindong Wu, Lucen Zhong, Mingdao Liu, Minlie H...
2024 arXiv
-
[17]
Kilem L. Gwet. 2014. Handbook of inter-rater reliability: The definitive guide to measuring the extent of agreement among raters. In 4th edition edition, pages 1--38. Advanced Analytics, LLC
2014
-
[18]
Hempelmann and Max Petrenko
Christian F. Hempelmann and Max Petrenko. 2015. https://doi.org/10.1007/978-3-319-20804-6\_59 An AI for humorously reframing interaction narratives with human users . In Distributed, Ambient, and Pervasive Interactions - Third International Conference, DAPI 2015, Held as Part ...
2015 doi
-
[19]
understanding
Jack Hessel, Ana Marasovic, Jena D. Hwang, Lillian Lee, Jeff Da, Rowan Zellers, Robert Mankoff, and Yejin Choi. 2023. https://doi.org/10.18653/V1/2023.ACL-LONG.41 Do androids laugh at electric sheep? humor "understanding" benchmarks from the new yorker caption contest . In Pro...
2023 doi
-
[20]
Wenyi Hong, Weihan Wang, Ming Ding, Wenmeng Yu, Qingsong Lv, Yan Wang, Yean Cheng, Shiyu Huang, Junhui Ji, Zhao Xue, et al. 2024. Cogvlm2: Visual language models for image and video understanding. arXiv preprint arXiv:2408.16500
2024 arXiv
-
[21]
Noman Islam, Zeeshan Islam, and Nazia Noor. 2017. https://api.semanticscholar.org/CorpusID:1040908 A survey on optical character recognition system . ArXiv, abs/1710.05703
2017 arXiv
-
[22]
Jentzsch and Kristian Kersting
Sophie F. Jentzsch and Kristian Kersting. 2023. https://doi.org/10.18653/V1/2023.WASSA-1.29 Chatgpt is fun, but it is not funny! humor is still challenging large language models . In Proceedings of the 13th Workshop on Computational Approaches to Subjectivity, Sentiment, & Soc...
2023 doi
-
[23]
Pu Jian, Donglei Yu, and Jiajun Zhang. 2024. https://aclanthology.org/2024.emnlp-main.613 Large language models know what is key visual entity: An llm-assisted multimodal retrieval for VQA . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Process...
2024
-
[24]
Liqiang Jing, Xuemeng Song, Kun Ouyang, Mengzhao Jia, and Liqiang Nie. 2023. https://doi.org/10.18653/V1/2023.ACL-LONG.635 Multi-source semantic graph-based multimodal sarcasm explanation generation . In Proceedings of the 61st Annual Meeting of the Association for Computation...
2023 doi
-
[25]
Justin Johnson, Andrej Karpathy, and Li Fei - Fei. 2016. https://doi.org/10.1109/CVPR.2016.494 Densecap: Fully convolutional localization networks for dense captioning . In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-...
2016 doi
-
[26]
Dayoon Ko, Sangho Lee, and Gunhee Kim. 2023. https://doi.org/10.18653/V1/2023.EMNLP-MAIN.176 Can language models laugh at youtube short-form videos? In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10,...
2023 doi
-
[27]
Shad Akhtar, and Tanmoy Chakraborty
Shivani Kumar, Atharva Kulkarni, Md. Shad Akhtar, and Tanmoy Chakraborty. 2022. https://doi.org/10.18653/V1/2022.ACL-LONG.411 When did you become so smart, oh wise one?! sarcasm explanation in multi-modal multi-party dialogues . In Proceedings of the 60th Annual Meeting of the...
2022 doi
-
[28]
Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Ziwei Liu, et al. 2024 a . Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326
2024 arXiv
-
[29]
Dongxu Li, Yudong Liu, Haoning Wu, Yue Wang, Zhiqi Shen, Bowen Qu, Xinyao Niu, Fan Zhou, Chengen Huang, Yanpeng Li, et al. 2024 b . Aria: An open multimodal native mixture-of-experts model. arXiv preprint arXiv:2410.05993
2024 arXiv
-
[30]
Junnan Li, Dongxu Li, Silvio Savarese, and Steven C. H. Hoi. 2023 a . https://proceedings.mlr.press/v202/li23q.html BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models . In International Conference on Machine Learning, ICML 20...
2023
-
[31]
Kunchang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. 2023 b . Videochat: Chat-centric video understanding. arXiv preprint arXiv:2305.06355
2023 arXiv
-
[32]
Bin Lin, Bin Zhu, Yang Ye, Munan Ning, Peng Jin, and Li Yuan. 2023. Video-llava: Learning united visual representation by alignment before projection. arXiv preprint arXiv:2311.10122
2023 arXiv
-
[33]
Chin-Yew Lin. 2004. https://aclanthology.org/W04-1013 ROUGE : A package for automatic evaluation of summaries . In Text Summarization Branches Out, pages 74--81. Association for Computational Linguistics
2004
-
[34]
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2024 a . https://doi.org/10.1109/CVPR52733.2024.02484 Improved baselines with visual instruction tuning . In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 2024 , p...
2024
-
[35]
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. 2024 b . https://llava-vl.github.io/blog/2024-01-30-llava-next/ Llava-next: Improved reasoning, ocr, and world knowledge
2024
-
[36]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023 a . http://papers.nips.cc/paper\_files/paper/2023/hash/6dcf277ea32ce3288914faf369fe6de0-Abstract-Conference.html Visual instruction tuning . In Advances in Neural Information Processing Systems 36: Annual Conference...
2023
-
[37]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023 b . Visual instruction tuning
2023
-
[38]
Yuanxin Liu, Shicheng Li, Yi Liu, Yuxiang Wang, Shuhuai Ren, Lei Li, Sishuo Chen, Xu Sun, and Lu Hou. 2024 c . https://doi.org/10.18653/V1/2024.FINDINGS-ACL.517 Tempcompass: Do video llms really understand videos? In Findings of the Association for Computational Linguistics, A...
2024 doi
-
[39]
Yuliang Liu, Zhang Li, Hongliang Li, Wenwen Yu, Mingxin Huang, Dezhi Peng, Mingyu Liu, Mingrui Chen, Chunyuan Li, Lianwen Jin, and Xiang Bai. 2023 c . On the hidden mystery of OCR in large multimodal models. CoRR, abs/2305.07895
2023 arXiv
-
[40]
Yanxin Long, Youpeng Wen, Jianhua Han, Hang Xu, Pengzhen Ren, Wei Zhang, Shen Zhao, and Xiaodan Liang. 2023. https://doi.org/10.1109/CVPR52729.2023.01462 Capdet: Unifying dense captioning and open-world detection pretraining . In IEEE/CVF Conference on Computer Vision and Patt...
2023
-
[41]
Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai - Wei Chang, Michel Galley, and Jianfeng Gao. 2024. https://openreview.net/forum?id=KUNzEQMWU7 Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts ....
2024
-
[42]
Ruipu Luo, Ziwang Zhao, Min Yang, Junwei Dong, Minghui Qiu, Pengcheng Lu, Tao Wang, and Zhongyu Wei. 2023. http://arxiv.org/abs/2306.07207 Valley: Video assistant with large language model enhanced ability
2023 arXiv
-
[43]
Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan. 2024. Video-chatgpt: Towards detailed video understanding via large vision and language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024)
2024
-
[44]
Abdelkader El Mahdaouy, Abdellah El Mekki, Kabil Essefar, Nabil El Mamoun, Ismail Berrada, and Ahmed Khoumsi. 2021. https://www.aclweb.org/anthology/2021.wanlp-1.42/ Deep multi-task model for sarcasm detection and sentiment analysis in arabic language . In Proceedings of the S...
2021
-
[45]
Minesh Mathew, Dimosthenis Karatzas, and C. V. Jawahar. 2021. https://doi.org/10.1109/WACV48630.2021.00225 Docvqa: A dataset for VQA on document images . In IEEE Winter Conference on Applications of Computer Vision, WACV 2021, Waikoloa, HI, USA, January 3-8, 2021 , pages 2199-...
2021
-
[46]
OpenAI. 2022. https://openai.com/blog/chatgpt Introducing chatgpt . CoRR
2022
-
[47]
OpenAI. 2023a. https://cdn.openai.com/papers/GPTV_System_Card.pdf Gpt-4v(ision) system card
-
[48]
OpenAI. 2024. https://openai.com/index/hello-gpt-4o/ Gpt-4o
2024
-
[49]
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff ...
2024 arXiv
- [50]
-
[51]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei - Jing Zhu. 2002. https://doi.org/10.3115/1073083.1073135 Bleu: a method for automatic evaluation of machine translation . In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, July 6-12, ...
2002
-
[52]
Yang Qiao, Liqiang Jing, Xuemeng Song, Xiaolin Chen, Lei Zhu, and Liqiang Nie. 2023. https://doi.org/10.1609/AAAI.V37I8.26138 Mutual-enhanced incongruity learning network for multi-modal sarcasm detection . In Thirty-Seventh AAAI Conference on Artificial Intelligence, AAAI 202...
2023 doi
-
[53]
Noam Shazeer. 2020. http://arxiv.org/abs/2002.05202 GLU variants improve transformer . CoRR, abs/2002.05202
2020 arXiv
-
[54]
The Mistral AI Team. 2023. https://huggingface.co/mistralai/Mistral-7B-Instruct-v0.2 Mistral-7b-instruct-v0.2
2023
- [55]
- [56]
-
[57]
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. 2024. Qwen2-vl: Enhancing vision-language mode...
2024 arXiv
-
[58]
Weihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong, Ji Qi, Yan Wang, Junhui Ji, Zhuoyi Yang, Lei Zhao, Xixuan Song, Jiazheng Xu, Bin Xu, Juanzi Li, Yuxiao Dong, Ming Ding, and Jie Tang. 2023. http://arxiv.org/abs/2311.03079 Cogvlm: Visual expert for pretrained language models
2023 arXiv
-
[59]
Chi, Quoc V
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. http://papers.nips.cc/paper\_files/paper/2022/hash/9d5609613524ecf4f15af0f7b31abca4-Abstract-Conference.html Chain-of-thought prompting elicits reasoning...
2022
-
[60]
Yubo Xie, Junze Li, and Pearl Pu. 2021. https://doi.org/10.18653/V1/2021.ACL-SHORT.6 Uncertainty and surprisal jointly deliver the punchline: Exploiting incongruity-based features for humor recognition . In Proceedings of the 59th Annual Meeting of the Association for Computat...
2021 doi
-
[61]
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jin Xu, Jingren Zhou, Jinze Bai, Jinzheng...
2024 arXiv
- [62]
-
[63]
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al. 2024 c . Qwen2. 5 technical report. arXiv preprint arXiv:2412.15115
2024 arXiv
- [64]
-
[65]
Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, et al. 2024. Minicpm-v: A gpt-4v level mllm on your phone. arXiv preprint arXiv:2408.01800
2024 arXiv
-
[66]
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. 2023. Sigmoid loss for language image pre-training. In Proceedings of the IEEE/CVF international conference on computer vision, pages 11975--11986
2023
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.