REVIEW 4 major objections 4 minor 59 references
Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor
T0 review · 4 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The paper claims a 1.8B multimodal model can be trained to identify visual commonsense errors in distractor options and generate explanations that help correct them, and that a new GPT-4-built benchmark, VCR-DF, exposes a gap in existing…
desk verdict A solid new benchmark and model for explaining VCR distractors, but the closed GPT-4 loop means the central 'significantly outperforms' claim is only as strong as the teacher. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by two artifacts. First, the VCR-DF dataset: GPT-4 receives the image's event description, objects, question, and correct answer, classifies the question by Bloom's taxonomy level, generates five image-relevant distractors (manually filtered to the top three), and annotates each distractor with a misconception and an explanation; over 90% of these preliminary samples pass manual checks for accuracy and clarity. Second, the PEIFG model: a visual feature extractor that fuses region-level features from a SAM-based visual marker perceiver (trained with OPT-350M to read object boxes) with global CLIP features; a learnable pool of ten expert prompts from which the top three are selected via Q-Former instruction-aware features under a cosine-similarity key-matching loss and a correlation loss; and a QWen1.5 text generator with LoRA that ingests image tokens, expert prompts, question, answer, and distractor to produce feedback. A final refinement stage samples multiple outputs, scores them with five GPT-4 diagnostic questions, and applies direct preference optimization. The machinery's work is to ground error analysis in fine-grained object location and global scene context, channeled through specialist prompts.
What would settle it
Take a random sample of VCR-DF test items and have independent human experts judge whether each GPT-4 'misconception' label actually identifies a genuine error relative to the image content; if human-expert agreement with GPT-4's labels falls well below the reported 90% internal check, the benchmark's validity collapses and PEIFG's advantage may be an artifact of training and evaluating on the same GPT-4 target.
Extended reading notes
Core claim
The central claim is that error correction in visual commonsense reasoning is a distinct, learnable capability: given a distractor that conflicts with visual commonsense, a model can be trained to name the misconception and explain the error in a way that guides a learner toward the correct answer. The paper's evidence is that on the newly constructed VCR-DF benchmark, the proposed PEIFG model achieves the highest scores on BLEU, METEOR, ROUGE-L, CIDEr, and BERTScore for both feedback and distractor generation, surpassing much larger open LMMs (3B–18B parameters), and it scores higher than GPT-4V on automatic metrics while ranking highest on human-rated helpfulness and logical consistency. The paper further asserts that existing LMMs, trained for forward reasoning, cannot perform this correction without specialized prompts, and that the VCR-DF benchmark exposes this gap.
Load-bearing premise
The entire benchmark and the measured superiority of PEIFG rest on the assumption that GPT-4's generated distractors and feedback are a valid ground truth for visual commonsense errors—that GPT-4's judgments about what counts as a misconception and what explanation is correct are themselves correct.
Editorial extensions
If this is right
- VCR-DF offers a reusable evaluation that does not reward entity-overlap heuristics: on the original VCR distractors, maximum entity overlap already gives over 60% accuracy, whereas the new distractors force genuine visual commonsense understanding.
- Error-correction ability becomes a reportable axis of LMM evaluation, distinct from answer accuracy, and the paper's ablations show that specialized expert-prompt guidance matters more than model scale for this axis.
- A single trained model serves both directions: PEIFG's feedback-generation skill transfers to generating new image-grounded distractors, enabling a loop of distractor design followed by misconception diagnosis.
- The DPO refinement with GPT-4 diagnostic scoring is a reusable recipe for improving the faithfulness of generated explanations in other explanation-generation tasks.
Reading between the lines
- Because GPT-4 supplies the training labels, the DPO reward, and the automatic evaluation references, the reported gap between PEIFG and other LMMs may partly reflect distillation toward GPT-4's own annotation style; an independent human-authored reference set would test whether the gap survives.
- A natural transfer test is to apply the VCR-DF pipeline to other multiple-choice domains (science QA, driving scenes, medical imaging) to see whether the expert-prompt architecture generalizes beyond VCR.
- The paper's entity-overlap diagnostic suggests other vision-language benchmarks may overstate reasoning ability; re-scoring them with image-relevant distractors could reveal similar inflation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces VCR-DF, a benchmark for explainable feedback generation on visual commonsense reasoning distractors, and proposes PEIFG, a 1.8B-parameter LMM with a visual marker perceiver, CLIP features, a learnable expert prompt selector, and a DPO refinement stage. The dataset is constructed by using language-only GPT-4 to generate distractors and feedback, followed by manual filtering. Experiments compare PEIFG against several open LMM baselines and GPT-4V using automatic metrics and a small human evaluation, reporting that PEIFG outperforms existing LMMs on both feedback and distractor generation.
Significance. If the benchmark and evaluation are valid, this is a useful new task: error-correction feedback for visual commonsense reasoning is underexplored, and the proposed PEIFG architecture with expert prompt selection and vision-grounded features is clearly described and thoughtfully ablated. The authors release code, and the ablations indicate that the visual branches, the expert prompt selector, and the refinement step all contribute. However, the main evidential value of the paper depends on breaking out of the GPT-4 loop: the training data, the DPO reward, and the automatic references all come from GPT-4, while the human evaluation is small and lacks agreement statistics. The claim of visual grounding is also not yet established because the benchmark generator receives only textual annotations and never the image. With stronger independent evaluation and statistical rigor, the contribution would be solid.
major comments (4)
- [§3.1, §4.3.1, §5.3] The evaluation is closed around GPT-4: GPT-4 writes the VCR-DF distractors and feedback (§3.1.1–3.1.2), GPT-4 scores the five diagnostic questions used for DPO refinement (§4.3.1 and Table 11), and the automatic metrics in Tables 1–2 are computed against those same GPT-4-written references (§5.3.1). Because PEIFG is trained to imitate this exact text style, high BLEU/CIDEr/BERTScore can indicate stylistic imitation rather than better error correction, and a stronger general LMM that produces useful but differently worded feedback would be penalized. The human evaluation in B.5 is the only independent check, but it covers only 200 samples with 5 raters, reports no inter-annotator agreement or significance tests, and Table 7 shows GPT-4V at or above PEIFG on Fluency (1.78 vs. 1.74) and Relevance (0.93 vs. 0.88). Please provide a larger independent human evaluation with agreement statistics and significance tests, and an analysis that separates imitation quality from correction quality (for example, by scoring whether the feedback identifies the actual visual contradiction rather than matching GPT-4 wording).
- [§3.1, Table 9] The benchmark generator is language-only: the GPT-4 prompt in Table 9 provides Event, Object boxes, Place, Question, Answer, and Educational Level, but not the image itself. Consequently VCR-DF can only contain misconceptions that are detectable from the textual annotations, and any image detail not captured in the event/place text is invisible to the benchmark. This undermines the claim that VCR-DF measures visual commonsense reasoning and raises the question whether PEIFG's visual branches (VMP and CLIP, §4.1) are actually necessary; a text-only ablation or a sanity check that feeds a mismatched image should be reported to show that the generated feedback depends on visual content beyond the provided annotations.
- [Tables 1–2, §5.4] No standard deviations, confidence intervals, or significance tests are reported for any automatic metric. Several differences are small relative to the likely sampling noise (e.g., BLEU-1 46.53 vs. 46.47 for CogAgent in Table 1; BLEU-4 17.69 vs. 17.32 for K=1 vs. K=3 in Table 4), so the phrase 'significantly outperforms' in the abstract and §6 is not supported by the reported evidence. Please report variances across runs or seeds and perform paired significance tests for the main comparisons, including the GPT-4V comparison in Table 2.
- [§3.1.2, B.5] The manual quality checks are not reported with enough detail to validate the ground truth: the paper states that 'over 90% of the preliminary samples' pass the Accuracy and Clarity criteria (§3.1.2), but gives no annotator count, no operational definition of the criteria, and no inter-annotator agreement; the same is true for the 200-sample human evaluation in B.5, which also lacks significance tests. Without this information, neither the claimed benchmark quality nor the human-evaluation advantage over baselines can be independently assessed.
minor comments (4)
- [§3, §5.4, §5.2.2] There are several typos: 'distracors' in §3, 'percevier' in §5.2.2, 'assisted bv GPT-4' in §5.4.2, and 'Q-Fromer' in §4.2.1; these should be corrected to 'distractors', 'perceiver', 'assisted by GPT-4', and 'Q-Former'.
- [Table 10] In Table 10, the third generated distractor is labeled 'Distractor1' instead of 'Distractor3', which makes the prompt example confusing.
- [§4.3, §5.1] The model name is written inconsistently as 'QWen1.5' and 'Qwen1.5'; the standard spelling is 'Qwen1.5'.
- [§5.4.1] The claim that VisualGLM and LLaVA-v1.5 predict the 'understand' level for almost all samples would be easier to verify if the per-level confusion matrix or per-level accuracy were reported.
Circularity Check
Closed GPT-4 loop: the training labels, DPO reward, and automatic evaluation references all come from GPT-4, so the reported superiority measures imitation fidelity rather than independently validated error-correction ability.
-
fitted input called prediction
[Abstract; Section 3.1.2; Section 4.3.1; Section 5.3.1; Eq. (8)]
"we employ GPT-4 as a ``teacher'' to collect the explainable feedback dataset VCR-DF for error correction, which serves as a benchmark to evaluate the ability of LMMs to identify misconceptions and clarify reasons behind the error in VCR distractors toward final answers... we formulate five diagnostic questions for GPT-4 to ascertain whether the generated feedback meets the specified criteria... We evaluate the performance with eight standard metrics, including BLEU-(1 to 4), ROUGEL, METEOR, CIDEr, and BERTScore."
The model is fit, via the language-modeling loss in Eq. (8), to reproduce GPT-4-written feedback tokens, and the automatic metrics in Table 1 score the generated text against those same GPT-4-written references. The DPO reward in Section 4.3.1 is also GPT-4's own answers to the diagnostic questions. Thus the automatic 'significantly outperforms' claim reduces to 'outputs are closer to the GPT-4 training distribution,' not to an independently established error-correction ability. The only non-GPT-4 evaluation is the small human study in Table 7 (200 samples), where GPT-4V actually scores higher on Fluency (1.78 vs 1.74) and Relevance (0.93 vs 0.88), and no inter-annotator agreement or significance tests are reported.
full rationale
The central derivation chain is a supervised-generation pipeline: GPT-4 creates the VCR-DF references, the PEIFG model is trained with cross-entropy to output those references, GPT-4 scores the model's own samples for DPO refinement, and the automatic evaluation measures n-gram/BERT overlap with the original GPT-4 references. This is a closed loop in which the 'quality' being optimized and the 'quality' being measured are both defined by the same GPT-4 distribution. Consequently, the paper's headline automatic result is a measure of imitation fidelity to the teacher that generated the training data, rather than a demonstration of independent error-correction capability. The human evaluation is the only external anchor, but it is limited in size, lacks statistical safeguards, and is partly favorable to GPT-4V, so it does not break the loop for the benchmark-validity claim. Because the loop is partial (humans did filter samples and provide a small independent evaluation), the appropriate score is 6 rather than 8-10; the central claim still has some empirical content, but its automatic support reduces by construction to matching the GPT-4 reference distribution.
Assumptions & free parameters
free parameters (4)
- lambda1 and lambda2 =
0.1
- prompt pool size S =
10
- selected prompts K =
3
- DPO sampling hyperparameters =
800 samples, top-p=0.95, temperature=0.8
assumptions (4)
- ad hoc to paper GPT-4-generated misconceptions and explanations are valid ground truth for visual commonsense error correction.
- ad hoc to paper Bloom's taxonomy levels assigned by GPT-4 are meaningful for VCR questions.
- domain assumption N-gram metrics (BLEU, METEOR, CIDEr) are appropriate for evaluating explainable feedback.
- domain assumption The VCR source data (object boxes, events, places) is accurate enough for the task.
Cite this review
Pith. "Pith review of Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor." pith.science (2026). https://pith.science/paper/T5XCMKSY
@misc{pith2026241207801,
author = {Pith},
title = {Pith review of: Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor},
year = {2026},
howpublished = {\url{https://pith.science/paper/T5XCMKSY}},
note = {Machine review of arXiv:2412.07801}
}
read the original abstract
Large multimodal models (LMMs) have shown remarkable performance in the visual commonsense reasoning (VCR) task, which aims to answer a multiple-choice question based on visual commonsense within an image. However, the ability of LMMs to correct potential visual commonsense errors in the distractor upon their occurrence is yet under-explored. Drawing inspiration from how a human teacher crafts challenging distractors to test students' comprehension of the concepts or skills and assists them in identifying and correcting errors toward the answer, we are the pioneering research for LMMs to simulate this error correction process. To this end, we employ GPT-4 as a ``teacher'' to collect the explainable feedback dataset VCR-DF for error correction, which serves as a benchmark to evaluate the ability of LMMs to identify misconceptions and clarify reasons behind the error in VCR distractors toward final answers. In addition, we propose an LMM-based Pedagogical Expert Instructed Feedback Generation (PEIFG) model to incorporate the learnable expert prompts and multimodal instruction as guidance for feedback generation. Experimental results show that our PEIFG significantly outperforms existing LMMs. We believe that our benchmark provides a new direction for evaluating the capabilities of LMMs.
Figures
Reference graph
Works this paper leans on
-
[1]
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023. Qwen-VL: A Frontier Large Vision- Language Model with Versatile Abilities. CoRR abs/2308.12966 (2023)
arXiv 2023
-
[2]
Benjamin S Bloom, Max D Engelhart, Edward J Furst, Walker H Hill, David R Krathwohl, et al. 1956. Taxonomy of educational objectives: The classification of educational goals. Handbook 1: Cognitive domain . Longman New York
work page 1956
-
[3]
Meyer, Yuning Chai, Dennis Park, and Yong Jae Lee
Mu Cai, Haotian Liu, Siva Karthik Mustikovela, Gregory P. Meyer, Yuning Chai, Dennis Park, and Yong Jae Lee. 2023. Making Large Multimodal Models Under- stand Arbitrary Visual Prompts. CoRR abs/2312.00784 (2023)
arXiv 2023
-
[4]
Delong Chen, Jianfeng Liu, Wenliang Dai, and Baoyuan Wang. 2024. Visual Instruction Tuning with Polite Flamingo. In Proc. of AAAI, Michael J. Wooldridge, Jennifer G. Dy, and Sriraam Natarajan (Eds.). AAAI Press, 17745–17753
work page 2024
-
[5]
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C. Lawrence Zitnick. 2015. Microsoft COCO Captions: Data Collection and Evaluation Server. CoRR abs/1504.00325 (2015)
arXiv 2015
-
[6]
Zhenfang Chen, Rui Sun, Wenjun Liu, Yining Hong, and Chuang Gan. 2023. GENOME: GenerativE Neuro-symbOlic visual reasoning by growing and reusing ModulEs. CoRR abs/2311.04901 (2023)
work page Pith review arXiv 2023
-
[7]
Rui Dai, Joseph C Fritchman, Qiaoyi Liu, Yang Xiao, Haibo Yu, and Lei Bao. 2019. Assessment of student understanding on light interference. Physical Review Physics Education Research 15, 2 (2019), 020134
work page 2019
-
[8]
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven C. H. Hoi. 2023. In- structBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning. In Proc. of NeurIPS, Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (Eds.)
work page 2023
Show all 59 references
-
[9]
Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, José M. F. Moura, Devi Parikh, and Dhruv Batra. 2017. Visual Dialog. In Proc. of CVPR. IEEE Computer Society, 1080–1089
2017
-
[10]
Denkowski and Alon Lavie
Michael J. Denkowski and Alon Lavie. 2014. Meteor Universal: Language Specific Translation Evaluation for Any Target Language. In Proc. of ACL Workshop . 376–380
2014
-
[11]
Qingxiu Dong, Ziwei Qin, Heming Xia, Tian Feng, Shoujie Tong, Haoran Meng, Lin Xu, Zhongyu Wei, Weidong Zhan, Baobao Chang, Sujian Li, Tianyu Liu, and Zhifang Sui. 2022. Premise-based Multimodal Reasoning: Conditional Inference on Joint Textual and Visual Clues. In Proc. of AC...
2022
-
[12]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recogn...
2021
-
[13]
Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. 2022. GLM: General Language Model Pretraining with Autoregressive Blank Infilling. In Proc. of ACL, Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.). Association for Computat...
2022
-
[14]
Jiaxin Ge, Sanjay Subramanian, Trevor Darrell, and Boyi Li. 2023. From Wrong To Right: A Recursive Approach Towards Vision-Language Explanation. In Proc. of EMNLP, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, 1173–1185
2023
-
[15]
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh
-
[16]
Tanmay Gupta and Aniruddha Kembhavi. 2023. Visual Programming: Composi- tional visual reasoning without training. In Proc. of CVPR. IEEE, 14953–14962
2023
-
[17]
Wenyi Hong, Weihan Wang, Qingsong Lv, Jiazheng Xu, Wenmeng Yu, Junhui Ji, Yan Wang, Zihan Wang, Yuxuan Zhang, Juanzi Li, Bin Xu, Yuxiao Dong, Ming Ding, and Jie Tang. 2023. CogAgent: A Visual Language Model for GUI Agents. CoRR abs/2312.08914 (2023)
2023 arXiv
-
[18]
Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. In Proc. of ICLR. OpenReview.net
2022
-
[19]
Kingma and Jimmy Ba
Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Opti- mization. In Proc. of ICLR
2015
-
[20]
Berg, Wan-Yen Lo, Piotr Dollár, and Ross B
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloé Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, and Ross B. Girshick. 2023. Segment Anything. In Proc. of ICCV . IEEE, 3992–4003
2023
-
[21]
Junyan Li, Delin Chen, Yining Hong, Zhenfang Chen, Peihao Chen, Yikang Shen, and Chuang Gan. 2023. CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding. CoRR abs/2311.03354 (2023)
2023 arXiv
-
[22]
Junnan Li, Dongxu Li, Silvio Savarese, and Steven C. H. Hoi. 2023. BLIP-2: Boot- strapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models. In Proc. of ICML (Proceedings of Machine Learning Research, Vol. 202), Andreas Krause, Emma Brunskill, K...
2023
-
[23]
Lei Li, Yuwei Yin, Shicheng Li, Liang Chen, Peiyi Wang, Shuhuai Ren, Mukai Li, Yazheng Yang, Jingjing Xu, Xu Sun, Lingpeng Kong, and Qi Liu. 2023. M3IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning. CoRR abs/2306.04387 (2023)
2023 arXiv
-
[24]
Yunxin Li, Baotian Hu, Xinyu Chen, Yuxin Ding, Lin Ma, and Min Zhang. 2023. A Multi-Modal Context Reasoning Approach for Conditional Inference on Joint Textual and Visual Clues. In Proc. of ACL, Anna Rogers, Jordan L. Boyd-Graber, and Naoaki Okazaki (Eds.). Association for Com...
2023
-
[25]
Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Proc. of ACL Workshop. 74–81
2004
-
[26]
Fuxiao Liu, Kevin Lin, Linjie Li, Jianfeng Wang, Yaser Yacoob, and Lijuan Wang
-
[27]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual Instruc- tion Tuning. In Proc. of NeurIPS, Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (Eds.)
2023
-
[28]
Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022. Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering. In Proc. of NeurIPS 2022 , Sanmi Koyejo, S. Mohamed, A....
2022
-
[29]
Pan Lu, Baolin Peng, Hao Cheng, Michel Galley, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, and Jianfeng Gao. 2023. Chameleon: Plug-and-Play Com- positional Reasoning with Large Language Models. In Proc. of NeurIPS , Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Morit...
2023
-
[30]
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi
-
[31]
Carl Ed Murchison. 1930. A history of psychology in autobiography Vol. I. Russell & Russell/Atheneum Publishers
1930
-
[32]
OpenAI. 2023. GPT-4 Technical Report. CoRR abs/2303.08774 (2023)
2023 arXiv
-
[33]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a Method for Automatic Evaluation of Machine Translation. In Proc. of ACL . 311–318
2002
-
[34]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. In Proc. of IC...
2021
-
[35]
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog 1, 8 (2019), 9
2019
-
[36]
Manning, Stefano Ermon, and Chelsea Finn
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning, Stefano Ermon, and Chelsea Finn. 2023. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. InProc. of NeurIPS, Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hard...
2023
-
[37]
Fawaz Sammani, Tanmoy Mukherjee, and Nikos Deligiannis. 2022. NLX-GPT: A Model for Natural Language Explanations in Vision and Vision-Language Tasks. In Proc. of CVPR. IEEE, 8312–8322
2022
-
[38]
Dustin Schwenk, Apoorv Khandelwal, Christopher Clark, Kenneth Marino, and Roozbeh Mottaghi. 2022. A-OKVQA: A Benchmark for Visual Question Answer- ing Using World Knowledge. In Proc. of ECCV (Lecture Notes in Computer Science, Vol. 13668), Shai Avidan, Gabriel J. Brostow, Mous...
2022
-
[39]
Dídac Surís, Sachit Menon, and Carl Vondrick. 2023. ViperGPT: Visual Inference via Python Execution for Reasoning. In Proc. of ICCV. IEEE, 11854–11864
2023
-
[40]
Lawrence Zitnick, and Devi Parikh
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015. CIDEr: Consensus-based image description evaluation. In Proc. of CVPR. 4566–4575
2015
-
[41]
Kankanhalli, and Ying Shan
Guangzhi Wang, Yixiao Ge, Xiaohan Ding, Mohan S. Kankanhalli, and Ying Shan
-
[42]
Jiayuan Xie, Yi Cai, Jiali Chen, Ruohang Xu, Jiexin Wang, and Qing Li. 2024. Knowledge-Augmented Visual Question Answering with Natural Language Ex- planation. IEEE Transactions on Image Processing (2024)
2024
-
[43]
Shiyu Xuan, Qingpei Guo, Ming Yang, and Shiliang Zhang. 2023. Pink: Un- veiling the Power of Referential Comprehension for Multi-modal LLMs. CoRR abs/2310.00582 (2023)
2023 arXiv
-
[44]
Da Yin, Feng Gao, Govind Thattai, Michael Johnston, and Kai-Wei Chang. 2023. GIVL: Improving Geographical Inclusivity of Vision-Language Models with Pre- Training Methods. In Proc. of CVPR. IEEE, 10951–10961
2023
-
[45]
What Makes for Good Visual Tokenizers for Large Language Models?CoRR abs/2305.12223 (2023)
2023 arXiv
-
[46]
Ayyubi, Kai-Wei Chang, and Shih-Fu Chang
Haoxuan You, Rui Sun, Zhecan Wang, Long Chen, Gengyu Wang, Hammad A. Ayyubi, Kai-Wei Chang, and Shih-Fu Chang. 2023. IdealGPT: Iteratively Decom- posing Vision and Language Reasoning via Large Language Models. In Proc. of EMNLP Findings, Houda Bouamor, Juan Pino, and Kalika Ba...
2023
-
[47]
Tianyu Yu, Jinyi Hu, Yuan Yao, Haoye Zhang, Yue Zhao, Chongyi Wang, Shan Wang, Yinxv Pan, Jiao Xue, Dahai Li, Zhiyuan Liu, Hai-Tao Zheng, and Maosong Sun. 2023. Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants. CoRR abs/2310....
2023 arXiv
-
[48]
Li Yuan, Yi Cai, Haopeng Ren, and Jiexin Wang. 2024. A Logical Pattern Memory Pre-trained Model for Entailment Tree Generation. In Proc. of COLING, Nicoletta Calzolari, Min-Yen Kan, Véronique Hoste, Alessandro Lenci, Sakriani Sakti, and Nianwen Xue (Eds.). ELRA and ICCL, 759–772
2024
-
[49]
Da Yin, Liunian Harold Li, Ziniu Hu, Nanyun Peng, and Kai-Wei Chang. 2021. Broaden the Vision: Geo-Diverse Visual Commonsense Reasoning. In Proc. of MM ’24, October 28-November 1, 2024, Melbourne, VIC, Australia Jiali Chen et al. EMNLP, Marie-Francine Moens, Xuanjing Huang, Lu...
2021
-
[50]
Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona T. Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer...
2022 arXiv
-
[51]
Shilong Zhang, Peize Sun, Shoufa Chen, Min Xiao, Wenqi Shao, Wenwei Zhang, Kai Chen, and Ping Luo. 2023. GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest. CoRR abs/2307.03601 (2023)
2023 arXiv
-
[52]
Weinberger, and Yoav Artzi
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi
-
[53]
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. From Recognition to Cognition: Visual Commonsense Reasoning. InProc. of CVPR. Computer Vision Foundation / IEEE, 6720–6731
2019
-
[54]
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023. MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models. CoRR abs/2304.10592 (2023). Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoni...
2023 arXiv
-
[58]
Ge Zheng, Bin Yang, Jiajin Tang, Hong-Yu Zhou, and Sibei Yang. 2023. DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multimodal Reasoning in Lan- guage Models. In Proc. of NeurIPS, Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (Eds.)
2023
-
[2017]
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering. InProc. of CVPR. IEEE Computer Society, 6325–6334
-
[2019]
OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge. In Proc. of CVPR. Computer Vision Foundation / IEEE, 3195–3204
-
[2020]
BERTScore: Evaluating Text Generation with BERT. In Proc. of ICLR . OpenReview.net
-
[2023]
CoRR abs/2306.14565 (2023)
Aligning Large Multi-Modal Model with Robust Instruction Tuning. CoRR abs/2306.14565 (2023)
2023 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.