REVIEW 3 major objections 5 minor 67 references
Generative Visual Commonsense Answering and Explaining with Generative Scene Graph Constructing
T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read The paper claims that generating location-free scene graphs from image patches and object prompts, then weighting the graph triplets by CLIP confidence during LLM generation, yields more accurate visual commonsense answers and…
desk verdict A sensible two-stage scene-graph pipeline with a clean internal ablation, undermined by an underspecified object-list input for VCR that could be the real source of the gains. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the location-free scene graph: a set of triplets $\langle subject, relation, object\rangle$ generated left-to-right as a text sequence, without bounding boxes, from CLIP patch tokens and a prompt that lists candidate objects. Training on Visual Genome teaches the generator to propose triplets; at inference the triplets are scored by CLIP image-text similarity, each triplet's tokens are multiplied by its confidence score before attention, and the weighted triplet text is prepended to the question as context for the answer-and-explanation generator. Cross-modal fusion between visual and text embeddings uses a single-head attention plus a gated combination, and the scene graph triplet weight $\alpha_{ij}$ in equation (6) is the mechanism that lets the model down-weight unreliable triplets.
What would settle it
Run G2 on the VCR test set with the object prompt $X_o$ restricted to noun phrases from the question, and compare CIDEr and BERTScore with the paper's full-object-prompt results; if the margin over the no-scene-graph baseline vanishes, the improvement is carried by unstated external object knowledge rather than by generated scene-graph triplets.
Extended reading notes
Core claim
On its own terms, the paper establishes that explicitly turning a scene into object-relation-object triplets before answering makes an LLM ground its reasoning in image content rather than in language priors. On the VCR benchmark, G2 improves filtered CIDEr from 47.3 to 57.7 and filtered BERTScore from 81.9 to 91.1 over the strongest earlier generative baselines, and human raters judged 63.1% of filtered explanations as well-justifying their answers. The scene graph is generated, not oracle-supplied: a Llama-3.2 model trained on Visual Genome emits triplets from CLIP patch tokens and an object prompt, and these triplets are then passed, with confidence weighting, into the answer-and-explanation generator.
Load-bearing premise
The load-bearing premise is that a usable list of object names is available for every VCR image, yet the paper never specifies where that list comes from: VCR has no ground-truth object annotations, Section 3.2 only describes combining subjects and objects from the question, and the example object lists include names like 'cup' and 'dining table' that do not appear in the question.
Editorial extensions
If this is right
- If the central claim holds, generative vision-language explanation systems can be improved without bounding-box annotations or region proposals: patch-level CLIP features plus an LLM are enough to supply relational context.
- Confidence-weighted token input provides a trainable, threshold-free way to filter noisy structured knowledge, suggesting that soft weighting can replace hard filtering whenever a pretrained scorer can rate generated facts.
- The two-stage design means the scene graph generator can be trained once on a large relation-rich dataset and then reused across downstream VCR, VQA-X, and e-SNLI-VE tasks, which the paper reports with improved overall e-ViL scores.
- Because the method constrains the model to attend to triplets, generated explanations should name concrete objects and relations rather than generic reasoning; the paper's human evaluation claims 63.1% of filtered explanations justify the answer well.
Reading between the lines
- A natural extension, not pursued in the paper, is to treat the confidence-weighting trick as the reusable idea: any structured text whose reliability varies token by token — retrieved facts, knowledge-base triples, OCR output — could be slotted into the same soft-weighting mechanism.
- The paper's object-prompt ambiguity suggests a concrete check for readers: if an external detector supplies the object lists, then the fair comparison to prior work holds that detector fixed, and the 'location-free' claim reduces to relation prediction without boxes rather than recognition without boxes.
- A testable successor would replace CLIP with an LLM-based triplet plausibility score and see whether the gains compound; the ablation in Table 4 only varies the threshold, not the scoring model.
- The two-stage design also points to a production recipe: train relation extraction once on a relation-rich corpus, freeze it, and reuse it as a plug-in context provider for any downstream question-answering LLM.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes G2, a two-stage generative framework for Visual Commonsense Reasoning (VCR) under the VL-NLE setting. In the first stage, a location-free scene graph generator is trained on Visual Genome using CLIP patch features and Llama-3.2, with a text prompt that includes an object list. In the second stage, the generated scene graph triplets are fed, together with the image and question, into a second Llama-3.2 model that produces the answer and explanation; a CLIP-based confidence score for each triplet is used to weight the input tokens during training. The authors report experiments on VCR, VQA-X, and e-SNLI-VE, and claim that the scene-graph-enhanced pipeline outperforms prior generative VL-NLE baselines and that an automatic confidence-based selection mechanism is superior to threshold-based selection.
Significance. If the claims hold, the paper would provide a practical demonstration that LLM-generated, location-free scene graphs can improve generative visual commonsense answering and explanation, and the confidence-weighted token selection idea is a useful mechanism for injecting structured knowledge into a decoder-only model. The internal ablation in Table 4 and the visualizations in Figures 5-8 give some support for the central mechanism. However, the significance is limited by two issues: the provenance of the object-list input for VCR is unspecified, and the headline comparisons against prior work are confounded by backbone differences. The scene graph generation comparison against Pix2SG is also not controlled for the object-list input. These issues affect the strength of the main claim rather than only the presentation.
major comments (3)
- [3.2, Figure 2] The provenance of the object list X_o used as input to the scene graph generator for VCR images is unspecified. The text defines X_o as 'subjects and objects from Q', but for VCR the only Q is the question, e.g., 'what are person1, person3, and person6 doing?', while Figure 2 and the case studies show object lists containing 'cup', 'dining table', 'tie', and 'handbag' that do not appear in the question. VCR provides no ground-truth object annotations, and the paper does not describe any detector, CLIP-based naming step, or other mechanism that produces these objects for VCR images. Because the generated scene graph is the only new signal introduced by G2, the reader cannot determine whether the improvement in Table 2 comes from the scene graph itself or from object information that is derived from an undocumented oracle or from the ground-truth answer/explanation. Please specify exactly how X_o is obtained at inference time for VCR, and provide an ablation that either removes the object-list input or obtains it from a described, non-oracle source.
- [Table 3, Section 5.1] The location-free scene graph generation comparison against Pix2SG is not apples-to-apples. Pix2SG is a location-free SGG method that does not receive an object list, whereas G2 is given X_o, a list of object names, as part of the text prompt. Since this object list fixes the node vocabulary of the generated scene graph, the large improvements at R@50 and R@100 (29.93 vs. 24.81 and 44.76 vs. 26.66) may reflect the provided object set rather than better relationship prediction. The paper should report a variant of G2 that does not receive the object list, or an equivalent setting for Pix2SG, before claiming superiority in location-free SGG.
- [Table 2, Sections 4.2 and 4.4] The headline comparisons on VCR, VQA-X, and e-SNLI-VE are confounded by backbone and pretraining differences. G2 is initialized from Llama-3.2-1B, while the baselines e-UG, OFA-X, NLX-GPT, and UMAE use GPT-2 or OFA backbones. A newer and larger decoder can explain a substantial part of the gains in n-gram and BERTScore metrics, so the current Table 2 does not isolate the contribution of scene graphs. The 'G2 (w/o SG)' row in Table 4 is a useful start, but it should be included in the main comparison table, and the authors should add a same-backbone scene-graph-free baseline that reproduces the full G2 fusion and training setup, in order to support the claim that scene graphs are the source of the improvement.
minor comments (5)
- [Figure 3 and Section 5.2] The human evaluation is reported only for G2 (filtered and unfiltered), with no comparison to the baselines or to the ground-truth explanations, and no inter-annotator agreement measure, which makes the 63.1% 'well demonstrated' figure difficult to interpret.
- [Section 4.3] The choice of 0.92 as the BERTScore threshold for filtering 'correct' answers is presented without justification or sensitivity analysis; please report unfiltered scores as well, or at least cite a precedent that uses the same threshold.
- [Table 2] Several entries are missing or unexplained, including the n-gram scores for UMAEVCR and some baseline scores on e-SNLI-VE; the paper should state explicitly which numbers are unavailable and why.
- [Throughout] There are numerous typos and inconsistencies: 'as shwon' in the Introduction, 'instancess' in Section 4.1, 'carring' and 'Selecction' in Figure 2, 'SSG' for SGG in Section 5.1, 'his generated results' in Section 5.2, and inconsistent 'GenGen' versus 'G2' labels in Figures 5-8.
- [Table 1] The caption of Table 1 says 'that contain goals and relationships'; this should read 'objects and relationships'.
Circularity Check
No significant circularity: scene graphs, confidence weights, and generated answers/explanations come from separately trained or frozen models, not from the target outputs.
full rationale
The derivation chain is self-contained at the equation level. The scene graph generator is trained on Visual Genome with ground-truth object sets (Sec. 3.2) and is then applied to VCR; the triplet confidence scores are produced by a frozen CLIP model (Eq. 5); and the VCR answer/explanation model is trained with cross-entropy on the question, scene graph, and image patch inputs (Eq. 7). No target quantity is defined as a function of itself, and no fitted parameter is renamed as a prediction. The only overlapping-author citation is [49] for the gated cross-modal fusion mechanism, which is an architectural component and not load-bearing for the central claim. The paper does leave unspecified how the object prompt Xo is obtained for VCR images (Sec. 3.2 vs. Figure 2 lists objects such as 'cup' and 'dining table' that are absent from the question), which is a significant reproducibility/fairness concern and could, if the objects were taken from ground-truth answers, create leakage; however, the paper does not state that, and the published derivation does not exhibit a concrete reduction of the prediction to its inputs. Therefore no circularity is established.
Assumptions & free parameters
free parameters (3)
- VG scene graph count cutoff =
50
- Threshold selection values for SG filtering =
0.7, 0.8, 0.9
- BERTScore answer filter threshold =
0.92
assumptions (5)
- standard math Attention and gated fusion equations (Eq. 1-3) correctly compute cross-modal alignment.
- domain assumption The scene graph generation model trained on Visual Genome transfers to VCR, VQA-X, and e-SNLI-VE images.
- domain assumption CLIP normalized similarity scores between triplets and images reflect triplet quality and are useful attention weights.
- ad hoc to paper An object list for VCR images is available at inference time as input to the scene graph generator.
- standard math Language modeling cross-entropy losses (Eq. 4 and 7) are the appropriate training objectives.
Cite this review
Pith. "Pith review of Generative Visual Commonsense Answering and Explaining with Generative Scene Graph Constructing." pith.science (2026). https://pith.science/paper/64MJQ725
@misc{pith2026250109041,
author = {Pith},
title = {Pith review of: Generative Visual Commonsense Answering and Explaining with Generative Scene Graph Constructing},
year = {2026},
howpublished = {\url{https://pith.science/paper/64MJQ725}},
note = {Machine review of arXiv:2501.09041}
}
read the original abstract
Visual Commonsense Reasoning, which is regarded as one challenging task to pursue advanced visual scene comprehension, has been used to diagnose the reasoning ability of AI systems. However, reliable reasoning requires a good grasp of the scene's details. Existing work fails to effectively exploit the real-world object relationship information present within the scene, and instead overly relies on knowledge from training memory. Based on these observations, we propose a novel scene-graph-enhanced visual commonsense reasoning generation method named \textit{\textbf{G2}}, which first utilizes the image patches and LLMs to construct a location-free scene graph, and then answer and explain based on the scene graph's information. We also propose automatic scene graph filtering and selection strategies to absorb valuable scene graph information during training. Extensive experiments are conducted on the tasks and datasets of scene graph constructing and visual commonsense answering and explaining, respectively. Experimental results and ablation analysis demonstrate the effectiveness of our proposed framework.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016. SPICE: Semantic Propositional Image Caption Evaluation. In Computer Vision - ECCV 2016 - 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part V (Lecture Notes in Computer Science, Vol. 9909) , Bastian Leibe, Jiri Matas, Nicu Sebe, and Max Wellin...
-
[2]
Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, and Isabelle Augenstein
-
[3]
Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization. 65–72
2005
-
[4]
Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Scott Wen-tau Yih, and Yejin Choi. 2019. Abductive commonsense reasoning. arXiv preprint arXiv:1908.05739 (2019)
arXiv 2019
-
[6]
Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom
-
[7]
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020. End-to-end object detection with trans- formers. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part I 16 . Springer, 213–229
work page 2020
-
[8]
Ting Chen, Saurabh Saxena, Lala Li, David J Fleet, and Geoffrey Hinton. 2021. Pix2seq: A language modeling framework for object detection. arXiv preprint arXiv:2109.10852 (2021)
arXiv 2021
-
[9]
Advances in Neural Information Processing Systems 31 (2018)
e-snli: Natural language inference with natural language explanations. Advances in Neural Information Processing Systems 31 (2018)
work page 2018
Show all 67 references
-
[10]
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020. Uniter: Universal image-text representation learning. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXX ....
2020
-
[11]
Bo Dai, Yuqi Zhang, and Dahua Lin. 2017. Detecting visual relationships with deep relational networks. In Proceedings of the IEEE conference on computer vision and Pattern recognition. 3076–3086
2017
-
[12]
Vincent S Chen, Paroma Varma, Ranjay Krishna, Michael Bernstein, Christopher Re, and Li Fei-Fei. 2019. Scene graph prediction with limited labels. InProceedings of the IEEE/CVF International Conference on Computer Vision . 2580–2590
2019
-
[13]
Helisa Dhamo, Fabian Manhardt, Nassir Navab, and Federico Tombari. 2021. Graph-to-3d: End-to-end generation and manipulation of 3d scenes using scene graphs. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 16352–16361
2021
-
[14]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv prepri...
2020 arXiv
-
[15]
Helisa Dhamo, Azade Farshad, Iro Laina, Nassir Navab, Gregory D Hager, Federico Tombari, and Christian Rupprecht. 2020. Semantic image manipulation using scene graphs. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 5213–5222
2020
- [16]
-
[17]
Zhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu, Yu Cheng, and Jingjing Liu. 2020. Large-scale adversarial training for vision-and-language representation learning. Advances in Neural Information Processing Systems 33 (2020), 6616–6628
2020
-
[18]
Radhika Dua, Sai Srinivas Kancheti, and Vineeth N Balasubramanian. 2021. Be- yond vqa: Generating multi-word answers and rationales to visual questions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion. 1623–1632
2021
-
[19]
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh
-
[20]
Justin Johnson, Ranjay Krishna, Michael Stark, Li-Jia Li, David Shamma, Michael Bernstein, and Li Fei-Fei. 2015. Image retrieval using scene graphs. InProceedings of the IEEE conference on computer vision and pattern recognition . 3668–3678
2015
-
[21]
Peng Gao, Jiaming Han, Renrui Zhang, Ziyi Lin, Shijie Geng, Aojun Zhou, Wei Zhang, Pan Lu, Conghui He, Xiangyu Yue, et al . 2023. LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model. arXiv preprint arXiv:2304.15010 (2023)
2023 arXiv
-
[22]
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al
-
[23]
Bei Li, Chuanhao Lv, Zefan Zhou, Tao Zhou, Tong Xiao, Anxiang Ma, and Jingbo Zhu. 2022. On Vision Features in Multimodal Machine Translation. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 6327–6337
2022
-
[24]
Yikang Li, Wanli Ouyang, Xiaogang Wang, and Xiao’ou Tang. 2017. Vip-cnn: Visual phrase guided convolutional neural network. In Proceedings of the IEEE conference on computer vision and pattern recognition . 1347–1356
2017
-
[25]
Maxime Kayser, Oana-Maria Camburu, Leonard Salewski, Cornelius Emde, Vir- ginie Do, Zeynep Akata, and Thomas Lukasiewicz. 2021. e-vil: A dataset and benchmark for natural language explanations in vision-language tasks. In Pro- ceedings of the IEEE/CVF international conference ...
2021
-
[26]
Wentong Liao, Bodo Rosenhahn, Ling Shuai, and Michael Ying Yang. 2019. Natural language guided visual relationship detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops . 0–0
2019
-
[27]
International journal of computer vision 123 (2017), 32–73
Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of computer vision 123 (2017), 32–73
2017
-
[28]
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin transformer: Hierarchical vision transformer us- ing shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision. 10012–10022
2021
-
[29]
Cewu Lu, Ranjay Krishna, Michael Bernstein, and Li Fei-Fei. 2016. Visual rela- tionship detection with language priors. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceed- ings, Part I 14 . Springer, 852–869
2016
-
[30]
Yikang Li, Wanli Ouyang, Bolei Zhou, Kun Wang, and Xiaogang Wang. 2017. Scene graph generation from objects, phrases and region captions. In Proceedings of the IEEE international conference on computer vision . 1261–1270
2017
-
[31]
Ege Özsoy, Felix Holm, Tobias Czempiel, Nassir Navab, and Benjamin Busam
-
[32]
Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out. 74–81
2004
-
[33]
Dong Huk Park, Lisa Anne Hendricks, Zeynep Akata, Anna Rohrbach, Bernt Schiele, Trevor Darrell, and Marcus Rohrbach. 2018. Multimodal explanations: Justifying decisions and pointing to the evidence. In Proceedings of the IEEE Conference’25, 2025, Yuan et al. conference on comp...
2018
-
[34]
Dong Huk Park, Lisa Anne Hendricks, Zeynep Akata, Bernt Schiele, Trevor Darrell, and Marcus Rohrbach. 2016. Attentive Explanations: Justifying Decisions and Pointing to the Evidence. CoRR abs/1612.04757 (2016). arXiv:1612.04757 http://arxiv.org/abs/1612.04757
2016 arXiv
-
[35]
Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022. Learn to explain: Multi- modal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems 35 ...
2022
-
[36]
Yue Qiu, Yoshiki Nagasaki, Kensho Hara, Hirokatsu Kataoka, Ryota Suzuki, Kenji Iwata, and Yutaka Satoh. 2023. VirtualHome Action Genome: A Simulated Spatio-Temporal Scene Graph Dataset With Consistent Relationship Labels. In Proceedings of the IEEE/CVF Winter Conference on App...
2023
-
[37]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learni...
2021
-
[38]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics . 311–318
2002
-
[39]
Fawaz Sammani, Tanmoy Mukherjee, and Nikos Deligiannis. 2022. NLX-GPT: A model for natural language explanations in vision and vision-language tasks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 8322–8332
2022
-
[40]
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2019. Vl-bert: Pre-training of generic visual-linguistic representations. arXiv preprint arXiv:1908.08530 (2019)
2019 arXiv
-
[41]
Björn Plüster, Jakob Ambsdorf, Lukas Braach, Jae Hee Lee, and Stefan Wermter
-
[42]
Kaihua Tang, Yulei Niu, Jianqiang Huang, Jiaxin Shi, and Hanwang Zhang. 2020. Unbiased scene graph generation from biased training. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 3716–3725
2020
-
[43]
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. 2015. Cider: Consensus-based image description evaluation. In Proceedings of the IEEE confer- ence on computer vision and pattern recognition . 4566–4575
2015
-
[44]
Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. 2022. Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. In International Conference on Machine Le...
2022
-
[45]
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems 28 (2015)
2015
-
[46]
Zhecan Wang, Noel Codella, Yen-Chun Chen, Luowei Zhou, Xiyang Dai, Bin Xiao, Jianwei Yang, Haoxuan You, Kai-Wei Chang, Shih-fu Chang, et al. 2022. Multimodal Adaptive Distillation for Leveraging Unimodal Encoders for Vision- Language Tasks. arXiv preprint arXiv:2204.10496 (2022)
2022 arXiv
-
[47]
Zhecan Wang, Haoxuan You, Liunian Harold Li, Alireza Zareian, Suji Park, Yiqing Liang, Kai-Wei Chang, and Shih-Fu Chang. 2022. SGEITL: Scene graph enhanced image-text learning for visual commonsense reasoning. In Proceedings of the AAAI Conference on Artificial Intelligence , ...
2022
-
[48]
Kaihua Tang, Jianqiang Huang, and Hanwang Zhang. 2020. Long-tailed clas- sification by keeping the good and removing the bad momentum causal effect. Advances in Neural Information Processing Systems 33 (2020), 1513–1524
2020
-
[49]
Zhiyong Wu, Lingpeng Kong, Wei Bi, Xiang Li, and Ben Kao. 2021. Good for Misconceived Reasons: An Empirical Revisiting on the Need for Visual Context in Multimodal Machine Translation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics a...
2021
-
[50]
Ning Xie, Farley Lai, Derek Doran, and Asim Kadav. 2019. Visual entailment: A novel task for fine-grained image understanding. arXiv preprint arXiv:1901.06706 (2019)
2019 arXiv
-
[51]
Danfei Xu, Yuke Zhu, Christopher B Choy, and Li Fei-Fei. 2017. Scene graph generation by iterative message passing. In Proceedings of the IEEE conference on computer vision and pattern recognition . 5410–5419
2017
-
[52]
Tzu-Jui Julius Wang, Selen Pehlivan, and Jorma Laaksonen. 2020. Tackling the unannotated: Scene graph generation with bias-reduced models. arXiv preprint arXiv:2008.07832 (2020)
2020 arXiv
-
[53]
Fei Yu, Jiji Tang, Weichong Yin, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang
-
[54]
Alireza Zareian, Svebor Karaman, and Shih-Fu Chang. 2020. Weakly supervised visual semantic parsing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 3736–3745
2020
-
[55]
Chenxi Whitehouse, Tillman Weyde, and Pranava Madhyastha. 2023. Towards a Unified Model for Generating Answers and Explanations in Visual Question Answering. arXiv preprint arXiv:2301.10799 (2023)
2023 arXiv
-
[56]
Rowan Zellers, Jiasen Lu, Ximing Lu, Youngjae Yu, Yanpeng Zhao, Moham- madreza Salehi, Aditya Kusupati, Jack Hessel, Ali Farhadi, and Yejin Choi. 2022. MERLOT RESERVE: Neural Script Knowledge through Vision and Language and Sound. In IEEE/CVF Conference on Computer Vision and ...
2022
-
[57]
Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, and Yejin Choi. 2021. Merlot: Multimodal neural script knowledge models. Advances in Neural Information Processing Systems 34 (2021), 23634– 23651
2021
-
[58]
Hanwang Zhang, Zawlin Kyaw, Jinyang Yu, and Shih-Fu Chang. 2017. Ppr- fcn: Weakly supervised visual relation detection via parallel pairwise r-fcn. In Proceedings of the IEEE international conference on computer vision . 4233–4241
2017
-
[59]
Shaotian Yan, Chen Shen, Zhongming Jin, Jianqiang Huang, Rongxin Jiang, Yaowu Chen, and Xian-Sheng Hua. 2020. Pcpl: Predicate-correlation perception learning for unbiased scene graph generation. In Proceedings of the 28th ACM International Conference on Multimedia . 265–273
2020
-
[60]
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675 (2019)
2019 arXiv
-
[61]
G2 w/o SG
Zhuosheng Zhang, Aston Zhang, Mu Li, Hai Zhao, George Karypis, and Alex Smola. 2023. Multimodal chain-of-thought reasoning in language models. arXiv preprint arXiv:2302.00923 (2023). A Visualization cases We showcase more visualization cases in Figure 5, 6, 7, and 8. Generativ...
2023 arXiv
-
[63]
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. From recognition to cognition: Visual commonsense reasoning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 6720–6731
2019
-
[67]
Renrui Zhang, Jiaming Han, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hong- sheng Li, Peng Gao, and Yu Qiao. 2023. Llama-adapter: Efficient fine-tuning of language models with zero-init attention. arXiv preprint arXiv:2303.16199 (2023)
2023 arXiv
-
[2017]
In Proceedings of the IEEE conference on computer vision and pattern recognition
Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition . 6904–6913
-
[2020]
Generating Fact Checking Explanations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel R. Tetreault (Eds.). Association for Computational Linguis...
2020
-
[2021]
In Proceedings of the AAAI Conference on Artificial Intelligence , Vol
Ernie-vil: Knowledge enhanced vision-language representations through scene graphs. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 35. 3208–3216
-
[2022]
arXiv preprint arXiv:2212.04231 (2022)
Harnessing the Power of Multi-Task Pretraining for Ground-Truth Level Natural Language Explanations. arXiv preprint arXiv:2212.04231 (2022)
2022 arXiv
-
[2023]
arXiv preprint arXiv:2303.10944 (2023)
Location-Free Scene Graph Generation. arXiv preprint arXiv:2303.10944 (2023)
2023 arXiv
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.